Cook-First Skills: cook with your agent first, write the skill later
Tabla de contenidos
Nobody writes a recipe before they have cooked the dish.
Yet with AI agent skills we do it all the time. We install someone else’s recipe without ever setting foot in their kitchen, or we sit down to write a SKILL.md imagining how something should turn out when we have never actually done it.
I want to propose a different way of working. I call it Cook-First Skills:
🔑 Cook-First Skills: a skill is not designed, it is cooked. First you solve the task with your agent, then you write the skill from what happened in that session, and you only call it good once you have cooked it again in a second real session.
That last part is the one that matters most, and it has a name of its own: the second-session rule. A skill that has not been used in a second real session is not validated. It is written, which is not the same thing.
In this article you will find:
- What Cook-First Skills is and why it goes against the grain of skill catalogs
- The second-session rule and why it is the only validation that really matters
- The working cycle: distill, repeat, refine
- Why the
descriptionfield is a trigger and not a description - Real examples, inside and outside of code
What is Cook-First Skills? ¶
Cook-First Skills is a way of working with AI agents in which skills are born from real work sessions, are validated by repeating the task and are refined by correcting the skill itself, not just the output.
The agent doesn’t matter: Claude Code, Codex, OpenCode, Pi, Cursor, Antigravity. Neither does the type of task. It can be a test suite, a module or a section of a website. But also a slide deck, a video edit, a batch of invoices or reconciling a bank statement.
All of those sessions follow the same pattern:
- You ask the agent for a specific task.
- There is some back and forth: you correct it, you give it context, you figure out together what works.
- The task turns out well.
- You know you’ll have to do it again.
At step 4 most of us close the session and forget about it. Two weeks later we’ll have to explain everything all over again. Cook-First Skills proposes using that moment to write the skill:
Create a skill from what we have done in this session.
Capture the steps that worked, the corrections I made
and the output format we agreed on. Use skill-creator.
So far, nothing many people aren’t already doing. In the skills mini-course I called it SDSD, Session-Driven Skill Development: you have the conversation once and you ask skill-creator to turn it into a skill. Cook-First takes that same gesture to the next level, with a complete cycle and a rule to know when a skill is truly validated.
What turns this into a method is what comes next.
Cook-First on /skills
The best skill isn't downloaded. It's cooked.
The cook, distill, repeat and refine cycle with copy-ready prompts, the 5 essential skills with their command, and how your agent decides which skill to load.
See the cycle on /skillsWhat is the second-session rule? ¶
A skill is not validated until you have used it in a second real session.
It sounds obvious, but it goes against almost everyone’s intuition. When the agent hands you a clean, well-structured, good-looking SKILL.md, the temptation is to call it done. You may even ask it to review it and it tells you it’s great.
That validates nothing. It’s like reading a recipe and declaring the dish delicious.
The second session is the one that tells you the truth, because the task never repeats exactly the same way. Another file, another client, another date, another nuance. That’s where you find out whether the skill captures the craft or just captures that particular day.
💡 If you take away only one thing from this article: a skill isn’t validated because it’s well written. It’s validated when you do the task again and have to correct the agent less than the previous time.
There’s a corollary I find just as important: when you correct the agent in that second session, fix the recipe, not just the dish. If you fix the output and move on, next time it will make the same mistake. If you fix the skill, it won’t.
Why cook first instead of installing other people’s recipes? ¶
Because someone else’s recipe can be excellent and still not be yours.
Catalog skills are written for someone else’s kitchen: their tools, their quirks, their conventions. Skills designed from scratch have a different problem: they describe how you think the task is done, not how you actually did it.
It’s the difference between a priori and a posteriori knowledge. The first two are a priori recipes. Cook-First produces a posteriori recipes: they are born from experience and corrected by it.
| Approach | Where the recipe comes from | Fit with your kitchen | Main risk |
|---|---|---|---|
| Catalog skill | Someone else’s work | Low at first | Instructions that don’t match your context |
| Skill designed from scratch | Your idea of how it should turn out | Medium | Writing for an imaginary case |
| Cook-First Skills | A real session that worked | High from day one | The recipe only works with that day’s ingredients |
I’m not telling you to stop installing other people’s skills. I’m telling you that the best skills you’ll ever have are the ones that come out of your own kitchen. And that you can make someone else’s skill your own with this same cycle: use it in a real session, correct it, and rewrite it.
The funny thing is that Anthropic’s own skill-creator already contemplates this path. In its intent-capture phase it anticipates that the current conversation may already contain the workflow you want to capture, and in that case it says to first extract from it the tools used, the sequence of steps and the user’s corrections. Cook-First doesn’t invent the technique: it turns it into a habit and gives it a validation rule.
How does it fit with Spec-Driven Development? ¶
If you follow Web Reactiva you know I spend a lot of time on SDD. And you may wonder whether this contradicts it. Quite the opposite: they complement each other because they look in opposite directions.
- The spec is the menu. It is written before the agent works and defines what has to leave the kitchen this time.
- The skill is the recipe. It is written after the agent has worked and captures how to do it next time.
One is prospective and the other retrospective. In fact, many of the skills I use to work with specs were born from sessions in which I was writing specs.
What is the Cook-First Skills cycle? ¶
The cycle has three verbs: distill, repeat, refine. Distilling happens in the first session; repeating and refining, in every session after that.
Distill: write the skill once you’re done cooking ¶
Solve the problem without thinking about the skill. Don’t try to make the session “look nice”: the valuable part is the stumbles and the corrections, because they are exactly what the agent didn’t know how to do on its own.
When you’re done, ask it to write the skill. It helps a lot to have skill-creator installed, or any skill that knows how to create skills. It’s not essential, but it knows the anatomy of a SKILL.md, progressive loading and the typical mistakes.
A trick that works for me: before asking for a specific skill, ask it which skills it sees in the conversation. Sometimes it finds patterns you hadn’t even put into words.
In this step you also decide where the recipe lives:
- At project level, versioned with the code, if the knowledge belongs to that project: how this app is deployed, how these tests are organized.
- At user level, available across all your projects, if the knowledge is yours: how you like to review a PR, how you plan your week.
Repeat: apply the second-session rule ¶
The next time that task comes around, do it with the skill loaded. Don’t reread the skill to see if it “looks right”. Use it and watch where you have to correct the agent.
Those corrections are gold. They tell you exactly what the recipe is missing.
Refine: fix the recipe, not just the dish ¶
When you have to correct it, ask the agent itself to update the skill:
We used the pending-work skill and I had to correct you twice:
I don't want to see tasks tagged "someday" and I want the
project name in front of each task. Update the skill so it
doesn't happen next time.
Markdown files aren’t set in stone. Every real session is an iteration: the recipe adjusts to your kitchen, one annotation after another.
And when should you split a skill into several? ¶
Sometimes a session doesn’t yield one recipe, it yields several. Just as a sofrito ends up being the base for ten different dishes, part of a skill may end up deserving a life of its own.
Deciding how to split a skill into different responsibilities is an engineering decision, like deciding when to split a class in two. My advice: split when it hurts, not before. When the skill gets heavy or when you see a part being reused in another context. Splitting before repeating is designing for imaginary cases.
Why is the description field a trigger and not a description? ¶
This is the technical point where most skills fail.
The description field in the frontmatter is not there to describe the skill. It’s there so the agent knows when to invoke it. It’s what lets you skip typing /skill-name every time and just say “I want to do such-and-such”, where “such-and-such” is a repeat of something you already did another day.
The skill-creator documentation is clear about it: the description is the main activation mechanism, and since models tend to under-invoke skills, they recommend writing it a bit “pushy”.
Cook-First gives you a huge advantage here: you have the exact sentence you started the session with. If the session kicked off with “what do I have pending today?”, that sentence has to be in the description.
---
name: pending-work
description: Review my pending work in the task manager and summarize what needs attention. Use this whenever I ask "what do I have pending", "what's on my plate today", "what's due today" or want to plan my day, even if I don't mention the task manager by name.
---
# Pending work
1. Fetch open tasks from the task manager, "Web Reactiva" project first.
2. Group them: overdue, due today, this week, no date.
3. Flag tasks untouched for more than 14 days.
4. Output a short list (max 10 items), most urgent first.
## Margin notes
<!-- Margin notes: they grow with every real session in which you correct the agent -->
- Skip tasks tagged "someday".
- Show the project name before each task.
Look at the last section. It’s not mandatory, but it’s the Cook-First pattern par excellence: the recipe’s margin notes. Like those family recipes with notes scribbled in pen (“less salt”, “better 10 more minutes”). They make visible how the skill has evolved and help you spot when a note has become important enough to move up into the body of the recipe.
⚠️ Watch what slips into the recipe. A real session can contain client names, internal paths, tokens or personal data. Review the skill before saving it, especially if it’s a project skill that will end up in a shared repository.
What validation levels exist beyond the second session? ¶
The second-session rule is the core, but it’s not the only quality check. There are four levels, from cheapest to most expensive.
Level 0: the first session. The task turned out well before the skill existed. You know the dish is possible.
Level 1: the second real session. You test it yourself, in your kitchen, with different ingredients. It’s the most valuable validation and the one that defines the method.
Level 2: the blind tasting. skill-creator lets you run evaluations: it generates test prompts, runs the task with and without the skill and helps you compare results. It can also optimize the description by testing which phrases trigger it and which don’t. It’s useful when the output is verifiable (file transformations, generated code, fixed-step workflows). For subjective outputs, such as a writing style, your palate still wins.
Level 3: the jury. Launch several agents, or the same one with different models, to cook with your recipe and send you back recommendations. Agents understand what a skill is and can tell you where they got stuck, which instruction was ambiguous or which step was unnecessary.
| Level | Cost | What it detects best |
|---|---|---|
| 0. First session | None | That the task is possible |
| 1. Second real session | Low | Whether the recipe holds up with different ingredients |
| 2. Blind tasting (evals) | Medium | Regressions and activation failures |
| 3. Jury of agents and models | High | Ambiguities and dependence on a specific model |
🛡️ Don’t skip level 1 to go straight to 2 or 3. A skill that passes synthetic evals but has never been used in a second real session is validated against made-up cases.
The skills that pay off most are the ones that come out of your own kitchen, and you learn that by sharing. Every Sunday we gather 12 resources on AI tools and productivity for developers (newsletter in Spanish). We're already 7,200+.
Apúntate gratis →What happens once the recipe is written? ¶
You no longer have to cook it yourself.
A skill with a good description can be invoked by any process that knows it exists: a scheduled agent that runs every morning, a loop that processes issues one after another, a subagent that reviews every PR. It’s the jump from the dish you prepare yourself to the daily menu that leaves the kitchen without you watching.
What started as “do this for me now” ends up as “do this every time that happens”. And without writing a new script, because the instructions are already written and tested in real sessions.
What real examples of Cook-First Skills are there? ¶
Let’s look at concrete cases. Some about code and some not, on purpose.
Pending work ¶
Every morning you ask the agent what you have pending. It has access to your task manager, or you can give it access. The first time you explain which projects to look at, which tags to ignore and how you want the summary. The second time you don’t want to explain it anymore.
That recipe is yours. You don’t need the one someone else published with their task manager, their tags and their quirks. You need the one that knows yours.
The trip planner ¶
This is one of my favorites. When I plan a road trip I want to know things no app gives me together: the best roadside stops, where the cheapest gas station of a specific brand is, and what the weather will be like right when I expect to pass each point, not the weather at the destination city.
The first time it was a long conversation, with searches, corrections and discarded formats. Out of it came a skill that now generates the route with times, weather per leg, stops and fuel. I didn’t design it. I cooked it first.
The review pattern ¶
In a development workflow you discover a pattern, or think you have one, or you ask the agent whether it sees one. A typical case: how you want the agent to review its own work before calling it done. Which criteria it applies, how it answers you, which tests it runs and what evidence it shows you.
You fine-tune it session after session until you realize you’re repeating it from memory. That’s the moment to write the skill. Some of my review skills started like that: a panel with different lenses (pre-mortem, architecture, over-engineering) that was born from asking the agent, many times in a row, to look at the same change from different angles.
The recipe that annotates itself ¶
There’s a variant I find to be the purest expression of the method: including the refinement inside the skill itself. I have a skill for solving GitHub issues whose last step is a reflection: what went wrong in this run and what should change in the instructions. And it applies that to its own SKILL.md.
It’s not magic and you should review those changes the way you’d review a PR. But the cycle no longer depends on you remembering to refine.
Beyond code ¶
The same rules apply to:
- Putting together the slides for a talk with your usual structure.
- Editing a video following your usual cuts, captions and pacing.
- Generating invoices from logged hours.
- Processing a bank statement and categorizing transactions the way your accountant does.
If you’ve cooked it once with an agent and you’re going to cook it again, it’s a candidate.
What mistakes should you avoid? ¶
These are the stumbles I see most often (and have made most often):
- The recipe that only works with that day’s tomatoes. The skill describes exactly what happened in that session, with that file and that client. The second session breaks it. skill-creator itself insists on generalizing from feedback rather than adding very specific patches.
- Writing recipes you never cook again. You create twenty skills in a week and don’t repeat any of them. That’s not Cook-First, that’s collecting.
- Vague description. “Helps with tasks” triggers nothing. Use the real phrases you use to ask for that task.
- Splitting preparations too early. Five ten-line skills that are always used together are one skill with an identity problem.
- Not reviewing what slipped in. Sensitive data, absolute paths from your machine, client names.
- Calling the recipe finished. A Cook-First skill is alive. If it hasn’t changed in three months and you use it often, either it’s perfect or you’ve stopped correcting the agent when it fails.
How do I start tomorrow? ¶
A two-week experiment:
- Install skill-creator (or whatever skill-creation skill your agent uses).
- Pick a task you do with an agent at least once a week.
- Next time you do it, when you’re done, ask: “which skills do you see in this session?”.
- Write the one that makes the most sense, with the
descriptionin your real words. - The following week, apply the second-session rule: repeat the task with the skill loaded and fix the recipe, not just the dish.
After three or four iterations you’ll have something no catalog can give you: instructions that know your way of working because they came out of it.
🔑 Cook-First Skills in one sentence: cook with the agent, write the skill, and don’t call it good until you’ve cooked it again.
Skills · Web Reactiva
Missing skill-creator to get started?
On /skills you'll find it with its command, alongside four other essentials, plus the anatomy of a SKILL.md explained line by line.
Go to /skillsFrequently asked questions ¶
What is Cook-First Skills? ¶
It’s a way of working with AI agents in which skills are born from real work sessions. First you solve the task with the agent, then you turn it into a SKILL.md and validate it by using it in a second real session, refining it every time you have to correct the agent.
What is the second-session rule? ¶
It’s the core principle of the method: a skill isn’t validated until you’ve used it in a second real session. Reviewing the text or asking the agent to assess it is no substitute for doing the task again with different data and checking whether it holds up.
Does it work with any AI agent? ¶
Yes, with any agent compatible with the Agent Skills standard: Claude Code, Codex, OpenCode, Cursor and others. The SKILL.md format is open, so the recipe you write in one usually works in the others with minimal tweaks.
Do I need skill-creator? ¶
It’s not essential, but it’s recommended. Any agent can write a SKILL.md if you explain the format. skill-creator brings best practices, automated evaluations and description optimization, which saves iterations.
How is it different from Spec-Driven Development? ¶
The spec is written before the agent works and defines what to build this time. The Cook-First skill is written afterwards and captures how to repeat it. One is the menu and the other the recipe, and they combine well.
When is it worth turning a session into a skill? ¶
When you know the task will repeat and during the session you had to give context or corrections you don’t want to give again. If the task is a one-off or the agent handles it well without help, there’s no need.
Where do I store the skill, in the project or at user level? ¶
If the knowledge depends on the project (its architecture, its deployment, its conventions), in the project. If it depends on you (how you review, plan or write), at user level.
How do I write a description so the skill triggers on its own? ¶
Describe when it should be used, not what it is. Include the real phrases you use to ask for that task, in the languages you ask in, and the contexts in which it should trigger even if you don’t name it.
When should I split a skill into several? ¶
When it becomes hard to maintain or when a part is reused in another context. Don’t do it in the first version: split when it hurts, not before.
Can a skill improve itself? ¶
It can include a final reflection step in which the agent proposes or applies changes to its own SKILL.md. It works, but you should review those changes as you would a PR to keep the skill from drifting toward odd instructions.
Sources ¶
12 recursos para developers cada domingo en tu bandeja de entrada
Además de una skill práctica bien explicada, trucos para mejorar tu futuro profesional y una pizquita de humor útil para el resto de la semana. Gratis.