Everyone Is Guessing About AI. You Can Stop.
Prompt tweaks never gave me a signal. Dialing in espresso did.
Have you ever experienced the magic of an LLM doing exactly what you wanted? Landing exactly where you planned. It feels game-changing in that moment. But your next session lands somewhere else, and you can’t say why.
"Prompt Engineering" promised magic words and clever phrases. That never happened — at least, it never happened consistently. So you reframe the conversation, add more context, more detail, a clearer expectation. You may have even explicitly given the exact steps. And yet, you find yourself ending up in the same place. A valley of the unknown, surrounded by a pile of keywords and phrases. The advice optimized the wording when the real problem was the guessing.
Iterative exploration yields stable feedback loops you can act on.
What guessing looks like
In the last three years, I got really deep into making coffee. My beverage of choice is an espresso drink called a Cortado. Two ounces of milk, two ounces of espresso. On the surface, this seems like an extremely simple recipe. This could not be further from the truth. I found myself swimming in a sea of options. Temperature, pressure, water, time, dose, flow, grind, bean, roast. The list goes on and on.
As any good novice does, I decided to pick something and go from there. Changing variables, hoping. Sometimes things were trending in the right direction. Other times, it was a complete miss. Before long, I burned my way through an entire $20 bag of coffee, learning nothing.
The entire time, I was never quite able to put my finger on what was causing the changes. I was changing everything at once, so when a shot improved, I couldn't say which dial had done it, and when it got worse, I couldn't say that either. You can't learn from this.
What changed?
The skill wasn't knowing the right settings. It was having a known starting point and changing one thing at a time. A finer grind, then taste. More dose, then taste. Hotter water, then taste.
If the shot got better, I knew why. If it got worse, I knew why. Over time, those small changes let me converge on what made a good shot. Now when I open a new bag, I don't guess. I start from the last good one and ask what's different. A lighter roast, so hotter and finer. A bag still isn't solved on the first shot, but I'm never starting from zero.
The same fix, applied to AI
The issue was never the prompt, the extra context, or the steps I spelled out. It was that I was trying to get everything right up front, in one big push, with nowhere specific to fix things when it went wrong.
The fix was the same one coffee gave me. Break the work into steps, give each step one job, and change one step at a time. The content changes with every project. The order doesn't. And because each step stands on its own, I can improve one without breaking the rest.
When developing my workflow with agents, I spent a lot of time reading other frameworks, interchanging their prompting, and seeing the effects. I could rapidly test my improvements side by side. Things I liked, I kept. Parts I found cumbersome, I removed. If I spotted a gap in how I wanted agents to behave, I created a skill or instruction to improve that part.
I want to emphasize that I didn't adopt someone else's entire process wholesale. Instead, I used their process to inform my own.
What it looks like in practice
The frameworks I studied were the popular ones, and I still look at new ones as they show up: Claude-mem, compound engineering, superpowers, mattpococks/skills. Each took a different approach, and some were more manual than others, but they all had similar parts. Research, design, proposals and specs, testing, and shipping.
In my own version, those became research, proposals and questions, a unified vision (a shared picture of what done looks like), and a named outcome. Then comes building, testing, and my own hands-on check before anything ships.
Take research as an example. Each of the frameworks solved for it, some by handing you a summary and others by handing you the evidence, but it was always part of the process. I wanted my research grounded in facts and shaped like a white paper: here's what we know, here's where it came from, and you decide what it means.
I did this for every step. I defined what I wanted from it, what it should produce, and something I could check on purpose. The research document, the set of options, and the checklist I use to test the finished work were each an artifact I could read and judge for myself. Those artifacts became my signal. When something was off, I didn't have to guess where, because the artifact told me which step to adjust.
The guessing stopped.
Every one of those first shots went down the drain. It felt like a waste. It was tuition. I was guessing at everything, so nothing could teach me anything.
One change, then taste. Every shot, every bag.
AI works the same way. I start from an order I trust, and every step leaves something I can read and judge for myself. When the result is off, I know where to look. The floor still moves. I'm still dialing in. But I'm not guessing anymore.
Iterative exploration yields stable feedback loops you can act on.
Everyone is guessing about AI. You don’t have to.