all writing
field note · 7 Sep 2026 · 5 min read

Same AI. Four results.

Most using AI all wrong: driving up costs, happy about low quality outcome.

techai for businessadoption curvecoststrategytoolsaiLLM

A knife will not fell a tree. A chainsaw will not trim a bonsai. Neither tool is bad. Each is simply wrong for the job in front of it, and the person holding it gets a poor result and blames the tool.

AI is the most capable tool most of us have ever been handed. And most people are using it like a knife on an oak: sawing away, frustrated, half-convinced the thing is overhyped.

Here is what I have come to believe after building an entire company with it. AI is not one tool. Depending on how you wield it, it is four different tools, and the gap between them is not linear. It is exponential.

The curve almost nobody is on

Picture a simple distribution. On the left, where the crowd stands, results are poor. On the right, where almost no one stands, results are extraordinary. The population thins fast as you move right. The quality climbs far faster than the population falls.

The shape is lopsided. Most people cluster on the left and get a fraction of what the tool can do. A smaller group does noticeably better. A smaller group still does better than that. And a thin band on the right gets results that look like a different product entirely, many times the bottom, not because they hold a better model, but because they use it differently.

You have seen this shape before. It is the technology adoption curve. The people getting the worst from AI are its laggards and late majority, using it grudgingly, the way people use any new tool they have not bothered to learn. The people getting the most are its innovators, a thin band on the right. Same curve, pointed at a single tool instead of a whole market.

Same model. Four results. Let me walk the four rungs.

Bad: a wish, not an instruction

“Write my CV.” “Summarise this.” “Make it better.”

No context, no objective, no material. A vague wish handed to a probabilistic machine, which does the only thing it can with a gap: it fills it. It invents the average of everything it has ever read and hands it back with total confidence. You get polished emptiness, and you conclude AI is not that clever after all.

This is the knife on the tree. Roughly 70% of all usage lives here.

Good: a precise instruction

The same request, done properly. A precise prompt. A clear objective. Real context and supporting material. “Rewrite this CV for this job description. Here is my history. Keep every fact. Flag anything you cannot support.”

Now the machine has something to work with, and the output is genuinely useful. This one move puts you ahead of most people alive. But you are still assembling the instruction by hand, from scratch, every single time.

Better: skills, not prompts

The next rung stops treating each task as a blank page. A skill is a reusable set of prompts with the right context built in, so the good version happens the same way every time without you rebuilding it.

You are no longer a clever person typing a clever request. You are running a small, repeatable process. The floor rises. Consistency arrives. This is the power tool, not the hand tool.

Best: the vertical application

The top rung is not a chatbot at all. It is a purpose-built application wrapped around the model: consistent, on-demand context pulled in automatically, and, the part that matters most, AI paired with deterministic code. The model does the judgement. Hard rules do the checking. Nothing important is left to a dice roll.

This is the workshop, not the tool. The saw, the jig that holds the wood square, the guard that stops the mistake before it happens.

What you actually gain by climbing

The four rungs are not simply “nicer output.” Three things improve together, all the way up.

Hallucination falls. More context and deterministic checks leave the model less room to invent.

Cost falls. A precise, well-scoped task burns a fraction of the tokens of a sprawling one you keep re-explaining.

Rework falls. The higher the rung, the less you fix, redo and babysit. At the top it is right the first time.

That is why the results curve bends the way it does. You are not just prompting better. You are removing the three things that make AI feel unreliable in the first place.

I built the same craft at all three levels

This is not theory for me. With JobMentis, my career tool, I shipped the exact same craft at three rungs, on purpose.

  • Good: copy-paste prompt guides. The prompts, free, for any chat model. They work, as long as you supply the context and check the output yourself, every time.

  • Better: open-source Claude skills. The same craft packaged as installable skills, with memory living in local files. Manual, one run per job, but repeatable.

  • Best: JobMentis itself. A persistent profile that already knows your history, context matched to every role automatically, and a groundedness reviewer, plain deterministic code, that flags every claim the model cannot source.

The prompt text barely changes between the three. What changes is leverage, and the leverage is exponential.

So where are you on the curve

Most people reading this are one deliberate step from a far better result. If you are giving AI wishes, start giving it instructions. If your instructions are already good, turn the best one into a skill. If you live in skills, ask what a deterministic check, or a real application, would do for the tasks you repeat every week.

AI is not magic. It is a fantastic tool with genuine power. But a chainsaw in the hands of someone trimming a hedge is just a dangerous way to do a small job badly. The whole game, the only game, is matching the tool to the work in front of you.

Pick the right one. Then climb a rung.