Article
Prompt debt: what happens after you build the library

Prompt debt is the accumulated cost of a prompt library nobody has gone back to: personas that stopped earning their keep, structures built around a model that has since moved on, ownership nobody claimed. KINTAL first wrote about this in April. Four model releases later, the cause is clearer than it was then: most creative teams built their prompting habits around the limits of an earlier generation of models, and those habits are still riding along in prompts now pointed at models that outgrew the need for them. The fix is matching the amount of scaffolding to the model reading it.
Prompting is usually the first skill people build when they start learning AI, and unlike most skills, it does not stay learned. The models change, best practice shifts, and what worked six months ago can quietly start producing results that are off, without anyone knowing why.
What is prompt debt and why does it matter?
With every team I work with, prompting practice sits somewhere on a spectrum. Some teams have a library and use it consistently. Others have one that individuals have quietly adapted to their own needs, so the shared resource and the actual practice have drifted apart. Some teams have no shared prompting practice at all, people do whatever works for them, and none of it gets captured. Prompt debt accumulates in all three situations. The difference is how visible it is.
The term borrows from software development's technical debt: the accumulated cost of shortcuts and decisions that made sense at the time but have not aged well. Prompts written twelve months ago were often built around techniques that have since been formally downgraded. Sander Schulhoff's research, co-authored with teams at OpenAI, Google and Stanford, found that role-based prompting, the "act as an expert copywriter" instruction most of us have used, shows statistically insignificant accuracy improvement on analytical tasks. It still works for expressive and stylistic tasks, which is most of what creative teams do, but the nuance matters. A library built on the assumption that persona framing is always the right move is now working harder than it needs to.
Is a persona the same as a role?
Not the same thing, and the gap between them is where most of the argument above lives.
A decorative persona, "you are a world-class business strategist with thirty years' experience", rarely earns its keep. A capable model already holds a reasonable model of how a strategist approaches a problem, and the claim hands it no information it did not already have. A functional role is a different instruction: "review this as the finance director responsible for protecting cash flow, prioritise payment terms and commitments, do not assess the creative quality of the proposal." That sentence sets the perspective, states what the model is responsible for, tells it what to prioritise and marks what falls outside its remit, four jobs decoration never does.
The test is simple. Remove the role and ask whether a reasonable model would produce materially different work without it. If the answer is no, it was decoration and can go. If the answer is yes, because the role changes the perspective, the responsibility, the authority or the professional standard being applied, it stays.
OpenAI's own guidance draws roughly the same line. Its GPT-5 prompting guide is almost entirely outcome-focused: goal, context, constraints, format. Its separate Realtime Prompting Guide, written for voice agents, still gives personality and tone a full section, because a customer-facing agent's warmth and pacing are part of the product a customer experiences. The role earns its place when the perspective changes the answer, not as a compliment paid to the model.
Why does a prompt library stop working?
Model behaviour keeps shifting under it. A prompt written for GPT-4 in early 2024 behaves differently on GPT-5.5, Opus 4.7 or whatever a team is running today. Register, structure, tone and level of inference all vary by model and by version, and teams managing multi-model workflows are dealing with an operational problem that still has very little practical guidance behind it.
Creative prompts degrade in ways enterprise governance frameworks tend to miss, because creative prompts are tied to things that move. Brand voice evolves, visual direction shifts, and a prompt that captured a client's tone accurately eighteen months ago may now be producing work that feels slightly off-register. If nobody owns the library, nobody is checking whether it still reflects where the brief has got to.
That question of ownership is where the published conversation goes quiet for creative teams specifically. CIO.com described prompts as "unowned infrastructure" earlier this year, and the description holds. In tech and product teams, governance frameworks are starting to assign prompt ownership to specific functions. In creative agencies, production houses and in-house teams, that conversation has not started. The closest thing most agencies have is an informal understanding that whoever built the library owns it, which works until they leave or get busy, and then it quietly stops working.
What four more model releases changed
Since April, KINTAL has run Nemotron 3 Super, GPT-5.4, GPT-5.5 and Claude Opus 4.7 through the same seven-task creative benchmark. The pattern that came back was more specific than "personas are finished": different models are asking for different things, and a library written for one tier will misfire on another.
Opus 4.7 needed a lighter hand to do its best work. On a brief-to-pitch strategy task and a financial planning task under constraint, it deployed extended reasoning without being asked to select a reasoning mode, something earlier versions needed explicit instruction to do. On a synthesis task, minimal scaffolding, an outcome and enough context, was what let it identify a buried insight and argue for it over the obvious alternative, scoring the highest mark on the benchmark. Give a model like this a rigid, twelve-step persona prompt and there is a decent chance you are getting in its way.
That lighter touch came with a limit worth noting. Told explicitly, across eight outputs, to drop the follow-up offer at the end of a piece of work, it dropped it once. The trained habit of ending with an offer of further help held on seven times out of eight, a reminder that instruction compliance and deeply trained conversational habits are not the same thing, and that no amount of context engineering switches off a behaviour that deep on its own.
GPT-5.5 sat further up the scaffolding needs than Opus 4.7. On structured variation work it produced five distinct executions from a loose brief. On tasks that needed an unstated tension surfaced, audience scepticism, a strategic reasoning chain, it read the brief's surface and stopped there, missing what was implied rather than written. That simply describes where it sits on the pyramid: strong on execution once the direction is given, needing more of that direction spelled out than a frontier reasoning model does.
OpenAI's own findings back this up from outside KINTAL's benchmark. Its GPT-5 prompting guide walks through a real example from Cursor: sections of their prompt tuned to get the most out of earlier, less capable models needed rewriting for GPT-5, because instructions written to compensate for weaker native tool use became counterproductive once the model's own judgement caught up. The guide is direct about the cost of skipping that rework: a highly steerable model is more damaged by contradictory or vague instructions than an older one, because it spends its reasoning hunting for a way to reconcile them rather than simply picking a plausible reading and moving on.
A closer comparison of the current frontier models is coming, and it was very nearly this piece instead. The plan was a straight model-by-model read on where each one has landed. What stopped it was realising the comparison does not stand up on its own: what a model needs from you changes with the tier it sits in, so ranking them side by side without saying which tier you are prompting into just moves the confusion somewhere else. That needed sorting first.
How much scaffolding does each model need?

That is the shape the pattern has settled into, the Prompting Pyramid. Four tiers, and the scaffolding a prompt needs drops as you move down them.
Small or cheap models want a narrow task, explicit steps, examples and a rigid output format. This is where persona framing still earns its place, since the model has less headroom to infer what "expert copywriter" is supposed to mean in practice, and a tight structure compensates for that.
General models want a clear goal, useful context, defined outputs and some step-by-step guidance. Less rigid than the tier above, but still benefiting from a reasonable amount of hand-holding.
Frontier and reasoning models, the Opus 4.7 tier, want the outcome stated first, relevant context, hard boundaries and success criteria, then less micromanagement of how they get there. This is where over-prompting shows up most: a persona script written for a smaller model gets left in place on one that can already reason about the brief unaided.
Agentic workflows are a different category of problem entirely. Tools, permissions, approval gates, verification and memory matter more than prompt wording, because the risk has moved from "does the output read well" to "did the system do the right thing without supervision."
This makes persona framing conditional on which tier is doing the work, a more useful instruction than either "always use a persona" or "personas are dead." The role has a narrower job now: perspective, responsibility and authority, not decoration.
| Tier | What it needs |
|---|---|
| Small or cheap models | Narrow task, explicit steps, examples, rigid output format |
| General models | Clear goal, useful context, defined outputs, some step-by-step guidance |
| Frontier and reasoning models | Outcome stated first, relevant context, hard boundaries, success criteria |
| Agentic workflows | Tools, permissions, approval gates, verification, memory and retrieval |
How do you audit a prompt library?
The audit does not need to be a large exercise. For most creative teams it comes down to five questions now, one more than KINTAL suggested in April.
Which prompts are still in active use? Check the view or edit history if the library lives in a shared doc or Notion page. Anything nobody has opened in three months is probably not in active use, or people have their own version saved elsewhere, which is its own finding.
Which have been modified beyond recognition? Ask someone to share the prompt they used for a recent piece of work and compare it with the library version. If they have added context, changed the tone instruction and dropped the format requirement, the library version is a first draft nobody is using, and the gap between that adaptation and the shared resource is the real finding.
Which were built for a technique that has since moved on? Rigid persona framing written in 2024 or early 2025 was best practice when it was written. It still works for stylistic tasks. Applied without distinction across a library, it is a blunt tool where a lighter touch now does the same job.
Which tier of model is this prompt talking to? This is the new question. A prompt built for a small or cheap model, heavy on persona and rigid format, applied unchanged to a frontier reasoning model, is fighting the model's own judgement rather than using it. The reverse is just as costly: a loose, context-only prompt aimed at a model that still needs explicit steps produces vague, unreliable output and gets blamed on the model rather than the mismatch.
Which have no identifiable owner or use case? These are the prompts nobody remembers writing, with no record of what they were for or which project they were built around. They look useful. Nobody quite knows if they are, and because nobody owns them, nobody updates or removes them. This category is usually larger than expected, and it is where the debt lives.
The structured frameworks, CO-STAR, RACE, RASCEF and the rest, remain a practical starting point for standardisation, particularly for copy, briefs and tonal consistency work. Treat them as scaffolding for the tiers that still need scaffolding, not a universal method.
The more durable fix splits what a prompt is doing into two layers. One is a canonical task specification that does not change with the model: the outcome, the required context, the constraints, the evidence rules and the success criteria. The other is a thin adaptation layer specific to whichever tier is running the job, rigid and example-heavy for a small model, outcome-led and unmicromanaged for a frontier one, wrapped in tools and approval gates for an agentic workflow. Anthropic calls the broader shift context engineering, describing it as "the natural progression of prompt engineering": less about the instruction, more about the complete set of information, tools and history the model has available when it answers. A canonical spec plus a model-specific layer is what that looks like inside an actual prompt library, rather than in a blog post about one.
The prompts a team built in 2025 are worth a second look, now that the tools have moved on and the practice has a sharper question available to it: what does this particular model need from you before you start typing.
Not sure how AI is feeling with your team? Find out how SIGNAL can help you.
A note on how this was made: the persona-versus-role distinction was worked through in conversation with ChatGPT. Research for this update was conducted using Claude's web search and fetch tools, cross-referencing KINTAL's own published benchmark findings on Opus 4.7, GPT-5.5, GPT-5.4 and Nemotron 3 Super against Sander Schulhoff's published research, Anthropic's engineering blog and OpenAI's GPT-5 and Realtime prompting guides. The framing, the audience, the editorial judgement and every significant decision along the way were mine.
More thinking
Have a project in mind?
Book a call and tell me what you're working on.
Book an intro call

