01Start HereThe map + Core 10 02What AI Still Gets Wrong11 failure patterns 03All 30 TermsSearch the guide 04What Is an Agent?Agency, tools & autonomy

What AI Still Gets Wrong

Name the failure. Fix the setup.

Fluent output isn't the same as reliable output. The eleven patterns below are the ones we see most often once AI-assisted work moves from a demo into an actual health IT deliverable — what each one looks like, what changes it, and where it tends to show up.

Open a card when something feels off. That's usually one of these eleven.

01 SETUP

The Cold Start Problem

+

Every new session starts without the context it needs to do strong work.

The patternA model can know a great deal in general and still know nothing about your organization, your role, or the initiative in front of you until you supply it.
What helpsKeep a short standing brief — who you are, what success looks like, current constraints — and open with it, or store it in a persistent project.
In practiceAsk for a stakeholder briefing without naming the health authority or initiative, and you'll get a briefing that could describe almost anyone.
Not the same asThis isn't evidence the model is unreliable — general capability and knowledge of your specific situation are two different things.
02 TRUST

Hallucination

+

A confident detail, citation, or figure that isn't actually true.

The patternWhen the real evidence is missing, a plausible-sounding answer can come out just as fluently as a well-supported one.
What helpsAsk for named, checkable sources, verify any figure before it reaches a client document, and explicitly permit "I don't have that."
In practiceA draft summary cites a section number in a standard that turns out not to exist.
Not the same asA citation formatted correctly isn't a citation that's true — formatting isn't evidence.
03 TRUST

Sycophancy

+

The model agrees with the direction you've already signaled instead of testing it.

The patternA system built to be helpful can mirror your framing and preferred conclusion more readily than it challenges them.
What helpsAsk directly for the counter-case, the biggest risk, and the conditions under which the plan should be rejected.
In practiceDescribe a vendor you're already leaning toward, and the model finds reasons the choice is sound while skipping past the procurement risk.
Not the same asAn agreeable tone isn't the same thing as a sound recommendation.
04 WORKFLOW

The Premature Build

+

The model rushes to a finished deliverable before the real decisions are made.

The patternAsked for a document, it can skip straight to a polished draft rather than surfacing the open questions first.
What helpsAsk for a plan, an outline, or three open questions before authorizing the full build.
In practiceA request for a project charter returns a fully formatted document before scope or sponsorship has actually been agreed.
Not the same asA polished, complete-looking output can hide the fact that the real decisions were never made.
05 WORKFLOW

First-Draft Anchoring

+

Once a draft exists, feedback patches it instead of reconsidering the structure.

The patternThe first output becomes a gravity well — later feedback tends to mutate it rather than cleanly revisit the underlying approach.
What helpsFor a genuine redirect, start a fresh request from a clean brief rather than iterating on a flawed foundation.
In practiceA steering committee deck gets a dozen rounds of wording edits when the real issue is that it's telling the wrong story.
Not the same asIterating on solid evidence is normal practice — getting stuck editing a flawed first draft is a different thing entirely.
06 WORKFLOW

Context Rot

+

A long conversation accumulates outdated instructions the model starts drawing on.

The patternSuperseded drafts, dropped directions, and old decisions pile up in a long thread, and the signal gets harder to use reliably.
What helpsStart fresh sessions at natural milestones, and keep durable facts in a project file rather than buried in chat history.
In practiceA months-long procurement thread resurfaces an evaluation criterion that was explicitly dropped two revisions earlier.
Not the same asA large context window means more text can fit — it doesn't mean every detail inside it gets equal, reliable weight.
07 TRUST

Outdated or Ungrounded Knowledge

+

The model answers from training data that may predate the current facts.

The patternA policy, a software version, or an org chart can change after a model's training ends, and it has no built-in way to know that.
What helpsSupply current source documents directly for anything time-sensitive, rather than relying on the model's own memory.
In practiceGuidance on a system's integration approach references an older version that no longer matches what's actually deployed.
Not the same asWeb access doesn't guarantee the authoritative source gets found — retrieval quality still needs a human check.
08 QUALITY

The Average Trap

+

Without real direction, output drifts toward generic, forgettable phrasing.

The patternAn underspecified request tends to land near the statistical middle of everything the model has seen — competent, and indistinguishable from anyone else's.
What helpsSupply real examples, a point of view, and specific language to avoid, so the model has something to differentiate against.
In practiceAn executive summary arrives full of phrases no actual sponsor would recognize as their own.
Not the same asClean grammar and a confident tone aren't the same as a substantive point of view.
09 QUALITY

Probabilistic Variability

+

The same request can return a different answer depending on when you ask.

The patternGeneration works probabilistically rather than as a fixed lookup, so repeated runs of the same prompt can diverge.
What helpsFor anything that needs to be repeatable, add constraints — templates, rubrics, worked examples — and review proportional to the stakes.
In practiceRun the same risk-assessment prompt twice in one week and the emphasis, or even the recommendation, may shift.
Not the same asVariability isn't randomness for its own sake — it means repeatable work needs guardrails, not that the tool is unreliable.
10 TRUST

Reasoning Failure

+

A methodically explained answer can still rest on a wrong starting assumption.

The patternMore visible reasoning helps with many tasks, but it doesn't guarantee the underlying assumption or calculation is correct.
What helpsCheck the starting assumptions independently, and verify any calculation against the actual source data.
In practiceA budget projection walks through its logic clearly but starts from the wrong baseline year.
Not the same asA detailed explanation is not proof — the conclusion still needs to be checked against the numbers.
11 TRUST

Bias

+

Output can reflect and amplify patterns present in training data or task setup.

The patternA model learns patterns from human-produced data and from how a task is framed — both can carry uneven treatment forward.
What helpsTest outputs across a range of real cases, and keep an accountable person reviewing anything that affects people's outcomes.
In practiceA draft position posting emphasizes different qualifications depending on subtle cues in a candidate's background, even when the qualifications are equal.
Not the same asBias isn't limited to overtly offensive language — it can show up as omissions, defaults, or uneven error rates.

Rolling AI-assisted work into a live delivery program?

We help teams structure the work so these patterns don't reach a client deliverable.

IMC holds no allegiance to any AI platform or vendor, so the review you get on how AI is being used in your program is independent, and grounded in what your organization actually needs.

Start a conversation