Prompting across models

What actually changes when you move a prompt between model families, and the three dials people confuse for one.

Updated 14 Sept 2026 · Current as of Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5

Three dials, not one

  • Which model runs the task
  • How hard it thinks
  • Whether it fans out into several agents

Most "prompting" problems are really a dial confusion.

Teams say "the model isn't smart enough" when what they changed was reasoning effort, or "we need a bigger model" when the task actually wanted decomposition into several agents. These are three independent controls. Name which one you are turning before you turn it, and most disagreements about model quality resolve on their own.

Name the role, not the model

  • Ask for "the reasoning rung", not a version string
  • One edit point when the lineup changes
  • Prefer a tier alias over a pinned ID

A pinned model ID always goes stale. An alias never does.

Write your prompts and tooling to name an intent level — the judgment tier, the execution tier — and map levels to models in exactly one place. When a new model ships you change one line, not fifty prompts. Pin an exact version only where something contractually requires it, such as a REST path or a check that must detect drift.

Spend capability where judgment lives

  • Judgment, planning, scoping, supervision: top rung
  • The execution leg of scoped work: one rung down
  • Genuinely mechanical work: two rungs down

Choose the model by the class of work. Choose the effort by the task.

Which class of work earns which model is worth writing down once and following. Effort is the setting to change task by task. I run everyday work at medium and turn it up for research and other open-ended work. On Artificial Analysis's September 2026 timing tests, the top setting adds minutes for a few points, so save it for work where time doesn't matter. A cheap model taking three times the turns can still cost more than one capable model taking one.

What transfers, and what does not

  • Clear success criteria transfer everywhere
  • Worked examples transfer well
  • Formatting scaffolds are model-specific
  • Jailbreak-ish coaxing transfers worst

Rewrite the scaffolding. Keep the intent.

The durable half of a prompt is the part describing what done looks like and what the constraints are. The fragile half is the part negotiating with a particular model's habits — the coaxing, the "think step by step" ritual, the formatting incantations. When you move families, expect to rewrite the second half and keep the first.

Never take a model's word for its own behaviour

  • A self-report is not evidence
  • Probe the behaviour, record the date
  • Re-probe when the lineup changes

Ask the transcript, not the model.

Models are unreliable narrators about their own configuration — which model is running, whether a flag took effect, what a sub-process actually did. If a behaviour matters, verify it from outside: the transcript, the logs, the API response. Record when you last checked, because the answer expires with the next release.

What would make your
work a little easier?

A recurring task, a half-formed idea, a system that could work better. That’s a good place to start.

Let’s talk it through