Choosing a model and effort level

Which model to pick in ChatGPT, Claude, and Gemini, how much effort to give it for each kind of task, and where turning it up stops paying.

Updated 14 Sept 2026 · Current as of Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, Gemini 3.8 Flash

Set the model once, move the effort

  1. Pick the top model your plan includes at no extra cost
  2. Everyday work: medium effort
  3. Research and brainstorming: turn it up
  4. Gemini 3.8 Flash: leave it on Extended, because it's fast

The model stays put. The effort moves with the task.

This applies if you use ChatGPT, Claude, or Gemini, on a company account, a personal plan, or both. Each app gives you two settings. The model is which AI answers. The effort is how long it works on your request before it answers. The apps call it effort, reasoning, or thinking level.

Most clients I work with pick a model once and forget about it. When an answer comes back weak, they retry or ask a different way. This rule keeps that habit for the model and adds one new one: change the effort to suit the task.

Effort moves the result more than the model choice does. GDPval-AA is a benchmark, a fixed set of tasks every model is scored on, built from work products such as reports and spreadsheets. On it, Claude's Opus 5 moves 363 points between its lowest and highest effort settings. At their highest settings, Opus 5 and ChatGPT's GPT-5.6 Sol are 111 points apart.

Pick the top model your plan includes

  • Claude: Opus 5
  • ChatGPT: GPT-5.6 Sol, or GPT-6 Astra in ChatGPT Work
  • Gemini: Flash, not Pro
  • Fable 5.1 costs extra on Pro and standard seats

Intelligence Index score against cost per task for Claude Opus 5 and Claude Sonnet 5 at each effort level. Opus 5 at low effort scores above Sonnet 5 at max for about a fifth of the cost.

Claude's smaller model, turned up, costs more and scores lower.

The top model is the most capable one your plan includes at no extra cost. In Claude, that's Opus 5. Fable 5.1, a higher Claude tier, scores higher but costs extra. Anthropic's help center says that on Pro and on standard company seats it runs on usage credits, billed on top of the plan.

In ChatGPT the model menu is mostly an effort menu. Instant, Medium, High, and Extra High all run GPT-5.6 Sol. On Plus, GPT-6 Astra is in ChatGPT Work.

In Gemini, pick Flash, even though Google calls Pro its most advanced model. The Artificial Analysis Intelligence Index is an independent benchmark that combines several reasoning and knowledge tests. On it, Gemini 3.8 Flash scores 41 and the current Pro, 3.1, scores 30.

The chart shows why the model is the setting to fix. Opus 5 at low effort scores 40 for $1.10 a task. Sonnet 5, the smaller Claude model, scores 38 at max and costs $5.09.

Find the effort setting

  • Claude: the menu next to the send button, Low to Max
  • ChatGPT: the model menu, Instant to Extra High
  • Gemini: Thinking level, Standard or Extended
  • Check what yours is set to now

Three apps, three names for the same setting.

These are the menus as of September 2026, from each vendor's help pages. Claude puts a model and effort menu next to the send button. The levels are Low, Medium, High, Extra high, and Max, and a change applies from the next reply.

ChatGPT folds effort into the model menu: Instant, Medium, High, and Extra High, with Pro above them on some plans. Gemini has a Thinking level of Standard or Extended, plus Deep Think on the Ultra plan only.

All three vendors say higher settings use up your plan's limits faster. Menus change often. If yours looks different, look for the word effort, reasoning, or thinking.

Medium for everyday work

  • Drafts, emails, summaries, quick questions
  • The top model at medium
  • Gemini 3.8 Flash on Extended, because it's fast
  • Low saves seconds and gives up quality

Medium is where I leave it for everyday work.

I run the top model at medium for everyday work. In my experience the results are better, and the cost in speed is small. The benchmark numbers line up with that.

Artificial Analysis times each model on the same standard request. At medium, the top models return a complete answer in under half a minute. Dropping Opus 5 to low saves between 1 and 8 seconds in those timings. Its scores fall from 45 to 40 on the Intelligence Index and from 1,525 to 1,372 on GDPval-AA.

Gemini 3.8 Flash is the exception. It writes about five times faster than the others, so even at high it answers in about the time they take at medium. I leave it on high. In the Gemini app, the higher of the two thinking levels is called Extended.

Turn it up for research

  • Research, brainstorming, open questions
  • Anything that searches and pulls sources together
  • Claude: Extra high. ChatGPT: High or Extra High
  • Max only when time doesn't matter

GDPval-AA office-work score at each effort level. Claude Opus 5 rises from 1,372 at low to 1,735 at max, and GPT-5.6 Sol from 1,355 to 1,624.

The more open the question, the higher the effort.

Higher effort gives me deeper research. The research for this playbook ran on Claude at extra high. It matched what I already knew about the topic and went deeper, with fresher information than I've had from the same model at medium.

GDPval-AA is the benchmark closest to this kind of work. Artificial Analysis runs it on tasks from 44 occupations. The model browses the web, uses tools, and produces a document or spreadsheet. The results are compared blind, two at a time.

Effort moves those scores a long way. Opus 5 goes from 1,525 at medium to 1,708 at extra high. Gemini 3.8 Flash tops out at high, where it scores 1,464, about level with GPT-5.6 Sol at medium. Anthropic's effort documentation lists extra high for exploratory work such as "detailed web search."

The top settings cost time

  • Opus 5: under 25 seconds at medium, 1 to 2 minutes at max
  • GPT-5.6 Sol: under 20 seconds at medium, 2 to 3 minutes at max
  • 6 to 8 points of score for the wait
  • Personal plans also burn through limits faster

Time to a complete answer at medium and at max. Claude Opus 5 goes from 45 in 13 seconds to 51 in 58 seconds. GPT-5.6 Sol goes from 39 in 12 seconds to 47 in 2.1 minutes. GPT-6 Astra goes from 50 in 14 seconds to 53 in 5 minutes. Gemini 3.8 Flash at high, its top setting, scores 41 in 18 seconds.

Medium takes seconds. Max can take minutes.

This is the trade: time for a little extra intelligence. Going from medium to max, Opus 5 gains 6 points on the Intelligence Index and GPT-5.6 Sol gains 8. GPT-6 Astra scores 53 at both extra high and max, and max takes 2 to 3 minutes longer.

These times move with server load. Between Artificial Analysis checks on September 11 and 14, Opus 5 at max went from about 2 minutes to 1. Medium stayed under half a minute.

On a company account, the wait is the main cost. On a personal plan, every vendor says higher settings also use your limits faster.

If time and cost don't matter for a task, go ahead and turn it all the way up. Max does add something on office work: 27 points over extra high for Opus 5 on GDPval-AA, and 39 for GPT-5.6 Sol. For everyday work, the extra minutes buy little.

What max does to simple tasks

  • Research: longer thinking can talk a model out of a right answer
  • Factual questions without search: more made-up detail
  • Today's top models: no drop found, only more cost
  • The sure cost is time

On simple tasks, max is at best a slower way to the same answer.

You may have heard that the highest settings make simple answers worse. There's research behind it, mostly on earlier models.

A study by Gema and colleagues found Claude models got more distracted by irrelevant details the longer they reasoned. Zhao, Hooi, and Ng tested 14 reasoning models on factual questions asked without web search, and more thinking often produced more made-up answers. Two 2026 studies, by Zhou and colleagues and by Caldarella and colleagues, caught models reaching a correct answer, thinking on, and dropping it. Anthropic's guidance for an earlier Opus model warns that max can overthink some tasks.

One recent test of today's top models points the other way. A developer publishing as Synthorai ran GPT-6 Astra and GPT-5.6 Sol on 11 simple math and counting tasks. Every level from low up got all of them right. On Astra, max cost 2.3 times as much per correct answer as low.

Where this stops

  • Coding and automated tasks: smaller models make sense
  • A new top model sends you back to step 1
  • Re-test effort after a model change
  • Every number here is from September 2026

Revisit the model when a new one ships. Choose the effort every task.

This rule is for office work. For coding and automated tasks I use smaller models, and most office workers don't need to worry about that. The Prompting across models playbook covers how to split that kind of work.

The model doesn't stay fixed forever. When your vendor adds a new top model to your plan, go back to step 1. Anthropic's guidance for Opus 5 says effort settings carried over from an earlier model should be tested again rather than reused.

The benchmark numbers here will date quickly. Google released three Flash models between late July and early September. When the next one reaches your menu, change the model. Until then, effort is the only setting to touch.

What would make your
work a little easier?

A recurring task, a half-formed idea, a system that could work better. That’s a good place to start.

Let’s talk it through