The most powerful AI models, like Anthropic’s Claude Mythos and OpenAI’s GPT-6 Astra, often cost five times as much to use, or more, as the everyday models most of us rely on, like Claude Sonnet, OpenAI’s GPT-5.6, or Google’s Gemini Flash. (Sources: Anthropic, OpenAI, and Google price lists.)
But the gap is smaller than the price suggests. Give a cheaper model a clear ask, the right facts, and a quick check, and it can get close. Below are the simple tips, with real before-and-after prompts you can copy today.
What do the top AI models do better?
- Staying on track. A top model can keep working on a long task without losing the thread. METR, an independent research group, estimated that an early version of Mythos succeeds about half the time on tasks that take a skilled person 16 hours or more, and said it cannot measure reliably beyond that. (Source: METR, as reported by The Decoder.) Smaller models tend to lose the thread sooner.
- Knowing more. A bigger model holds more facts, especially about niche topics. A smaller one knows less, so it is more likely to guess, and a guess can sound just as confident as a fact.
- Hard problems. On Anthropic’s own tests, Mythos beats the company’s earlier top models on difficult coding, science tasks, and expert-level exam questions. (Source: Anthropic’s launch announcement.)
- Checking its own work. Anthropic says Mythos is better at verifying its own work before handing it back. (Source: Anthropic’s launch announcement.)
Where the gap barely matters
For a lot of daily work, a smaller model is already enough: rewriting an email, summarizing a page, fixing grammar, turning notes into a list, pulling names and dates out of a document. These jobs are short and clear, and the answer is easy to check.
The companies say so themselves. Anthropic advises Haiku for simple tasks, Sonnet for most work, and its larger models for the most complex reasoning. Google describes its Pro model as the one that excels at coding and complex reasoning, and sells cheaper Flash and Flash-Lite models alongside it. (Sources: Anthropic’s pricing guide, Google’s price list.) A good rule: start with the smaller model. Switch to the big one when the job is long, hard, or a mistake would be expensive.
Tips for better prompts
- Say exactly what you want. Don’t leave it to guess. Include who it is for, how long it should be, and what it should sound like. Anthropic’s first general rule of prompting is to be clear and direct, and to think of the AI as a brilliant new employee who lacks context on your work. (Source: Anthropic’s prompting guide.)
- Give it the facts. Don’t make it rely on memory. Paste in the few facts it needs, like prices, dates, or the details of your situation. Google’s guide puts it plainly: include the information the model needs instead of assuming it already has it. (Source: Google’s prompting guide.)
- Tell it what to do when it doesn’t know. Add one line: “If the answer isn’t in what I gave you, say so.” Anthropic says giving the AI permission to admit it doesn’t know can drastically reduce made-up answers, which AI experts call hallucinations. (Source: Anthropic’s guide to reducing made-up answers.)
- Show an example. If you want a certain style or format, paste one you like and say “like this.” Anthropic calls examples one of the most reliable ways to steer the format and tone of an answer, and Google recommends always including them. (Sources: Anthropic’s prompting guide, Google’s prompting guide.)
- Give all the details up front. Put the full picture in your first message instead of adding details bit by bit. In one study by Microsoft and Salesforce researchers, AI answers got 39 percent worse on average when the same task was spread over a back-and-forth conversation, and once a model took a wrong turn, it rarely recovered. (Source: Microsoft and Salesforce research paper.)
Tips for a better workflow
- Break big jobs into small steps. Don’t ask for a whole report in one go. Ask for an outline, then one section at a time. Small steps are where small models do their best work, and OpenAI and Google both recommend splitting complex tasks this way. (Sources: OpenAI’s prompting guide, Google’s prompting guide.) For each step, still give all the details it needs in one message.
- Let it think first. For harder questions, add “Think this through carefully before you answer,” and turn on your app’s “thinking” option if it has one. Anthropic’s guide says a simple instruction like “think thoroughly” often works better than spelling out every step yourself. (Source: Anthropic’s prompting guide.)
- Start a fresh chat when it drifts. Long conversations get worse over time, because early mistakes stick around. That is the same study again: models that go wrong early tend to stay wrong. Ask for a short summary of where you are, open a new chat, and paste it in. I wrote more about this in context rot.
- Keep your important material outside the chat. If the only copy of an important fact, link, or screenshot lives in an old conversation, you will put off starting fresh, even when you should. Keep it somewhere you can grab it and paste it back in.
Tips for better results
- Ask it to check its answer. “Now check your answer against the facts I gave you. List anything that doesn’t match.” Anthropic recommends this kind of second pass for catching mistakes. (Source: Anthropic’s guide to reducing made-up answers.)
- Check with something that can’t be fooled. A calculator for numbers. Clicking the link. Running the code. Reading the source it quoted. Don’t just ask the AI if it is sure. Even Anthropic says to always check important information yourself.
- Ask more than once. Ask the same question two or three times, or ask for three versions. If the answers disagree, treat that as a warning sign that it may be guessing. If you are writing, just pick the best draft. (Source: Anthropic’s guide to reducing made-up answers.)
- Know when to switch up. If the smaller model gets the same thing wrong twice, stop rewording. Hand that one step to a bigger model, then come back.
See it in action
Does this really work?
Yes. Researchers have shown that a much smaller model, given the right help, can beat a giant on a clear task.
There are limits. A smaller model still knows less, and it still struggles with something truly new or a job that runs for hours with nobody watching. For those, the big model earns its price.
The checklist
- Did I say exactly what I want? Who it is for, how long, and what tone. Why: the AI can’t read your mind, so it fills any gaps with guesses.
- Did I give all the details up front? Why: when details arrive bit by bit, the AI locks onto its first guess and rarely lets go.
- Did I paste in the facts it needs? Why: if it has to rely on memory, it may make something up.
- Did I tell it to say “I don’t know”? Why: otherwise it may guess, and a guess can sound just as sure as a fact.
- Is this one small job? If not, split it into steps. Why: AI does its best work on one clear task at a time.
- Is this chat getting long? Ask for a summary and start a new chat. Why: old mistakes linger in long chats and quietly affect new answers.
- Did I check the answer myself? With a calculator, the original source, or by trying it. Why: AI can sound confident and still be wrong.
Your copy-and-fill prompt template
Where Tansei fits
Almost every tip above comes down to one thing: context. The facts, the example, the summary you carry into a fresh chat, the source you check the answer against. Give a model the right context and even a small one does good work. Leave it guessing and even the biggest one can go wrong.
That is why I built Tansei. It is a simple shelf for Mac and Windows that sits at the edge of your screen and keeps your working context in one place: the screenshots, links, notes, snippets, and files you are using right now. When you write a prompt, drag in exactly what the AI needs. When a chat drifts, start a fresh one without losing anything. It works the same with any AI, big or small, because it lives beside every app.
Everything stays on your computer, and there is no account.
Get Tansei for Mac and Windows
Frequently asked questions
What is Claude Mythos?
Claude Mythos is Anthropic’s top AI model. Its latest version, Mythos 5.1, came out on September 1, 2026. It is invitation-only, for approved organizations doing work like cybersecurity and life sciences. Claude Fable is the same model with extra safety limits, and anyone can use it.
What is the difference between Claude Mythos and Claude Sonnet?
Mythos is Anthropic’s top model, built for long and complex work. Sonnet is a smaller, faster model that costs about a fifth as much and handles most everyday tasks well. The gap shows up mostly on long, hard, or multi-step jobs.
Do I need the most powerful AI model for everyday tasks?
Usually not. Emails, summaries, rewriting, and simple research work well on smaller models. Anthropic, for example, suggests Haiku for simple tasks and Sonnet for most work. Use the biggest model when the job is long, hard, or a mistake would be costly.
How do I get better answers from a cheaper AI model?
Be specific about what you want, give all the details up front, paste in the facts it needs, and tell it to say when it doesn’t know. Break big jobs into small steps, start a fresh chat when it drifts, and check the result with something other than the AI.
Why does AI make things up (hallucinate)?
When a model doesn’t know the answer, it often produces something that sounds right instead of saying so. This is called hallucinating, and the made-up answer is called a hallucination. It happens with every AI model, big or small, but models that know less have more gaps to fill. Giving it the facts and telling it to admit when it doesn’t know both reduce hallucinations.
How can Tansei help me get better results from AI?
Tansei keeps your working context in one place: the screenshots, links, notes, snippets, and files you are using right now. That makes it easy to give any AI model, big or small, the facts and examples it needs, to start a fresh chat without losing your material, and to check answers against the original source.
Back to top