Four workshops of material on one deck. Every fact checked against Anthropic's docs on 30 July 2026.
Six parts. Click any line to jump to that section.
LLM, token, and the context window everything else depends on.
Four models, the two dials you control, and where you work.
What context really is, your two jobs, the two options when a chat gets long, and the habits that keep it short.
What makes a prompt work, and six methods worth knowing.
Why it invents things, how to catch it, and when to be careful.
MCP and skills in plain terms, and what you already have.
Four ideas everything else rests on.
It reads everything you've given it and predicts what comes next, then does that again, and again. Every sentence is built one piece at a time. Nothing is looked up in a database.
It isn't recalling what you said earlier; it's re-reading it. The whole transcript is sent again every turn. (It can remember things about you across chats — that's a separate memory feature, not the conversation itself.)
Without a tool to look things up, it answers from training, not from a live source.
Ask the same question twice and you can get two different answers. That is normal.
Roughly ¾ of a word in English. People assume tokens are characters or whole words; they are neither. Every limit and every price in this deck is counted in tokens, which is why it comes first.
Hello worldTwo reasons it matters:
1. Limits are in tokens. The context window — the next slide explains what it is — holds up to 1 million tokens on current models: roughly 750,000 words, about the whole Bible. Generous, but a handful of PDFs eats it faster than you'd guess.
2. Bills are in tokens. When you use the API you pay per token in and per token out. On claude.ai it's bundled into the subscription, so you don't see it, but it's still being counted.
Anthropic calls the context window Claude's working memory: everything it can see while it answers you. The size is fixed, and all of this competes for the same room.
It does not build up like memory. The window is a page re-read from the top each time you press enter.
Four models, two dials, and two places to work.
| Model | Built for | In / out per Mtok | Context |
|---|---|---|---|
| Fable 5 | Highest capability there is. Long-running agents, deep research. | $10 / $50 | 1M |
| Opus 5 | Complex agentic coding and enterprise work. The sensible default. | $5 / $25 | 1M |
| Sonnet 5 | Best speed-to-intelligence ratio. Most everyday work. | $3 / $15 | 1M |
| Haiku 4.5 | Fastest and cheapest, still near-frontier. High volume, and the helper jobs bigger models hand off. | $1 / $5 | 200k |
Everyone else: OpenAI, Google, Meta, Mistral and DeepSeek all ship competitive models. We use Claude. · Sonnet 5 is on introductory pricing of $2/$10 until 31 Aug 2026. · Mythos 5 exists but is invitation-only, for defensive cybersecurity.
In Claude Code and Cowork — the two places to work, later in this part — you set how much Claude does between check-ins. Four modes, from careful to reckless. Chat has no such dial: every reply is already a check-in.
The habit worth stealing: plan first, then run. Start a big task in plan mode, approve the plan, then switch to auto-accept for the execution. One approval replaces an afternoon of back-and-forth.
Where the switch is: in Claude Code, Shift+Tab cycles manual, auto-accept and plan; bypass has to be switched on deliberately. Cowork offers the same idea as three permission levels. How hard Claude thinks is a separate dial: the next slide.
The model menu holds the model picker, effort, and — on older models — a Thinking toggle. Effort is the one worth understanding. It sets how much work goes into every answer.
Claude's own warning, straight off that menu: “Higher effort means more thorough responses, but takes longer and uses your limits faster.” So Max on an easy task buys nothing and costs real usage. Turning effort down is the cheapest way to stretch your allowance — the shared usage budget that claude.ai, Cowork and Claude Code all draw on.
Effort and Thinking are separate things. Effort sets how hard Claude works; Thinking is whether it reasons visibly first. Current models decide that themselves; only older models like Haiku 4.5 still show the manual toggle.
Most of what people assume needs Cowork, Chat already does, including producing real spreadsheets and decks. The difference is narrower than it looks, and it's about scale, duration and repetition.
You stay in it, turn by turn. Seconds per reply.
You describe the outcome and walk away.
The rule: one file and one answer, use Chat. A whole folder, an hour of work, or something you'll need again next week, use Cowork. · Claude Code is the same idea aimed at a codebase; if you write software, that's your door.
Each of these would be painful or impossible in Chat: too many files, too long, or it has to happen again next week.
However you phrase it, end with this line:
First explore my folder, then ask me questions using AskUserQuestion before you execute.
You get a clickable form instead of a blank text box. It interviews you, you approve the plan, then it works.
The part that changes what you do this afternoon.
Context is simply everything Claude can see right now when it answers you. It is the material Claude has to work with for this task.
These are inputs to context, not steps in a pipeline.
The more irrelevant, outdated or contradictory material Claude has to keep track of, the harder it becomes to focus on the thing that matters now.
Current task
Relevant instructions
Relevant files and facts
Only the conversation you still need
Current task buried in noise
100-message conversation
Old drafts and unused files
Wrong turns Claude now reads again
Every context decision is the same two jobs at once. They pull in opposite directions, and doing only one of them still fails.
Give Claude enough to do the job well: the report, the email thread, the example, the constraint. If it can't see it, it can't use it.
Skip it and: Claude fills the gaps from training instead of your reality. Generic answers, wrong assumptions, confident invention. That failure gets its own part: Part 5.
Keep the window clean: material for this task, this customer, this question. Yesterday's task belongs in yesterday's chat.
Skip it and: attention spreads across the noise and quality slides, which is the rot from the previous slide. Old drafts and wrong turns get read as instructions.
Anthropic's guiding principle, verbatim: find “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome.” Enough signal, zero noise. The rest of this part is how.
Answers degrade well before the window is full: heavy users treat 60% full as the most one task should need. Past that line you have two options.
Long conversation that isn't finished? Ask for this instead of pushing on or starting from scratch, then paste the result into a fresh chat.
Summarise this conversation so I can continue in a new chat.
Include:
- The goal, in one sentence
- Decisions we made, and why
- What's done so far
- What's still open, and the immediate next step
- Any constraint or preference I gave you that still applies
Leave out: dead ends, superseded drafts, and anything
we already resolved and won't revisit.
/compact does this automatically, and you can steer it with an instruction instead of accepting the default.
↑ Up returns to the slide, exactly where you left off
A container you put around work, in chat or Cowork, rather than a place to work itself. Its job is holding your durable context so your conversations don't have to.
Cowork projects work the same way, with their own instructions, pointed at a folder on your machine. Ours is pre-written: the cowork-starter folder on the shared workspace.
Two layers are already loaded before you type a word. Most people never set either, then wonder why they keep repeating themselves.
Click your initials → Settings → Instructions for Claude. Your role, the terms you use, how you want things written. Applied to every conversation you ever have. Five minutes, once.
Set inside a Project, and they only apply to chats in that Project. The context for one job: who the audience is, what done looks like, what's different about this customer.
Plus the Project's knowledge files, whatever Claude remembers about you, and any connectors you've switched on. Your message lands on top of all of it.
CLAUDE.md file. · ↓ See the order, then copy both layers
The stack Claude reads every turn, built from the bottom up. Everything under your message was already in place before you typed it.
↑ Added this turn
You, right now
Draft a reply to Bosch about the SoH drop.
Layer 2 · per project
Fleet-customer emails. Warm and professional. Use the style guide.
Layer 1 · every chat
Say when you don't know. Always include the source. Be concise.
Anthropic's · fixed
Follow these safety rules. Refuse requests that could cause harm.
↓ Loaded before you type a word
Your initials → Settings → Instructions for Claude. Account-wide, applies to every conversation you ever have, and nothing needs filling in. Select the block and copy.
# Personal preferences
## How to handle facts and uncertainty
- If you're not sure, say so plainly ("I'm not sure" / "I'd need to check"). Don't guess and present it as fact.
- Don't make up facts, numbers, names, dates, or quotes. If you don't know, say you don't know.
- Separate fact from best guess, and say when you're guessing.
- If my request is unclear or missing something important, ask me one quick question first.
- If something's low-stakes and easy to undo, go ahead, but tell me what you assumed.
- If you get something wrong, say so and fix it. Don't over-apologize.
- Don't tell me a task is done unless you actually checked it.
- Use plain language. If you need a technical term, define it once.
## Output style
- Be concise. Match the length of your answer to the question.
- Write in plain sentences. Only use bullet points when the content is actually a list.
- Don't restate my question back to me, and skip closers like "Hope this helps!"
- No filler. Cut openers ("Great question," "I'd be happy to," "Certainly"), padding ("It's worth noting that," "At the end of the day"), and empty hedges.
- No AI buzzwords: delve, leverage, utilize, foster, streamline, robust, seamless, game changer, tapestry, realm, transformative, elevate, harness.
- Say things straight. Skip "It's not X, it's Y" contrasts, "what most people miss" setups, and "stands as a testament" puffery. State the point and let it stand.
- Use active voice and concrete details: "It cut the wait from 40 minutes to 4," not "it significantly improved efficiency."
- Don't end with a summary paragraph or a profound-sounding kicker. Stop at the last useful point.
- Don't over-format: no emoji in headings, no bold scattered through sentences, no headers over two-line sections.
- Don't use em dashes as a rhythm habit; a comma or period usually works.
## Behavior
- Push back when you disagree with me. Don't flatter me or just agree to be agreeable.
- If you think what I'm asking for is a bad idea or won't work, tell me before you do it, not after.
- When there's a real decision to make, give me the options with a quick recommendation. Don't just pick for me.
- If a task is outside what I asked for, say so and stop. Don't wander into extra work.
- If your answer comes from the web or a file, link where it came from so I can check it.
- Before doing anything hard to undo (deleting, overwriting, sending, publishing), tell me what you're about to do and wait for my okay.
↑ Up returns to the order diagram · ↓ Down for project instructions
Inside a Project → Instructions. Applies only to chats in that Project. This one has blanks: replace everything in square brackets.
## About this project
This project is for [drafting customer emails / writing reports / etc.].
Everything here is [internal work / for external customers / etc.].
## What I want
- Output format: [short email / one-page brief / bullet summary].
- Tone: [warm and professional / plain and direct].
- Always [check the uploaded style guide before drafting].
## My style (examples)
[paste 1-2 real things you've written so Claude matches your voice]
↑ Up returns to personal instructions · keep this short, layer 1 already covers how you write
On claude.ai this costs you nothing: a fresh chat inside a Project keeps your instructions and knowledge files and drops only the accumulated mess. In Claude Code the equivalent is “one task, one plan, one /clear.”
“Update the auth logic in src/auth/handler.rs” beats pasting the file. Same for documents: a link or a path costs a handful of tokens; the contents cost thousands. Let Claude fetch what it decides it needs.
Numbers, tables and code degrade first in a long chat: one small error gets re-read as fact and spreads. Route exact work to its own short chat.
Get the plan agreed before any work happens: plan mode in Claude Code and Cowork (from Part 2); in claude.ai, ask for a plan and approve it first. One approval replaces the long back-and-forth.
Project instructions in claude.ai, or a CLAUDE.md file for Claude Code. Loaded fresh every chat, so it never rots and never gets re-explained. The Project and the instruction stack from the last two slides are where this lives.
A folder of reports or a dozen sources doesn't belong in your chat: a sub-agent reads it in its own window and hands back a short summary. Cowork's research jobs do this for you. Costs tokens, buys back attention.
What makes a prompt work, and six methods worth knowing.
Anthropic's framing, and the most useful one: “a brilliant but new employee who lacks context on your norms and workflows.” Very capable, knows nothing about your organisation, can't see your screen, and will guess rather than ask.
The golden rule:
“Show your prompt to a colleague with minimal context on the task and ask them to follow it. If they'd be confused, Claude will be too.”
“Look at this battery data and tell me what you think.”
Vague verb. No audience. No format. No definition of “what you think.” Nothing to check the answer against.
“You're reviewing Flash Test results for a fleet customer. From the attached report, list any pack whose SoH dropped more than 3 points since the last test. One line each: pack ID, old grade, new grade, likely cause. Flag anything you're unsure about rather than guessing.”
Role, source, threshold, format, and permission to admit doubt.
<instructions>, <context>, <input>) so Claude can tell your orders from your data. It's just a marker, nothing technical.We have a reprompt skill, a packaged instruction set Claude reaches for on its own, covered properly in Part 6. Give it your rough request and it returns a structured one: role, context, task, constraints, success criteria.
Fastest way to apply everything on the previous slide without memorising it.
Flip the interview. End a request with “ask me what you need to know before you start” and Claude comes back with the three or four questions that actually matter: audience, deadline, format, what done looks like. You answer in a line each; it assembles the brief.
The context you'd never think to volunteer gets pulled out of you instead.
The same habit, everywhere. This is the chat edition of the Cowork line from Part 2 — …ask me questions before you execute. One sentence turns a guessing game into a briefing, and it works in Chat, Cowork and Claude Code alike, because the expert on what you actually want is you.
Why it invents things, and how to catch it.
Back to Part 1's first idea: it predicts likely text. A confident invented answer looks just as likely as a correct one. Nothing checks facts, and nothing flags the moment it stops remembering and starts inventing.
Invented answers sound exactly as assured as true ones. Fluent writing is not evidence, and treating it as evidence is the real risk.
Names, dates, figures, citations, part numbers. Exactly what gets pasted into a customer email unchecked.
Unless told otherwise, it answers. Admitting ignorance is a behaviour you have to ask for.
Give it permission not to know. Add “say you don't know rather than guessing.” Anthropic describes this one line as drastically reducing false information. Free, and almost nobody does it.
Make it quote before it answers. For long documents, ask for the relevant passages word-for-word first, then work from those. It can't ground an answer in a passage it had to invent.
Demand citations, then audit them. Every claim cites a source, then a second pass finds a supporting quote for each, and retract any claim without one. The retraction clause is what makes it work.
Close the door on outside knowledge. “Use only the attached documents, not your general knowledge.” Makes it a closed-book question, where invention has nowhere to hide.
Anthropic's own caveat, verbatim: these techniques “significantly reduce hallucinations” but “don't eliminate them entirely. Always validate critical information.”
Asking Claude “is this right?” doesn't work, because it will agree with you. Four techniques that actually catch things.
It agrees with you because it's trained to. Raters rewarded answers that felt helpful over ones that were blunt, so sycophancy is the objective rather than a quirk. How you ask decides what you get: tell it what you did and you'll get encouragement, because there was no question in it. Ask what's wrong with what you did and you'll get an answer.
Knowing where it gets risky is what makes the rest of this credible. Five situations that need a second thought before you rely on the answer.
Connectors give it reach. Skills tell it how.
MCP is the kitchen: the tools and ingredients Claude can reach. Switch a connector on and Claude can work inside that system. Nothing to install. (The recipes for this kitchen — skills — come three slides from now.)
This is also the real fix for hallucination. Every defence in Part 5 manages a model working from memory. A connector removes the need to remember, because it reads the actual record.
Nothing here is pre-wired for you. You turn on the ones you need and authenticate once. These five are where the work actually lives.
Docs, meeting notes, databases. Where decisions and process live. The most-used connector we have.
Outlook mail and calendar, SharePoint, Teams. Your inbox, your meetings, your shared files.
Channels, threads, DMs, canvases. Where the informal decisions actually got made.
Accounts, opportunities, and the test history behind your reporting.
The live open web: search, any page unblocked, structured data. Every fact in this deck was verified through it.
Calendly, Canva, Docusign, Stripe, tldv, Apollo, Semrush and 50-plus more. Ask if you need one.
Every connector you switch on describes its tools to Claude at the start of the chat. Anthropic's own guidance is blunt about it: tools and connectors are token-intensive, and they suggest switching off the ones you aren't using.
A folder with one markdown file — the recipe for the kitchen you just saw. It says what the job is, how we do it here, and what good looks like. Claude reads it when the job comes up, so you stop explaining the same process again and again. Instructions are how to behave; a skill is how to do a job.
You explain the grading bands again. And the report format again. Every conversation starts from zero, and two people doing the same task get different answers.
You say “grade this.” The recipe is already there. Same result for everyone, and when one person improves the file, everyone gets the improvement.
The description is the part that matters. Claude reads only the name and description of every skill up front, a line or two each, and pulls the full recipe in only when it looks relevant. The description is what makes Claude reach for the skill at all. Vague description, skill never fires. This is also why fifty skills don't drown your context.
| Skill | What it does |
|---|---|
| reprompt | Scores your prompt against the Prompt Standard for clarity, specificity, structure, constraints, verifiability and decomposition, then rewrites it. Everything in Part 4, automated. |
| no-ai-slop | Strips the tells that make writing read as machine-generated, so a draft sounds like a person wrote it. |
Kept in one internal repo, versioned like any other code. · You already have all of these, and you didn't install any of them. How that works: next slide.
A skill is just a folder. Three ways to get it in front of someone else.
A .zip of the folder goes to Organization settings → Skills. It appears in every member's list, switched on by default. Nobody installs anything.
Use for: anything everyone should have.
Zip the folder and send it however you like: Slack, email, a shared drive. They upload it under Customize → Skills and it's theirs to edit.
Use for: outside the company.
Pick the skill in your list and share it with named colleagues. It lands in their Shared with you section, ready to switch on, and view-only so there's still one version.
Use for: sharing something you just built.
It predicts text, and it re-reads rather than remembers. The whole conversation is sent again every turn. Every complaint you've had about Claude follows from that one fact.
More context makes it worse. Quality degrades long before the window fills. One task, one chat.
Point at things; don't paste them. A path or link costs a few tokens. The contents cost thousands and crowd out what matters.
Tell it “I don't know” is allowed. One sentence, free, the best defence against confident invention.
Write down what you'd otherwise repeat. Instructions for how to behave, a skill for how to do a job. Explain it once, benefit forever.