Cursor vs Claude Code vs Codex (2026): Pick by Work Shape
Cursor vs Claude Code vs Codex in October 2026: features, pricing, value for money and a simple rule for choosing by the shape of your work.
· By Guilherme Salgueiro
Put the Cursor, Claude Code and Codex websites side by side this autumn and you will get the uneasy feeling of marking three students' homework after someone left the answer sheet on the desk. All three run agents in the cloud while your laptop sleeps, all three split big jobs across swarms of subagents, all three ship a command-line tool, and all three park a review bot on your pull requests. On 18 September 2026 the last difference on paper vanished, when Claude Code version 2.1.277 started reading AGENTS.md, the instruction file the other two already read. You can spend a whole weekend on comparison tables and finish exactly where you started, only more tired, because the tables measure the one thing that changes every month.
- Cursor, Claude Code and Codex now share almost every headline feature, so pick by the shape of your work, not the feature list.
- Cursor suits developers who want their hands on the code; Claude Code suits work judged by feel, like design and writing; Codex suits long, well-specified jobs across apps and browsers.
- Prices match at $20, $100 and $200 a month; since Opus 5.5, Claude Code and Codex give balanced value, while Cursor burns fast on Claude and GPT.
- Keep one AGENTS.md as your instructions file, run only as many agent threads as you can review, and log two weeks of real tasks before you trust the feeling of speed.
Tools will come and tools will go
In May 2025 Rick Rubin, the producer behind records for Jay-Z, Kanye West, Adele and Justin Timberlake, and the man Justin Bieber went into the studio with when he wanted to sound grown up, published The Way of Code with Anthropic: Lao Tzu's Tao Te Ching rewritten, chapter by chapter, for people who build software by talking to machines. Most of it is gentle and strange, full of valleys, water and uncarved wood, and then chapter 28 lands a line that reads like a verdict on every tool war since: "The Vibe Coder knows the tools, yet keeps to the block. Thus, he can use all things. Tools will come and tools will go. Only The Vibe Coder remains." The block is the uncarved wood that every tool is cut from, which in plain terms means your judgement, your taste and your sense of what the thing should become, and the advice is to know the tools well while keeping your weight on the part of you that doesn't depreciate.
Rubin is the right person to say it, because he has lived the argument for forty years. When Anderson Cooper asked him on 60 Minutes in January 2023 whether he played instruments, he said "barely", and when asked whether he could work a soundboard he went further: "I have no technical ability. And I know nothing about music." What he does know, he explained, is what he likes and what he doesn't, and he is decisive about it, which happens to be the exact job description of someone directing an AI agent. The studios he worked in moved from tape to Pro Tools to laptops, the engineers and the instruments changed with them, and none of it changed what artists were paying him for: "the confidence that I have in my taste." In The Creative Act, the book that interview was promoting, he puts it even more simply: "No matter what tools you use to create, the true instrument is you."
That is the lens for this guide. You still have to pick a tool on Monday morning, and the differences between Cursor, Claude Code and Codex are real enough to cost you weeks if you pick badly, so the sections that follow compare them with dated facts and without mercy. But the comparison is arranged around the part that stays, because a vibe coder who knows what good looks like will get good work out of any of the three, while someone without that ear will get confident nonsense out of all of them, only at different speeds and prices.
Seen from the producer's chair, the three tools stop looking like rival products and start looking like three different ways of making a record. Cursor is playing the instrument yourself with a gifted session player at your elbow, who finishes your phrase before you've found it and lets you keep or erase every note, so your hands never leave the keys. Claude Code is producing a band in the room, the way Rubin does it from the couch: you describe the feeling you're after, the band plays it, and you stop the take, ask for it "again, but slower and sadder", and listen once more. Codex is sending your tracks to a mastering house with a precise brief, then getting back a finished master and a spec sheet proving it met the brief, without ever hearing the work in progress. Each is a legitimate way to make a great record, and each produces disasters when you choose it for the wrong song, because you don't send a half-written demo to mastering and you don't book a full band to fix one wrong note.
So the questions worth asking are about you rather than the tool. What shape is the work in front of you, and how much listening can you honestly do before your ears get tired? The second question matters more than any benchmark, because in Stack Overflow's 2025 Developer Survey 66% of developers named AI answers that are "almost right, but not quite" as their biggest frustration, which makes your attention the scarcest instrument in the room. And before we walk through the three studios, it's worth seeing how fast chapter 28 is already coming true, because in the last three months alone one of these tools was sold, another was folded into a bigger product, and the third swapped its engine.
Feature by feature: what Cursor, Claude Code and Codex share, and where they still differ
Before taste, the facts. The table below compares the three tools as they stand on 8 October 2026, checked against each vendor's own docs and pricing pages, and it is worth reading slowly, because about two thirds of it is the same in every column. That shared ground is the good news for anyone choosing: whichever you pick, you get an agent that edits across files, runs commands, works in the cloud, reviews pull requests and reads an open instruction file. The rows where the columns diverge are the ones that should drive your choice.
| | Cursor | Claude Code | Codex | |---|---|---|---| | Started life as | An AI code editor (a VS Code fork) | A terminal agent | A cloud agent plus an open-source CLI | | Where you use it | Editor and Agents Window, CLI (agent), web, iOS and iPad, Slack, JetBrains | Terminal, VS Code and JetBrains, desktop app, web, mobile, Slack | ChatGPT desktop app, CLI, IDE extension, web, mobile, Slack, Linear | | Models | Composer 2.5 and Grok (latest: Grok 4.7, 21 Sep 2026) in the cheaper in-house pool; Claude and Gemini at API rates; OpenAI models due to leave (proposed shutoff 12 Nov 2026) | Claude only: Opus 5.5 (default), Sonnet 5.5, Fable 5.1, Haiku 5.5 (since 7 Oct 2026) | OpenAI only: GPT-6.1 Sol (recommended default), GPT-6 Astra, GPT-6 Luna | | Reads AGENTS.md | Yes, including nested folders | Yes since 18 Sep 2026, but by default only when there is no CLAUDE.md | Yes, natively (32 KiB combined cap) | | Reads CLAUDE.md | Yes, always applied | Yes, natively | Only if you add it as a fallback name | | Agents in the cloud | Cloud Agents on their own VMs | Cloud sessions on the web, scheduled routines | Codex cloud tasks, scheduled tasks | | Parallel work | Cloud subagents, Projects coordinator (beta) | Subagents, worktrees, workflows, agent teams (experimental) | Subagents, worktrees, Ultra reasoning that splits work across agents | | Pull request review | Bugbot (usage-billed on Pro, included on Teams) | Code Review and the GitHub Action | @codex and automatic PR review | | Browser and computer use | Built-in browser; cloud agents can use a desktop | Desktop app browser, Chrome extension, computer use on macOS and Windows (CLI: macOS only) | In-app browser, computer use on macOS and Windows | | Extend it with | Rules, skills, hooks, MCP, plugins (also loads Claude Code skills and hooks) | Skills, hooks, plugins, mods, MCP | Skills, plugins, hooks, MCP, /import from Claude Code and Cursor | | Safety by default | Sandbox toggle, review every diff in the editor | Permission modes; auto mode is the default since 28 Sep 2026 | OS-level sandbox; workspace-write with approvals on request | | Entry plan | Pro, $20/month | Pro, $20/month | Plus, $20/month (Free and Go get Luna only) | | Heavy-use plan | Ultra, $200/month | Max 20x, $200/month | Pro, $100, $200 or $500/month |
Three rows explain most of the real-world differences. The models row matters most, because each tool is now married to a model family: Claude Code runs only Claude, Codex runs only OpenAI, and Cursor, once the neutral ground where you could switch between everyone, is tilting towards its new owner's models. The safety row tells you how each tool expects to be supervised, with Codex fencing the agent in at the operating-system level, Claude Code trusting a permission system and a classifier, and Cursor assuming your eyes are on the diff. The browser and computer-use row looks close on paper, since both Claude Code and Codex can now drive a browser and a desktop, but it is where Codex's polished app and plugin ecosystem make the bigger difference for work beyond code, which matters if half your week is research, spreadsheets and admin rather than features.
What changed this summer
Most comparisons still ranking on Google were written before three changes that reshaped the table above, which is chapter 28 playing out in a single season.
- Cursor was sold. SpaceX's $60 billion all-stock acquisition closed on 14 August 2026 and made Cursor a wholly owned SpaceX subsidiary, a sister company to SpaceXAI, formerly xAI. Grok joined Cursor's in-house model pool, with Grok 4.7 (released 21 September) the latest, and SpaceXAI's model card says Grok 4.6 was trained partly on anonymized Cursor workflow data. On 28 August OpenAI announced it would stop supplying its models to Cursor, with a proposed shutoff on 12 November.
- Codex moved into ChatGPT. On 9 July 2026 the Codex desktop app became the ChatGPT desktop app, with Codex as one view next to Chat and Work, and GPT-6.1 Sol has been OpenAI's recommended model for Codex since 29 September.
- Claude Code changed engines. Opus 5.5 became the default model on 22 September 2026, four days after Claude Code started reading AGENTS.md, a feature requested in a GitHub issue that had gathered over 5,000 thumbs-up since August 2025.
None of this makes any of the three a bad choice, but it does change what is worth investing in. Hours spent memorising one tool's quirks or one vendor's billing unit depreciate on somebody else's schedule, while an instruction file in an open format, a habit of reviewing what comes back and a clear sense of what good looks like move with you from one studio to the next. With the table in hand, the next three sections take each tool in turn: what it does best, where it lets you down, and which work I give it.
Cursor: the best choice when you want your hands on the code
Cursor is the session player at your elbow. It began as a fork of VS Code, the most popular code editor in the world, and its whole design assumes that you are reading and touching the code yourself while the AI suggests, completes and edits alongside you. That is no longer the whole story, though, because 2026 has been the year Cursor learned to work when your hands are off the keys, and most of what it shipped since April is about agents running without you watching.
Where Cursor is strongest
- Tab completion. Cursor's own autocomplete model predicts your next edit, often several lines ahead, and is still the one developers single out as the best in any editor.
- Reviewing changes. Every agent edit arrives as a visual diff you can keep or undo chunk by chunk, which is the most comfortable way of the three to stay in control of each line.
- Cloud Agents. Each one runs on its own virtual machine with a full desktop, so it can open what it built and check it before handing it back, and you can start one from Slack, Linear, GitHub, the web or your phone and close the laptop while it works.
- Parallel agents in one window. The Agents Window, launched with Cursor 3 on 2 April 2026, runs several agents side by side, locally or in the cloud, each in its own worktree, so three attempts at a problem don't trip over each other.
- Projects. In beta since 10 September 2026, Projects adds a coordinator agent that breaks a larger piece of work into parts and hands them to subagents, which is the closest Cursor gets to running a small team rather than a single assistant.
- Remote control from your phone. Since 6 October 2026 the iPhone app can steer agents running on your own computer, not only the ones in the cloud.
- Compatibility. It is the most open of the three to other tools' setups: it reads AGENTS.md, CLAUDE.md and Claude Code's skills and hooks without any conversion, so a repository set up for Claude Code or Codex works in Cursor as it is.
Where it falls short
- Models. The cheap, generous pool is now Cursor's own: Composer 2.5 and xAI's Grok models, the latest being Grok 4.7 from 21 September. Claude and Gemini are still available, but billed at API rates, while Claude Code and Codex subscriptions bundle their frontier models into the plan, and OpenAI's models are due to leave Cursor on 12 November.
- Pricing you can't read. Cursor moved Pro from 500 requests to a usage allowance in June 2025, apologised and refunded users in July, and has changed its billing unit several times since. In August 2026 it stopped publishing how much frontier-model usage each plan includes, so you find your limit by hitting it.
- Weight. It is a full desktop IDE, and reviewers still complain about memory and CPU use on large projects.
My take
I don't use Cursor, and the reason says more about the reader than about the tool. Cursor's main selling point used to be access to every model in one place, but I already pay for Claude Code and Codex, so I get Anthropic's and OpenAI's best models directly, and the in-house alternatives, Composer 2.5 and Grok, were noticeably less reliable in my work. Setting up parallel worktrees also took more fiddling than in Claude Code or Codex, where it works out of the box. As a surface I like it less than Claude Code and much less than Codex. If you are a developer who wants to stay inside the code and collaborate with the AI line by line, Cursor may well be the most natural home. For a non-developer like me, though, an editor built around reading code adds nothing that the other two desktop apps don't already do better. The one exception is Grok Bot, the persistent cloud assistant that now comes with paid Cursor plans, which in my experience is probably the best product in its category. In studio terms, the session player is wonderful if you play the instrument, and an expensive piano in the corner if you don't.
Claude Code: the best choice when you want to direct and still hear every take
Claude Code is the band in the room, with you on the couch. It started in February 2025 as an agent that lived in the terminal, and its design still assumes that you describe what you want in plain words, let the agent do the playing, and stop the take whenever it drifts. The terminal is now only one of its doors, because the same agent runs in VS Code and JetBrains, in a desktop app, on the web and on your phone. What changed this autumn is bigger than another door, though. In the space of three weeks the band got a new lead singer, a manager who runs several sessions at once, and a set designer, and Claude Code quietly stopped being a coding tool and became a studio for making almost anything.
Where Claude Code is strongest
- The strongest model right now. Opus 5.5 became the default on 22 September 2026, and in my experience it is the best model available today, in any tool. Anthropic's own claim is that it matches the bigger Fable 5.1 on most work while costing 40% less to run than Opus 5, and Sonnet 5.5 (28 September) and Haiku 5.5 (7 October) renewed the rest of the lineup within three weeks. What I notice most is not a benchmark but reliability: it does what I asked, without the detours.
- Projects: a manager for your sessions. Relaunched in beta on 17 September 2026, a project is one conversation where you drop work as it comes, and Claude splits it into threads, runs them in parallel in the cloud, keeps shared instructions and memory across them, and tells you which are finished, which opened a pull request and which are waiting on you. It is the difference between playing every instrument yourself and briefing a producer who books the players, and it keeps working after you close the laptop.
- Design, inside the session. Claude Design launched in April 2026 as a separate site, has worked inside every Claude conversation, Claude Code included, since 16 September, and left beta on 8 October. From a coding session you can ask for a design, sync your existing design system with
/design-sync, refine it on a canvas and hand it straight back to code, which closes the gap between "what should this look like" and "build it" that used to mean a round trip through another tool. - Dashboards and Motion, from 8 October 2026. Dashboards, in beta on paid plans, connects to a data source such as BigQuery, Snowflake or Salesforce and builds a dashboard that stays current. Motion, in beta on Team and Enterprise plans only, turns a report or a chart into a short animated explainer you can export as an MP4. Neither is about code, which is exactly the point: the same agent that fixes your bug can now present the result.
- Shaping how it works. No tool of the three gives you more control over the agent's habits. CLAUDE.md and its own notes carry your preferences between sessions, skills teach it repeatable jobs, hooks can block an action before it happens, and plugins package all of that for a team.
- Less babysitting by default. Since 28 September 2026 new sessions start in auto mode, where a classifier approves routine actions and stops the risky ones, so you answer far fewer permission questions without handing over the keys entirely.
Where it falls short
- Limits within limits. Claude Code shares one pool with Claude chat, Design and everything else, with a five-hour window and a weekly cap on top, and Fable has its own ceiling inside that pool: on Max it can use at most half of your weekly limit, burns it faster than other models, and still counts towards the total, while on Pro it is billed separately as usage credits from the first message. Projects make this sharper, because every thread is a full session drawing on the same meter. The pool itself has moved a lot, settling on 14 September 2026 at 25% above the 2025 baseline, which Anthropic admitted was a 17% cut compared with the temporary boost before it.
- A trust dent that took a while to fix. In April 2026 Anthropic published a postmortem admitting that three product changes made between March and April had quietly made Claude Code worse, including a lower default effort level and a bug in how it remembered its own reasoning. It reverted them and reset everyone's limits, but for a few weeks many users had good reason to wonder whether the tool or their prompts were to blame.
- New features arrive in stages. Projects started with select Pro and Max users, Motion is Team and Enterprise only, and several of the best parts are still labelled beta, so what you read about this month may not be in your account yet.
- AGENTS.md with conditions. It reads AGENTS.md since 18 September, but by default only when there is no CLAUDE.md in the project, and it ignores the shared
.agents/skillsfolder the other tools are starting to use.
My take
Claude Code was the first coding tool I picked up after Lovable, and I started using it the moment it came out. For many months it was the only tool I used, and almost everything I built until about four months ago was built in it. Like every model family, though, Anthropic's has had its hits and misses, and some of its releases were disappointing: in the months before this autumn Codex was ahead more than once, and I started splitting my coding and AI work between the two. Opus 5.5 is what won me back. It is the strongest model I have used, it feels faster in my hands than Codex on GPT-6.1 Sol, because it writes nearly twice as fast even if Sol tends to finish whole tasks sooner, and it talks to me in plain English instead of filling its reports with "contracts" and "receipts" that I have to translate before I can judge them. That plain speech matters more than it sounds, because trust comes from understanding what happened, and you can only direct a band whose playing you can follow. Today it is where I go for anything that needs design sense or creativity; I run Projects to keep several pieces of work moving at once, and I reach for Design when an idea needs a face. The shared limits are the price of admission. Tools will come and tools will go, and I have watched this one fall behind and come back, but right now, when the work needs taste, this is the band I want in the room. If your work is mostly design, writing or anything where you judge the result by feel rather than by a test, that is the case for starting here too.
Codex: the best choice when the work goes beyond code
Codex is the best work surface of the three, and the only one that mixes coding and knowledge work so well that you stop noticing the switch. It still does what made its name, taking a precise brief and coming back with a finished change and proof that it met the brief, but it now lives inside the ChatGPT desktop app, next to Chat and ChatGPT Work, and that app is the easiest and the most beautiful place to work with an agent today. It rules at computer use and browser use, and its MCP connections simply hold: where Claude Code still drops a server and asks you to reconnect, Codex keeps working. If the mastering house was where you sent finished tracks, Codex has become the whole studio complex, with the best equipment in every room.
Where Codex is strongest
- The best surface of the three. The desktop app is the most polished way to work with an agent today: projects in a sidebar, parallel threads, built-in worktrees, a review panel for pull requests and an in-app browser where you can comment directly on a page. It feels like a product designed around people rather than around a terminal, which matters when you spend whole days inside it.
- Computer and browser use that actually works. Since April 2026 Codex can drive its own browser and your desktop, first on macOS and then on Windows, so it can test what it built, fill in a form or pull numbers from a site that has no API. Claude Code can do this too, but Codex does it more smoothly and with less setup.
- Plugins and MCP that stay connected. MCP servers in Codex are noticeably more stable than in Claude Code, with far fewer reconnects, and plugins, added in March 2026, bundle skills, MCP servers and app integrations into one install, and since 25 August scheduled tasks can start when an email lands in Gmail, a message arrives in Slack or something happens on GitHub. That is what makes Codex good at work that has nothing to do with code: research, reports, inboxes, spreadsheets.
- Long, unattended jobs. Goal mode, out of experimental since 21 May, lets Codex work towards one objective for hours or days, and Ultra reasoning splits a hard problem across several agents working in parallel. Since 25 June you can start or steer all of it from your phone.
- Cheaper and quicker per finished task. On Artificial Analysis' Coding Agent Index (read 1 October 2026), Codex with GPT-6.1 Sol at xhigh effort scores 63 and finishes an average task in 15.5 minutes for $1.04 at API prices. Claude Code with Sonnet 5.5 reaches the same 63 at xhigh in 27 minutes for $3.33, and its top scores, 68 with Sonnet 5.5 and 66 with Opus 5.5 at max effort, take 87 and 65 minutes at $14.19 and $13.04 a task. Those are pay-per-token prices, not what your subscription costs, but the gap in time is real.
- Safe by default. Codex fences the agent in at the operating-system level, so out of the box it can only write inside your project, and independent reviewers rate it the safest of the three before you change a single setting.
- An easy way in.
/importbrings your Claude Code or Cursor settings, skills, MCP servers and recent sessions across in one step, and nobody else offers the same in the other direction.
Where it falls short
- The models talk too much. OpenAI's models are verbose, their reports are harder to read than they should be, and they have a habit of adding things nobody asked for, an extra layer of abstraction here or a defensive check there, which is fine in a demo and expensive when you have to review every line.
- A bumpy model history. Codex has been through more model changes in 2026 than any other tool: the dedicated GPT-5.x-Codex line ended in February, GPT-5.4, 5.5 and 5.6 followed in quick succession, GPT-6 Astra arrived on 3 September and GPT-6.1 Sol took over as the recommended default on 29 September. GPT-5.5 leaves Codex on 14 October, so the model you tuned your habits around this summer is already on its way out.
- Limits that move. OpenAI has reset and rebalanced Codex usage several times this year, and since Astra arrived users report that it drains a weekly allowance far faster than earlier models did.
- A coding tool inside a bigger product. Codex is now one mode of OpenAI's everything-app. That is great if you live in ChatGPT, but developers who only want a coding agent get a lot of product around it.
- CLAUDE.md is invisible by default. Codex reads only AGENTS.md, capped at 32 KiB, unless you tell it to fall back to other file names.
My take
Codex is the better tool, and Claude Code is the better collaborator, which is why I pay for both. When Codex pulled ahead of Claude this summer I started splitting my work, and that split has stuck: Codex is where I go for knowledge work, for anything that needs a browser, a desktop or a long chain of plugins and MCP servers, and it has the nicest interface of any agent I have used. The models are the catch. GPT-6 was a disaster for me, GPT-6.1 Sol is a real improvement on 5.6, but even at its best it feels slower in my hands, wordier and less clear than Opus 5.5. The numbers explain the feeling: Opus 5.5 writes about 96 tokens a second against Sol's 54, so it reads faster while you watch, even though Sol finishes the whole task sooner because it takes fewer, shorter steps. And Sol still likes to add things that are not critical to what I asked. In studio terms, Codex has the best building in town and the best equipment in every room, and I book it for the sessions where the equipment matters most. If your work is more operations than craft, more inbox, browser and spreadsheet than design and prose, that is exactly the case for making Codex your first studio rather than your second.
One instructions file can serve all three tools
Every session musician who walks into a studio gets the same thing on the music stand: a lead sheet with the chords, the key and a few notes on feel, written so that any good player can pick it up and play. AGENTS.md is that lead sheet for coding agents, and since 18 September 2026 it is the first one all three tools can read, which is the most practical gift this autumn has given anyone who works with more than one of them. You write down once how your project works, what to avoid and how to check the work, and the same file briefs Cursor, Claude Code and Codex alike, so you stop maintaining three slightly different versions of the truth and wondering which one each agent believed.
The catch is that each tool reads the sheet slightly differently, and one of those differences can quietly switch it off:
- Claude Code reads AGENTS.md only when the project has no CLAUDE.md. Add a CLAUDE.md for a Claude-only setting, or even a private
CLAUDE.local.md, and AGENTS.md goes silent unless you import it. - Codex reads AGENTS.md natively, up to 32 KiB in total, and ignores CLAUDE.md unless you add it as a fallback name in its settings.
- Cursor reads both files and applies CLAUDE.md to every conversation, so anything you write in both files reaches Cursor twice, and anything Claude-specific in CLAUDE.md reaches Cursor too.
The setup that works in all three is a single AGENTS.md as the source of truth and, only if you need Claude-specific extras, a short CLAUDE.md whose first line imports it:
Anthropic's own docs recommend exactly this pattern, and say that keeping the import never makes Claude read AGENTS.md twice. Because Cursor reads CLAUDE.md as well, keep that file to the import and a few Claude-only lines, and never copy the shared rules into it. Keep AGENTS.md itself short, too: an ETH Zurich study from February 2026 found that context files did not generally make coding agents more successful but raised their running cost by over 20%, whether an AI or a developer wrote them, and that repository overviews in particular did not help. Where the files earned their place was in spelling out the unusual rules of a project, the things an agent could not guess on its own. I have written about the file itself in Start with CLAUDE.md. The short version is that the lead sheet is the one part of your setup that moves with you when the tools change, which makes it the best hour you will spend on any of them.
What Cursor, Claude Code and Codex cost in October 2026
Studio time has always been sold in two ways: by the hour, where you pay for exactly what you use and wince at every overrun, or by the block, where you pay up front and the clock stops mattering until you hit the end of the booking. All three tools now sell a block with an hourly meter hidden behind it, and the sticker prices look almost identical, which is precisely why they are misleading.
| | Cursor | Claude Code | Codex | |---|---|---|---| | Free tier | Hobby: limited agent requests | No Claude Code on Free | Free and Go ($8): GPT-6 Luna only | | Entry plan | Pro, $20/month | Pro, $20/month ($17 billed yearly) | Plus, $20/month | | Middle plan | Pro+, $60/month | Max 5x, $100/month | Pro, $100/month | | Heavy-use plan | Ultra, $200/month | Max 20x, $200/month | Pro, $200 or $500/month | | Team seat | Teams, $40/user/month | Team Standard, $25/user/month ($125 Premium) | Business, $25/user/month ($125 Premium) | | What the plan buys | A usage allowance, generous on Composer and Grok, at API rates for Claude and Gemini | A five-hour window plus a weekly cap, shared with all of Claude; Fable up to half the weekly cap on Max | Usage shared with ChatGPT Work, measured in messages per five hours and weekly limits | | When you run out | On-demand usage, billed afterwards | Usage credits at API rates, if you turn them on | ChatGPT credits |
Prices are monthly in US dollars before tax, checked on each vendor's pricing page on 9 October 2026.
The $20 tier in all three is a test drive rather than a working plan: enough for a few focused sessions a week, and Cursor's own docs admit that daily agent users typically spend $60 to $100 a month and power users $200 or more. The real comparison starts at $100 to $200, and there the difference is less the price than the meter. Codex publishes the clearest numbers, a range of messages per five hours for each model, although OpenAI has reset and rebalanced them several times this year. Claude Code publishes only multiples, Max being "5x or 20x" Pro, and its pool is shared with everything else you do in Claude. Cursor's limits are the hardest to see, because it stopped publishing what each plan includes in August.
Value for money is where my own view has moved most this year. For a long stretch Claude Code burned through its limits fast, and Codex was the better value for money, comfortably. Since Opus 5.5 arrived on 22 September, that has changed. Anthropic says it uses fewer tokens per task than Opus 5 and raised five-hour limits on the same day, and in normal, reasonable use I no longer hit mine, and today I would call the two balanced. Both give good value for what they cost, so the choice between them comes down to the shape of your work rather than the price. Cursor is the odd one out. Its allowance resets monthly rather than weekly, and on its own models, Composer and Grok, it lasts a very long time; switch to Claude or GPT and even the top plan burns fast, because those run at API rates. The shame is that the models that last are not the ones you most want doing the work.
What no price table shows is how much of the booking you will actually use. Theo Browne made the bluntest version of this point in a video on 1 October 2026: once you have paid, "your ability to work and put out even bigger things is only limited by your token usage." His test is how many agent threads you have running right now, and his bar is five. I would keep the idea and drop the number. A plan earns its price when it is busy on work you had already decided was worth doing, ideally while you sleep, and an expensive plan sitting idle is the worst deal of the three. Five threads of the wrong song are still the wrong song, though, which is why the cheapest hour of studio time is the one you don't waste on a take you never needed, and where the next section comes in.
Which one to pick: start from the work, not the tool
A good song can still be ruined by the wrong room. An intimate ballad drowns in a hall built for an orchestra, and a loud band sounds boxed in inside a vocal booth, even when the musicians and the gear are exactly the same. The same is true here, which is why "which tool is best?" is the wrong first question. The better ones are about the work itself: what shape is it, and how much of what comes back can you honestly listen to before it ships?
Shape comes first, and it is easier to read than it sounds:
| If the work is mostly… | Start in | Because | |---|---|---| | Code you want to read, shape and own line by line | Cursor | You stay at the keyboard, with an agent one keystroke away | | Judged by feel: design, writing, product decisions, a blank page | Claude Code | The strongest model right now, and the clearest collaborator while the brief is still forming | | Well specified, long, or spread across apps, browsers and documents | Codex | The best work surface, computer use that works, and jobs that keep running while you're away | | A bit of everything, and you are not a developer | Claude Code or Codex | Cursor's edge lives in the editor, the one place you won't spend your day |
The second question is the one no pricing page asks. Every agent thread hands you something to listen back to before it ships, and your ears do not scale the way the threads do. This is why I keep Theo's idea and drop his number: the right number of threads is however many you can review properly in the time you actually have, and for most people on most days that is fewer than their plan allows. A thread that finishes and then sits unheard until tomorrow was never saving you time. It was only queueing it, which is the same lesson I found in running several agents without losing judgement.
Do you need two studios? I pay for both, Claude Max and ChatGPT Pro at $200 a month each, because my work splits between them, and each does its half better than the other would. If yours doesn't split, one is enough, and the saving is the smallest part of it: one set of instructions, one set of habits, one place to look when something breaks. If you do add a second, give each one a job rather than asking both the same question, which I've written about in Stop Being Loyal to One AI Stack.
Whichever you pick, you will feel sure about it within a week, and that feeling is exactly what the last section is about.
Only the vibe coder remains, so measure like one
The line this guide opened with began as a joke. In February 2025 Andrej Karpathy coined "vibe coding" for building software by describing what you want and letting the AI write it, and within weeks the internet had chosen its face: a photo of Rick Rubin, eyes closed, headphones on, one hand resting on a mouse. Rubin had never seen the picture and assumed it was AI-generated, but it was real, taken at a hi-fi show in Germany where the mouse was only turning the volume up. The week after that, he said, "a company called Cursor, which I don't really know what they do," published fifteen rules of vibe coding with his photo at the top. So he did what a good producer does with an accident and leaned into it. He posted the first joke of his life, "Tools will come and tools will go. Only the vibe coder remains," watched it pass a million views when his daily posts usually drew tens of thousands, and turned the joke into a book, The Way of Code, with the gloriously deadpan subtitle The Timeless Art of Vibe Coding, for a term that was, in his words, "maybe ten weeks old."
He wasn't only joking, though. Before vibe coding, he said, you had to be a virtuoso coder to make something great, "and now everybody can do it." That is what these three tools really are under the pricing tables: they hand the designer with an app in a sketchbook, the writer with a product in their head and the founder who never learned to code everything they need to put an idea live. I'm not an engineer either, and the site you're reading this on exists because describing it turned out to be enough.
That freedom has a catch. Making something is now easy, but easy doesn't tell you whether it is any good, and the honeymoon with a new tool makes that harder to hear: the first big job lands in minutes, and from then on every good result confirms the switch while every bad one gets filed under bad luck. That is outcome bias, the habit this blog is named after, judging a decision by one result rather than by whether it was sound. Don't expect the researchers to settle it for you either. METR, the lab that caught developers working slower with AI in 2025 while they believed they were faster, tried to run its study again and said in February 2026 that the new data was "unreliable", partly because so many developers now refuse to work without AI, and partly because nobody could time a task cleanly once several agents were running at once.
So listen to your own work the way Rubin listens to a take. When he describes how he makes things with AI, it is never a single prompt and a verdict: he asks for different versions, compares them, narrows them to two he likes, often combines the best of both, and stays "open to being wrong." For two weeks, do the same with your tools. Give each one the kind of work the table in the last section sends its way, and keep a one-line note per task: what you asked for, how long until you had something you could use, how many minutes you spent checking it, and whether any of it had to be redone. Ten tasks a tool is enough to see a pattern, and what you are looking for is not the fastest take but the lowest total cost of reaching something you would ship, because that is where the honeymoon hides its bill.
By this time next year at least one of these three studios will look nothing like it does today, and every comparison table, this one included, will need rewriting. Your notes won't, because what they train is the thing Rubin has been selling for forty years without playing a note: knowing what you like, knowing what you don't, and checking the one against the other. Pick a studio, write down what happens, and make the thing only you would make. Tools will come and tools will go. The joke was right.
Frequently asked questions
Which is better, Cursor, Claude Code or Codex?
None wins outright in October 2026. Cursor is best when a developer wants to stay in the code, Claude Code when the work is judged by feel (design, writing, product decisions), and Codex for long, well-specified work across apps, browsers and documents.
Is Claude Code or Codex better value for money?
Since Opus 5.5 launched on 22 September 2026, the two are roughly balanced. Codex used to give more usage per dollar, but Anthropic cut Opus 5.5's token use and raised five-hour limits, so at $100 to $200 a month both give good value.
Can one instructions file work for Cursor, Claude Code and Codex?
Yes. Since 18 September 2026 all three read AGENTS.md. Keep it as the single source of truth and, if you need Claude-only settings, add a short CLAUDE.md whose first line is @AGENTS.md.
Do I need to pay for more than one AI coding tool?
Only if your work splits cleanly, for example craft in Claude Code and operations in Codex. If it doesn't, one tool is enough, with one set of instructions and habits.
How do I know if an AI coding tool actually makes me faster?
Log two weeks of real tasks: what you asked for, time to a usable result, minutes spent reviewing, and whether anything had to be redone. Even METR said in February 2026 that agent work has become hard to time, so your own log beats the feeling of speed.