Blog
Insights
Grok vs ChatGPT: Grok 4.7 vs GPT-6 (2026)
Grok 4.7 vs GPT-6 Astra compared on benchmarks, pricing, speed, and true cost per task - plus why neither new model is in the chatbot you actually use.

Nafis Amiri
Co-Founder of CatDoes

TL;DR: Both labs shipped new flagships in the same week. Grok 4.7 landed September 21, 2026; GPT-6 Sol and GPT-6 Luna landed September 22. Neither one is in the chatbot you actually use. Grok 4.7 is API, Cursor, and Grok Build only, and SpaceXAI's own plan pages still say Grok 4.6. GPT-6 Sol and Luna run in ChatGPT Work and Codex but not in Chat, where Plus still gets GPT-5.6 Sol. On the one index that measures both flagships the same way, GPT-6 Astra scores 53 to Grok 4.7's 46. Grok is 5.7x cheaper per token but costs 15% more per finished task, because it writes three times as much.
Table of Contents
Grok vs ChatGPT at a Glance
What Changed in September 2026
Which Model You Actually Get
Pricing and Plans
Benchmarks: GPT-6 Astra vs Grok 4.7
Coding and Development
Real-Time Data and Web Search
Writing and Creative Tasks
AI Agents: ChatGPT Work vs Grok Bot
Image and Video Generation
Which Should You Choose?
Beyond Chatbots: Building Apps With AI
Frequently Asked Questions
Grok vs ChatGPT at a Glance
Grok vs ChatGPT got harder to answer in September 2026, not easier. Both labs shipped new flagship models one day apart, and in both cases the model you read about is not the model your subscription runs.
ChatGPT is OpenAI's chatbot. Its top model is GPT-6 Astra, released September 3, 2026, joined on September 22 by the cheaper GPT-6 Sol and GPT-6 Luna. OpenAI reported 900 million weekly active users and 50 million paying subscribers as of February 2026.
Grok is built by SpaceXAI, a trading name of xAI LLC after SpaceX acquired the company in February 2026 and rebranded it that July. Its newest model is Grok 4.7, released September 21, 2026. SpaceX's IPO filing put Grok at roughly 117 million monthly active users in March 2026, though that figure counts Grok features used inside X rather than standalone app usage.
Feature | ChatGPT | Grok |
|---|---|---|
Newest model | GPT-6 Sol and Luna (Sept 22, 2026) | Grok 4.7 (Sept 21, 2026) |
Top model | GPT-6 Astra | Grok 4.7 |
In the consumer chat app | GPT-5.6 Luna or Sol, by plan | Grok 4.6 |
Maker | OpenAI | SpaceXAI (formerly xAI) |
Context window | 1,050,000 tokens | 500,000 tokens |
Intelligence Index (v4.3.2) | 53 (Astra), 48 (Sol) | 46 |
Flagship API price per 1M tokens | $10 in / $50 out | $2 in / $6 out |
Entry paid plan | $8/mo (Go) | $10/mo (SuperGrok Lite) |
Standard paid plan | $20/mo (Plus) | $30/mo (SuperGrok) |
Web traffic share (Aug 2026) | 55.5% | 2.4% |
Market share figures come from Similarweb's August 2026 data and measure website traffic only, so they exclude mobile apps. On app-based measurement the gap narrows: Sensor Tower put ChatGPT at 46.4% of the assistant app market through May 2026. Grok also peaked far higher in one specific slice. Apptopia data reported by Reuters showed Grok reaching 17.8% of US app usage in January 2026, up from 1.9% a year earlier, before falling back over the following months.
What Changed in September 2026
Three releases landed inside 36 hours, and the ordering matters for how you read the benchmark claims.
Grok 4.7 arrived September 21. It uses a larger base model than Grok 4.6, trained with a longer reinforcement learning run weighted toward tasks that take hours rather than minutes. Pricing did not move: $2 per million input tokens and $6 per million output, exactly what Grok 4.6 cost. This is a capability upgrade at an unchanged price, not a price cut.
Claude Opus 5.5 arrived September 22, roughly 90 minutes before OpenAI's launch. It matters here only as context: it currently tops both Artificial Analysis's index at 58 and the WebDev Arena leaderboard, above anything either lab in this comparison ships. Neither OpenAI's nor SpaceXAI's launch table includes it, because both were locked before it shipped.
GPT-6 Sol and GPT-6 Luna arrived September 22. These extend the GPT-6 family downward rather than replacing Astra. Sol costs $2 in and $10 out per million tokens, half what GPT-5.6 Sol cost. Luna costs $0.10 and $0.50, half the input price and 58% less on output. OpenAI has confirmed these are permanent prices, not introductory rates.
Worth noting for anyone reading older coverage: Grok 4.7 spent months as a rumour while Elon Musk's public timeline slipped, so comparisons written before September 21 correctly said it did not exist. That is now resolved. Grok 5 remains unreleased.
Also unchanged since June 2026: SpaceXAI replaced per-product daily caps with a single weekly usage pool shared across Chat, Imagine, Voice, and Build, shown as a percentage in your settings. Any article quoting a fixed "10 prompts every two hours" limit for Grok is describing a system that no longer exists. SpaceXAI still publishes no numeric allowance for any tier, so treat specific figures you find elsewhere as invented.
Which Model You Actually Get
This is the part almost every comparison gets wrong right now. Neither September flagship is available in the consumer chat product, and both companies say so in their own documentation.
SpaceXAI's Grok 4.7 model card states that the company "plans to add Grok 4.7 to its consumer surfaces (web, mobile apps, and Grok-in-X on the X platform) at a later date." As of publication, x.ai's pricing page still describes the $30 SuperGrok tier as running the "Grok 4.6 model," and grok.com's plan page headline still reads "Unlock the full power of Chat with Grok 4.6." Grok 4.7 currently runs in the SpaceXAI API, Grok Build, Cursor, GitHub Copilot, and the Grok Office add-ins. No subscription tier puts it in the chat app.
OpenAI is equally explicit. Its model documentation says that in ChatGPT, "GPT-6 Sol and GPT-6 Luna are available in Work and Codex. They aren't available in Chat." The chat picker is a generation behind across the board.
Plan | Model in the chat window | Newest model you can reach |
|---|---|---|
ChatGPT Free | GPT-5.6 Luna | GPT-6 Luna, desktop app only |
ChatGPT Go | GPT-5.6 Luna | GPT-6 Luna, desktop app only |
ChatGPT Plus | GPT-5.6 Sol | GPT-6 Astra, Sol, Luna in Work and Codex |
ChatGPT Pro / Business | GPT-6 Astra | GPT-6 Astra in chat and Work |
Any Grok plan | Grok 4.6 | Grok 4.6 (4.7 is API and tools only) |
The practical consequence: if you are comparing the two chatbots as chatbots, you are comparing GPT-5.6 Sol against Grok 4.6, not GPT-6 against Grok 4.7. Every benchmark below describes models you reach through a developer surface, a coding agent, or a $100-plus plan.
Sam Altman called the Astra rollout "messy" on September 4 after Plus subscribers went looking in the chat picker for a model their plan appeared to include. Three weeks later, OpenAI's help documentation still lists GPT-6 Astra in chat for Pro, Business, and Enterprise only, with Plus limited to Work and Codex.
Pricing and Plans
ChatGPT has six consumer and business tiers. Grok has seven, and they overlap confusingly because some are billed through X.
ChatGPT plan | Price | What you get |
|---|---|---|
Free | $0 | GPT-5.6 Luna, unlimited text chats, ads |
Go | $8/mo | Higher upload, image and memory limits, still ads |
Plus | $20/mo | GPT-5.6 Sol in chat, ChatGPT Work, Codex, ad-free |
Pro (lower) | $100/mo | 5x Plus limits, GPT-6 Astra in chat |
Pro (upper) | $200/mo | 20x Plus limits — new sign-ups paused since Sept 10 |
Business | $20-25/seat/mo | $20 annual, $25 monthly, 2 seat minimum, Astra access |
Three details trip people up. Go is a paid plan that still shows ads, and the ad-free toggle is a Free-plan option, so Go subscribers pay $8 and cannot switch ads off. Free and Go both now include unlimited everyday text chats, so the gap is in uploads, images, memory, and data analysis rather than raw message volume. And Pro is documented as two prices, but only the $100 tier is currently purchasable.
Grok plan | Price | What you get |
|---|---|---|
Free | $0 | Chat and image generation, no video |
X Premium | $8/mo | Higher Grok limits inside X, plus X perks |
SuperGrok Lite | $10/mo | Image and video creation, longer conversations |
SuperGrok | $30/mo | DeepSearch, Voice, Grok Bot, 720p 30-second video |
X Premium+ | $40/mo | SuperGrok access plus ad-free X |
SuperGrok Plus | $100/mo | 1080p video, priority access, early features |
SuperGrok Heavy | $300/mo | Largest agent team, highest usage, X Premium+ included |
SuperGrok Plus at $100/mo is the tier most comparisons miss. It appeared in mid-2026 without an announcement post and is the first individual plan that unlocks 1080p video. SpaceXAI publishes the $30 and $100 prices on its pricing page but not the Lite and Heavy prices, which appear only as unpriced columns in a comparison table, so treat $10 and $300 as well-corroborated rather than official.
ChatGPT Go vs SuperGrok Lite
The budget matchup is $8 against $10. ChatGPT Go buys headroom: more uploads, images, memory and data analysis than Free, but no flagship model and no Deep Research. SuperGrok Lite buys media, restoring image and video creation alongside longer conversations and one-prompt app building.
Go is the better deal for people who mainly type. Lite is the better deal for people who mainly generate. Note that "Grok Lite" is not a real product name, only shorthand for SuperGrok Lite.
ChatGPT Plus vs SuperGrok
At $20 against $30, this is where most buyers land. Plus includes GPT-5.6 Sol in chat, ChatGPT Work, and the full Codex coding agent, and it is the cheapest plan that reaches GPT-6 Astra at all. SuperGrok includes DeepSearch, voice mode, Grok Imagine, Grok Bot, and native access to X's live firehose.
Plus is 33% cheaper and carries the deeper toolset. SuperGrok's advantage is real-time social data, which no ChatGPT plan can match at any price.
ChatGPT Pro vs SuperGrok Heavy
At the top, SuperGrok Heavy runs $300/mo against ChatGPT Pro at $100. Heavy unlocks the largest agent team and bundles X Premium+. ChatGPT's $100 Pro tier is the cheapest plan that puts GPT-6 Astra in the normal chat interface, which is the single clearest reason to go past $20.

One pricing trap applies to both sides. Grok 4.7 costs $2 in and $6 out below 200,000 prompt tokens, then doubles to $4 and $12 for every token in the request once you cross that line. GPT-6 Astra does the same thing at 272,000 input tokens, doubling input and adding 50% to output. Long-context work costs materially more than either headline rate suggests.
Benchmarks: GPT-6 Astra vs Grok 4.7
Most benchmark comparisons circulating right now are unusable, for two reasons worth explaining.
First, neither lab published a SWE-bench Verified score for its current flagship. OpenAI stopped reporting the benchmark after finding flawed test cases in a majority of its hardest unsolved problems, and SpaceXAI's Grok 4.7 launch table has no SWE-bench row. Any article giving you a head-to-head SWE-bench number for these models is quoting older models or inventing it.
Second, SpaceXAI's own launch table does not compare Grok 4.7 to GPT-6 Astra. It compares against GPT-5.6 Sol, a model OpenAI has since replaced, and it runs each model at a different reasoning effort: Grok 4.7 at xHigh, Grok 4.6 at High, GPT-5.6 Sol at Max. Those rows are not like-for-like and SpaceXAI does not claim they are.
The reliable comparison comes from Artificial Analysis, which runs both models through the same harness on the same index version.
Benchmark (AA v4.3.2) | GPT-6 Astra | Grok 4.7 |
|---|---|---|
Intelligence Index | 53 | 46 |
Terminal-Bench 4.0 | 59% | 26% |
Humanity's Last Exam | 55% | 43% |
GDPval-AA v2.1 (Elo) | 1,542 | 1,695 |
AA-Briefcase v1.1 (Elo) | 1,569 | 1,657 |
AutomationBench-AA | 68% | 66% |
SciCode | 56% | 57% |
AA-Omniscience | 43 | 32 |
CritPt (physics) | 32% | 18% |
AA-LCR v1.1 | 81% | 77% |
Across the ten evaluations in the index, Astra wins seven and Grok 4.7 wins three. Astra leads the overall index by seven points. Grok's wins are meaningful, though: GDPval-AA and AA-Briefcase both measure real-world agentic knowledge work, and Grok 4.7 takes them by wide margins.
Grok 4.7 improved on Grok 4.6, just modestly. The index rose two points, from 44 to 46. AA-Briefcase gained 111 Elo and GDPval-AA gained 90. Terminal-Bench rose 4.5 points and AA-LCR fell 3.7. Its hallucination rate on AA-Omniscience dropped from 34% to 29%.
Two version traps are worth flagging. Scores quoted at "61" for Astra come from an older index generation that Artificial Analysis has since rebuilt twice; the current v4.3.2 figure is 53. And Terminal-Bench scores only mean something within a single version. SpaceXAI reports Grok 4.7 at 37.6% on Terminal-Bench 4.0 at xHigh effort while Artificial Analysis measures 26% on the same benchmark version, and older Grok scores in the 80s come from version 2.1, effectively a different test.
Cheaper per token, more expensive per task
This is the finding that changes the buying decision, and it reverses what was true of Grok 4.6.
Grok 4.7 is 5.7x cheaper than Astra on blended token price, $1.35 against $7.70 per million. But it emits roughly 81,000 output tokens per benchmark task against Astra's 27,000, three times as many, and 125% more than Grok 4.6 needed. Multiply it out and Artificial Analysis measures Grok 4.7 at $3.74 per task against Astra's $3.26. The cheaper model costs about 15% more to finish the same work.
Speed inverts the same way. Grok 4.7 returns its first token in 0.92 seconds against Astra's 352 seconds, which looks decisive. But it generates at 40 tokens per second against Astra's 52, and because it writes so much more, it takes about 17.7 minutes per task against Astra's 8.8. Grok starts sooner and finishes later.
On WebDev Arena, where humans vote on head-to-head outputs, GPT-6 Astra ranks second with 1,792 and Grok 4.7 ranks twelfth with 1,632, a hair above Grok 4.6's 1,624. GPT-6 Sol slots in fifth at 1,686. Claude Opus 5.5 leads at 1,818.
Coding and Development
On coding evaluations run under the same harness, Astra leads clearly. The Terminal-Bench 4.0 gap is the widest in the index, 59% to 26%.
Grok 4.7 does better when you give it its own agent. Artificial Analysis's Coding Agent Index, which scores a model inside its native harness rather than bare, puts Grok 4.7 plus Grok Build at 56, up from 47 for Grok 4.6, ranking fourth behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. Inside that harness its Terminal-Bench score rises from 18% to 33% and DeepSWE from 65% to 73%. The harness is worth a lot to Grok specifically.
The cost argument for Grok is weaker than it looks, for the reason above: identical input pricing to GPT-6 Sol at $2 per million, cheaper output at $6 against $10, but two to three times the token output per task. If you are optimizing an agent loop for cost, measure cost per completed task rather than per token.
Both ship serious coding agents. OpenAI's Codex is built into the ChatGPT desktop app and available on every plan, and it now runs GPT-6 Sol and Luna. SpaceXAI's Grok Build is a command-line agent, open sourced under Apache 2.0, that runs parallel subagents in separate git worktrees; it defaults to Grok 4.7 and gained persistent memory in September. Grok also ships a Grok 4.7 Fast variant at double the token price, available only inside Cursor and Grok Build, never on the public API.
Neither one deploys your code. They generate it, and version control, hosting, testing, and release remain your problem. If you want AI that goes past generating code to building and shipping a complete app, that is a different category of tool.
Real-Time Data and Web Search
This remains Grok's clearest advantage, and it is structural rather than technical.
Grok has native access to X's live post stream. Ask what is trending right now and it reads the platform directly instead of waiting for search engines to index the conversation. For breaking news, social sentiment, and anything that happened in the last hour, nothing in ChatGPT's lineup competes.
ChatGPT browses the web well and cites its sources. But it pulls from indexed pages, which introduces lag that Grok's direct feed does not have.
Knowledge cutoffs favour ChatGPT slightly and are closer than the model numbers suggest. GPT-6 Astra is trained to April 30, 2026, GPT-6 Sol to April 20, and GPT-6 Luna to May 18. SpaceXAI's two published sources disagree with each other on Grok 4.7, with the developer docs saying May 2026 and the model card saying June 2026.
For multi-source research where accuracy beats recency, the two are close. For live social data, Grok wins outright.

Writing and Creative Tasks
The writing difference is about personality, and it is more consequential than benchmark gaps for most people.
ChatGPT writes in a measured, structured register. For business emails, reports, and documentation it produces output you can send with light editing.
Grok is blunter and funnier. SpaceXAI tuned it to have opinions, and for social posts, brainstorming, and first drafts that need energy, it usually gives you something more interesting to react to.
The tradeoff runs both ways. Grok's looser moderation means it will occasionally produce something you cannot use at work. ChatGPT's tighter guardrails feel restrictive but keep output consistently safe. Grok 4.7 did tighten noticeably: SpaceXAI shipped what it calls an entirely new safeguard stack and reports letting only 3.3% of risky dual-use prompts through on its internal HackerBench test.
AI Agents: ChatGPT Work vs Grok Bot
Both companies shipped autonomous agent products this year, and this is where they diverge most. Neither is a subscription tier, despite how often they get described as one.
ChatGPT Work launched July 9, 2026. You give it an outcome, it gathers context from your connected apps and files, breaks the job into steps, and works for hours before returning finished spreadsheets, slides, documents, or web apps. It is bundled into Plus and above with usage metered rather than flat-rate, and it is now the main place a Plus subscriber can reach GPT-6 models.
Grok Bot launched August 11, 2026. Instead of one task-scoped agent, you get persistent AI teammates, each with a cloud virtual machine, browser, filesystem, and terminal. They sign into your real SaaS tools through the interface, which means they can drive legacy software that has no API, and they message each other and share context.
Grok Bot's access rules changed fast, and most write-ups are stale. It is no longer restricted to SuperGrok Heavy. Since August 26, 2026 it has been included with all SuperGrok plans and every paid Cursor plan including Cursor Teams, and an enterprise version followed on September 3. It is also no longer macOS-only: SpaceXAI now ships desktop builds for macOS, Windows, and Linux, plus iPhone, iPad, and Android apps. There is still no standalone web app.
One security detail deserves more attention than it gets. Per SpaceXAI's own documentation, every Bot on an account shares a single persistent cloud computer, including files, browser sessions, and logins. Separate Bots are not a security boundary.
Grok's Skills feature, launched May 2026, generates Word documents, presentations, spreadsheets, and PDFs, and it officially supports seven connectors: SharePoint, Outlook, OneDrive, Google Workspace, Notion, GitHub, and Linear, plus any custom MCP server. OpenAI's connector library is considerably larger, with more than 60 first-party integrations alongside a public directory of roughly 2,300 apps.
Image and Video Generation
OpenAI shipped ChatGPT Images 2.5 on September 8, 2026, cutting generation latency by up to 50% versus Images 2.0 and adding a Sketch feature for drawing directly in the interface. It went to every tier at once, including Free.
Grok Imagine has not shipped anything new since early August. Imagine Image 2.0 handles targeted editing of existing images and accepts up to five source images per edit, at up to 2K resolution. Imagine Video 1.5 does text-to-video, image-to-video, and reference-to-video, generating clips of up to 15 seconds.
The resolution rules are more restrictive than most summaries admit. 1080p output is limited to text-to-video and image-to-video, while reference-to-video caps at 720p, and on the consumer side 1080p is a SuperGrok Plus unlock at $100/mo. The $30 SuperGrok tier tops out at 720p and 30 seconds.
Grok has the stronger video story. ChatGPT has the stronger access story, since ChatGPT's image generation reaches every tier while Grok gates video behind a paid plan and its best resolution behind a $100 one.
Which Should You Choose?
The answer depends on what you do most days.
Use case | Winner | Why |
|---|---|---|
General productivity | ChatGPT | Higher index score, deeper ecosystem |
Production coding | ChatGPT | 59% vs 26% on Terminal-Bench 4.0 |
Cost per finished task | ChatGPT | $3.26 vs $3.74 despite 5.7x pricier tokens |
Cheap high-volume API calls | Grok | $2/$6 per million vs $10/$50 |
Agentic knowledge work | Grok | Wins GDPval-AA 1,695 to 1,542 |
Real-time social data | Grok | Native X firehose access |
Professional writing | ChatGPT | More consistent, work-safe tone |
Creative writing | Grok | Sharper personality, better first drafts |
Long documents | ChatGPT | 1,050,000 tokens vs 500,000 |
Video generation | Grok | Up to 15 seconds, 1080p on higher tiers |
Budget under $10/mo | ChatGPT | Go at $8 vs SuperGrok Lite at $10 |
If you want the most capable model and the widest set of tools around it, ChatGPT Plus at $20/mo remains the default that is hard to argue against. Just know that $20 gives you GPT-5.6 Sol in the chat window and GPT-6 models only inside ChatGPT Work and Codex. Chat access to Astra starts at $100.
If you care about cost per token, first-token latency, or live social data, SuperGrok at $30/mo earns its premium, as long as you accept that the chat app runs Grok 4.6. And if you are weighing other tools in this space, our roundup of the best free app builders covers the adjacent landscape.
Beyond Chatbots: Building Apps With AI
ChatGPT and Grok both write code, debug functions, and answer technical questions. Then they stop. You still set up hosting, wire a backend, handle app store submissions, and maintain the thing after launch.
AI app builders remove that middle step. Rather than copying code out of a chat window into your editor, you describe what you want and the tool builds it.
CatDoes works that way. You explain the app in plain language, and the agent handles frontend, backend, database, auth, and deployment, then ships to the App Store, Google Play, or the web with a custom domain. No local dev environment, no manual deploys.
For coding questions and general AI work, ChatGPT and Grok are the right tools. For going from an idea to a live app without writing code, that is what CatDoes is built for.
Frequently Asked Questions
Is Grok better than ChatGPT?
On capability, no. GPT-6 Astra scores 53 on the Artificial Analysis Intelligence Index v4.3.2 against Grok 4.7's 46, and wins seven of the ten evaluations in that index. Grok 4.7 wins three, including both agentic knowledge work benchmarks, and it is 5.7x cheaper per token. It also has real-time access to X, which no ChatGPT plan matches.
Can I use Grok 4.7 in the Grok app?
Not yet. SpaceXAI's model card says Grok 4.7 will reach the web, mobile apps, and Grok-in-X "at a later date," and both x.ai and grok.com still advertise Grok 4.6 on their plan pages. Grok 4.7 currently runs in the SpaceXAI API, Grok Build, Cursor, GitHub Copilot, and the Grok Office add-ins.
Are GPT-6 Sol and GPT-6 Luna in ChatGPT?
Only partly. OpenAI's documentation states they are available in ChatGPT Work and Codex but "aren't available in Chat." Free and Go users can try GPT-6 Luna in the desktop app. In the chat window itself, Free and Go run GPT-5.6 Luna and Plus runs GPT-5.6 Sol.
Is Grok cheaper than ChatGPT?
Per token, yes: Grok 4.7 costs $2 in and $6 out per million against GPT-6 Astra's $10 and $50, a 5.7x blended difference. Per finished task, no. Grok 4.7 emits about 81,000 output tokens per benchmark task against Astra's 27,000, so Artificial Analysis measures it at $3.74 per task against Astra's $3.26. Note that Grok's rate doubles above 200,000 prompt tokens.
Which AI has the bigger context window?
ChatGPT. GPT-6 Astra, Sol, and Luna all handle 1,050,000 tokens, with 922,000 of input and up to 128,000 of output. Grok 4.7 handles 500,000. This reverses the position from earlier in 2026, when Grok 4.3 led at 1M tokens against ChatGPT's 128K.
Is Grok free?
There is a free tier with chat and image generation, but no video generation. Paid plans start at $10/mo for SuperGrok Lite, with SuperGrok at $30/mo, SuperGrok Plus at $100/mo, and SuperGrok Heavy at $300/mo. ChatGPT's free tier includes image generation but shows ads, and its ad-free toggle is only available on the Free plan.
Does ChatGPT Plus include GPT-6 Astra?
Partly. Plus subscribers get Astra inside ChatGPT Work and Codex, but not in the standard chat model picker, where OpenAI lists it for Pro, Business, and Enterprise. Sam Altman called the rollout "messy" on September 4, 2026, and it has not changed since.
Which AI is best for coding?
GPT-6 Astra, on the evaluations run under a common harness: 59% against 26% on Terminal-Bench 4.0. Grok 4.7 closes much of the gap inside its own Grok Build harness, where it scores 56 on Artificial Analysis's Coding Agent Index, fourth overall. Ignore any SWE-bench Verified comparison between these models, because neither lab published one.
What is the difference between ChatGPT Work and Grok Bot?
Both are autonomous agents, not subscription plans. ChatGPT Work is included with Plus and above and returns finished documents, spreadsheets, and slides. Grok Bot creates persistent AI teammates that each run on their own cloud machine and can operate software with no API. Since August 2026 it is included with all SuperGrok plans and every paid Cursor plan, not just SuperGrok Heavy, and it runs on macOS, Windows, Linux, iOS, and Android.

Nafis Amiri
Co-Founder of CatDoes


