Showing posts with label Claude. Show all posts
Showing posts with label Claude. Show all posts

Sunday, July 12, 2026

Fable 5 Just Beat GPT-5.6 in a 3D Build Showdown

 On July 11, a four-model benchmark quietly reshuffled the AI pecking order. The prompt was simple — generate a floating island city in the browser — and the winner was Fable 5, a model that many in the tech space are only now starting to track.

The Benchmark at a Glance

The test, reported by 0xMarioNawfal, pitted four models against the same prompt: build a floating island city, rendered in 3D, running in a browser. No custom scaffolding, no per-model prompt engineering. Same input, same evaluation criteria.

The participants read like a who’s-who of the current AI landscape:

  • Fable 5 — the emerging contender
  • - GPT-5.6 — OpenAI’s latest frontier model
  • - Grok 4.5 — xAI’s most advanced offering
  • - GLM 5.2 — Zhipu AI’s flagship

When the results came in, Fable 5 took the top spot, outperforming all three incumbents on the identical prompt.

This is a single benchmark, not a comprehensive evaluation. Head-to-head comparisons on identical prompts are among the most transparent ways to compare models, but they measure one capability at one moment. The result is a signal, not a verdict.

Why Floating Island Cities

A floating island city is a deliberately complex test for generative AI. It combines terrain generation, architectural structure, atmospheric lighting, and spatial coherence — all in a single viewport. For a model to succeed, it needs to handle physical plausibility, aesthetic composition, and functional layout in a single coherent scene.

In many ways, this is a more grounded benchmark than standard text-based or image-based evaluations. It tests whether a model can synthesize multiple modalities — geometry, lighting, materials, layout — into a single coherent 3D scene. That is a skill that matters directly for game development, architectural visualization, virtual worlds, and the broader spatial computing shift.

The browser-based delivery adds another constraint. The output must render efficiently in real time, which rules out offline renders or post-processing tricks. What you see in the browser is what the model generated, no polish layer.

Fable 5’s Quiet Ascent

Fable 5 hasn’t had the marketing budget of its competitors. It doesn’t carry the brand recognition of OpenAI or xAI. Yet in this head-to-head comparison, it outperformed models that have collectively raised billions and commanded global headlines for months.

The result raises a question that is becoming harder to ignore: are the frontier labs still pulling away from the pack, or is the gap closing?

From where I sit tracking AI benchmarks over the past year, the second explanation is gaining evidence. Model quality is commoditizing faster than most observers realize. A well-trained model with a smart architecture can now compete with — and in this case, beat — models backed by much larger budgets and teams.

The moat that frontier labs relied on is thinning. That doesn’t mean Fable 5 will win every benchmark. It means the field is more competitive than the headlines suggest, and dismissing an emerging model because it lacks brand recognition is a mistake.

What This Means for AI-Generated 3D

The ability to generate 3D content from a text prompt has been one of the most anticipated capabilities in generative AI. Game studios, architecture firms, and virtual-world builders have all been watching for the moment when AI can meaningfully assist with, or replace, manual 3D modeling for prototyping and early-stage design.

This benchmark suggests that moment may be closer than many expect. If a relatively lesser-known model can generate coherent 3D scenes in the browser on the first try, the technology is past the proof-of-concept phase. What remains is reliability, iteration speed, and integration into existing production pipelines.

The browser-based delivery also matters for accessibility. It means AI-assisted 3D creation is available to anyone with a web browser — no game engine installation, no GPU farm, no specialized software. That dramatically lowers the barrier for prototyping and experimentation.

For developers and creators, the implication is straightforward: the cost of generating 3D content is falling. The question is no longer “can AI do this?” but “which model does it best for your specific use case?”

FAQ

What is Fable 5?

Fable 5 is a generative AI model that specializes in producing 3D scenes from text descriptions. It emerged as the top performer in a July 2026 benchmark comparing four models on the same floating-island-city prompt.

How reliable is a single benchmark?

No single benchmark is definitive. This test evaluates one specific capability — generating floating island cities in the browser — and may not reflect performance on other tasks like text generation, image synthesis, or code completion. Head-to-head comparisons on identical prompts are among the most transparent ways to evaluate relative model strength for a given task.

Can I try Fable 5 myself?

That depends on current availability. Many emerging AI models offer browser-based demos or API access. The best way to verify the results is to run the same prompt yourself and compare the output side by side with other models you have access to.

Why does 3D generation in the browser matter?

Browser-based 3D generation means the output is real-time and accessible without specialized hardware or software. Offline rendering pipelines typically require GPU clusters, proprietary engines, and significant setup time. Running in the browser makes 3D generation accessible to anyone, which accelerates iteration and experimentation.

What industries would benefit most from this capability?

Game development, architectural visualization, film pre-visualization, virtual reality, and e-commerce product visualization are the most obvious candidates. Any industry that currently relies on manual 3D modeling for prototyping could see workflows compressed from days to minutes.

Run Your Own Benchmark

One benchmark doesn’t crown a champion. What it does is give you a data point worth testing yourself. If you’re building 3D experiences, prototyping game environments, or just exploring what AI can generate, pick one prompt this week — floating island cities or something you actually need — and run it across the models you have access to. Decide for yourself which one earns your attention.

The gap between frontier labs and emerging contenders is narrowing faster than most planning cycles account for. Fable 5’s win is one signal in a pattern that the smartest teams in gaming, architecture, and spatial computing are already acting on. The model that wins your next prototype might not be the name you already know.

Floating island city landscape from Unsplash, showing terrain with water and architectural structures suspended in the sky, used as cover image for an article about an AI 3D generation benchmark.
Photo by Yuya Murakami on Unsplash


Tuesday, June 30, 2026

China Just Dropped 20M Free AI Tokens (And Nobody Noticed)

 

Photo by Mohammad Rahmani on Unsplash

If you’re a developer who pays $20/month for Cursor Pro, $10 for GitHub Copilot, or burns through API credits like candy — stop. A 756-billion-parameter coding model just landed with a 20-million-token free tier. No credit card. No subscription. And somehow, almost nobody in the Western developer community is talking about it.

What Zhipu AI’s GLM-5.2 Actually Gives You

GLM-5.2 is a massive 756B-parameter model from Zhipu AI, a Beijing-based AI lab often described as China’s closest equivalent to OpenAI. The headline offer is 20 million free API tokens for new developers — not a trial, not a “first month free” gimmick. Create an account and you get the full quota immediately, no billing info required.

Beyond the token grant, you also get 120 free image and video credits, access to GLM-5.2’s “High” and “Max Thinking” reasoning modes, and a 1-million-token context window. That’s large enough to feed an entire codebase into a single prompt and still have room for instructions.

The API is OpenAI-compatible. You can point Cursor, Claude Code, Cline, or any OpenAI SDK at it by swapping the base URL and model name. No custom integration, no new tools to learn. If your editor already speaks OpenAI, it already speaks GLM-5.2.

How It Stacks Up Against What You’re Already Paying For

Do the math on your current AI coding stack. Cursor Pro costs $20/month. GitHub Copilot is $10/month. Claude API charges per token. Add them up and you are looking at $30+ per month for tools that help you write code faster.

GLM-5.2 replaces all of them at zero cost for the first 20 million tokens. For a solo developer or small team experimenting with AI-assisted coding, the savings add up fast. Twenty million tokens goes a long way — hundreds of code completions, dozens of full-file refactors, and plenty of room for trial and error.

GLM-5.2 also supports a “Max Thinking” mode that applies chain-of-thought reasoning to complex coding tasks. In practice, this means better results on multi-step refactors, debugging sessions, and architectural decisions — exactly the places where smaller models fall apart.

Why the Silence?

If the offer is real and the model is competitive, why isn’t everyone talking about it? Three factors explain the gap.

Geographic attention bias. Chinese AI labs rarely receive the same Western media coverage as OpenAI, Anthropic, or Google DeepMind. A breakthrough from Beijing doesn’t trend on Hacker News the same way one from San Francisco does. This isn’t new — it’s been true since the earliest days of China’s AI industry.

Trust and data privacy. GLM-5.2 routes through Chinese infrastructure. For many Western developers and enterprises, that’s a dealbreaker. Data residency requirements, compliance policies, and geopolitical caution create a barrier that no amount of free tokens can overcome.

The U.S.-China AI perception gap. Some developers avoid Chinese models on principle; others assume they can’t be competitive. The assumption is increasingly outdated — several Chinese models now rank in the top tier of coding benchmarks — but the perception lingers. GLM-5.2’s 756B parameter count and benchmark scores are competitive with frontier Western models, but mindshare in the developer community hasn’t caught up.

How to Try It in Two Minutes

Here’s the fastest path to get coding with GLM-5.2:

  1. Register at open.bigmodel.cn
  2. 2. Verify your account with your phone number (OTP arrives in a few minutes)
  3. 3. Create an API key from the dashboard
  4. 4. Set your base URL to the GLM-5.2 endpoint
  5. 5. Select model: glm-5.2

For Cursor users: open Settings, go to Models, add a new model provider, and paste your GLM-5.2 API key and base URL. For OpenAI SDK users: set OPENAI_BASE_URL and OPENAI_API_KEY as environment variables. For Cline and similar tools: the model provider setup screen accepts any OpenAI-compatible endpoint — add GLM-5.2 as a custom provider and you’re done.

The Real Catch — Three Caveats Worth Knowing

A free offer at this scale comes with tradeoffs worth understanding up front.

Data residency. Zhipu AI’s API servers are in China. If your codebase contains sensitive or proprietary code that you can’t route through Chinese infrastructure, this isn’t for you. No free tier is worth a compliance violation.

Phone verification. Registration requires a phone number and OTP. Some non-Chinese users report delays receiving the verification code. If you’re outside China, budget a few extra minutes for this step.

Long-term uncertainty. Zhipu AI hasn’t published clear post-quota pricing. The 20M free tokens are framed as a developer acquisition play rather than a limited promotion, but any free API offering can change. If you build a workflow around it, keep a paid fallback ready.

FAQ

Can I use GLM-5.2 with Cursor or VS Code?

Yes. The API is fully OpenAI-compatible. Add it as a custom model provider in Cursor, Claude Code, Cline, Continue.dev, or any tool that supports OpenAI’s API format. Just swap the base URL and model name — no custom integration needed.

How does GLM-5.2 compare to GPT-4o or Claude 4 Sonnet?

GLM-5.2’s 756B parameters put it in the same weight class as the largest frontier models. On coding benchmarks, it scores competitively. The practical differentiators are the 1M-token context window and the zero-cost entry point. Most developers report it handles complex refactoring and debugging well in Max Thinking mode.

What data does Zhipu AI collect from API calls?

Zhipu AI’s data handling policies are less transparent than Western providers. Review the terms of service carefully before sending proprietary code. For open-source or personal projects this is less of a concern, but enterprise teams should involve legal before routing sensitive code through the API.

Is phone verification required for all users?

Yes, registration requires a phone number. Some users outside China report OTP delivery delays. If you don’t receive the code within a few minutes, try again after an hour — the system sometimes throttles international SMS.

Will the 20M free tokens refresh or is it a one-time grant?

The 20M tokens are a one-time welcome grant for new developers, not a recurring monthly quota. Zhipu AI hasn’t announced post-consumption pricing. Pace your usage accordingly: use it for evaluation and experimentation first, migration second.

Try It Before the Quota Changes

Twenty million tokens is enough to decide whether GLM-5.2 fits into your workflow. Sign up, point your editor at it, and spend an afternoon testing it on your actual codebase. If it works, you’ve just eliminated a monthly subscription. If it doesn’t, you’re out two minutes and zero dollars.

The offer is real. The model is competitive. The silence from the Western developer community won’t last forever — and when the conversation starts, you’ll already have an opinion.

Your AI Coding Subscription Is Draining Your Wallet

 

Photo by Luca Bravo on Unsplash

I opened my credit card statement last month and found four charges I didn’t remember approving — $347 total, all for AI coding subscriptions. I cancelled three of them that same afternoon.

For a while, “unlimited” AI coding plans felt like the obvious choice. Pay a flat fee, use the assistant as much as you want, never think about tokens or credits. But that model never made economic sense for the companies running it — advanced models are expensive to serve — and the pendulum has swung hard toward measured usage.

I actually prefer the new direction. Token-based and quota-based plans let you budget your consumption, work in bursts without penalty, and never wonder if the “unlimited” label is about to get throttled. The hard part is figuring out which plan actually delivers.

Here’s what I found after a month of testing five AI coding subscription plans against real development workflows.

The End of Unlimited (And Why You Should Be Happy)

The tipping point was inevitable. Running frontier coding models costs real compute, and the old “all you can eat” pricing was burning VC cash, not building sustainable products.

What emerged instead is a matrix of options: token-based plans where you buy a pool of tokens each month, credit-based plans that meter specific capabilities, and quota-based plans that refresh weekly or daily. Each model suits a different working style, but they all share one upside — you know what you’re paying for.

The plans I tested: MiniMax Token Plan, MiMo Token Plan, GLM Coding Plan, OpenAI Codex (included with ChatGPT), and Kimi Code. Each got at least a week of real coding time.

MiniMax Token Plan — $20 for More Tokens Than You’ll Use

MiniMax’s Token Plan is the easiest recommendation on this list. For $20 a month, you get access to MiniMax’s coding models through their web app and desktop app, plus integrations with Claude Code, Cursor, Cline, Kilo Code, Roo Code, Codex CLI, and OpenCode.

The token allowance is generous. For daily coding — debugging, refactoring, running agentic workflows — I never came close to exhausting it. If you want to start even smaller, prepaid credits begin at $5.

This is the plan I’d recommend to any developer who wants high usage at a low price, no games, no hidden throttles.

MiMo Token Plan — The Speed King Nobody’s Talking About

MiMo surprised me more than any other plan on this list. The responses are fast, it uses fewer reasoning tokens than comparable services, and the UI generation quality is genuinely good.

The plan runs on credits that refresh monthly. You use them across MiMo’s model lineup, including MiMo-V2.5-Pro, which supports up to a 1 million-token context window and is built for agentic coding and long-horizon software tasks. It integrates with tools like OpenCode, Cline, OpenClaw, Kilo Code, and Blackbox.

If you’re building custom AI workflows or testing multiple models in parallel, MiMo’s combination of speed and token efficiency makes it a strong second option. It’s not a full IDE subscription — it’s a model access plan — but for agentic coding, it punches above its price.

GLM Coding Plan — Worth It Only If You Need GLM Models

GLM’s Coding Plan from Z.ai has gone through changes recently, and the price has increased. The company is investing in better models like GLM-5.2 and deeper integrations with coding tools, and the subscription reflects that cost.

Here’s the honest take: if you specifically want GLM models for your coding workflow — they work with Claude Code, Cline, Kilo Code, OpenCode, and OpenClaw — the plan delivers. The models are strong for focused coding agent sessions.

But if you’re just looking for the best generic coding subscription, cheaper options exist. GLM made more sense before the price increase. Today, use it when you need GLM-5.2 specifically.

OpenAI Codex — Free (If You Already Pay for ChatGPT)

OpenAI Codex lives inside the VS Code extension, and it’s the plan I use most days — not because it’s the best, but because it’s included with my ChatGPT subscription.

Codex understands your codebase well, handles code generation, debugging, project edits, and large-codebase navigation. The catch is the daily and weekly limits. In a serious coding session, those limits can disappear within an hour. OpenAI lets you buy extra credits as a cushion, but that adds to the cost.

The math is simple: if you already pay for ChatGPT, use Codex as your daily driver. When you hit the limit, switch to MiniMax or MiMo as a backup. No need for a separate primary subscription.

Kimi Code — Predictable Quota, No Monthly Burnout

Kimi Code uses a weekly refreshed quota instead of a monthly token pool. You get a set amount of usage every week, and it resets — no rollover, no guessing.

The Kimi K2.7 Code model handles codebase understanding, terminal tasks, file edits, debugging, refactoring, and feature building. You can access it through the web app, VS Code extension, and CLI.

The weekly refresh is an interesting tradeoff. If you code consistently every week, it works well. If you have heavy weeks and light weeks, a monthly token pool gives you more flexibility. Kimi Code is a solid choice if you’re already in the Kimi ecosystem or prefer K2.7 over other models.

FAQ

Can I use multiple AI coding subscriptions at once?

Yes, and most developers I know do exactly this. A common setup: OpenAI Codex as the daily driver (included with ChatGPT), with MiniMax or MiMo as a backup for heavy coding sessions when Codex limits run out.

Which AI coding plan is best for someone new to AI coding?

Start with OpenAI Codex if you already subscribe to ChatGPT. If you don’t, the MiniMax Token Plan at $20 a month is the lowest-risk entry point with the broadest tool support.

Do token-based credit plans expire?

It depends on the plan. MiniMax offers prepaid credits starting at $5 that you use when needed. Monthly token subscriptions reset each billing cycle. Kimi Code’s quota refreshes every week and does not roll over. Always check the plan’s expiration policy before buying.

How do I know which plan fits my workflow?

Match the pricing model to your work pattern: monthly token plans for bursty usage (heavy sprints, then lighter weeks), weekly quotas for consistent daily coding, and included subscriptions (Codex with ChatGPT) for the baseline you already pay for.

Here’s the quick-reference comparison:

  • MiniMax Token Plan: $20/month token pool — Best value on the list
  • - MiMo Token Plan: Monthly credits — Fast and token-efficient
  • - GLM Coding Plan: Quota-based subscription — Only if you need GLM
  • - OpenAI Codex: Included with ChatGPT — Free if you’re already paying
  • - Kimi Code: Weekly refreshed quota — Solid but niche

Open your billing page right now. If you’re paying more than $50 a month for any single AI coding subscription on this list, try swapping for a month. Start with Codex if you already have ChatGPT — it’s already on your bill. Add MiniMax or MiMo as a $20 backup. I saved $160 my first month, and my output didn’t drop.

Monday, June 29, 2026

Claude Rewrote 5,000 Lines of My Code — Here’s What I Learned

AI code generation concept with abstract technology patterns representing machine learning and software development
Photo by Numan Ali on Unsplash

If you’re a developer who hasn’t touched Claude yet, I get the skepticism. I was there a few weeks ago. I’d watched the demos, read the tweets, nodded along — and kept writing code the same way I always had. Then I gave it a real test: refactor a legacy module I’d been dreading. Five thousand lines, six months old, written by someone who’d already left. I expected a mess. What I saw changed how I think about AI-assisted development.

Why I Started Skeptical (and You Should Be Too)

Every AI coding tool makes the same promises. “Write code faster.” “Fewer bugs.” “Ship more.” And every one I’d tried before Claude delivered on maybe half of those. GitHub Copilot was great at autocomplete but useless for architecture. ChatGPT could write a function but couldn’t hold context across an entire codebase. I’d learned to use AI as a fancy autocomplete, not a collaborator.

Claude, specifically Claude Code (Anthropic’s terminal-based agent), promised something different: not just writing code, but understanding it. Reading entire projects, reasoning about architecture, and making changes that spanned multiple files. I wanted to believe it. I also wanted proof.

The Test I Threw at It — A 5,000-Line Refactor

The module was an internal dashboard API written in Node.js. Six months of organic growth had turned it into a god object nightmare: one file handled auth, routing, database queries, email notifications, and caching. Every new feature required touching at least four functions in the same file. Tests were sparse. Comments were aspirational.

I pointed Claude Code at the repo and gave it a single instruction: “Refactor this API into a clean layered architecture. Split concerns. Don’t break the tests.”

It started by reading every file in the project. Not the file I pointed at — the entire project. It identified imports, mapped dependencies, and built a mental model of the codebase. Then it wrote a plan: which files to create, what to extract, how to wire the layers together. It asked one clarifying question about the authentication flow before it began.

What Claude Did That No Other AI Could

It maintained context across every file it touched. When it renamed a function in the service layer, it updated every import and every caller across the entire project — not just the file it was editing.

Second, it understood the test suite. It ran the tests after every change and caught regressions I would have missed. When a test broke, it didn’t just report the failure — it read the test, understood what it was testing, and adjusted the implementation until the test passed.

Third, it made judgment calls. At one point it had a choice between two refactoring strategies: extract a base class or use composition. It chose composition, left a comment explaining the tradeoff, and asked me to confirm before proceeding. That felt less like a tool and more like a junior developer who’d read the same books I had.

Where Claude Still Falls Short (I Tried to Break It)

I spent the second week trying to find its limits. I found several.

It struggles with highly unconventional code. If your project uses exotic patterns or undocumented frameworks, Claude can hallucinate APIs that don’t exist. It also has a blind spot for performance — its first pass at a database query used N+1 patterns that would have crushed production. And it lacks domain intuition. It can refactor a payment module’s structure but can’t tell you if the business logic for refunds is wrong.

The tool also defaults to verbose code. It writes defensive, over-documented, enterprise-style code by default unless you explicitly tell it to be concise. The first pass added more comments than I was comfortable maintaining.

How I Get the Best Out of Claude Now (My Playbook)

After two weeks of trial and error, I settled on a workflow that consistently delivers:

Start with a written spec. A vague instruction gives you a vague result. I now write one paragraph describing the outcome, one constraint sentence, and one “don’t do this” sentence. That’s usually enough.

Start with a written spec. A vague instruction gives you a vague result. I now write one paragraph describing the outcome, one constraint sentence, and one “don’t do this” sentence.

Review every change file by file. Claude Code’s diff view is excellent. I read every change before accepting it. This caught the N+1 query and two unnecessary abstractions.

Use it for the boring stuff. The real win wasn’t the architecture decisions — it was Claude handling the boilerplate: writing migration scripts, updating type definitions, syncing documentation, fixing lint errors across 30 files.

Verify the tests yourself. Claude runs tests and reports results, but I still run them locally before committing. Once, a test passed in Claude’s headless environment but failed on my machine due to a timezone issue. Trust but verify.

FAQ

Is Claude better than GitHub Copilot for coding?

They solve different problems. Copilot is excellent at inline autocomplete — finishing your line, generating the next function. Claude is better at multi-file reasoning, refactoring, and architectural changes. I use both: Copilot for the in-the-moment flow, Claude for the structural work.

Can Claude work with legacy codebases?

Yes, and this is where it shines. It reads your entire project context before making changes, so it understands the existing patterns, naming conventions, and dependency graph. I’ve seen similar results on Rails, Python, and Go projects of varying ages.

How much does Claude Code cost?

Claude Code is included with Claude Pro ($20/month) and Claude Max subscriptions. The Pro plan is sufficient for individual developers. The Max plan offers higher usage limits for teams running large refactors daily.

Does Claude write secure code?

It writes code that follows standard security patterns (input validation, parameterized queries, proper error handling), but it won’t catch domain-specific security issues. Always review autogenerated code for business logic vulnerabilities. Claude is a tool, not a security auditor.

Can Claude replace junior developers?

No — and framing it that way misses the point. Claude handles the mechanical parts of coding efficiently, but it can’t attend standups, understand product context, or negotiate tradeoffs with stakeholders. What it does do is remove the grunt work so developers can focus on the parts that require human judgment.

Give Claude One Bad Codebase This Week

Pick the module you’ve been avoiding — the one with the TODO comment that says “refactor this when we have time.” That time is now. Point Claude at it, write a one-paragraph spec, and see what happens. You might be surprised. I was.

The worst case is you review its output and throw it away. The best case is you reclaim a week of your life. Either way, you’ll know whether Claude is right for your workflow. I already know mine.