Why I Switched (And Why I Almost Switched Back)
Look, I was that person. The one with ChatGPT pinned to my browser tabs, the one who’d reflexively opened it for everything from debugging code to untangling family drama in email drafts. Then in February, Anthropic released Claude 3.7 Sonnet, and the AI discourse got loud enough that I couldn’t ignore it anymore. The big claim? This new version shows you its actual reasoning before giving you an answer, which sounded either revolutionary or like marketing nonsense, depending on which thread I was reading.
I decided to do what I always do when I’m skeptical: go all in for a month and actually feel the difference instead of just reading hot takes. No hedging, no “I’ll use both,” just switch completely and sit with the discomfort. Thirty days. Full commitment. I’ll be honest though—the first week had me reaching for ChatGPT like muscle memory reaching for coffee.
Here’s what actually happened when I stopped running away from the learning curve.
The Extended Thinking Feature Changed How I Work (In Ways I Didn’t Expect)
The headline feature of Claude 3.7 Sonnet is extended thinking, where you can see the model’s reasoning chain before it gives you the final answer. When Anthropic’s official Claude 3.7 Sonnet announcement dropped, I expected this to feel like watching someone think out loud. It does feel that way, but the practical impact was weirder than I anticipated.
The first time I used it for a coding problem, I got genuinely disoriented. Claude showed me forty-something lines of its reasoning process before the actual code. My instinct was to skip it, get to the answer. But I forced myself to read through it, and something clicked. I could see where it was considering different approaches, where it was second-guessing itself, where it was choosing between solutions. When the final code didn’t work (it happens), I could actually trace back and see which assumption led to the wrong path. That’s not nothing. That’s the difference between getting an answer and understanding why the answer is shaped that way.
After about two weeks, I stopped seeing the reasoning as extra friction. I started seeing it as the actual product. You know that feeling when someone explains something complicated and you go from confused to “oh, I see why they chose that”? This is what that looks like in AI form. The code itself wasn’t always better than ChatGPT’s code. But the path to understanding was.
The Hard Numbers (Where Claude Actually Pulls Ahead)
I could tell you my subjective feelings, but let’s talk concrete stuff. According to Anthropic’s internal benchmarks, Claude 3.7 Sonnet scored 70.3% on SWE-bench Verified, a pretty rigorous test for coding tasks. GPT-4o hit 38.8% on the same test. That’s not a small gap. That’s the kind of gap where you notice the difference when you’re actually using it.
But here’s where I’m going to be the friend who levels with you: the Stanford HAI 2025 AI Index Report noted something that doesn’t make headlines. The gap between leading AI models on reasoning benchmarks compressed to under 5% across the top competitors by mid-2025. Translation? The days of one model being obviously better across the board are mostly over. We’re in the territory of specific strengths and tradeoffs now, not clear winners.
I saw this play out in real work. Claude 3.7 was genuinely better at hard coding problems and working through complex logical tasks. But when I switched to ChatGPT for a creative writing project, it felt more natural. When I needed information from a specific recent event, ChatGPT had fresher training data. The numbers tell part of the story. They don’t tell all of it.
Where This Model Actually Stumbled (Because Honest Means Honest)
I’d be doing you a disservice if I pretended this was a clean win. The 200,000 token context window—basically a 500-page novel in a single prompt—sounds incredible until you hit the edge cases. I tried uploading a massive research document to get Claude to synthesize it, and the extended thinking feature sometimes got so deep in reasoning that it spiraled on ambiguous questions. Not stuck in loops exactly, but generating reasoning that felt like it was going in circles.
There’s also the speed factor. ChatGPT gives you a response in seconds. Claude, especially with extended thinking enabled, takes noticeably longer. For quick questions, this feels irritating. For complex problems, the extra time makes sense. But if you live in an attention economy where a five-second delay feels like a failed experience, this matters.
And I’m going to say it: the interface isn’t quite as polished as ChatGPT’s. It’s good, it’s clean, but ChatGPT feels like a tool refined through millions of hours of user feedback. Claude feels like a tool refined by talented engineers who know what matters and haven’t quite had time to sweat the small details yet. That’s not a death sentence. It’s just reality.
What This Means for My Actual Life (The Verdict That Matters)
Here’s the truth: I didn’t switch back to ChatGPT as my primary tool. But I also didn’t delete ChatGPT. I’m using Claude as my main squeeze for coding problems and serious reasoning tasks, with ChatGPT around for speed and for those moments when I need something slightly more intuitive.
Anthropic raised 2.75 billion dollars in early 2025, valuing the company at 61.5 billion dollars. That money isn’t just sitting there. It’s going to keep pushing Claude forward. The funding matters because it signals that despite ChatGPT’s market dominance, something real is happening here. Claude isn’t a scrappy underdog anymore. It’s a well-funded competitor that’s genuinely strong in ways that matter for specific use cases.
The unpopular opinion? Claude 3.7 Sonnet isn’t going to replace ChatGPT for most people, and it shouldn’t. But for people who write code, analyze data, or work through complex logical problems, it’s worth the friction of switching. The extended thinking feature isn’t a gimmick. It’s a real shift in how you interact with AI assistance.
If you’ve been curious but hesitant, my honest take is this: spend a week with it. Get past the muscle memory of ChatGPT. Let the reasoning chains feel weird until they feel useful. Then decide if it fits your actual work. That’s how you’ll know if this matters for you, not from reading reviews or benchmarks or hot takes from people trying to prove a point.
What’s been your experience if you’ve tried it? I’m genuinely curious whether this landed differently for you than it did for me.