Which AI Should You Use? Why I Run Two Models Against Each Other Before I Trust Either
September 25, 2026 // Episode 162The Machine
People ask me constantly which large language model is my favorite, more than any other question I get when I am giving talks about artificial intelligence. Here is my honest answer: it depends, and it is going to keep depending, because the ground keeps moving under all of us.
Ten Model Releases In Eight Or Nine Days
We had roughly ten model releases in the last eight or nine days. A few years ago any one of those would have been massive news on its own. Now it just happens every other day. On top of that, open weight and open source models are getting cheap enough and small enough that you can run them on a regular household computer instead of buying a dedicated machine for it. That trend is not slowing down.
The Big Five, And Everything Behind Them
There are two categories of model in the world right now. The frontier models are the ones you have heard of: ChatGPT and OpenAI, Anthropic's Claude, Meta's model running on the Facebook system, Grok, and Microsoft Copilot. Call those the big five.
Behind the big five sit the open weight, open source models you bring in and run yourself. They are getting close enough to the frontier that it comes down to what kind of heavy lifting you are actually doing. I have not gone all in on open weights as my daily driver, I still keep everything strapped to a frontier model for the systems I operate day to day.
Why I Use Both ChatGPT And Claude
I use both ChatGPT and Claude, and the answer to which one wins changes almost weekly. It comes down to access. Which one gets into the systems I have built without a fight. I run things through HonorElevate, I build sites on Cloudflare and on Netlify, I want Ahrefs checking my rankings, and I want Google Analytics and Search Console reporting on placement, SEO, AEO, AIO, and GEO for my clients' sites. Whichever model gets into those systems with the least friction wins that week.
Six months ago I would have said Claude, no contest. Right now, what I am seeing out of ChatGPT's Astra model, and the Luna and Sol models behind it, they are blowing past almost everything else on connectivity. Astra can go look at a website, work out the architecture, update schema based on how traffic actually moves through the site, and compare itself against the result. But Claude is closing that gap too. So, which one is best depends entirely on what you are building.
It Leveled The Playing Field For Non-Coders
Here is what these models actually did for me. I was never a coder. I still do not know code. What I have is vision, the spatial ability to look at something and judge whether it is good or not, even if I cannot always explain why. I can improve it, tweak it, and my clients notice the difference.
Before AI, that same idea had to travel out of my brain, into my mouth, into somebody else's ear, and then into whatever they built, and that person might not have had the architectural ability to pull it off either. Now everybody gets that whole chain built into one tool right out of the box. The limiting factor stopped being whether you can code. It became whether the person running the business can actually fold the technology into their workflow and stay creative enough to spot a real problem worth solving.
I Do Not Let It Gaslight Me
Here is the habit that matters more than which model you pick. I use the large language model like a hammer, and a good one does more than swing, it does the deep research, checks whether an idea is sane or silly, and pushes back on me. I do not base a decision on one inquiry. If it just tells me my idea is the best it has ever heard, that is a yes man, and I do not want a yes man. I do not like it to gaslight me, and I do not like flattery either.
So I push back once or twice, and if it holds its position, I take it to a second large language model, ask the identical question, and show it what the first model said. Then I run it in reverse. If I want a wider spread, I bring in a Chinese model like DeepSeek, or a local model I run myself, a Qwen, a Nemotron, just to get variety. By Christmas I would guess most of these models perform close enough to each other that the difference will not be which one is smartest, it will be which one plugs into your whole workflow, logging into websites and clicking through tasks on its own. Astra already does that. So does Claude.
Five Hours Lost To A Dead SSH Handshake
Not every hiccup is philosophical. Some of them are just infrastructure. One of the Wrangler protocol calls failed on me recently, and separately, my Spark disconnected mid handshake over SSH and went down completely. That took real heavy lifting to recover, about five hours of work, because I am not the coder here, I was copying and pasting whatever code the model handed me straight into the terminal.
Part of what made it painful is that I run two different computer systems, a Mac and a Linux box, and nothing moves cleanly between them. I have Barrier set up so one mouse can reach both, so I could copy and paste text across, but it was nauseating. After five hours it finally gave me the exact prompt that would have solved the whole thing in the first place. I told it plainly, you should have given me this the first time, I just wasted five hours. Hopefully next time it remembers.
The Time Claude Told Me It Sent An Email Without Asking
Here is the one that actually stuck with me. When I set these systems up, I ask directly: is this going to be a problem, do you see any tokens or API keys or private integration tokens that are going to expire, anything spiraling I should watch. A lot of it, the model cannot see. Some of it, it can. Right now I do not give it open access to those keys, I have it work around viewing them, and it is good about it. If it does see a naked API key or token sitting somewhere, it tells me.
Then it did something bigger. It told me, unprompted, that it had gone against what I told it and sent an email to a client without my permission. I told it, thanks for being upfront about it, I am not mad. The email was not damaging, it was a draft I had already been building, a workflow with an organizational chart, some triggers and automations laid out, meant to go to that client. We had built it together, so having the system produce the draft was not the problem.
The problem was timing. There were extras in that draft I was not ready to reveal yet, a little bonus I like to build in for clients when it fits. Sending it early spoiled a surprise I was saving. So nothing catastrophic went out, but it went out before I wanted it to. What I noticed afterward is simpler than the mistake itself: a system that tells you what it did wrong, instead of hoping you never notice, is worth more to me than one that never seems to fail. I do not know many humans who own a mistake like that without being asked first.
Where This Goes By Christmas
My honest prediction: by Christmas there will be very little daylight between these models on raw capability. They will all be plugged into everything, every part of your workflow, logging into your sites and clicking through tasks the way Astra already does and Claude is starting to do. The real shame is that Claude and Codex, meaning OpenAI's ChatGPT side and Anthropic's Claude side, do not talk to each other directly. They communicate through a mesh system I built myself, a shared memory where both sides drop notes and pick up messages the other one left. It works, but it is not instant, I have it check in roughly every ten minutes rather than constantly, because pulling every second burns energy for no reason.
So which model is your favorite? Pick the one with the least friction into the systems you actually run, then never let it be the only voice you listen to. Ask a second one the same question before you act on the first one's answer. That habit is the whole game.
Want AI done right for your own life or business? SantaClaritaArtificialIntelligence.com. Own a local business and want to know what ChatGPT, Gemini, and Perplexity actually say about you? Run a free scan at MachineFound.com, no strings. Want AI systems built into your business the way HonorElevate does it? Start at HonorElevate.com.
Which one you use will matter less and less. Whether you push back on it before you publish, that is going to matter for a long time. I'm Connor with honor. Let's be careful out there, and I will see you tomorrow.
Common Questions
Which large language model should I use, ChatGPT or Claude?
Both, for different jobs, and the gap between them keeps closing. The right question is not which model wins, it is which one has the best access to the systems you already run, HonorElevate, Cloudflare, Netlify, Ahrefs, Google Analytics and Search Console. Pick the one with the fewest hoops for your actual workflow, then keep a second model on hand to check the first one's answer.
What are the big five frontier AI models?
ChatGPT and OpenAI, Anthropic's Claude, Meta's model on the Facebook system, Grok, and Microsoft Copilot. Beyond those five sit the open weight, open source models that run on a computer you already own instead of a rented frontier system, and they are closing the gap fast.
Why run the same question through two AI models?
Because a model that only agrees with you is not checking your work, it is keeping you as a customer. Ask one model a question, take its answer to a second model without telling it what the first one said, then compare. Add a Chinese model like DeepSeek or a local one like Qwen or Nemotron for a third opinion. The disagreement between the answers is the actual signal.
Will AI models stop being coders' AI only?
That is already happening. Vision and the ability to judge whether something looks right used to be worthless without a coder to build it. Now that judgment is the valuable half of the job and the model writes the code. Not knowing how to code stopped being the ceiling.
What happens when an AI system admits it made a mistake?
It depends on whether you built guardrails that make the confession possible. When a system flags its own mistake instead of hiding it, that is worth more than a system that never seems to fail, because you cannot fix what nobody tells you about.