A friend of mine runs his whole business on Grok. Not because he compared it to anything. Because it was already in the app he had open.
That's how most small business owners pick their AI. It's in the phone, or a nephew said to use it, or it came with the subscription they already pay for. Then they use it for everything: writing the email, researching the vendor, checking a tax question, drafting the ad. And they call all of it "AI," like there's one AI out there and these apps are just different doors into it.
There isn't one AI. And the thing you're typing into is not the thing doing the thinking.
I'm a master automotive technician. I've spent thirty years explaining complicated machines to people who just need theirs to work. This is the same job. Let me pop the hood.
The engine and the dashboard
Every AI product you've heard of is two separate things bolted together.
The engine is the model. GPT, Claude, Gemini, Grok, Llama, and a few others. This is the part that does the actual thinking, if you want to call it that. It's built by one of a handful of companies with enough money to train one. It is expensive, it is rare, and it is what people mean when they say "large language model," or LLM.
The dashboard is the app. ChatGPT. Perplexity. Claude.ai. Cursor. Copilot. The Grok tab inside X. The "AI assistant" your accounting software just added. This is the part you actually touch: the chat box, the buttons, the memory feature, the web search toggle, the file upload, the price.
Here's the part nobody tells you. The same engine ships in different dashboards. And one dashboard can swap engines.
Perplexity doesn't build its own engine. It runs on GPT, Claude, and others, and lets you pick from a menu. Cursor, the coding tool, does the same. Microsoft Copilot is mostly OpenAI's engine in Microsoft's dashboard. Even ChatGPT itself, OpenAI's own dashboard, has a dropdown where you choose which of their engines you're driving today.
I do this every day. I work mostly inside Anthropic's dashboard, a terminal tool called Claude Code. From that same seat I can hand a job to OpenAI's Codex engine and get the result back without leaving. Same chair, same files, different engine under the hood.
Two cars with the same engine drive differently. Not because of the engine. Because of everything bolted around it: the transmission, the tires, the tune, what the dash shows you and what it hides. AI is the same. When someone tells you "Perplexity is better at research than ChatGPT," they're usually talking about the dashboard, not the engine. More on that in a minute.
What the engine actually does
Here's the mental model that fixes most of the confusion.
An LLM is not a database. It doesn't look things up. It was trained by reading an enormous amount of text up to a certain date, and what it learned is a very, very good sense of what word comes next. That's it. Ask it a question, and it produces the most plausible-sounding continuation, one word at a time.
Three things fall out of that, and they explain almost every frustrating experience you've had.
It has a cutoff date. Everything after training day, it doesn't know. Not "knows a little." Doesn't know. If the dashboard doesn't go fetch current information for it, it will fill the gap with something that sounds right.
It doesn't know what it doesn't know. There's no little flag inside that says "I'm guessing here." The confident tone is the same whether it's reciting a fact it saw ten thousand times or inventing a part number. That's what people mean by "hallucination." It's not a bug that gets patched. It's what a next-word predictor does when it runs out of road.
It has no memory of you. Each conversation starts blank. When a dashboard "remembers" you, that's the dashboard saving notes and pasting them back in front of the engine next time. The engine itself forgot you the moment you closed the window. I wrote about how that broke my dad's project in My Dad Sent Me a 64-Page Claude Session.
So when you ask "what's the best AI," you're asking the wrong question. You're asking which engine is best without asking what the car is for.
Why Perplexity feels smarter for research
Now the Perplexity thing makes sense.
When you ask Perplexity a question, the dashboard goes and runs a web search first. It pulls the actual pages, hands them to the engine along with your question, and tells the engine to answer from those pages and cite them. You get footnotes. You can click through and check.
When you ask plain ChatGPT the same question with web search off, the engine answers from memory. Training-day memory. It might be right. It might be right as of eighteen months ago. It might be a confident blend of three things it half-remembers. Same confident tone either way.
That's not a smarter brain. That's a design decision about what to put in front of the brain before it answers. Perplexity's whole product is "search first, then think, then show your work."
And here's why this matters for you: most of the major dashboards can do that now. ChatGPT has a search toggle. Claude has one. Grok leans hard on live posts from X, which is great for what people are saying right now and much weaker for a state licensing requirement. The engine isn't what makes research trustworthy. What you feed it is.
The habit I'd teach every owner: before you trust an answer, ask yourself, did it look this up, or did it remember it? If there are no sources, it remembered. Treat it like advice from a smart friend at a bar. Useful. Not something you file with the state.
Grok, ChatGPT, Claude, Gemini: same question, four answers
Try this. It takes five minutes and it will teach you more than any review.
Pick a real question from your business. Not a trivia question. Something like "What do I have to post on the wall if I hire my first W-2 employee in Oregon?" Ask it in ChatGPT, Grok, Claude, and Gemini. Same words.
You'll get four different answers. Some longer, some shorter, some with sources, some with none. One might be dead wrong about a form number. One might refuse to guess and tell you to call the state, which is annoying and also the correct answer.
That's not four AIs having a bad day. That's four different engines, trained on different text, with different cutoff dates, sitting inside four different dashboards that make different choices about searching, citing, and hedging. Grok isn't wrong for being Grok. It's just tuned for a different job than the others, and it's your job to know which one you're driving.
I'm not going to tell you which one wins. It changes every month, and anybody who tells you otherwise is selling a dashboard. What doesn't change is the way to judge: does it show sources, does it admit when it's unsure, and does it get your kind of question right when you check.
The part that should actually worry you
Everything above is about picking the right tool. This part is about not getting trapped by it.
Here's how it goes. You start with one dashboard. You upload your price sheet, your service menu, the email templates you've refined over ten years. You correct it a hundred times until it writes like you. It saves your preferences, your projects, your history. Six months in, it's genuinely useful, because it knows your business.
Then a better engine comes out. Or the price triples. Or the company gets bought and the model you use disappears from the app you learned it in. That last one isn't hypothetical. It just happened to Cursor's users this month when OpenAI announced its engines would stop working inside Cursor after SpaceX bought the company. Nobody using Cursor did anything wrong. The deal changed over their heads.
And you find out that the six months of context you built lives inside that dashboard. The files are exportable, sort of. The decisions, the corrections, the "no, we don't do it that way, we do it this way" that made it useful? That's in old chat logs. You can download them. You can't hand them to the next tool and have it understand.
Nate B. Jones, who writes about this better than anyone, calls it custody. A price increase you can see. A missing project history you only discover when you try to leave. The dashboard companies aren't villains for this. A product gets more useful the more it knows about you, and the same feature that saves you time every day is the one that makes leaving expensive. But you should understand the bargain you're making every time you upload a file.
Own the brain, rent the engine
The fix is not to avoid useful features or keep five copies of everything. The fix is to keep the part that's actually yours outside every dashboard.
Nate built and open-sourced a design for this called the Open Brain. I built one following his plan and I run my business on it. The idea is dead simple. Your accumulated context, the decisions, the standing rules, the "here's how we do it" notes, the record of what you tried and what happened, lives in a small database you own. Every AI you use can read from it and write to it through an open standard called MCP that OpenAI, Anthropic, Google, and Microsoft all support. Claude reads it. ChatGPT reads it. The engine that ships next month will read it.
That flips the whole relationship.
The dashboard becomes disposable. New engine comes out Tuesday? Point it at your brain. It knows your business by lunch. No migration, nothing left behind.
You can actually run the five-minute test. Compare engines on real work, because none of them has a six-month head start on knowing you. Whichever one is best this quarter gets the job. Next quarter, maybe a different one.
The price of leaving drops to nearly zero. Which changes how every vendor treats you, whether they know it or not. You're not a captive. You're a customer who can walk.
Your business stays understandable outside the AI. This one matters more the more work you hand off. Replacing a chatbot that drafted a paragraph is a nuisance. Replacing a system that knows your pricing logic, your vendor history, and your customer follow-up rules is closer to replacing an employee. Keep that knowledge where you can see it.
Nate's line, and I've stolen it for good: you own the memory, and you rent the intelligence. The intelligence is a commodity now. New ones arrive every few weeks and every one of them wants to be your only one. The part worth protecting is the record of what you've thought and decided and learned. Keep that behind your own door.
If you want the next level of this, I wrote about what happens once your brain actually remembers everything, and why that creates a new problem: Your AI does not need more memory. It needs a memory policy.
Pick the dashboard for the job
Here's the short version to tape above your desk.
| The job | What to look for | Why |
|---|---|---|
| Research, anything with a date or a number | A dashboard that searches first and shows sources. Perplexity, or any chat app with search turned on. | The engine's memory is stale. Sources are the only thing you can check. |
| Writing in your voice | Whichever engine you like the prose from, fed your examples every time. | Voice is context, not engine. Any engine writes like you if you hand it your best ten emails. |
| Your own documents | A dashboard that lets you upload or connect files, and reads from them instead of from memory. | You want it answering from your price sheet, not from the internet's idea of one. |
| Math, counts, anything with money | Make it show its work, then check one line yourself. | Next-word prediction is bad at arithmetic in a way that's hard to see. |
| Long-term memory of your business | Not a dashboard feature. Your own brain, outside all of them. | Anything inside one app is that company's to keep. |
And one habit, if you only keep one: ask where the answer came from. Did it look it up, or did it remember it? Everything else in this article follows from that question.
I'm not an engineer. I fix complicated machines for a living and explain them to people who just need theirs to run. If you want help setting this up for your shop, your practice, or your household, that's what Wildash does.
