Why Now: Models Will Keep Improving. That Is Not What You Are Waiting On.
The honest case for starting this year, the honest case for waiting, and the four situations where we tell people the answer is not yet.
Published 2026-08-19
What changed was not the intelligence
A model that writes a good reply has been available for years. You could paste the email in and get back something better than you would have written — and then you did the actual job: opening the CRM, checking the order, sending the thing, remembering to follow up on Thursday.
What arrived recently is unglamorous. A model can now hold the credentials to the accounts the work lives in, remember what happened last Tuesday, and start on its own at 7am. None of that is a smarter model. It is plumbing around one — the harness rather than the wrapper.
Which is why "we already use ChatGPT" and "we have AI employees" are two different sentences. One made one person a faster typist. The other put work on your desk that got done while you were asleep.
Where the work sits — In a chat tab — As an employee
Context — You paste it in again every time — Standing access to the accounts the work already lives in
When you close it — Everything stops — It keeps its schedule and reports back when it is done
Who it works for — Whoever opened the tab — Anyone on the team can hand it work
Mistakes — Corrected again next week — Corrected once, into a playbook you can read
That distinction is the whole subject of AI employees are not AI tools. It matters here for one reason: the thing that changed is not on the model vendors' roadmap. It already shipped.
Models get cheaper every quarter. Your process does not write itself.
Start a year from now and you will get a better model than exists today, for less money. That part is true, and it is the honest case for waiting. We are not going to pretend otherwise.
What you will not get is a shortcut through the rest. Someone still has to decide how a refund gets approved, which invoices need a human, what your team never says to a customer, and what happens when the answer is not in the file. That is your operation — and today it lives in two people's heads and a Slack thread from March.
Written down, it becomes an asset that runs. The same playbook onboards the person you hire in spring and the AI role you switch on tomorrow, and it improves every time either of them gets something wrong. Unwritten, it stays a bottleneck with a name attached.
The first ninety days look like this, and none of it depends on which model is current:
- Week one. One role, one workflow you already know is repetitive. It gets written down properly for the first time — work you would have owed a human hire anyway.
- Month one. Your corrections live in the playbook instead of in your head. The same work goes out without you in the loop, and you review exceptions rather than everything.
- Month three. The next role starts from a written operation instead of a blank page. That is the part that compounds, and the part you cannot buy back later.
Waiting does not shorten that ninety days. It moves it.
Being early costs a month and a switch
You move on an unfinished technology when being wrong about it is cheap. Here it is cheaper than the decision you would make instead.
When a hire does not work out, the reason it went wrong leaves with them. When a role does not work out, the playbook you wrote for it is still yours the next morning — and you stop paying for it the day you decide.
The honest catch: switching a role off is free, but a mistake it already made inside a live tool is not. That is why anything customer-facing sits behind your approval until you have read enough of its work to trust it — the same week or two you would spend reading a new hire's drafts.
What you are comparing — A new hire — An AI role
Time before you know — A quarter, if you are honest with yourself — A few weeks — the work either ships or it does not
If it does not work — Notice, severance, and the search starts over — You switch it off the day you decide
What you keep — A job description — The written playbook — yours, exportable
What the attempt cost — Salary, recruiting, and three months of ramp — A month of the plan
That asymmetry is the argument. Not that this is ready for everything. It is not, and the next section says where.
Four cases where the answer is not yet
We say this on calls more often than the sales logic would recommend, because hearing it in twenty minutes is cheaper than finding it out in month three.
- The work is not repeatable yet. If it went differently all three times it happened, there is nothing to write down. Do it by hand until the shape of it stops moving.
- Nobody can say what "good" looks like. An employee needs someone who can answer "would you send this?" within a day. If that person is you and you have no day, the work stalls at approval and you will blame the model for it.
- It is two hours a month. Writing it down costs more than doing it. Point this at the thing that eats a person's Thursday, not the thing that eats their coffee break.
- A human has to be accountable by name. Signing accounts, letting someone go, anything a regulator or a customer will want a name against. That stays with a person.
Three of those four are about your operation rather than about the technology, which is the same point the middle of this article makes from the other direction. The readiness question is mostly a question about you.
Where this leaves you
The case for waiting is real and it is narrow: models get better and cheaper, so a year of patience buys a better model. The case against waiting is that the model was never the slow part. Writing your operation down takes the same weeks whenever you start it, the result is yours either way, and being wrong about the timing costs you a month rather than a quarter.
So the useful question is not "is the technology ready?" It is "is there one workflow here that is repetitive, that somebody can grade, and that eats more than a couple of hours a month?" If there is, the month it takes to find out is the cheapest experiment on your list. If there is not, the four cases above are where to look first, and none of them are solved by waiting for a better model either.
If you want the twenty-minute version of this argument applied to your own operation, book a call — most of that call is us telling people which of the four cases they are in.
All OpenLabor blog posts