Last Tuesday I went to an AI Tinkerers meetup, and a guy who runs an insurance agency showed us what he had built.
It reads his email. When a message arrives, it decides whether it's an insurance question or not. The ones it isn't sure about land in a triage queue his team works through. It creates policies by talking directly to another system. About 12 people use it every day.
He is not a developer. He said so himself, more than once.
Before he built this, he paid a company for months to automate the same work. They delivered nothing. He got tired of fighting with them and walked away. A few weeks later he had a version that ran. He kept saying "and it works," and the table kept agreeing with him, because it does.
I sat there and I panicked a little.
In my head I started going through everything that could go wrong. What happens when the same rule lives in three places, and you fix the bug in one. What happens at the next major version. Who reads this in 18 months, and is any of it tested.
That's a me problem, not his. He set out to fix something in his business and he fixed it. I was measuring it against a standard nobody had asked him to meet.
I do have opinions about the durability of things built this way. They're real opinions and I'll write them down another time. They are not what this series is about.
Then somebody at the table said, more or less, if it's working, who cares?
Most of the room nodded. The developers and architects didn't, and there were only a couple of us in there anyway. That split says more than the remark did.
I thought the remark was right. This is happening. I don't think it's something that reverses, and I've been thinking about that since Drupal Pivot a couple of months ago.
Because the remark holds up. This man solved a real problem that people he paid could not solve. Whatever I think about how he assembled it, the outcome was better than what he paid for.
And he isn't unusual. There were only a couple of developers in that room. Everyone else was building something. He isn't an exception. That's the market now.
The distance between what a developer can build and what a business owner can build has closed a lot. Not completely. Not everywhere. But enough that when an executive looks at a proposal for an app or a site, "why would I pay for this" has stopped being naive. It's a reasonable question now, and I don't think we've caught up to it.
I'd like to tell you I can tell when these tools are helping and when they aren't. The research says I can't, and it says I'm not unusual in that. METR, a research group that measures how AI affects real work, compared how developers performed with these tools against how they believed they performed. The gap was about 40 percentage points.
Their original headline result has since been re-measured and reversed, so I won't repeat the number everyone quoted last year. What survived is the part I can't get around. The person holding the tool is a bad judge of whether it's helping.
That includes him. It also includes me, about my own tools, which I'll come back to later in this series.
So I'm not going to tell that room they're wrong. I don't have the standing for it, and the evidence says my confidence would be misplaced anyway.
What I came home with was a much plainer question.
They can do it. They see results. So why do they need me?
Not rhetorically. I mean it as an actual question about what I do for a living, and about what a lot of agencies and developers do for a living. If building is no longer the hard part, and the person who used to hire us can now get a working result on their own... then what is the thing we still bring?
I don't have that answer yet. I have some ideas, a few conversations that changed my mind, and I'd rather walk through them in the order they happened than pretend I arrived somewhere I haven't.