HASSANHABIB

On Spec

12 min read   aisoftware-engineeringagents

On June 19, 1905, a converted storefront in Pittsburgh began charging five cents to sit on a hard wooden chair and watch pictures move against a sheet on the back wall. Harry Davis and John Harris called it the Nickelodeon, joining the common word for a five-cent piece to the Greek word for a theater. Within five years, something like twenty-six million Americans were visiting places like it every week. If you wanted to see a moving picture, you went to where the picture was, at an hour somebody else had chosen, and you watched whatever happened to be running.

The arrangement is worth stating plainly, because every improvement since has been an argument with it. The experience existed first. It sat in a fixed place on a fixed schedule, and the person rearranged their evening around it. Software has rested on the same assumption for as long as there has been software, which is where I am going with this.

0/ The Schedule

Television carried the screen into the house and left the schedule exactly where it was. Families still gathered at a particular hour for a particular program, and the listings were printed in the newspaper so you could plan your week around other people’s decisions. The place had been solved. The time had not.

The home video recorder broke the time. A person could capture a broadcast and watch it after the children were asleep, which sounds unremarkable now and was contested enough to reach the Supreme Court. In 1984, in the case brought against Sony over the Betamax, the Court held that recording a broadcast at home in order to watch it later was fair use. What was being argued over, in the end, was whether an ordinary person was allowed to decide when.

Then the catalog arrived in a pocket. Online video removed the place and the schedule together, and it removed something else that had been invisible while the other two were in the way: you no longer had to know in advance that you wanted the thing. You could want it at eleven at night and have it at eleven at night.

Each of those steps handed a little more control back to the person watching. Every one of them was an answer to the same two questions. Where, and when.

1/ On Spec

Underneath all four arrangements sits a condition nobody thought to name, because there had never been an alternative. The thing you watched had to already exist. Somebody wrote it, cast it, shot it and cut it, months or years earlier, for an audience they could only imagine. The nickelodeon and the phone differ enormously in convenience and not at all in this: both hand you a finished object made in advance by people who did not know you.

Software is made the same way, and the habit is so old that it gets mistaken for the nature of the thing. A team decides what a product will do, builds it, and ships it to people who have not arrived yet, to perform tasks nobody has stated yet, through an interface chosen before anyone’s actual problem was visible. I call this Spec Software, after the way builders describe a house put up before there is a buyer. A house built on spec is a real house and often a good one. It is also a guess, and the person who eventually lives there takes the kitchen the developer guessed at.

That is what using an application is. You find one that roughly fits your purpose, you learn where its designers put things, and you bend your work until it fits the shape of their guess. The better the guess, the less bending, and the industry has gotten remarkably good at guessing. The bending never goes away, because the guess was made before you existed.

2/ The Request

The next step in that progression is not about where or when. It is about what.

I expect an increasing share of what people watch, read and use will be assembled after they ask for it rather than selected from what was made before. Someone says they want a romantic comedy with a particular pair of characters, set somewhere cold, slower through the middle than films of that kind usually are, and the experience is put together as they watch it. Someone with a forty-minute drive asks for a conversation about a subject they are curious about, in a register they find easy to listen to, lasting exactly as long as the drive. Neither existed the moment before it was requested.

I want to be careful about how strongly I say this. I am describing a direction I find persuasive, not a scheduled arrival, and the interesting question is not whether any of it is possible but how quickly the cost of generating something falls below the cost of maintaining a catalog of everything a person might want. That is an economics question at least as much as a capability one.

3/ Layers of Manifestation

If experiences are going to be assembled on request, something has to do the assembling, and the shape of that is already visible in three layers. I think of them as Layers of Manifestation, because each one sits a step closer to the moment a purpose becomes something a person can see, use or hold.

Three stacked layers. At the bottom, deterministic systems, owned by software engineering. In the middle, non-deterministic systems, the agents, owned by AI engineering. At the top, the box, owned by user interface design. A person sends intent down and receives an experience back, and the authoritative records never leave the bottom layer.

The bottom layer is the systems that hold what is true. Records of accounts, of appointments, of orders, of medical history, of who is allowed to do what, along with the rules that keep those records consistent. This layer is not going away and it is not getting smaller. It is getting more exposed, through APIs, through MCP servers, through data resources that a program can reach rather than a person. I would resist one comfortable assumption here, which is that putting an operation behind an API makes it dependable. An endpoint can be as ambiguous, as racy and as poorly specified as anything else. Substantial and unglamorous engineering still lives in this layer, and the value of the layers above depends on it being boring and correct.

The middle layer interprets what somebody wants. It reads the instruction, consults whatever skills and policies apply, decides which tools to call and in what order, coordinates with other agents when the work is larger than one of them, and decides when it is out of its depth and a person needs to be asked. This is the layer that improvises, and it should be the only one that does.

The top layer is the thing the person actually talks to. A harness, an agent application, a super app. I have been calling it the box, which is a working metaphor rather than a product, and I do not mean one company operating one box for everybody. I expect many boxes from many providers, differing in taste and in what they are connected to, built around the same idea: you state a purpose, and the box connects that purpose to systems that can serve it, while making the progress, the results, the permissions and the controls legible to you.

Those three layers are also three jobs, and that is the part I find worth sitting with. The bottom layer is software engineering as it has always been, and none of its obligations soften because something clever sits above it. The middle layer is the new one, and it is where most of what gets called AI engineering actually lives: not training models, but deciding what an agent may attempt, what it must read before it decides, and what happens when it is wrong. The top layer is still design, applied to a different object. Instead of laying out the screens, it specifies the vocabulary the screens are assembled from, the controls that let a person look inside, and the moments where the system has to stop and ask.

4/ Tomorrow’s Meeting

Consider a sentence a person might say to such a thing. Help me prepare for tomorrow’s client meeting.

What comes back is not an application. It is a workspace built around that meeting and nothing else: the recent correspondence with those people, the commitments still open from last time, the two documents that matter and not the forty that do not, the agenda beside the calendar entry, a draft of what might be said. The email system still holds the mail and the calendar still holds the appointment. Nothing was copied out from under its owner. What was assembled is a view with a purpose, and when the meeting is over most of it can dissolve, because its reason for existing has passed.

The organizing unit of that experience is the task, not the product. This is the part I find genuinely different. For as long as I have worked in software, the unit has been the application, and a person’s day has been spent moving between applications and carrying context across the gaps by hand. If the task becomes the unit, that carrying is the work the machine does.

None of this implies talking instead of looking. A conversation is a poor way to compare six options or move an appointment by twenty minutes. The box should put a table on the screen when a table is the right answer, a calendar you can drag when that is the right answer, and an editor when there is something to write. What changes is not that interfaces disappear, but that they are assembled around a purpose instead of chosen from a catalog, and that the ones you return to every day can be kept exactly as you last arranged them.

5/ What Holds

A prediction is only useful if it says what it does not cover, so let me be direct about what I think survives all of this intact.

Authored work survives. There is a difference between a film someone made because they had something to say and a film assembled to match my mood on a Tuesday, and the difference is not resolution or pacing. It is that someone meant it. Shared work survives for a related reason: a large part of what a film is worth is that other people saw the same one, and an experience generated for an audience of one cannot be argued about over dinner. Recordings of things that actually happened survive absolutely, because their whole value is that they were not assembled to please anyone.

The line I would hold hardest is between presentation and record. How a story ends can bend to what somebody prefers. What is in their account, when their appointment is, what dose they took and what they agreed to in writing cannot bend at all. A system that blurs those two has not become more helpful, it has become less trustworthy, and the bottom layer exists precisely to make that line enforceable rather than aspirational.

There is one more thing that holds, and it is the objection I find hardest to answer. An application can be inspected. It has a version, a history of changes, a place to report that it is wrong. An interface assembled for one person for one afternoon has none of that, and when it shows a number that is not right, it is not obvious what anyone would even file a complaint against. I do not think that is fatal. I do think it is unsolved, and that solving it is more of the work than generating the interface was.

6/ Whose Box

Which leads to the question I would want asked of any box, including one I built.

Something is choosing which tools get called, which results are surfaced first, which of several possible answers is the one you see, and what appears when you have not asked for anything in particular. Those are not neutral mechanics. They are the same choices that decided what was on television at eight o’clock, made faster and much closer to the person. A box that assembles your work, your entertainment and your record of what you agreed to has more influence over an ordinary day than any application ever had, and the honest question is not what it can do but whose purposes it serves when those purposes and yours are not the same.

I would also watch what sits underneath. Many boxes competing for people’s attention is a healthy picture, and it stops being one if all of them are thinking with the same few systems owned by the same few companies. Plurality at the layer you can see is not plurality if the layer you cannot see is held by three parties. That is a question about ownership rather than capability, and it will be settled by decisions rather than by progress.

What I want from all of this is narrow and I think worth saying plainly. People have spent forty years learning the shapes of other people’s guesses, and getting good at software has partly meant getting good at accommodation. If less of that is necessary, the time comes back. Whether that is progress depends entirely on what fills it, and the measure I would use is not how little a person has to do, but how much more of what they meant to do they actually get to.

7/ Mutating Matter

Everything above happens behind glass. The box assembles a workspace, a film, a conversation, and all of it lives on a screen, in a headset, or in a voice in your car.

The last step is the one where it stops being light. You say, hey AI, make me a chair, and what answers is not a picture of a chair or a link to buy one, it is a chair. Not one from a catalog, but one that fits the way you actually sit, in the room you actually have.

None of this exists, and I offer it as a direction, not a forecast. It is not a new direction, either.

In August 2008, Intel’s chief technology officer, Justin Rattner, described programmable matter at the Intel Developer Forum: millions of tiny robots called catoms, each carrying its own processor, holding on to one another to take whatever shape they were told. The work came out of Intel Research Pittsburgh and Carnegie Mellon under the name Claytronics, and it never left the lab. I suspect it was not wrong so much as early, and I find it fitting that it came from the same city as the nickelodeon.

Follow that line far enough and intelligence stops sitting beside matter and moves into it, woven into material at the smallest scale we can build. I call this Mutating Matter: substance that changes its shape because somebody asked. The three layers do not disappear when that happens, they get heavier, because the bottom layer becomes physics and boring and correct becomes the only thing between a person and the floor.

How the chair looks can bend to taste. Whether it holds you cannot bend at all.

Then something stranger happens. Matter that can understand an instruction can also say something back, and a chair that knows its own shape knows when a leg is cracking. That would be the end of silent matter, because for as long as there have been people, the only feedback the world gave us was whatever broke.

Think of a tree. It has been telling us things all along, in curling leaves and thinning bark, and soil sensors already translate a little of it. Carry that far enough and the tree simply says it: I need water, I need light, please do not cut here.

This is where the idea turns against the way I began it. I opened with a command, make me a chair, but you do not command family, and a world that can tell you it is hurting is very hard to treat as furniture.

The feeling itself is old. Many cultures have long held that rivers and forests are kin rather than resources, and in 2017 New Zealand recognized the Whanganui River as a legal person. It is also my answer to the question of whose box: if matter takes orders, the frightening question is who else can give them, and matter you are in a relationship with is not a device anyone can simply instruct.

Every arrangement in this essay has been an argument about where, when and what. The last question is with whom. The first moving pictures asked people to sit on hard wooden chairs and watch whatever was running, and it would be a fitting end to that story if the chair, the room and the tree outside the window were no longer things we used, but company we kept.


Hassan

Comments

← All posts