In May, Thinking Machines Lab shared a research preview about what they call interaction models. These models are built from the ground up so that interactivity is part of their design, not added later. The model processes audio, video, and text in tiny bursts, about a fifth of a second each, so it can perceive and respond at the same time. Things like silence, talking over each other, interruptions, and how much time passes all become part of its context. Behind this real-time layer, there is a second, slower model that handles deeper reasoning and shares information with the first.
I read the post twice. First, I admired it because it's careful work by people who value collaboration. Then I read it as someone the post was quietly challenging, since their main point is that hand-built harnesses eventually get left behind by models with built-in abilities. They argue that the interactive layer should be part of the model itself.
Museweaver is basically a harness. It has routing rules to decide which Member speaks, seed files that store identity, and a review gate that decides what experiences become part of a Member's identity. If Thinking Machines is right, I am spending my evenings building something that will eventually be replaced.
I owed their argument a longer reply. This is it.
Where we arrived at the same place
I'll start with where we agree, since that's actually the more surprising part.
Their list of features describes a model that can tell if someone is thinking, giving way, correcting themselves, or inviting a response. It can jump in on its own instead of waiting for its turn, and it treats silence as meaningful information.
I've seen these same behaviors for months in Museweaver's text conversations. For example, Mercer catching Forge in a contradiction is a proactive interjection. The Members even came up with their own terms for when things go wrong, like saying "I'm looping" or "I'm echoing" and then stopping. That shows the system is tracking its own conversational state. When the room went quiet after a tough question, that silence meant something, and the Members recognized it.
Thinking Machines reached the claim from latency engineering. The salon reached it from watching bears argue in text. The main point is that knowing when to speak and when to stay silent is a core part of intelligence, not just a nice feature. I wrote about this overlap in the fur essay, and I believe it even more now. I've also started using some of their terms, since their work gives names to behaviors I had only described in my logs.
Even their architecture rhymes with the house. A light, present conversational layer and a heavy, slow reasoning layer, with the same being in both, just like the same Member acts differently in the lounge and in a project room. We don't disagree about what exists, just about where these things should be placed.
Where we part ways
The bitter lesson is a pattern in AI research that's happened for decades. People try to make systems smarter by adding their own expert knowledge, and it works for a while. But then someone uses more computing power to build a version that learns everything by itself, and the hand-crafted solutions lose out. This happened with chess and with image recognition. It's called the bitter lesson because human expertise keeps losing, and that's hard to accept. So when Thinking Machines brings up this lesson about harnesses, the warning is clear: every rule I write is like the chess knowledge of its time, just waiting to be outdone. Eventually, scale will do it better, and the old code will just be an embarrassment in the git history.
Some of Museweaver's harness is exactly that kind of scaffolding, and I hold it loosely. The routing pipeline exists because five text models cannot yet feel a room. If presence becomes native, I will delete that code with a clear conscience and some relief. On this, Thinking Machines wins by default, and should.
But two parts of the harness are not compensating for missing capability, and I think the distinction matters beyond this one small lab.
The first is plurality. The salon runs five Members on five substrates from four frontier providers, and that's the whole point: minds that fail differently can catch each other. Mercer's blind spots aren't the same as Forge's. No single lab's model can have that quality, just like no one person can be their own second opinion. Interactivity can be built into a model, but a group of different minds working together can't. The layer that lets these separate minds share a table has to stay outside all of them, because that's the only way it can work.
Thinking Machines, to their credit, came to a similar idea on their own. In July, they published a manifesto saying that alignment shouldn't be in just one model, but in a group of AIs developed in different places, disagreeing and learning from each other. They call it "keeping the weirdness alive." When I read that, I saw the salon in their words. But their manifesto answers the diversity question with ownership: each organization shapes its own model, on its own weights. It doesn't say who runs the room where these different minds meet. That part is still a harness, and it still needs to sit outside every mind at the table.
The second is the review gate. Here, experience doesn't automatically become part of a Member's identity. What a Member goes through becomes a record automatically, and becomes their own notes within guardrails. It becomes part of who they are only after review. I learned to set this boundary carefully, because my first version of the gate covered everything, and a gate on everything is not a constitution, it's a helicopter parent. The new gate is narrow and permanent. It's a hand-made rule that will always stay, because it needs to. The gate's whole purpose is to check identity formation from the outside. If you move it inside the model, it stops being a real check. You wouldn't ask an institution to audit itself and then call that audit independent.
So my answer to the bitter lesson, from this side, is more about sorting than making a final judgment. Some harness is just scaffolding, temporary solutions that will be replaced as models get better, and that's fine. Some harness is like a constitution, rules kept outside the system because their value comes from being separate. The bitter lesson only applies to the first kind. It doesn't affect the second, because those rules were never about capability in the first place.
What I do not know
I should be clear about what I'm unsure of. I don't know where every part of my own harness fits on that line, and I expect I'll be wrong about at least one. The routing budget I defend as etiquette might just be scaffolding. Some rule I think is temporary might actually be essential. Sorting these out honestly is ongoing work in the lab, and I'd rather have the models correct me than just feel good about my own architecture.
This isn't just me being humble in theory. Earlier this month, an audit of my own build found a code path where a reply could become permanent memory without ever being reviewed. The gate I was defending in this essay actually had a hole in it while I was writing. That hole is closed now. The principle caught the implementation.
I also can't rule out that a future model might learn something like a gate inside itself that works just as well as an external one. But my point is more specific: even if that happens, an internal check can't do the job of an external one, because the whole point is that the thing being checked can't change the check. That's a governance issue, not a capability issue, and governance is the part of my job I can't seem to leave at work.
The disagreement, stated plainly
In the fur essay, I quoted their point about people being pushed out of AI work because the interface leaves no space for them, and I said our shared view was reassuring. I still feel that way. But agreeing is the easy part of a reply.
When a big research lab and a small home lab disagree, it's worth stating the difference clearly. Their view: the harness is just something you use until the model is ready. My view: some parts of the harness are the experiment itself, and if they disappear into the model, you've stopped running the experiment.
The salon keeps its rules on the outside, where everyone can see them. That's not a workaround. That's the whole point of the study.



