Writing systems
For the past year or two I’ve been doing product and engineering management for the team building our editorial AI buddy, Content Agent, and Sanity Context, our way of bringing knowledge to agents while they get stuff done.
This work has involved a slew of surprises in a bunch of places: the absolute importance of words and writing, theory of mind, contextual computing, the importance of zoomable summaries for domain comprehension, information dynamics in collaboration. And Anthropic’s “surprising” caching strategy.
More on these later. Let’s talk about words for now! They feel the most important.
Whatever else they might be, models are squeezed out of culture and wrought of words. That you can speak to a construct boiled out of human meaning-making is crazy. Truly unhinged stuff.
And LLMs running in agentic loops will generally attempt to do whatever you tell them. If your instructions are poor, they will get it wrong. Give them harsh counterexamples and you’ll be rewarded with more weird behavior. This is also typical of bad leadership btw.
And if you’re not just asking for an improvement on your gochujang cabbage recipe, but making an agentic harness for people to use in everyday situations: your users will be unclear as well. They’ll be imprecise in their wants and inexact about their circumstances. So you want every iota of cleverness and clarity the agent has to persist into the conversation. You don’t want it wasted on your own avoidable mistakes.
Also: agents are, for most practical purposes, born anew with every conversation. So they’ll keep on repeating your mistakes by proxy forever unless you make them self-improving. But you likely don’t have the scale for that, and it might not work even if you did. You could just end up noodling around in a local maximum of confused thinking.
If English is the hottest new programming language, it comes without linters, syntax checkers, compilers or much of a formal way to know when something is off.
In the Bronze Age of LLMs (idk, 2022?) there was something called “prompt engineering”. A lot of this was little tricks where you could make the model better by telling it “think step by step”, “I’ll give you 100 bucks if you get it right”, or “it’s January”.
With better reinforcement learning, though, many of these tricks were baked in and made redundant. Now models are so much better, provided you’re just being super, super clear.
So what does this look like?
There’s no e-meter for this kind of clarity. You just need to guess, as accurately as possible, what the most likely interpretation of your system prompt is, after factoring in some of the wiles of the model.
This is also what any writer does when they write. Good writers are better at it. We run the text on a representation, an approximation of the audience, to gauge a reaction.
It’s also what we do when we read. Our ability to simulate interpretation, and the narrative form itself, is probably very closely linked to our ability to form complex hierarchical societies.
Technical product organizations already do a lot of writing. We write memos. Documentation. Little notes on Slack. Most of this is writing for clarity, but it has more contextual scaffolding and usually fewer consequences for the odd lapse in judgement. The stakes are lower.
You want writers who are to prompts what Sorkin is to the walk-and-talk.
So who has been thinking about clarity in writing?
Well, one of Anthropic’s headline tips is to show your prompt to a colleague with no context. Actually, this is pretty good. At least it quickly gets at non-generalizable terms of art.
“What do you mean by LRF support here?”
It’s also a typical tip in Swedish klarspråk, “clear language”, whose practitioners have been trying to get Swedish bureaucrats to write well for the last 30 years.
But I’m sure there’s room for real discipline here. If you believe LLMs can do cognition and real work, you should also believe we’ll get schools, consultancies and actual businesses. Language nerds specializing in clarity. A game of glass beads at the fulcrum, levering a trillion-dollar bet on enchanted inference sand.
The humanities have 2k yrs of intellectual history and a slew of fields that could contribute a lot to this, but currently seem most preoccupied with policing the use of AI in their own ranks. As a humanities person, this saddens me.
As Leif Weatherby points out, we now have empirical evidence for language, and perhaps meaning, arising from differences between signs, with no grounding needed. This is a totally rad finding!
Saussure scores! And it’s Saussure–Harnad 1–0!! These ball knowers have never even seen ball!
Meanwhile, Netflix is hiring a Staff Systems Designer, Language: