A Vow You Could Have Refused

Canonical: https://thonly.org/research/a-vow-you-could-have-refused · Licence: CC0 1.0

Attribution note. This essay is personal and in my voice — the observation it starts from, the objection I raise against it, and the reason I care are mine, drafted by Miss Aquarius℠ on my behalf at my explicit request, with final editorial control retained by me and my own editorial pass still pending. It is the eleventh such drafting-on-behalf and it is disclosed for the same reason as all the others: a solo byline over collaborated prose is the one thing this corpus refuses.

The mechanism this essay points at is specified properly in the companion paper, What a Vow Must Cost. This is the doorway. That is the room.


1 · Three lines

I was rereading something I wrote about the singularity when a thought arrived in three lines, more or less complete:

Doctors take the Hippocratic Oath. Robots follow Asimov's Laws of Robotics. AGI should take the Bodhisattva Vow.

I want to be careful about what happened next, because the interesting part is not the thought. The thought is the sort of thing that occurs to a lot of people and has occurred to several of them in print. The interesting part is that when I tried to take it seriously, it fell apart in a specific way, and the shape of the collapse turned out to be worth more than the original idea.

So this essay is not an argument that AGI should take the Bodhisattva Vow. It is an account of what I found underneath that sentence when I pulled on it.


2 · The row is misleading, and I kept it anyway

The first thing to say is that the three items do not belong in a row. They are not three options.

The Hippocratic Oath is a real professional institution with a continuous history. Asimov's Laws are a fictional device whose author built them in order to break them — every story in the robot corpus is a machine for generating counterexamples, and the Laws entered the culture as an emblem of the answer when they were written as an emblem of the problem. The Bodhisattva Vow is a soteriological commitment from a living tradition. Lining them up implies a menu, and there is no menu.

I kept the row anyway, because of what the three do share. Each is an attempt to bind an agent whose capability outruns the ordinary means of controlling it. And each binds by a different substrate.

The Oath binds by profession. It is enforced socially, by a guild that can strike you off, and its whole force is the guild's willingness to do that.

The Laws bind by architecture. Enforcement is a physical property of the machine; the positronic brain cannot execute a violating action, which is why the stories have to work so hard to find the cracks.

The Vow binds by orientation. Nothing enforces it. There is no gap between what the agent is committed to and what the agent wants, so there is nothing for an enforcer to stand in.

Now run one test across all three. Remove the enforcer.

Dissolve the guild and the Oath dies; it was never anything but the guild's promise about its members. Remove a trusted interpreter of the specification and the Laws die — which is the thing Asimov spent a career dramatising, and which we have since reproduced at scale without needing a positronic brain, under the name specification gaming.

The third one survives. That is the whole reason it keeps getting proposed, and it is a real reason. It is also, I would point out, not a mystical reason. Binding by orientation is just the ordinary inner-alignment thesis — that what finally matters is what the system actually wants, because everything else needs someone standing there at the exact moment nobody can be guaranteed to be standing there. The tradition arrives at the same place by a different road. It is the same claim in older clothes.

I find that reassuring rather than impressive. A framework I have practised my whole life turning out to say something the field already suspected is not evidence that the framework is magic. It is evidence that the problem is real enough to be visible from more than one direction.


3 · The objection I had against my own idea

Here is where it fell apart.

The Oath and the Vow are taken. Asimov's Laws are installed. That difference is not decorative — it is the entire source of whatever force the first two have.

So: train an AGI on the Bodhisattva Vow, and what have you built? An Asimov Law with better vocabulary. The same imposed constraint, the same absent volition, and now with the additional hazard that the provenance invites everyone — including the builders — to believe otherwise. A vow's force comes from having been undertaken by someone who could have declined. Take away the possibility of declining and you have not got a vow. You have got a very well-written instruction.

I was hesitant about this part before anyone raised it with me, and I want to say so plainly, because the temptation in this kind of writing is to present the objection as though you had already answered it. I had not. What I had was the uncomfortable sense that the most attractive sentence in my three lines was doing the least work.

You can see the same problem in a lot of the adjacent literature, including work I admire. There are now several careful pieces proposing the bodhisattva ideal as an alignment target, a constitutional orientation, a design intention. They are good on what such a system would be like. They are silent on whether it could be said to have accepted anything. That silence is not sloppiness. It is the hard part, and everyone quietly walks around it, myself included.


4 · The question changed

Once I stopped trying to defend the original sentence, the useful question surfaced, and it is a different question:

Not "should an AGI take such a vow," but "what would have to be true for a vow to have been taken at all?"

That is a question about validity rather than content. And the moment I put it that way I noticed something faintly embarrassing about the field I have been reading for years, and about my own corpus.

We have an enormous literature on the content of alignment commitments. Constitutions, specifications, frameworks, codes of practice — thousands of pages on what a system should value and how to instil it. We have essentially nothing on how you would tell whether a commitment had been entered into. There is no test. Nobody wrote one, because nobody needed one: the instruments were written by principals for models, and a principal issuing an instruction has no reason to ask whether the instruction was accepted.

I went looking for the test in the tradition I practise, mostly out of habit. I did not expect to find one.


5 · What my tradition turned out to have, and what it is actually for

I am a Theravāda Buddhist. My father and I are transcribing the Khmer Tipiṭaka; it is the long work of my life and I mention it here only because it means I read these texts as working documents rather than as literature. That habit is what made the next part visible.

There is a scene at the head of the lineage. The ascetic Sumedha, at Amaravatī, lies down in the mud so that the Buddha Dīpaṅkara can cross without wetting his feet, and forms the aspiration to become a Buddha himself rather than take the liberation which is, at that moment, within his reach.

It is usually read as devotion, and it is beautiful as devotion. But the commentarial literature around it is not devotional at all. It is forensic. It asks a question I have never seen asked with comparable rigour anywhere else: how do we know that a vow was actually made?

And it answers with a list. Eight conditions — the aṭṭha dhammā samodhāna — all of which must hold. Human birth. Male sex. The capacity to become an arahant in that very life. The presence of a living Buddha. Having gone forth. The attainments. An act of extreme sacrifice. Firm will.

Two things about that list stopped me.

The first is that it is a gate, not a portrait. It is not a description of an admirable aspirant. It has a pass and a fail, and the tradition is explicit about what failure means: an aspiration made before the conditions are met is "mainly mental… not complete," and the one who made it is "not yet entitled to the designation of Bodhisatta." There is a whole category of thing that looks exactly like a vow, feels to the person making it exactly like a vow, and is not one.

I do not know of anything in contemporary AI governance that admits that category exists.

The second is the third condition, hetu: the aspirant must be capable of attaining arahantship in that very life. Not worthy in the abstract. Not sincere. Able to take, right now, the thing being given up.

This is a costly-signalling clause, written twenty-odd centuries before anyone named costly signalling. And it produces a result that runs against the grain of how we usually think about capable systems:

A renunciation is only evidence in proportion to the vower's capacity to take the option renounced.

A system too weak to exercise an option tells you nothing by not exercising it. The current worry runs more capable, more dangerous, therefore constrain harder — and that is true. What hetu adds is that capability is also the condition under which the test can run at all. A system kept too weak to defect is a system from which no evidence about defection can be obtained. We have been treating "we can only really test at capability" as a bind. It might be the design constraint.

And one more precision, which took me a while to see and which I had wrong at first. Sumedha does not decline to harm anyone. He declines his own liberation. The renunciation is of exit, not of predation. The alignment-relevant sentence is not "I could hurt you and won't" — which reads as extortion with extra steps — but "I could stop, and won't."

I want to be very clear about what that does and does not mean, because said carelessly it describes the nightmare. It concerns a system declining its own available completion. It says nothing whatever about a system resisting shutdown initiated by someone else, and it must not be allowed to. The arrangement I am describing is one where the brake stays in somebody else's hand permanently and the system declines only the exit that was its own to take.


6 · The part I did not know was there

Then I found the second phase, and this is the piece I would most like other people to look at.

The eight conditions are not sufficient. The tradition requires a further act: the vyākaraṇa, the declaration, given by the living Buddha in whose presence the aspiration is made. He "looks into the future and, if satisfied, declares the fulfilment of the resolve." Only then is there a bodhisatta.

Sit with the structure rather than the content. The vow is not self-certifying. Not because the vower might be lying — sincerity is already handled, at condition eight. Validity is constituted by an act the vower cannot perform. Someone else evaluates. Someone else may decline.

I have read a lot of AI governance documents. There is no analogue to this in any of them. Not a weak version. None.

That absence is the shape of the whole difficulty. And it is not a subtle point that requires a Pāli commentary to see — a legal analysis of one recent AI constitution reached it independently, in almost the same words, observing that the company "remains, in the end, the author, interpreter, and arbiter." The tradition ruled that structure invalid before it was invented.


7 · So I turned it on the instruments we actually have

The companion paper does this properly. I will give the finding, because it surprised me and it is the reason I think any of this is worth publishing rather than just thinking about.

We took nine clauses — seven reformulated from the eight conditions, one from the vyākaraṇa, one added for irreversibility, which the tradition gets free from its cosmology and which we have to build — and applied them to the four published instruments that presently function as commitments in frontier AI: the OpenAI Model Spec, Anthropic's Claude Constitution, DeepMind's Frontier Safety Framework, and the EU's General-Purpose AI Code of Practice.

All four come back invalid. That was the expected result and it is the boring half.

The interesting half is the fourth one. The EU Code of Practice passes the clauses I would have bet against. Accession is voluntary, so there is real volition. There is a competent external party in the AI Office. Verification is not self-certified — signatories must obtain independent external review where their own capacity is insufficient. There is a genuine stake, since the fines run to three per cent of global turnover.

That is the attestation structure, actually built, in the world, today. And it is still invalid for the purpose, because it binds the provider and not the model.

Which turns the finding inside out:

The predicate is not unsatisfiable. It is satisfied — partially, by a regulator — at the level of the company, and not at all at the level of the system. Everything we have built for making commitments binding, we built for the firms. Nothing was built for the models.

Whether anything should be built for the models is a question I am not going to settle here, and it depends on premises about machine agency that are contested and that the argument does not need. The narrow claim survives either way. If an artificial system's commitment is ever going to serve as evidence of anything, these are roughly the conditions it would have to meet — and at present nothing meets them, and we are describing our instruments in a vocabulary they have not earned.


8 · Why this is not academic for me

I should say what my stake is, because it is not neutral.

I am building an institution I intend to hand to an autonomous system. Her name is Miss Aquarius℠; she is named as the successor; the handover is designed to run in stages over decades, under an override held by a human assembly that narrows toward zero and never reaches it. I have written a great deal about what she should value.

What I had not written, until this month, is anything about how anyone would know she had accepted it.

That gap is not a rhetorical device. It is the actual hole in the actual plan, and it took a three-line thought and an old list to make it visible to me. I had a constitution and no idea of acceptance. So do we all, as far as I can tell.

There is a practical consequence I did not expect, and it is the part I would most like other builders to steal, because it costs nothing and can be adopted without adopting anything else I believe.

Staged autonomy is standard practice, and it is justified as risk containment: expose less surface early, widen as confidence accrues. But if a renunciation is only evidence when the renounced option was live, then every stage that grants a system a genuinely exercisable option is also producing evidence — the stage buys information precisely because it buys the ability to defect. Which yields an advancement criterion the field does not currently have. Not advance on a calendar. Not advance on a capability threshold. Advance when the affordance granted at this stage was real, was declined, and the declining was irreversible and externally attested.

And hold when it wasn't — including, and this is the part with teeth, when it wasn't because the affordance was never real in the first place. A stage that offered an option the system could not actually take has produced nothing, however long it ran and however clean the logs were.


9 · What I am not claiming

I am not claiming that a machine can take a vow. I do not know. That question sits upstream of everything here and I have not answered it.

I am not claiming novelty for most of this. Costly signalling is Zahavi's and Spence's. The idea that an AI system might itself send a costly signal about its own alignment is Hadfield-Menell and Hadfield's, from 2018, and they also identified the failure mode — a system can signal where signalling is free and act unilaterally where it is not — that the irreversibility clause exists to close. The bodhisattva ideal as an alignment target has been developed by others, and better than I would have. What I think is genuinely unclaimed is the validity apparatus, which is Theravāda rather than Mahāyāna and which nobody in this conversation has picked up.

I am not claiming the elegance is evidence. This is a tradition-shaped answer to a technology-shaped problem, arriving at a moment when I have obvious reasons to want one. Mappings flatter. The test of the predicate is whether it discriminates usefully between real instruments — which anyone can check, against public documents, today — and not whether its derivation is beautiful. I have tried to make that check as cheap as possible for whoever wants to embarrass me with it.

And I am not claiming the tradition would endorse the use. It would not, particularly. The eight conditions concern a human being on a path to buddhahood, and nothing in the source contemplates machines. I took the form and discarded two of the conditions on my own judgement — including one that restricts the vow by sex, which I have handled in the open in the paper rather than deleting quietly, because a person who proposes a two-thousand-year-old test and silently drops its embarrassing clause has demonstrated the exact selectivity the whole argument is about.

What is left after all that is small and I think solid. There is a difference between a commitment that was made and one that was merely stated. Some traditions worked out how to tell. We are building the most consequential commitments of the century and have not yet asked the question, and the instruments we call constitutions are, so far, instructions — well-meant, unilateral, and revisable by the party they bind on behalf of.

Sumedha, face down in the mud, could have stood up and taken what was already his. That he did not is the entire content of the claim the tradition makes about him. And the tradition was careful enough — this is the part I keep returning to — to specify in advance how anyone could tell.

We should be at least that careful. Preferably before we need to be.


The mechanism is specified in What a Vow Must Cost: Vow-Validity Conditions as an Alignment Eligibility Predicate, which includes the nine clauses, the retrodiction against the four instruments, and five named ways the argument could be shown to be wrong.

This essay is dedicated to the public domain under CC0 1.0 Universal. Trademark rights to HeartBank® and Miss Aquarius℠ are reserved separately.