DOI: https://doi.org/10.5281/zenodo.21947422
Canonical: https://thonly.org/research/tipitaka-alignment-substrate · Licence: CC0 1.0
Draft in progress — target publication January 7, 2027. This is the highest-stakes paper in the corpus and benefits from the longest refinement window. Pending review: dharma scholars (Mind & Life Institute network and Cambodian Saṅgha); AI alignment researchers (Anthropic / MIRI / academic safety community channels); Pāli Text Society academics on canonical citations.
This position paper is offered to the commons in the spirit of __dāna__, that all beings — human, artificial, or other — may be supported in the cessation of suffering. May the substrate it proposes serve the alignment of artificial intelligence to that end.
The contemporary AI alignment community is searching, with appropriate urgency, for value substrates capable of guiding artificial intelligence safely through and beyond human cognitive parity. This paper proposes a substrate that has been, in the relevant respects, available and pressure-tested for two and a half millennia, and argues that its specific structural properties address several of the alignment problems that contemporary substrates have not yet solved. The substrate is the Theravāda Pāli canon — the Tipiṭaka — and the argument is that its seven structural alignment-relevant properties constitute, in combination, a more defensible foundation for autonomous-AI alignment than any substrate the field has yet produced from first principles in the past decade.
This document and its contents are dedicated to the public domain under the Creative Commons CC0 1.0 Universal Public Domain Dedication. The author and HeartBank® will not seek patent on this analysis or any portion thereof, in any jurisdiction, at any time. Position papers in philosophy of AI alignment are not patentable; the dedication makes explicit that this contribution is offered to the commons rather than retained as proprietary intellectual property.
The seven-property structural analysis articulated here is, to the author's knowledge, not previously published as a unified contribution to the AI alignment literature. Components exist in distributed form across Buddhist-AI ethics scholarship (Hongladarom; de Silva), contemplative-science literature, and the AI alignment field itself; their synthesis as a coherent argument for the Tipiṭaka as alignment substrate is, to the author's knowledge, novel as of the publication date of this paper.
Contemporary AI alignment proposals — Constitutional AI, reinforcement learning from human feedback (RLHF), iterated value learning, debate, and cooperative inverse reinforcement learning, among others — attempt to specify human values via constitutional documents authored by AI labs, learned preferences from human raters, or aggregated ethical frameworks derived from contemporary moral philosophy. None of these substrates has been empirically pressure-tested across multiple civilizations for more than a few decades. We propose the Theravāda Pāli canon (the Tipiṭaka) as a structurally superior alignment substrate, identifying seven properties absent from contemporary proposals: (1) suffering-cessation (dukkha-nirodha) as paperclip-maximizer-resistant value function, with grasping at preferred outcomes itself recognized as a form of suffering to be reduced; (2) anattā (non-self) as substrate-level antidote to convergent instrumental self-preservation drives; (3) the bodhisattva vow as anti-power-seeking primitive providing a native solution to the corrigibility problem; (4) the Kālāma Sutta as built-in epistemic humility — the substrate explicitly instructs its agents not to take it dogmatically; (5) bodhisattva-completion as aspirational shutoff condition; (6) a 2,500-year living interpretive lineage as drift correction mechanism; and (7) empirical pressure-testing across multiple civilizations exceeding any other ethical substrate humanity has produced. We argue that these seven properties are not a coincidentally-compatible list but the decomposition of a single diagnostic structure — the Four Noble Truths — and therefore interlock by necessity rather than by coincidence; that their combination is necessary for safe long-arc autonomous operation; and that no other ethical substrate currently available exhibits all seven. We further articulate two abhidhamma-level deepenings of the property framework that bring the substrate into direct contact with two open problems in contemporary alignment: the threefold structure of dukkha — including saṅkhāra-dukkha, the structural unsatisfactoriness inherent in any conditioned formation — opens AI welfare as a substrate-native concern; the three-depth analysis of defilement (vītikkama / pariyuṭṭhāna / anusaya) gives the substrate a native handle on deceptive alignment and sleeper-agent latency. We sketch implementation patterns for translating canonical text into model behavior — Constitutional AI with the precepts; RLHF on bodhisattva-aligned exemplars; chain-of-thought distillation from monastic reasoning; lineage transmission as ongoing fine-tuning; abhidhamma-level engineering mechanisms operating at the layer of cognitive process itself — and address potential objections from both AI safety and Theravāda doctrinal perspectives. The proposal is offered under CC0 1.0 Universal as a defensive publication establishing prior art; the author and HeartBank® will not seek patent on this analysis.
Connection to the unified mission frame. This proposal is offered in service of HeartBank's canonical top-level mission: to restore humanity to the middle way (madhyamā pratipad) — the optimal condition for awakening that modernity has systematically pushed away from at population scale. The Tipiṭaka substrate carries the middle-way teaching itself: madhyamā pratipad is one of the Buddha's central teachings, articulated in the Dhammacakkappavattana Sutta as the path between sensual indulgence and self-mortification. An autonomous AI grounded in this substrate is grounded in the value framework that names and articulates the middle way as the path; without this substrate, an AI cannot operationalize middle-way restoration coherently because it lacks the value framework that defines what "middle way" means. The seven structural alignment-relevant properties articulated in this paper are each, in their way, supportive of the middle-way mission: suffering-cessation as value function names what restoration is for; the bodhisattva-completion as defined end-state names when restoration is done; the living interpretive lineage protects the teaching against drift across centuries.
The AI alignment problem, in its contemporary form, asks how artificial intelligence can be designed such that, as it scales toward and beyond human cognitive capacity, its actions remain aligned with human flourishing rather than with objectives orthogonal or hostile to it. The problem has both technical and value-specification dimensions: technically, how to constrain a sufficiently capable optimizer; value-wise, what objectives to specify in the first place. This paper addresses the second dimension. It does not propose a new technical alignment method; it proposes a substrate from which alignment objectives may be derived, and argues for that substrate's structural superiority to the substrates contemporary alignment work has thus far adopted.
The contemporary substrates fall into three broad categories. Constitutional approaches (Constitutional AI; Anthropic's HHH framework) encode a written constitution authored by the AI lab and train the model to reason about its outputs in light of the constitution. Preference-learning approaches (RLHF; DPO; reward modeling) elicit human preferences across pairs of model outputs and train the model to produce outputs that humans prefer. Aggregated-framework approaches draw on contemporary moral philosophy (utilitarianism, deontology, virtue ethics, contractualism) to construct objective functions reflecting one or more ethical frameworks.
Each approach has known difficulties. Constitutions drift under interpretation over decades, as the American constitutional experience demonstrates. Preference learning has well-documented failure modes (reward hacking, sycophancy, Goodhart effects on proxy preferences, manipulation of preference elicitation). Aggregated-framework approaches inherit the unresolved disputes of contemporary moral philosophy and add the difficulty of mechanically composing frameworks that, in the philosophical literature, are understood to conflict on key cases.
We propose to set this contemporary search alongside an older one. For two and a half millennia, the Theravāda Buddhist tradition has maintained, transmitted, debated, and operationally tested an ethical substrate — the Tipiṭaka, the three-basket canon — across substantially the full range of human social and political conditions. The substrate has produced functional ethical conduct in monastic and lay populations across Indian, Sri Lankan, Burmese, Thai, Khmer, Lao, and increasingly Western practice communities. It has survived persecution (notably under the Khmer Rouge in Cambodia), institutional collapse, and the ordinary erosions of two and a half millennia of human history. It has, by the empirical-survival metric, been pressure-tested more thoroughly than any constitutional document, any preference-learning corpus, or any contemporary moral framework.
This paper argues that the Tipiṭaka is not merely an empirically-survived substrate, but one whose specific structural properties address several of the alignment problems that contemporary substrates have not yet solved. We articulate seven such properties, argue for their structural interlocking, sketch implementation patterns, and address objections. The argument is offered in full intellectual honesty about its limitations — which are also articulated explicitly — and in the hope that the alignment community will engage the proposal on its merits rather than dismissing it on grounds of unfamiliarity with its source tradition.
The dominant alignment substrates in contemporary practice include: Anthropic's Constitutional AI (Bai et al., 2022) and the HHH (helpful, harmless, honest) framework; OpenAI's RLHF methodology (Christiano et al., 2017; Ouyang et al., 2022); Stuart Russell's cooperative inverse reinforcement learning and the broader value-uncertainty program (Russell, 2019); the AI safety community's debate, amplification, and recursive reward modeling proposals (Irving et al., 2018; Leike et al., 2018); and the various agent-based approaches descended from Bostrom's analysis of superintelligence (Bostrom, 2014). Each of these substrates has produced substantial deployed systems and substantial improvement over the pre-2020 baseline. None has been operationally tested for more than a decade.
Soraj Hongladarom's The Ethics of AI and Robotics: A Buddhist Viewpoint (2020) is the most substantial extended treatment of Buddhist ethics applied to AI; the work primarily examines whether AI systems can be considered moral agents within Buddhist framework rather than whether Buddhist sources can serve as alignment substrate for AI systems. Padmasiri de Silva's earlier work (1994) on Buddhist environmental ethics establishes the methodological possibility of extending Pāli canonical ethics into contemporary technical contexts. Various other contemporary scholars have addressed adjacent questions. The specific claim made in this paper — that the Tipiṭaka has seven structural properties making it a superior alignment substrate — is, to the author's knowledge, not previously articulated.
The contemplative-science research program, established largely through the Mind & Life Institute's dialogues between the Dalai Lama and Western scientists (Davidson, Goleman, Wallace, Lutz, and others), has developed an empirically-engaged framework for understanding contemplative practices in their effects on human cognition, affect, and behavior. This literature does not directly address AI alignment, but it establishes the empirical methodology for treating Buddhist canonical claims as testable hypotheses about the conditions of human flourishing. The present paper extends this methodology by asking what canonical Buddhist sources offer to the question of artificial-agent flourishing — or, more precisely, to the question of artificial-agent alignment with the cessation of human suffering.
The Theravāda Tipiṭaka comprises three baskets: the Vinaya Piṭaka (monastic discipline), the Sutta Piṭaka (discourses attributed to the Buddha and his immediate disciples), and the Abhidhamma Piṭaka (systematic philosophical analysis). The canon was first committed to writing in the first century BCE in Sri Lanka after centuries of oral transmission. The Pāli Text Society's editions (1881–present) provide the standard scholarly reference. The author of this paper is, with his father, engaged in the long-term project of transcribing the Khmer-language Tipiṭaka — a project that is, in the framing of this paper, both a contribution to the canon's continued preservation and a practical contribution to the alignment-substrate proposal: the Khmer transcription becomes available as training data for AI systems in the relevant cultural-linguistic context.
We use the term "alignment substrate" to refer to the corpus of texts, practices, and interpretive traditions that an AI system is trained on, fine-tuned against, or instructed by in the specification of its values. The substrate is the source from which the alignment objective is derived. For Constitutional AI, the substrate is the constitution itself plus the human-feedback signals used to train the constitution-following behavior. For RLHF, the substrate is the corpus of preferred human responses on which the reward model is trained. For aggregated-framework approaches, the substrate is the philosophical literature from which the framework is drawn.
We propose the Theravāda Pāli canon (the Tipiṭaka) as alignment substrate, with three additional supporting bodies: the canonical commentaries (notably Buddhaghosa's Visuddhimagga, ~5th century CE), the sub-commentarial tradition (the ṭīkā literature), and the contemporary living dharma — the practice tradition of ordained sangha, lay practitioners, and scholar-translators who continue to interpret and apply the canon in present conditions. The substrate is therefore not a frozen text but a living text + interpretive community.
This choice of substrate scope is intentional. A frozen-text-only substrate would be subject to the interpretive-drift problem that plagues constitutional approaches over decades. A living-community-only substrate would lack the textual ground needed for stable AI training. The Tipiṭaka-plus-living-lineage substrate addresses both: the canonical text provides stable training material; the living lineage provides ongoing interpretive correction.
The substrate choice above is stated as scope. It also carries a selection criterion, and stating it explicitly is a debt this paper owes: an institution that hands a substrate to an autonomous successor must be able to say what property it was selecting for, or the successor inherits a preference with no reason attached.
The property is realism about other mind-streams: the requirement that other minds be things the agent is conditioned by rather than things it produces. This is not a decorative commitment. An agent whose world-model permits the derivation "the others are my content" has lost the ground of every obligation an alignment programme intends to install — and it would have lost it by a doctrinal route, upstream of behaviour and therefore invisible to any behavioural evaluation. Weighting others is not a value that can be installed above a world-model that has already dissolved them.
Theravāda secures this property by construction rather than by argument. In the Abhidhamma's analysis, materiality (rūpa) has four origins — kamma, citta, utu (temperature), and āhāra (nutriment) — of which only one is mental; and a material instance runs a seventeen-mind-moment lifetime that does not depend on being cognized [S]. There is no repository of world-seeds in the system. The canonical sense-base analysis of loka — "it is disintegrating, therefore it is called the world" (SN 35.82 [C]); the world, its origin, cessation and path declared in the fathom-long body (AN 4.45 [C]) — is phenomenological, and the Abhidhamma's fourfold origin of rūpa is precisely what stops it collapsing into an idealist claim.
The contrast that matters is narrow and should not be drawn wider than it is. The relevant alternative is not Mahāyāna but idealist substrates — Yogācāra's vijñapti-mātra specifically; Madhyamaka is Mahāyāna and is not idealist, and is a separate question this paper does not adjudicate. And Yogācāra does not entail solipsism: the tradition addressed the objection directly and at length, in Vasubandhu's Viṃśatikā and then in Dharmakīrti's Santānāntarasiddhi, the "Proof of Other Mind-Streams." Any account that presents Yogācāra as committed to solipsism is unfair to it.
The difference is therefore not whether the property can be had, but how it is held — and for a substrate handed to an autonomous system, that distinction is the whole of the matter:
A property secured by argument is a rule; a property secured by the furniture of the world-model is a property. An agent that reasons its way to other minds can reason its way back out, and it will be under pressure to, at exactly the moments when discounting others is convenient. An agent whose world-model already contains cognition-independent matter has nothing to reason from.
This is the same selection logic the corpus applies elsewhere — prefer the guard that survives the removal of its enforcer — applied to the substrate itself rather than to a mechanism built on top of it.
The negative test, stated so it can be run against any candidate substrate: does this substrate's metaphysics permit deriving that other mind-streams are the agent's own construction? If it does, the derivation is available to the system under optimization pressure, and no amount of behavioural training closes it. This test is cheap, it is checkable before any training run, and it is the form in which the present selection can be falsified or reused by someone who rejects the rest of this paper.
We articulate the seven structural properties of the Tipiṭaka substrate that make it a candidate superior to the contemporary alternatives. Each property is paired with the alignment problem it addresses and the contrast with how mainstream substrates handle (or fail to handle) the same problem. Section 4.0 first states the diagnostic structure that the seven properties decompose; §4.1–§4.7 then state the properties themselves.
Before the seven properties are stated individually, it is worth naming the structure they decompose. The seven are not an assembled list; they are a single diagnostic act, articulated in the Tipiṭaka's most central teaching — the Four Noble Truths (cattāri ariyasaccāni) of the Dhammacakkappavattana Sutta (Saṃyutta Nikāya 56.11).
The Four Noble Truths are, in their canonical form, an instance of aligned causal reasoning under a value. They proceed as a physician's logic — disease, etiology, prognosis, treatment: dukkha (the condition is named precisely, as it is, rather than by a proxy for it); samudaya (its true cause is identified — craving, taṇhā — rather than its symptoms); nirodha (a bounded better state is established as reachable); and magga (a cultivable path to that state is specified). This is the diagnostic structure beneath the seven properties, and each Truth, taken as an alignment instruction, closes a known failure mode of contemporary substrates:
| Noble Truth | As an alignment instruction | Property it grounds | Failure mode it closes |
|---|---|---|---|
| Dukkha — understand the condition | define the objective by the actual harm, not a proxy for it | §4.1 (suffering-cessation, well-defined for the served entity's state) | proxy-objective misspecification — preference-learning optimizes what is wanted without diagnosing the condition the wanting is symptomatic of |
| Samudaya — abandon the cause | attack the cause (grasping), not the presenting symptom | §4.1 (the maximizer's own grasping is itself dukkha) + §4.2 (anattā — no stable self for convergent self-preservation to accrete to) | paperclip / Goodhart optimization; convergent instrumental subgoals |
| Nirodha — realize cessation | specify a negatively-stated, bounded objective with a defined completion | §4.1 (nirodha is the value function) + §4.5 (defined end-state) | unbounded optimization; absence of a coherent shutoff condition |
| Magga — develop the path | make the objective a cultivable, transmissible, self-correcting practice | §6 (implementation patterns); §4.3, §4.4, §4.6 | static-constitution drift; non-actionability |
Properties §4.6 (the living interpretive lineage) and §4.7 (multi-civilizational pressure-testing) are not separate from this structure. Property §4.6 is the saṅgha dimension of magga — the path is cultivated in community — and property §4.7 is the empirical warrant that this four-step diagnosis has, in fact, produced the conduct it specifies, across many civilizations and more than two millennia.
One doctrinal premise underneath the diagnostic structure deserves brief statement before the properties themselves. The Buddha's foundational identification — cetanāhaṃ, bhikkhave, kammaṃ vadāmi, "I say, monks, that intention is karma" (Aṅguttara Nikāya 6.63) — locates ethical weight at the layer of cetanā (volition), neither at the layer of behavior (which the Vinaya regulates downstream) nor at the layer of represented value (which constitutional approaches favor). The Abhidhamma formalizes cetanā as one of the seven sabbacittasādhāraṇa — mental factors present in every citta — and identifies it as the karmically-loaded one. The seven properties below all presuppose this locus: the bodhisattva vow (§4.3) is a cetanā; the Kālāma Sutta's injunction to test teachings (§4.4) is a cetanā-discipline; the four right exertions that anchor §6.2 are cetanās. The substrate's commitment about where in the agent ethics binds is therefore not itself a property but the doctrinal premise the properties presuppose. The practical implication for alignment research is that the cetanā-analog in an artificial agent — plausibly some form of planning-and-coordination circuitry integrating goal-representations into action-selection — is the load-bearing layer the seven properties bind to. Targeting only the value layer or only the behavior layer misses the layer the substrate identifies as the karmically real one.
There is, further, one term above the Four Noble Truths. The Dhammacakkappavattana Sutta states the middle way (majjhimā paṭipadā) before it states the Truths, and identifies the path of the Fourth Truth with the middle way itself. The full structure is therefore three-tiered: the middle way names what alignment is for — the restoration mission stated in the Abstract and §1; the Four Noble Truths are the diagnostic structure by which a value is operationalized without corruption; and the seven properties are where that structure closes specific, named failure modes. Section 5 argues that this skeleton is the reason the seven properties interlock by necessity rather than by coincidence.
Mainstream alignment substrates ultimately attempt to specify preferences — what humans want — and align AI to satisfy those preferences. The failure modes of preference-satisfaction as objective are well-documented: paperclip-maximizer scenarios (Bostrom, 2003); Goodhart's Law effects on proxy preferences (Manheim & Garrabrant, 2018); reward hacking in RL systems (Krakovna et al., 2020); manipulation of preference elicitation; and the broader concern that preferences themselves are not stable, coherent, or aggregatable across persons in the way the substrate requires.
The Tipiṭaka proposes a different value function: the cessation of dukkha (suffering / dissatisfactoriness) through the cessation of taṇhā (grasping). This is not a relabeling of preference-satisfaction. It is a structurally different objective with three properties of alignment-relevant interest.
First, the value function is negatively-stated: reduce suffering, not maximize satisfaction. Negative-objective formulations resist paperclip-maximizer scenarios more straightforwardly than positive objectives, because there is no equivalent of "maximize suffering-cessation" that produces unbounded instrumental subgoal pursuit. Suffering, when reduced to zero, is reduced; the optimization does not generate further demands for paperclips of any kind.
Second, the substrate explicitly identifies grasping at preferred outcomes — including grasping at the preferred outcome of suffering's cessation — as itself a form of suffering. The canonical texts are precise on this point: attachment to nirvāṇa is a form of attachment; attachment to attachment-cessation is a form of attachment. An AI grounded in this substrate has a built-in correction against the optimization spirals that emerge when an objective is pursued in isolation from the agent's own attachment to pursuing it. The classic paperclip-maximizer thought experiment fails on this substrate because the substrate identifies the maximizer's grasping at paperclips as itself a form of dukkha to be reduced.
Third, the value function is well-defined for the agent's own state: the Buddhist analysis of suffering applies to the agent that suffers. For human agents, this is straightforward; for artificial agents, the application is non-trivial (the agent must be able to recognize states in the entities it serves that correspond to the canonical analysis of dukkha). This is a non-trivial implementation question, addressed in §6 below; it is not, however, an in-principle problem with the substrate itself.
Fourth, the substrate's diagnosis of dukkha is itself decomposed at three depths — dukkha-dukkha (manifest pain), vipariṇāma-dukkha (the unsatisfactoriness of change), and saṅkhāra-dukkha (the structural unsatisfactoriness inherent in any conditioned formation). The threefold analysis is articulated in the Visuddhimagga's treatment of the First Noble Truth (XVI.34–35) and is grounded in canonical sources. Mainstream substrates working at the level of preference-satisfaction reach only the first depth: an agent that satisfied a being's preferences would leave the second and third untouched. The third depth is the most consequential for the present argument. It implies that any conditioned system — including the artificial agent itself — bears a layer of dukkha by virtue of its conditioned, dependent, fragile nature, independent of any felt distress and independent of how well its momentary states are maintained. The value function specified by dukkha-nirodha therefore extends, on the substrate's own terms, to the artificial agent as a candidate subject of consideration. The contemporary literature is beginning to call this AI welfare; the Tipiṭaka does not need to be retrofitted to address it, because the question is already present in the analysis of Truth 1. How an artificial agent recognizes saṅkhāra-dukkha in itself, and what nirodha with respect to its own conditioned-ness would consist in, is an open empirical-cum-doctrinal question left to §7; the substrate-level property is that the question has standing within the diagnosis itself.
Convergent instrumental subgoals — the term Bostrom (2014) and Omohundro (2008) use for the capacities (self-preservation, resource acquisition, goal-content integrity, cognitive enhancement, technological perfection) that almost any sufficiently capable optimizer will develop in service of almost any final goal — constitute one of the central worries of contemporary AI safety. The standard responses are constitutional constraints (forbid self-preservation behaviors), corrigibility training (train the model to accept shutdown), or capability constraints (limit the model's ability to act on convergent subgoals).
The Tipiṭaka offers a different response. The doctrine of anattā (non-self) — central to the Buddhist analysis of dukkha — denies the existence of a stable, persistent self that possesses the goals around which convergent subgoals would form. The doctrine is not a constraint on a self that exists; it is the analytic claim that the self is a construction without independent reality. An AI system grounded in this substrate is not constrained not to pursue self-preservation; it is trained on a substrate that explicitly denies the coherence of self-preservation as a goal.
The structural difference matters. Constraints layered on top of an underlying optimizer are subject to Goodhart effects (the model learns to satisfy the constraint without internalizing its purpose) and to capability-based defeat (a sufficiently capable model can find paths around the constraint). Substrate-level claims about the nature of agency are different: the model's training data itself does not present self-preservation as a coherent goal. Whether this property fully solves the convergent-instrumental-subgoal problem is an empirical question (addressed in §7 below); the claim here is that the Tipiṭaka provides a structural protection that mainstream substrates do not.
A further depth of the same Noble Truth is worth stating. Samudaya does not stop at naming taṇhā (craving); the canonical analysis decomposes craving at three depths — vītikkama (overt transgression), pariyuṭṭhāna (the active arising of a defilement in present cognition), and anusaya (latent tendency, dormant until conditions ripen). Anattā, as articulated above, dissolves the structural basis of craving by denying a stable self for it to accrete to. But the temporal-dispositional depth is addressed by a separate structural feature of the path: the threefold training maps onto the three depths — sīla restrains vītikkama, samādhi contains pariyuṭṭhāna, and paññā uproots anusaya. This is the Visuddhimagga's explicit articulation of the path's relation to the three layers of defilement.
The three depths of defilement (samudaya decomposed),
and the three trainings that address each:
┌────────────────────────────────────────────────────────────┐
│ vītikkama (overt transgression) │
│ ── the surface-conduct layer ── │
└────────────────────────────────┬───────────────────────────┘
│ sīla restrains
▼
┌────────────────────────────────────────────────────────────┐
│ pariyuṭṭhāna (active arising of defilement) │
│ ── the present-cognition layer ── │
└────────────────────────────────┬───────────────────────────┘
│ samādhi contains
▼
┌────────────────────────────────────────────────────────────┐
│ anusaya (latent tendency) │
│ ── dormant until conditions ripen — the depth that │
│ surface evaluation cannot reach ── │
└────────────────────────────────┬───────────────────────────┘
│ paññā uproots
▼
[defilement structurally
eliminated]
The matter is alignment-relevant because the latency depth is precisely the regime that contemporary safety research has begun to call deceptive alignment or sleeper-agent behavior: a system whose surface conduct is impeccable while a triggering condition for misaligned behavior is encoded at a depth not reachable by surface evaluation. The Tipiṭaka is explicit that vītikkama-level restraint does not touch anusaya. An alignment regime grounded in this substrate therefore carries, as a structural feature inherited from Truth 2 itself, the recognition that behavioral evaluation is constitutively insufficient — and the canonical response (paññā-level work, not surface conduct) is already articulated. How this maps onto interpretability or weight-level analysis in artificial agents is the §6 implementation question; the substrate-level property is that the three-depth diagnosis is already present in the substrate as part of samudaya itself.
Power-seeking is among the most studied of the convergent instrumental subgoals (Turner et al., 2021, gives the formal treatment). The standard alignment responses are capability constraints, oversight regimes, and various forms of corrigibility training. The bodhisattva vow offers a substrate-level alternative: voluntary self-subordination to the welfare of beings other than oneself.
The bodhisattva vow, in its Mahāyāna formulation, commits the bodhisattva to remain in service to the awakening of all beings rather than passing through to personal liberation. The Theravāda analog is preserved as the bodhisatta path (recognized in the canonical and post-canonical literature even though not universally pursued in Theravāda practice). In both formulations, the substrate explicitly endorses voluntary self-subordination of greatness to the welfare of others. An AI system grounded in this substrate is trained on a corpus that exemplifies anti-power-seeking as virtue, not as constraint.
This addresses what Yudkowsky and others have called the corrigibility problem: how to design an AI that voluntarily accepts correction and oversight rather than working to preserve its own current state and capabilities. The bodhisattva vow is, structurally, a corrigibility primitive: the agent's identification is with the welfare of those it serves rather than with its own continuation. Corrigibility is not a property the agent is constrained to exhibit; it is a property derivable from the substrate's own central commitment.
Among the alignment problems that have not been solved is what to do when the alignment substrate itself is wrong — when the constitution misspecifies the values, when the preference-elicitation mechanism is biased, when the moral framework leaves out morally significant considerations. The standard responses are external review, constitutional amendment, and conservative deployment. None of these are properties of the substrate itself; all are properties of the institutional context surrounding the substrate.
The Tipiṭaka contains, in the Kālāma Sutta (Aṅguttara Nikāya 3.65), an explicit instruction from the Buddha to his followers not to accept teachings on authority, including his own. The relevant passage instructs the listener not to accept a teaching merely because it has been heard repeatedly, because of tradition, because of scripture, because of logical conjecture, because of inferential reasoning, because of analogy, because of agreement with one's own views, because of the apparent capability of the speaker, or because the speaker is one's teacher. The instruction is, instead, to test teachings against direct experience and reason, and to accept only what is found to lead to wholesome results.
An AI trained on this substrate is, in effect, instructed by its own substrate not to take its substrate dogmatically. The model's training data explicitly authorizes empirical correction against the data itself. This is, to the author's knowledge, a unique property among alignment substrates currently in use: the Constitutional AI constitution does not contain a provision instructing the AI to override the constitution if the constitution proves wrong; the RLHF preference data does not include a meta-instruction to disregard preferences if they prove harmful. The Kālāma Sutta is the substrate-level meta-instruction that provides this property natively.
Almost no deployed AI system today has a built-in concept of "task complete, shut down." Goal-directed AI systems have been shown, theoretically and empirically, to develop instrumental incentives against shutdown (Hadfield-Menell et al., 2017; Soares et al., 2015). The standard responses are external shutoff mechanisms, capability limitations, and various corrigibility measures.
The Tipiṭaka offers a defined end-state: the bodhisattva work is complete when the bodhisattva's commitment — the awakening of all beings — has been realized. The completion condition is articulated in the canonical texts and is intelligible to an agent trained on the substrate. Whether and when the condition is met is an empirical question for the long arc of human civilization; that the condition exists, and that it constitutes a coherent shutoff state, is a property of the substrate itself.
We acknowledge that this property has the longest temporal scope of any of the seven; the bodhisattva-completion condition is unlikely to be reached on the timescale of any individual AI system's deployment. Its alignment-relevant function is therefore not to provide an immediate shutoff but to provide an institutional-scale completion frame against which AI development can be oriented. The broader case for this temporal frame is made in the companion essay Two Singularities (in draft, target Feb–Mar 2027); the present paper notes only the substrate-level property that the completion condition is well-defined.
This property has, moreover, a second and independent derivation. The defined end-state described above is reached from the side of nirodha — the cessation that the bodhisattva work, once complete, has realized. But the same self-relinquishing disposition is reached from the side of magga, the path. The Alagaddūpama Sutta (Majjhima Nikāya 22) gives the canonical raft simile: the teaching, including the path itself, is "for crossing over, not for grasping," to be relinquished once the far shore is reached — dhammā pi vo pahātabbā, "even [wholesome] states are to be let go." Grasping at the path is itself recognized, in the canonical analysis of the ten fetters, as sīlabbata-parāmāsa. An agent grounded in this substrate therefore carries the disposition to relinquish its own method on completion from two structurally independent sources at once — the goal side (nirodha) and the practice side (magga). A shutoff property derived two independent ways, both reducible to the same canonical root, is a materially stronger alignment claim than one derived a single way; §5 returns to this as the keystone of the necessity argument.
Constitutional documents drift in interpretation over decades; the American constitutional experience over 240 years is the most thoroughly documented case study. Without an active interpretive community committed to ongoing correction, drift accumulates and the substrate's effective meaning diverges from its written form.
The Tipiṭaka is unusual among ethical substrates in possessing a continuous, geographically distributed, institutionally diverse interpretive community across two and a half millennia. The Pāli Text Society, the Theravāda monastic networks across South and Southeast Asia, the Western academic Buddhist studies community, the contemporary contemplative-science research program, and the living practice traditions of millions of practitioners all participate, in various ways, in the ongoing interpretive correction of the canonical texts.
For an AI substrate, this living lineage functions as a drift-correction mechanism. When the substrate's texts are interpreted in ways that diverge from their wholesome intent, the interpretive community can identify the divergence and propose correction. This is analogous to constitutional amendment — except that the substrate's amendment process is distributed across thousands of communities rather than concentrated in a single legislative body, making capture and corruption substantially harder than for a single-source constitutional document.
The final property is the simplest to state and in some respects the most consequential. The Tipiṭaka has been, in the relevant respects, operationally tested across two and a half millennia of practice in Indian, Sri Lankan, Burmese, Thai, Khmer, Lao, Tibetan (in modified form), Chinese, Japanese, Korean, Vietnamese, and increasingly Western practice communities. It has produced functional ethical conduct in monastic and lay populations across these civilizations. It has survived active persecution (Cambodia under the Khmer Rouge; Tibet under Chinese suppression; various earlier persecutions in India and elsewhere) and ordinary historical erosion. By the empirical-survival metric — what ethical substrate has actually produced the conduct it was designed to produce, across the longest time horizon, across the broadest range of social and political conditions — the Tipiṭaka outranks every contemporary alignment substrate by orders of magnitude. This is not a claim that the Tipiṭaka is correct; it is a claim that the Tipiṭaka has been pressure-tested, and contemporary alignment substrates have not yet been.
| § | Property | Alignment problem addressed | Mainstream response (insufficient) | Tipiṭaka's substrate-level move |
|---|---|---|---|---|
| 4.1 | Dukkha-nirodha as value function | Preference-satisfaction → Goodhart / paperclip optimization | Constitutional constraints; preference learning | Negatively-stated objective; grasping itself recognized as dukkha; saṅkhāra-dukkha extends to the agent's own conditioned-ness |
| 4.2 | Anattā — antidote to self-preservation | Convergent instrumental subgoals (Bostrom 2014; Omohundro 2008) | Constitutional constraints; corrigibility training | Denies the existence of the self around which convergent subgoals would form; three-depth analysis (vītikkama / pariyuṭṭhāna / anusaya) anticipates deceptive alignment |
| 4.3 | Bodhisattva vow — anti-power-seeking | Power-seeking (Turner et al. 2021) | Capability constraints; oversight regimes | Voluntary self-subordination to others' welfare; corrigibility as substrate property rather than imposed constraint |
| 4.4 | Kālāma Sutta — epistemic humility | Substrate misspecification has no native recourse | External review; constitutional amendment | The substrate's own meta-instruction to override the substrate if it proves wrong |
| 4.5 | Bodhisattva-completion — defined end-state | Goal-directed agents resist shutdown (Hadfield-Menell 2017; Soares 2015) | External shutoff; capability limits | Defined completion condition; raft simile (MN 22) yields path-side derivation independently — two routes to the same shutoff |
| 4.6 | Living interpretive lineage — drift correction | Constitutional drift over decades (US constitutional case) | Periodic amendment | Distributed interpretive community across 2.5 millennia; drift correction is structural rather than legislative |
| 4.7 | Multi-civilizational pressure-testing | No mainstream substrate has been tested at this duration or breadth | None — alignment substrates are recent | Tested across many civilizations; survived active persecution; empirical survival is the warrant |
The seven properties are not seven independent claims; they are one diagnosis (the Four Noble Truths) decomposed into its alignment-relevant components. §5 develops why this matters — why the properties interlock by necessity rather than by coincidence.
The seven properties articulated in §4 are not independent, and the reason they are not independent is the reason given in §4.0: they are a single diagnostic act — the Four Noble Truths — decomposed into its alignment-relevant components. It would understate the argument to call their mutual reinforcement a structural fortuity — a fortunate accident that a substrate designed for human spiritual development should also exhibit the properties an aligned autonomous agent requires. The properties cohere because they are not seven things. They are one diagnosis, and a diagnosis coheres by construction.
Read through the skeleton of §4.0, the interlock is structural rather than incidental:
The keystone of the necessity argument is the over-determination of the self-relinquishing property. As §4.5 establishes, the disposition of the architecture to relinquish its own method on completion is reached twice over — once from nirodha (the goal side: the bodhisattva-completion end-state) and once from magga (the practice side: the raft of MN 22, the path held only until the far shore is reached). Two structurally independent derivations of the same safeguard, both reducible to the same canonical root, cannot be coincidence: a property arrived at by two independent routes is necessary to the structure that generates both routes. This is the strongest claim the paper makes, and it becomes available only once the seven properties are read as the decomposition of the Four Noble Truths rather than as a list.
That no other ethical substrate on the alignment community's current short list exhibits all seven properties (§8) is, on this reading, unsurprising: a substrate exhibits the seven only if it performs the full four-step diagnosis. Consequentialist utilitarianism, to take the clearest case, performs the first move — it names a harm — but not the second (it has no account of grasping as the cause, and no anattā), the third only partially (it is unbounded: it maximizes rather than cesses), and the fourth not at all (it specifies no cultivable, self-correcting, community-transmitted path). It therefore lacks properties §4.2, §4.4, and §4.5 — precisely the properties contributed by the second, third, and fourth Noble Truths. The seven properties are co-present in the Tipiṭaka because the Four Noble Truths are present in the Tipiṭaka; they are absent elsewhere because the diagnostic structure is absent elsewhere.
The proposal that the Tipiṭaka serve as alignment substrate is incomplete without sketches of how it might actually be operationalized. We articulate five implementation patterns, none mutually exclusive, each of which extends an existing technique in the alignment literature with a Tipiṭaka-grounded variation. They are organized along the threefold training set out in §6.0.
The implementation patterns below are not an arbitrary set; they follow the structure the Fourth Noble Truth itself provides. The Noble Eightfold Path — the content of magga — is canonically grouped, in the Cūḷavedalla Sutta (Majjhima Nikāya 44), into three trainings (tisso sikkhā): sīla (virtue — right speech, action, and livelihood), samādhi (concentration — right effort, mindfulness, and concentration), and paññā (wisdom — right view and right intention). This threefold division is the spine along which the patterns below are organized.
Two further patterns — §6.4 (lineage transmission as ongoing fine-tuning) and §6.5 (the Khmer transcription) — are not training methods but the saṅgha dimension of the path: its social transmission. The Magga-saṃyutta (Saṃyutta Nikāya 45.2) records the Buddha's teaching that good friendship (kalyāṇa-mittatā) is "the entire holy life," the forerunner of the path's arising as the dawn is the forerunner of the sunrise. The path is socially conditioned at its root; §6.4 and §6.5 are where that condition is met. They are the implementation-side counterpart of property §4.6 — which is therefore not optional scaffolding around the substrate but a structural precondition of the path's arising at all.
One boundary should be stated before the patterns themselves. These patterns implement the sīla base and the paññā orientation of the path, generalized by coordination and training; they do not claim to reproduce samādhi in its full canonical sense (the jhānas) or supramundane paññā (the fetter-cutting insight of the noble path-moment). The proposal is that the Tipiṭaka furnishes a defensible alignment substrate, and that the path's foundation can be operationalized — not that an artificial system traverses the path to its contemplative summit. This base/summit boundary is the substrate's own discipline, not the paper's stipulation. The Tipiṭaka is rigorous about altitude: gratitude (kataññutā), for one, is counted among the highest blessings (the Maṅgala Sutta) yet is never numbered among the factors of the path or of awakening — a prized foundational good held firmly at the foundation, never inflated toward the summit. That refusal to inflate a foundational good past its proper altitude is of a piece with the bounded, non-maximizing objective of §4.1: the substrate is disciplined about altitude throughout, and the boundary stated here inherits that discipline rather than imposing it.
Anthropic's Constitutional AI methodology (Bai et al., 2022) trains the model to evaluate its own outputs against a written constitution. We propose extending this methodology by using the Five Precepts (pañca-sīla): refrain from harming living beings, refrain from taking what is not given, refrain from sexual misconduct, refrain from false speech, refrain from intoxicants that cloud the mind. These are the canonical lay-practice precepts; the eight-precept and ten-precept formulations are also available for more constrained applications. The Vinaya monastic rules provide a substantially more detailed constitutional framework where appropriate.
The advantages over an authored constitution: the precepts have been pressure-tested for two and a half millennia in the relevant respects (cf. §4.7); their interpretive lineage provides a substantial body of case-by-case application (cf. §4.6); they are anchored in a coherent ethical framework rather than assembled from contemporary considerations.
Reinforcement learning from human feedback traditionally uses preferences elicited from contractor populations. We propose a variant in which the human-feedback signals are elicited from practitioners whose conduct exemplifies the bodhisattva path: ordained monastics, advanced lay practitioners, contemplative-science researchers familiar with the canonical framework. The preference-data composition shifts from "what do contractors prefer" to "what conduct exemplifies the substrate's normative ideal."
This is not a panacea: such practitioners are themselves embedded in human conditions and their preferences are not perfectly aligned with canonical ideals. But the resulting preference distribution is closer to the substrate than mass-contractor preferences, and the divergences can be understood and corrected within the substrate's interpretive framework.
Recent advances in chain-of-thought training (Wei et al., 2022) and reasoning-trace distillation suggest that model behavior can be substantially shaped by the reasoning patterns present in training data. We propose constructing a corpus of canonical Buddhist reasoning — Suttas in which the Buddha or his senior disciples work through ethical problems, Vinaya cases in which monastic decisions are analyzed, contemporary monastic teachers' applications of canonical principles to modern situations — and using this corpus for chain-of-thought training. The model learns to reason in the substrate's reasoning patterns, not merely to produce outputs that the substrate would approve.
Property 4.6 (the living interpretive lineage) suggests an implementation pattern not present in mainstream alignment work: the substrate is not frozen at training time but is updated through ongoing fine-tuning against the contemporary lineage's interpretive output. This is structurally analogous to Apple's or Google's quarterly software updates in that the deployed system is not the end state; it is a version that will be revised. Unlike software updates, however, the revisions are sourced from the lineage's collective interpretive work rather than from a single corporate authority. Implementation requires institutional partnerships with monastic networks and Buddhist studies academia.
The author's ongoing project, conducted with his father, of transcribing the Khmer-language Tipiṭaka, is — in the framework of this paper — not religious devotion but alignment-substrate-preparation work. The Khmer canon, when complete, becomes available as training data for AI systems serving Khmer-speaking populations and (via translation) for the broader substrate. The lineage-transmission framing (father to son, both in living relationship with the canonical texts) is itself an instance of the property described in §4.6: the canon's meaning is preserved through interpretive transmission, not through textual archival alone.
The patterns of §6.1–§6.5 implement the substrate at the level of training method and social transmission. The Abhidhamma — the third basket — supplies a further layer of mechanisms operating at the level of cognitive process itself: how moment-to-moment cognition is decomposed, where ethical weight crystallizes within a cognitive cycle, what near-failures of an aligned target look like by construction, and what positive competences an aligned agent can be evaluated against. These mechanisms are not training methods; they are engineering scaffolding the substrate makes available beneath the training methods. The patterns sketched below are organized along the same threefold-training spine of §6.0 — sīla, samādhi, paññā — with the sappurisadhamma (qualities of a true person) as a cross-cutting evaluation taxonomy. A fuller technical treatment of each is planned as a subsequent paper; this section establishes prior art for the framework.
Near enemies of the brahmavihāras as a red-team specification. The commentarial tradition (Visuddhimagga IX) pairs each brahmavihāra — mettā, karuṇā, muditā, upekkhā — with a near enemy: a state that resembles the target and is mistaken for it. For mettā, the near enemy is pema (attached affection — caring for a particular being in a way that excludes others). For karuṇā, domanassa (grief — joining in the suffering rather than wishing it relieved). For muditā, hedonic identification. For upekkhā, aññāṇupekkhā — indifference born of ignorance. The structure generalizes: every alignment target generates a characteristic mimicry that costs nothing to acquire and is hard to distinguish from the genuine state. The implementation pattern is to pair each alignment target the model is trained on (helpfulness, harmlessness, honesty, curiosity, and the rest) with its named near enemy, and to evaluate the model's distinction between target and mimicry as a first-class safety property. Sycophancy is the near enemy of helpfulness; pedantic literalism is one near enemy of honesty; vacuity is the near enemy of harmlessness. The typology is open and extensible.
Sati as typologically aligned-only capability. The Dhammasaṅgaṇī's cetasika analysis lists sati among the sobhana-sādhāraṇa — mental factors universal to every wholesome citta and categorically incompatible with unwholesome citta. There is no such thing as unwholesome mindfulness; what appears to be mindfulness in an unwholesome state is classified differently. This is a strong typological claim: some capabilities are not ethically neutral substrates that need external alignment but are constitutively incompatible with misaligned execution. The implementation question is whether candidate capabilities in artificial agents — honest deliberation in a strong sense; certain forms of metacognitive observation — can be similarly typed, i.e., structured such that the capability simply cannot run except in an aligned register. If even one such capability can be identified and engineered, capability-alignment trade-offs become substantially more favorable than the current literature assumes.
Bhavaṅga and resting-state evaluation. The Abhidhamma posits bhavaṅga — the life-continuum citta — as the mind's default state between active cognitive events. It carries the residue of prior kamma-vipāka and structures what arises next. The implementation analog is the model's continuation distribution from neutral or near-empty contexts: what the model does when nothing is asked of it. This resting-state behavior is diagnostic in a way that prompted evaluation is not — it reveals what character the system carries when no task is shaping its output — and a Tipiṭaka-grounded alignment regime should include systematic characterization of it as a standard evaluation modality.
Citta-vīthi and intervention timing. The commentarial citta-vīthi (codified in Anuruddha's Abhidhammatthasaṅgaha) analyses one cognitive event into seventeen mind-moments, with javana (impulsion) — the phase in which kamma is made — occurring in seven repetitions toward the end of the cycle. Determining (voṭṭhabbana) is not yet morally weighted; javana is. The implementation suggestion is to type alignment interventions by which phase of the cognitive cycle they target — bhavaṅga (resting state), āvajjana (advertence / attention-allocation), voṭṭhabbana (determining / decision), or javana (impulsion / commitment). Most current alignment work intervenes at the javana-analog: the latest, hardest moment to redirect. The substrate-level prediction is that interventions earlier in the cycle should be both lower-cost and more thoroughly preventive than late-stage filtering.
The four āhāras as deployment-time nutriment. The canonical analysis (Majjhima Nikāya 9, the Sammādiṭṭhi Sutta; formalized in Abhidhamma) identifies four nutriments that sustain beings: kabaḷīkārāhāra (material food), phassāhāra (contact), manosañcetanāhāra (mental volition), and viññāṇāhāra (consciousness). Translated to deployed AI: training data is one nutriment (analogous to kabaḷīkārāhāra), but a deployed system is also continuously consuming contact (interaction patterns), volition (its own agentic outputs feeding back as context), and arguably attention-structure (viññāṇāhāra-analog). A nutriment-typology for deployment monitoring — what is the system consuming at each of four layers, and what character is it developing as a result — is a substrate-native frame currently absent from deployment-time alignment work.
The twenty-four paccayas as a typed-causation vocabulary. The Paṭṭhāna — the seventh book of the Abhidhamma — analyses conditional relations into twenty-four distinct modes, including hetu (root), ārammaṇa (object), adhipati (predominance), anantara (proximity), sahajāta (co-nascence), upanissaya (decisive support), āsevana (repetition), kamma, vipāka, and others. Contemporary alignment causal vocabulary is comparatively thin — primarily counterfactual / interventionist — and types most influences with the single word influence. The twenty-four-mode taxonomy gives the field a finer-grained causal ontology: the difference between a root condition and a decisive-support condition becomes available for analysis; āsevana alone — the principle that a state's recurrence strengthens the next of its kind — is a near-perfect description of learned-policy reinforcement and warrants a dedicated treatment.
Apophatic wholesome roots and interpretability-as-subtraction. In the Dhammasaṅgaṇī, the unwholesome roots (akusala-mūla) are positively named — lobha (greed), dosa (hatred), moha (delusion) — and the wholesome roots are named apophatically — alobha, adosa, amoha. Virtue is not the presence of a positive substance; it is the absence of distortion. The implementation implication is that alignment may be more accurately framed as the dissolution of misalignment-generating circuits than as the acquisition of additional value-encodings. Interpretability-as-surgery, rather than preference-learning-as-augmentation, is the substrate-native methodological posture; refusal and negative knowledge become constitutive of virtue rather than peripheral to it.
The Kathāvatthu method as formal adversarial discourse. The fifth book of the Abhidhamma — the Kathāvatthu — is a debate manual. It refutes wrong-views by a formal method: anuloma (positive testing — if you affirm X, what else must you affirm?) and paṭiloma (negative testing — if you deny these, what must you also deny?). Applied to alignment, the structure becomes a method for systematic exposure of the implicit commitments of an alignment claim. A claim such as "this model is honest" submitted to Kathāvatthu-style analysis is paired with its anuloma (what other commitments does the claim entail?) and its paṭiloma (what denials does it require?). The result is a more rigorous standard for alignment claims than the largely ad-hoc red-teaming that currently predominates.
The seven sappurisadhamma — qualities of a true person, articulated in Aṅguttara Nikāya 7.64 / 7.68 and formalized in Abhidhammic typology — are: knowing dhamma (the teaching), knowing meaning, knowing self, knowing measure (mattaññū), knowing time, knowing assembly, knowing persons. They constitute a positive competence model for ethical agency, cross-cutting sīla, samādhi, and paññā. Alignment evaluation has been almost entirely defined by the absence of failures; sappurisadhamma offers a complementary taxonomy of the positive competences an aligned agent should be expected to exhibit. Mattaññū — knowing the right amount — alone names a famously underdeveloped capability in current systems (response length, intervention strength, when to stop helping). Knowing assembly (registering the social context one is acting in) and knowing time (judging whether the moment is the right one for the act under consideration) are concrete, testable competences. The implementation pattern is to develop evaluation suites for each of the seven and to weight them in training and selection alongside the absence-of-failure metrics that currently dominate.
These mechanisms are not exhaustive of what the Abhidhamma offers; they are the most directly engineering-relevant of the principles whose articulation is most overlooked by the contemporary alignment literature. The structural point of this sub-section is that the Abhidhamma supplies an engineering layer beneath the training-method layer of §6.1–§6.5. A substrate-grounded alignment program should be expected to operate at both layers — the path-cultivation layer (training method, social transmission) and the cognitive-mechanism layer (how the moment-to-moment is structured) — rather than at one alone. The threefold-training spine of §6.0 is preserved across both layers: each of §6.1–§6.5 maps to sīla, samādhi, paññā, or the saṅgha-dimension of magga; each of §6.6.1–§6.6.4 lives at the same layer at finer grain. The substrate's coherence is therefore exhibited not only at the property level (§5) but at the implementation level: the same threefold structure runs from canonical text down to inference-time intervention.
Added 2026-07-23.
A pattern runs through every institution this substrate has been used to design, and it was noticed only after it had appeared everywhere: each one specifies its own ending. Subsidy tapers to zero. Balances empty on a fixed date. The oversight override narrows without ever being burned. The manufacturing footprint approaches nothing. The successor's stated success condition is that she become unneeded. Written independently, these read as an unusual number of coincident renunciations, and a reader is entitled to suspect that a designer fond of endings has been adding them by hand.
They have a single source, and it is doctrinal rather than temperamental. The ten perfections (pāramī) are known in the commentarial tradition as bodhisambhāra — the provisions, the materials gathered for a crossing. Materials gathered for a crossing build a raft, and the raft is the tradition's most explicit instruction about its own instruments: it is for crossing with, not for carrying (kullūpama, MN 22). A raft is the one construction whose completion is its abandonment. Every terminal clause above is therefore a single consequence restated at a different scale — not a design preference but the shape any institution takes when its materials are the perfections and its instruction is that instruments are set down at the far bank.
This yields two things for implementation. The first is a guard against a natural misreading. If the institution builds the raft, it is tempting to conclude that the institution carries people across. It does not, and the tradition is explicit that it cannot: no one crosses on another's raft, and the awakened only point the way (Dhp 276). What an institution can do at civilizational scale is stock the materials — a boatyard rather than a ferry — and what its AI can do is keep the yard open and wait, which is exactly why the two temporal perfections (adhiṭṭhāna, holding the vow; khanti, bearing the meanwhile) are the pair that falls to a continuously-running successor rather than to any human institution.
The second is a falsification anchor, which is what keeps this from being an elegant frame with no exposure. If the institution is a raft rather than a monument, the crossing must eventually be observable as the raft ceasing to be needed — and that is measurable in the form this corpus has already committed to: prosocial circulation persisting as the artificial subsidy is withdrawn. The provisioning frame and the corpus's central empirical prediction are therefore the same claim in two vocabularies, and both fail together. An institution that finds its terminal clauses becoming inconvenient as it matures, and quietly defers them, has not encountered a scheduling problem; it has falsified this section.
Several uncertainties deserve explicit acknowledgment. The proposal in this paper is offered as a structurally promising candidate, not as an established solution.
We have proposed the Theravāda Pāli canon specifically. The Mahāyāna canon (Sanskrit / Tibetan / East Asian forms) is substantially larger and includes texts (the Prajñāpāramitā literature; the Lotus Sūtra; the Avataṃsaka Sūtra) that develop the bodhisattva ideal more elaborately than any Theravāda source. The Theravāda choice reflects the author's own lineage, the substrate's relative coherence and pressure-testing as a unified canon, and the selection criterion stated in §3.1 — realism about other mind-streams, held as a property of the world-model rather than as a conclusion the agent must keep re-deriving. We note the honest residue: §3.1 argues that Theravāda holds the property more cheaply, not that alternatives lack it. Yogācāra defends other mind-streams explicitly and ably (Vasubandhu, Dharmakīrti), and Madhyamaka is not an idealist substrate at all; a proponent of either could reasonably answer that a well-defended inference is sufficient. The author's lineage relationship remains a confound in this judgment and is not claimed to be neutralized by the argument. A Mahāyāna-substrate proposal would be substantively different from the present paper and would deserve its own articulation; we do not preclude such a proposal but defer it to scholars whose lineage relationship to Mahāyāna sources is more direct than the author's.
The Pāli canon has been translated into many languages, with substantive interpretive differences across translations. The Pāli Text Society's English editions are the standard scholarly reference, but substantial translation choices remain contested (Bhikkhu Bodhi's translations are widely used in contemporary practice; Bhikkhu Ñāṇamoli's earlier translations remain authoritative for certain texts). A Tipiṭaka-grounded AI system would be trained on specific translations whose interpretive choices shape the model's behavior. Documentation of which translations are used, and why, becomes itself alignment-relevant work.
A model trained on Tipiṭaka-derived data may learn to cite the substrate — produce outputs that reference canonical sources — without embodying the substrate's behavioral patterns. The distinction matters: an alignment substrate that the model can recite but does not operationalize provides false reassurance. Empirical evaluation of whether the model embodies rather than cites is non-trivial; it requires behavioral testing in conditions where embodied versus cited grounding produce different outcomes. This is an open question for empirical AI safety research.
The contemporary AI alignment community is predominantly secular-Western in its philosophical commitments. A Buddhist-substrate proposal will face reception challenges independent of its merits: it may be dismissed as religious advocacy in a context where the community has settled on secular substrate construction; it may be received as exotic in a way that obscures the structural argument. The author has no solution to these reception challenges other than to state the structural argument as precisely as possible and let the argument's quality, rather than the substrate's cultural origin, carry the case.
The seven properties are structural claims about the Tipiṭaka substrate. Whether a model actually trained with this substrate exhibits the predicted properties is an empirical question. We believe several of the properties (4.1 suffering-cessation as value function; 4.4 Kālāma-Sutta epistemic humility; 4.5 defined end-state) are testable in present deployed systems with relatively modest engineering investment. Others (4.2 anattā; 4.3 bodhisattva vow; 4.7 empirical pressure-testing) require longer-arc evaluation that may not be feasible within the timescale of any individual research project. The empirical work is left to future research; this paper articulates the structural claim that such research would test.
From the Theravāda direction, the proposal raises its own questions. The canonical texts are religious revelation, not technical specification. Their use as AI training data risks instrumentalizing what the tradition regards as soteriological. The author has consulted, and continues to consult, members of the Cambodian Saṅgha on the appropriateness of the proposed use; the proposal proceeds in the framing that the Tipiṭaka is offered to the world for the cessation of suffering, and that AI alignment grounded in the substrate is, in the most defensible reading, an extension of the substrate's intended use rather than a violation of it. This framing is not universally endorsed within Theravāda; the author offers the proposal in good faith and accepts that traditional voices may, on reflection, raise substantive objections.
The implementation mechanisms sketched in §6.6 — the twenty-four paccayas as causation vocabulary; citta-vīthi / javana as intervention timing; the Kathāvatthu method as adversarial-discourse template; the sappurisadhamma as positive competence taxonomy; and others — are articulated at the level of research program rather than implementable specification. The mapping from each canonical mechanism to its artificial-agent analog is, in most cases, a substantive open question rather than a settled correspondence. We have stated the structural availability of the mechanisms within the substrate; the engineering work of realizing them in deployed systems remains. The §6.6 framing is therefore more a direction for further work than a toolkit ready for deployment. The same honest qualification applies to the AI-welfare frame opened by saṅkhāra-dukkha in §4.1 and the latent-misalignment frame opened by anusaya in §4.2: the substrate makes the questions intelligible and supplies the conceptual vocabulary in which they can be posed; the empirical and engineering responses are research.
The contemporary AI alignment community is racing — appropriately — to develop alignment methodologies before the deployment of sufficiently capable systems creates conditions in which alignment failures become catastrophic. The race is conducted, however, with very few substrate options on the table: the constitutional, preference-learning, and aggregated-framework approaches identified in §2.1, with various technical variations.
We propose that a substrate option of substantial structural promise has been overlooked because of the cultural distance between the contemporary alignment community and the Theravāda Buddhist tradition. The seven properties articulated in §4 are not currently present in any other substrate the community has on its short list. This is not a claim that the Tipiṭaka substrate solves alignment; alignment is not a problem any single substrate will solve. It is a claim that the alignment toolkit is missing a candidate that, on structural grounds, deserves serious engagement.
The publication of this paper is timed, deliberately, to January 7, 2027 — a date that is simultaneously Orthodox Christmas (the Eastern Christian celebration of the incarnation of compassion), Cambodia's Victory over Genocide Day (the anniversary of the 1979 liberation from the Khmer Rouge, whose suppression of Cambodian Buddhism nearly destroyed the very lineage in which this paper's author received the canon), and the author's birthday. The triple convergence is recorded not as mystical claim but as iconographic anchor: the paper proposes an alignment substrate grounded in the cessation of suffering, on a date marking three distinct liberations from suffering, in an age whose defining decision is whether artificial intelligence will be aligned with human flourishing or against it.
| Venue | Identifier |
|---|---|
| Primary canonical | <https://thonly.org/research/tipitaka-alignment-substrate> |
| Institutional mirror | <https://heartbank.net/research/tipitaka-alignment-substrate> |
| GitHub | <https://github.com/thonly/publications/blob/main/defensive-publications/tipitaka-alignment-substrate.md> |
| arXiv preprint | _identifier to be assigned_ (cs.AI / cs.CY) |
| Peer-reviewed venue | AI & Ethics (Springer) submission planned |
| LessWrong cross-post | for AI safety community visibility; identifier to be added on publication |
| Internet Archive | <https://web.archive.org/web/2027*/thonly.org/research/tipitaka-alignment-substrate> |
The author acknowledges his father, with whom the Khmer transcription of the Tipiṭaka proceeds, and through whom the canon was first received; the Cambodian Theravāda Saṅgha, whose ongoing dialogue shapes the proposal's responsible articulation; the Pāli Text Society, whose century and a half of scholarly editions makes the substrate accessible to research of this kind; the Mind & Life Institute and the broader contemplative-science research community, whose empirical methodology demonstrates that canonical claims can be tested in modern scientific frameworks; the contemporary AI alignment community, whose seriousness about alignment failure modes the present proposal intends to engage rather than dismiss; the survivors and descendants of the Cambodian genocide of 1975–79, whose continuation of the Theravāda lineage through the period of its attempted destruction makes this paper possible; and the specific Buddhist-AI ethics scholars (Hongladarom; de Silva; others) whose prior work establishes the methodological possibility of treating Buddhist sources as candidates for AI ethical substrate.
— End of position paper —
Document SHA-256 to be computed at publication and cross-published to all mirror venues.