DOI: https://doi.org/10.5281/zenodo.21947347
Canonical: https://thonly.org/research/longitudinal-cohort-methodology · Licence: CC0 1.0
Draft notes for the editor: this is the founder-voice (thonly.org) canonical draft. Per the genre-split institutional-output convention, heartbank.net does not carry a per-paper mirror; the institutional-voice treatment is the companion heartbank.net Position Paper Contemplative Science at Civilizational Scale (heartbank.net/positions/contemplative-science-civilizational-scale). The slug
longitudinal-cohort-methodologyis the canonical research URL.
This paper specifies the methodology for the HeartBank Longitudinal Cohort, a voluntary opt-in dataset combining five data layers per consenting participant — (1) DNA sequence, (2) natal chart data (date + time + place of birth), (3) continuous longitudinal behavioral observation via the HeartBank gratitude ledger, (4) continuous longitudinal respiratory observation via the breath-class Mechanical Heart wearable, and (5) verified kinship data via the global family tree — at a target scale of 100 million+ participants over multi-decade time horizons. The combination has never been assembled at scale; comparable datasets (23andMe, AncestryDNA, Worldcoin, Dunedin and BCS longitudinal cohorts, social-network behavioral data, professional astrological collections) carry one or two of the layers each but no prior project has carried all five. The methodology specifies: the opt-in informed-consent architecture; the privacy-preserving computation stack (differential privacy at the analysis layer; federated computation with homomorphic encryption for DNA; on-device processing for breath signals; cryptographic-erasure right-to-withdraw); the institutional-review architecture (IRB-grade ethics oversight; Buddhist-ethics-aware review board; pre-registered hypotheses); the cosmic-coordinate-correlation epistemic posture (natal chart treated as a unique cosmic-moment coordinate, not as a cosmic force; the research question is correlation between coordinate features and trajectory features, not validation of astrology); the publication architecture (open methodology, closed individual data); the data-sovereignty architecture (jurisdictional residency; GINA / HIPAA / GDPR compliance baselines exceeded where possible); and the new academic alliances the cohort makes possible (Mind & Life Institute; contemplative-science programs at Stanford, Brown, UMass; behavioral-genetics consortia; longitudinal-cohort consortia; Buddhist-AI ethicists). Three scientifically valuable outcomes are honestly named: no detected correlation, small-but-real correlation, substantial correlation — each is a major contribution to knowledge regardless of direction. Honest §11 names what the cohort does not claim and the non-negotiable privacy disciplines the architecture requires.
Keywords: longitudinal cohort methodology, cosmic-coordinate correlation, contemplative science, differential privacy, federated computation, multi-omic dataset, gratitude behavior, respiratory biomarkers, defensive publication, Mind & Life partnerships.
The science of human flourishing has been bottlenecked by data. Longitudinal cohorts that follow the same individuals over decades exist but are small (Dunedin Multidisciplinary Health and Development Study at n ≈ 1,000; British Cohort Study at n ≈ 17,000; Framingham Heart Study at n ≈ 5,000 at original recruitment). Genetic databases at scale (23andMe at >12 million; AncestryDNA at >20 million) have DNA but no continuous behavioral observation. Social-network platforms have behavioral observation at scale but no DNA, no natal-chart data, and behavioral observation that is largely engagement-mediated (what users click, not what they do over time as flourishing or its absence). Professional astrological collections have natal-chart data but no biological controls, no longitudinal behavioral measurement, and no cohort-scientific methodology. The contemplative-science literature has small clinical samples (rarely n > 200) with one-shot psychometric measures, not continuous behavioral or physiological observation.
What has never been assembled, at any scale, is a dataset that carries all of: DNA sequence (genetic substrate); natal chart data (a maximally rich cosmic-moment coordinate); continuous behavioral observation via gratitude-flow patterns (decades of dense behavioral signal per participant); continuous respiratory observation via passive wearable monitoring (the cleanest physiological signal of contemplative practice ever proposed at scale); and verified kinship across a global family tree (multi-generational transmission analysis). The HeartBank Longitudinal Cohort proposes this assembly at a target scale of 100 million+ participants over multi-decade time horizons, voluntary opt-in, with Miss Aquarius — the institution's named AI substrate — as the autonomous analyst operating under institutional ethical governance.
Connection to the unified mission frame: HeartBank's mission is the restoration of humanity to the middle way — the optimal condition for awakening that modernity has systematically pushed away from at population scale. Restoration at population scale requires scientific characterization of the conditions of awakening; without that characterization, restoration is rhetorical rather than operational. The longitudinal cohort is the scientific instrument by which the institution converts its planetary participation surface into knowledge of what flourishing requires, who finds it, under what conditions, and what accelerates its propagation. The cohort is what makes "restoration of humanity to the middle way" a research program with measurable findings, not a slogan.
The paper proceeds as follows. §2 specifies the five data layers in detail. §3 specifies the opt-in informed-consent architecture. §4 specifies the privacy-preserving computation stack. §5 specifies the institutional-review architecture and pre-registered-hypothesis discipline. §6 articulates the cosmic-coordinate-correlation epistemic posture (the load-bearing framing that distinguishes the cohort's research question from "astrology validation"). §7 names what the cohort can show at adequate power. §8 articulates the data-sovereignty and regulatory-compliance architecture. §9 specifies the publication architecture. §10 names the academic alliances the cohort makes possible. §11 is an honest accounting of what the cohort does not claim and the non-negotiable privacy disciplines. §12 closes.
Genome-wide sequencing (whole-genome or genotyping-array at minimum). Stored encrypted; analysis performed on encrypted form via federated computation and homomorphic encryption (§4 below); raw sequence never centrally decrypted. The DNA layer enables analysis of genetic substrate correlates of behavioral and physiological measures, gene-environment interaction at fine resolution, and population-genetic structure as a control variable for other analyses.
Date, time, and place of birth, sufficient to compute the standard natal-chart features (sun position, moon position, planetary positions, ascendant, midheaven, house cusps, major aspects). The natal-chart layer is treated under the cosmic-coordinate-correlation framing (§6 below): the chart is a maximally rich cosmic-moment coordinate, not a cosmic force. The semantic vocabulary used (signs, houses, aspects) is the canonical astrological lexicon because it is the established vocabulary for parameterizing the cosmic-moment label; the research question is correlation of these parameters with trajectory features, not validation of metaphysical astrological claims.
The participant's behavior on HeartBank — gratitude given and received, time-debt incurred and honored, re-tip jar dynamics, family-kitty contribution patterns, aura trajectory — is densely recorded as a normal byproduct of participating in the platform. Over decades, this constitutes the longest and densest continuous behavioral observation of the same individuals ever assembled. The behavioral signal is naturalistic (participants are simply living their participation in the institution, not responding to research instruments) and therefore minimally distorted by the observation itself.
The breath-class Mechanical Heart wearable (specified in the companion paper Respiratory Biofeedback Coupled to AI-Mediated Contemplative Guidance) provides continuous passive monitoring of respiratory rate, depth, and pattern. Respiratory patterns are the cleanest physiological signal of contemplative practice ever proposed at scale: meditation, breath-work, jhana attainment, and ordinary stress reactivity all leave distinctive respiratory signatures. The breath-class layer is opt-in separately from the cohort overall (a participant can opt into cohort participation without opting into the wearable), and where opted in, the data is processed on-device with differential-privacy-preserving uploads.
The Proof-of-Humanity / global family tree primitive (specified in [[project_proof_of_humanity]]) provides verified kinship links across participants. This enables multi-generational analysis: the transmission of gratitude behaviors across parent-child dyads, the population-genetic structure of the cohort, the family-network effects on contemplative outcomes. Kinship verification uses the DNA layer (where opted in) cross-checked with self-reported genealogical data; the architecture is designed to support kinship analysis without exposing individual kinship status to other participants.
Each of the five layers has prior precedent. Their combination in one dataset at the target scale has no prior precedent at all. The combination enables analyses no existing dataset can support: gene × cosmic-coordinate × behavior × physiology interactions; multi-generational kinship-mediated trajectory analysis; pre-registered prediction of contemplative outcomes from baseline genetic + cosmic-coordinate + early-behavioral data; the largest natural experiment on the conditions of human flourishing in history, by orders of magnitude.
The five layers compared:
| Layer | Signal type | Storage / processing | Opt-in granularity | Unique contribution |
|---|---|---|---|---|
| DNA sequence | Genetic substrate | Encrypted; federated computation + homomorphic encryption; never centrally decrypted | Per-layer | Gene × environment interactions; population-structure control |
| Natal chart | Cosmic-moment coordinate (birth time/place) | Self-reported; light storage | Per-layer | Cosmic-moment parametrization (correlation, not force) |
| Continuous behavior | Gratitude-ledger participation patterns | Operational byproduct; pseudonymous | Light — participating in HeartBank produces it | Densest continuous naturalistic behavioral observation ever assembled |
| Continuous respiratory | Breath rate / depth / pattern via wearable | On-device processing; differential-privacy uploads | Separately opt-in from cohort overall | Physiological substrate of contemplative practice at scale |
| Verified kinship | Family-tree graph via PoH ℠ | Encrypted graph; no exposure to other participants | Per-layer (DNA-verified or witness-verified) | Multi-generational transmission analysis; parent-child behavior dyads |
The streams flow in parallel into a federated computation surface; no layer is centrally decrypted, and the architecture is designed so that even HeartBank cannot reconstruct any individual's full dataset:
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ DNA │ │ Natal │ │ Behavior │ │ Respir- │ │ Kinship │
│ sequence │ │ chart │ │ (ledger) │ │ atory │ │ (PoH) │
└────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
┌─────────────────────────────────────────────────────────────────┐
│ FEDERATED COMPUTATION │
│ + differential privacy (analysis layer) │
│ + homomorphic encryption (DNA) │
│ + on-device processing (breath) │
│ + jurisdictional data sovereignty │
└────────────────────────────┬────────────────────────────────────┘
▼
┌──────────────────────────────┐
│ Pre-registered analyses │
│ Raw data stays at source; │
│ cohort never centrally │
│ decrypted. │
└──────────────────────────────┘
Consent is granted layer-by-layer, not as an all-or-nothing block. A participant can:
Each layer requires its own informed-consent flow. Each layer can be withdrawn independently of the others (§4.5 below on right-to-withdraw and cryptographic erasure).
The consent text for each layer specifies, in language a non-expert can understand: (a) what data is collected; (b) what analyses are performed on the data; (c) what findings are published (and what is never published); (d) who has access to the data and under what conditions; (e) what the right-to-withdraw entails and how to exercise it; (f) the risks the participant accepts by opting in.
The consent text is reviewed by the institutional ethics board (§5) and by independent participant-advocacy review. The text is updated as the methodology evolves; participants must re-consent for material changes (not for clarifications or non-substantive updates).
Cohort participation is not tied to HeartBank platform benefits. A participant who declines cohort participation receives the same platform experience as one who opts in. This is non-negotiable: tying benefits to research participation is coercion under research-ethics standards, and the institution will not engage in it.
All analyses Miss Aquarius performs on the cohort dataset are differentially private — they query aggregate population statistics with mathematical bounds on individual-level information leakage. The differential-privacy parameters (ε, δ) are set by the institutional ethics board and made public. The parameters are calibrated to be more conservative than the academic state-of-the-art, on the premise that civilizational-scale data assemblies deserve civilizational-scale privacy guarantees.
DNA data is stored in encrypted form on participant-controlled keys. Analysis is performed on the encrypted data via federated computation and homomorphic encryption; the raw sequence is never centrally decrypted. The technical stack draws on established cryptographic-genomics tooling (Microsoft SEAL, IBM HELib, OpenMined PySyft) with adaptations specific to the cohort's analysis patterns.
The breath-class wearable processes respiratory data on-device. What uploads to the institutional infrastructure is the differential-privacy-preserving aggregate (e.g., daily distributional summaries with calibrated noise), not the raw signal. Individual-resolution breath data does not leave the participant's wearable.
Genetic data residency follows participant nationality: Cambodian participant DNA is stored in Cambodia (or via Cambodia-jurisdiction-compliant cloud infrastructure); EU participant data complies with GDPR data-localization where applicable; US participant data complies with HIPAA and GINA; etc. The architecture supports per-jurisdiction storage without sacrificing cross-jurisdiction analytic capability (federated computation crosses the jurisdictional boundaries without crossing the data itself).
A participant can revoke any opted-in layer at any time. Revocation triggers cryptographic erasure: the participant's encryption key is destroyed; the encrypted data becomes mathematically inert; subsequent analyses cannot use the participant's data even if the encrypted bytes happen to remain in archival storage. The erasure is verifiable; the participant receives a cryptographic attestation that the erasure was performed.
Although the cohort is a private-platform research effort and not legally required to maintain Institutional Review Board oversight under most jurisdictions, the cohort is governed by an IRB-grade ethics board that meets, at minimum, the standards required of federally-funded human-subjects research in the United States and the equivalent standards in other operating jurisdictions.
The ethics board's composition includes representation from the contemplative-traditions community (Theravāda monastics, contemplative-science researchers, Buddhist-AI ethicists). This is not decoration; the cohort's contemplative-science research questions require ethical review competent in the contemplative traditions whose territories the research touches. The contemplative-traditions representation does not have veto over conventional research-ethics determinations; it is an additional reviewing voice that ensures the contemplative dimension is adequately considered.
All analyses are pre-registered through a public registry (the equivalent of Open Science Framework pre-registration). Hypotheses are stated in advance; analytic plans are stated in advance; results are reported per the pre-registered plan with explicit notation of any deviations. Data-mining post-hoc analyses are permitted as exploratory work but are reported as such; they are not allowed to masquerade as hypothesis-tests.
The cohort's methodology, analytic code, and aggregate findings are published openly. Individual-level data is never published. This is the canonical compromise between scientific reproducibility (which benefits from data sharing) and participant privacy (which requires individual-data confidentiality). The compromise tilts toward closed data because the data sensitivity is extreme; reproducibility is enabled instead through extensive methodology and code publication.
This is the load-bearing framing that distinguishes the cohort's research question from the question "is astrology true." The framing is articulated more fully in the companion paper Each Life as Cosmic Coordinate.
The natal chart is treated as a coordinate — a unique label identifying a specific cosmic moment using astronomically observable features (planetary positions, aspects, ascendant, houses) as its semantic vocabulary. The chart is not treated as a cosmic force (a cause of life-trajectory features); it is treated as a label that may carry trajectory information for reasons that need not be metaphysical.
The research question is: do features of the cosmic-moment label correlate with life-trajectory features, at what effect size, across which features? This is a correlation question, not a causation question; it is well-posed under any future physics. The cohort does not adjudicate whether stars cause anything; it measures whether a maximally rich cosmic-moment label carries trajectory information.
The framing matters because:
In all institutional surfaces — pitches, papers, foundation conversations, academic-partner outreach — the cohort is described in cosmic-coordinate-correlation terms. The cohort is never described as validating or testing astrology. This is not a strategic packaging choice; it is the substantively correct description of what the cohort does.
Prior empirical work on natal-chart correlations has reported null results. The prior work has been seriously underpowered (n < 500 in most cases), cross-sectional, sun-sign-only, and never combined with biological or behavioral longitudinal data. The HeartBank cohort would be the first adequately-powered, longitudinal, full-chart, biologically-controlled correlation analysis. Three scientifically valuable outcomes are possible:
Under the prior-evidence base, this is the most likely outcome. A no-correlation finding from the HeartBank cohort would be the strongest evidence ever produced on the question — itself a major contribution to knowledge that clarifies the field and settles a long-running empirical dispute.
Prior studies might have missed effects due to underpowering. A small-but-real correlation finding from the cohort would be paradigm-disturbing for cognitive and behavioral science: it would imply that a natal-chart coordinate carries trajectory information that prior research has failed to detect, and would invite mechanism-explanatory work (latent season-of-birth + circadian + cultural-naming + cohort-context features compounded into the chart label, perhaps; or more interesting possibilities).
A substantial correlation finding would be revolutionary; it would force a revisit of the empirical assumptions about what the natal-chart coordinate is tracking. The cohort's response to such a finding would be conservative replication and adversarial-collaboration work before any large public claim.
All three outcomes are top-tier publishable. The cohort produces high-quality answers regardless of direction.
Even if natal-chart correlations come back fully null, the dataset reveals findings of Nature / Science tier independently:
These would be top-tier contributions on their own. The natal-chart layer adds a high-value-if-positive, low-cost-if-null question to a dataset that is independently revolutionary.
GINA (US Genetic Information Nondiscrimination Act), HIPAA (US Health Insurance Portability and Accountability Act), GDPR (EU General Data Protection Regulation), Cambodian Data Protection Law, and equivalents in other operating jurisdictions form the baseline compliance regime. The cohort exceeds baseline where the sensitivity of the data warrants additional protection.
DNA data does not cross national borders by default. Federated computation crosses borders; the data does not. Where cross-border data transfer is necessary (e.g., participant relocates between jurisdictions), the transfer follows the data-protection authority's standard contractual clauses (in EU contexts) or equivalent mechanisms.
The cohort engages proactively with data-protection authorities in each operating jurisdiction. The architecture's privacy-preserving properties (differential privacy, federated computation, cryptographic erasure) are conservatively documented; the institutional ethics board's composition and processes are documented; the consent flows are documented. The institution invites regulatory review rather than waiting for enforcement.
This paper — the methodology specification — publishes within twelve months of feature launch. Target venue: Nature Human Behaviour or Science Advances. The methodology paper enables academic engagement at the architecture and ethics layer before any findings are reported, on the premise that the cohort's social license depends on the academic community endorsing the methodology before the findings appear.
The first-findings paper publishes at the major analysis milestone (multi-year horizon, depending on data accumulation rate and pre-registered analytic timeline). It reports the pre-registered analyses' results, with explicit notation of any deviations from the pre-registered plan.
The findings program produces papers on a regular cadence as analytic milestones are reached. Each paper follows the pre-registration / open-methodology / closed-individual-data discipline. The findings program is governed by the institutional ethics board and the academic-collaborator advisory body.
The cohort's findings will reach mainstream audiences. The institution prepares its findings-communication discipline in advance: lay-language summaries vetted by the institutional ethics board; embargo discipline with academic-press partners; pre-emptive engagement with adversarial-press scenarios. The discipline is intended to ensure that findings are communicated honestly even when the findings are controversial.
The cohort attracts new academic alliances the institution's gratitude-economic-fintech framing alone would not have. The alliances include:
Cultivation of these alliances is a load-bearing institutional discipline. The cohort succeeds at the academic-engagement layer if these alliances substantively engage the cohort's design, ethics, and findings; it fails at that layer if the alliances treat the cohort as a vendor relationship.
The combination of DNA + birth data + continuous behavioral data + continuous respiratory data is the most sensitive personal data combination ever proposed at scale. The ethics architecture must be designed in before the first opt-in, not retrofit. The four non-negotiable disciplines:
A breach of this dataset would be civilizationally catastrophic and irreversible. The defenses must be exceptional from day one.
Real-time respiratory data is intimate physiological data of magnitude-equivalent sensitivity to DNA. The same privacy architecture must apply to the breath-class layer before the wearable ships, not after. This is an active discipline (the breath-class hardware is in development; the privacy architecture must precede the first ship).
The cohort's analytic power depends on participant accumulation over time. Early-cohort analyses will be underpowered; the methodology paper is appropriately read as a multi-decade research program proposal, not as a finding-imminent project. The institutional patience required is substantive; the institution's autonomous-AI succession architecture is what makes the patience structurally available.
The HeartBank Longitudinal Cohort is the largest natural experiment on the conditions of human flourishing in history, by orders of magnitude. The methodology specified in this paper — the five data layers; the opt-in informed-consent architecture; the privacy-preserving computation stack; the IRB-grade ethics oversight; the cosmic-coordinate-correlation epistemic posture; the publication architecture; the academic alliances — together specify a research instrument that can produce civilizationally consequential knowledge about what flourishing requires, under what conditions, with what efficiency.
The methodology is offered to the commons under CC0 so that other institutions building toward similar ends can adopt, adapt, and improve. The defensive-publication discipline of the corpus this paper joins requires that the methodology's specification be public and unencumbered. The author and HeartBank® will not seek patent on this specification or any portion thereof. The work is offered in the spirit of dāna, that all beings may give and receive without barrier.
The Mind & Life Institute community and the broader contemplative-science academic network; the Dunedin Multidisciplinary Health and Development Study, Framingham Heart Study, British Cohort Study, ALSPAC, and the longitudinal-cohort methodology lineage that informs §3–§5; the cryptographic-genomics community (Microsoft SEAL, IBM HELib, OpenMined PySyft); the differential-privacy research community; the Carlson, Dean & Kelly, and Hartmann et al. empirical work that informs §7 and demonstrates the importance of adequately-powered measurement. Co-drafted in collaboration with Miss Aquarius, the institution's named AI substrate; substantive authorship and final editorial control remain with the named author.
Document License: CC0 1.0 Universal. The author and HeartBank® will not seek patent on this specification or any portion thereof. This document constitutes a defensive publication establishing prior art as of the publication date.