Published research

Eight papers published on Zenodo. Each represents a complete, citable contribution to the Living Framework. Together they constitute the theoretical and empirical foundation on which RITAM and NIYOM build.

01
Foundations of the Living Framework — Reliability Architecture
2024 · Structural reliability conditions for extended AI collaboration
02
Memory Governance in AI Collaboration
2024 · Structured memory architecture and context integrity
03
Structural Failure Modes in AI Collaboration
2024 · Six named failure modes — taxonomy, mechanisms, and signals
04
The Living Framework
2024 · Philosophical foundations of sustained human-AI partnership and epistemic trust
05
Language Reliability in AI Systems
2025 · Linguistic patterns as reliability indicators
06
Intervention Architecture in AI Collaboration
2025 · Structural interventions for reliability maintenance
07
Distributed Cognition in Human-AI Systems
2025 · Extended mind theory applied to AI collaboration
08
The Verification Layer in AI Collaboration Architecture
2025 · Structural verification as the missing architectural component
ML
The Mahdi Ledger
2024 · Applied reliability framework in complex information environments · AI-generated, separately published

Currently in progress

Two active research projects extending the published framework. Both are pre-publication — findings will be incorporated into future papers.

Project RITAM — Reliability Information and Technical Architecture Model
Active · 2025–2026
Foundational substrate research for governed AI cognition. Nine architectural primitives — each justified by a failure mode. v1.1.1 runtime published open source, 146 tests, Apache 2.0. Founding principle: governance must precede persistence.
→ Project RITAM
Project NIYOM — Verification Harness
Active · 2025–2026
A structured execution harness for AI, in active development. Every task goes through Plan, Execute, Verify, and Repair — with a separate adversarial verification pass between generation and acceptance. Built on the LC-OS paper series. v1.0.0 · Stage 7 in progress.
→ Project NIYOM
Failure Library — ongoing documentation
Continuous
The Failure Library published six named failure modes from Paper 03. New failure modes and new instances of documented modes continue to be identified from active research. The library is maintained as a living document alongside the published papers — new entries will be added as the empirical base accumulates.
→ Failure Library

Under active investigation

Questions the research has identified but not yet answered. These are not speculations — they are gaps that the existing papers point to and that current research is working toward.

Assessment tools for reliability architecture evaluation
The published assessment (20 questions, 5 domains) is a functional diagnostic but it is primarily structured around failure mode exposure rather than architecture evaluation. A more systematic assessment methodology — one that could evaluate an AI collaboration architecture against the RITAM model — is under development. This requires RITAM to reach a stable enough state to define the evaluation criteria.
Multi-agent reliability in distributed AI systems
The published framework addresses human-AI collaboration. As multi-agent AI systems become more common — where multiple AI models collaborate, pass information between themselves, and produce outputs without continuous human review — the reliability architecture questions change significantly. Paper 07 (distributed cognition) provides a theoretical basis, but the empirical work on multi-agent failure modes has not yet been done.
Automated verification — structural conditions for delegation
Paper 08 argued for the verification layer as a dedicated architectural component. A key open question is what conditions must hold for verification to be delegated to an automated process rather than requiring human review. This is not an AI-specific problem — it has engineering precedent — but the conditions in human-AI collaboration systems have characteristics that may require their own analysis.
Reliability degradation rates across domain types
The empirical work underlying Papers 01–08 was observational rather than controlled. A more rigorous investigation of how reliability degrades at different rates across domain types — quantitative work, legal analysis, creative work, strategic planning — and what structural factors explain these differences, would significantly strengthen the empirical foundation of the framework.

Open research questions

Unresolved questions that the current research cannot answer — where the framework's limits are visible and future work is needed.

How much of the reliability architecture needs to be built before each session begins, versus how much can be constructed adaptively during a session?
The framework describes architectural conditions but is underspecified on timing. Some conditions (like initial context structure) are clearly antecedent. Others (like verification) might be adaptively triggered. The practical trade-off between pre-session architecture cost and in-session reliability benefit is not yet understood quantitatively.
Do the structural failure modes interact? If Context Drift occurs, does it increase the probability of File Divergence or Domain Erosion?
Paper 03 documented the six failure modes as distinct. The research did not examine whether they cascade or interact — whether one failure mode raises the probability of others, or whether some failure modes are prerequisites for others. Understanding the failure cascade structure would substantially improve diagnostic prioritisation.
Is the reliability architecture domain-general, or does it require domain-specific adaptation?
The framework claims to describe structural conditions for reliability across domains. But domain-specific research (legal AI, medical AI, financial AI) applies reliability concepts differently and emphasises different failure modes. The question is whether the Living Framework's architecture is domain-general with domain-specific configurations, or whether some domains require fundamentally different architectural models.
What is the relationship between model architecture changes and framework validity?
The framework argues that reliability is an architectural property of the human-AI system rather than the model. But model architecture changes (extended context windows, better long-term coherence, built-in tool use) may change the nature of the structural conditions required. The framework needs to be evaluated against significantly more capable models to test whether its core claims remain valid or require revision.
Can the verification layer itself be a source of failure, and if so, what makes verification reliable?
Paper 08 argued for a dedicated verification layer. The assumption is that verification is more reliable than generation. But if verification is itself performed by an AI process, it inherits its own reliability conditions. A verification process that has the same failure modes as the process it is checking does not improve system reliability. The conditions for reliable verification are currently unexamined.
How should the framework be evaluated? What would falsifying evidence look like?
The Living Framework makes claims about what structural conditions are necessary for AI collaboration reliability. For these claims to be scientifically meaningful, it must be possible to specify what evidence would disconfirm them. This is currently underspecified — the research is descriptive and theoretical more than hypothesis-testing. The transition to a more falsifiable form is a methodological priority.
On roadmaps in research

Why "Living Framework" is not a marketing phrase

Most roadmaps in product or technology contexts describe planned features. This roadmap describes active empirical inquiry — questions being investigated, gaps being closed, and limits that the research currently cannot reach past. The items under "open" are not backlog items; they are limits of the current knowledge state.

The framework is called "living" because it is not a fixed model. Each paper revised and extended what came before. RITAM and NIYOM will revise and extend again. The open questions above are places where the framework's account of reliability is currently incomplete — where the evidence is insufficient, the analysis is underspecified, or the empirical work has not been done.

Researchers who work on adjacent problems and have evidence that bears on any of the open questions are welcome to make contact. The Advisory page has the right channel.