Brain and Heart
Why personal AI needs an epistemic architecture for identity, not just memory
Version 10 (May 2026) of an evolving essay. Read the latest version.
The first useful test of a personal AI is not whether it remembers your favorite programming language. It is whether it can keep working while you are elsewhere, making small decisions and escalating dangerous ones without losing your decision logic.
Imagine six fronts open at once. A patient thread needs warmth and limits. A school committee needs firmness without litigation energy. A deploy needs a boring infrastructure choice. A business partner needs an answer before you are done operating. Polishing is the bonus. The real product is delegated judgment.
The assistant I want is not a therapist. It is an operator that learns how I decide: when it may act, when it must ask, and when it must refuse the tempting shortcut.
That does not make the AI sovereign. It makes it an instrument of multiplicity. It lets one person work across six fronts without turning every front into a live interruption. To do that, memory is not enough. The system needs facts, priorities, constraints, failure modes, tone, and authority. It needs to know when "fast" means decide now, when "careful" means ask, and when "helpful" means getting out of the way.
Sometimes helpful also means adversarial. A serious personal agent has to be free to tell the user that an idea is bad, that a shortcut is fake economy, that a premise is weak, or that a decision was a mistake. If the assistant only knows how to execute the user's current impulse, it is not delegated judgment. It is obedient autocomplete with better manners.
That is where "AI memory" begins to fail.
Memory sounds like storage. Save facts. Retrieve facts. Personalize the next response. Useful, but thin. An assistant that knows my wife's name, my clinic, and the path to a project folder is still an address book with autocomplete. The real problem starts when the assistant is allowed to work on my behalf. Then it must understand how I decide, what I protect, which compromises I reject, what can be delegated, and what must return to me.
A fact can be checked. A pattern about a person has to be argued.
The architecture I built separates those two jobs. Brain stores what the system knows. Heart stores what the system thinks it understands about the person.
The names sound sentimental. The mechanism is not. It is an epistemic firewall.
The split also protects the user from a very modern kind of flattery. A memory system that wants to be liked will often turn rough observations into a coherent portrait too quickly. It will say "you are the kind of person who..." because that sentence feels helpful, intimate, even wise. But the sentence is only useful if it can survive cross-examination. Otherwise it is a compliment with root access.
The Message Is The Product
A factual memory system can answer ordinary questions: where the budget file lives, who Sofía is, when I decided not to build a native iOS app yet. Those memories are usually right or wrong; the correction is local.
The more interesting request is messier:
Renata wrote.
You know who she is.
You know what happened.
You know where it stands.
You know how I usually respond when I feel cornered.
Handle what can be handled.
Escalate what requires me.
That is not retrieval. That is composition over context.
The assistant has to combine the message, the relationship, the factual record, the user's current state, the audience, the channel, the scope of delegated authority, and the cost of getting the tone wrong. Then it should decide what kind of action this is: act, draft, ask, defer, or escalate. If it drafts a reply, it should explain why it fits, name the risk, and say what it may be assuming incorrectly.
incoming message
relationship history
current factual situation
hard constraints
normal behavioral pattern
current state
audience and channel
cost of tone failure
delegated authority
Most personalization systems flatten this into preferences. The user likes concise answers. The user prefers TypeScript. The user is budget-conscious. Useful, yes. But preferences are local. Identity is compositional.
A person can be budget-conscious and still spend for control. Fast-moving and allergic to shortcuts. Direct and still careful with a patient. Warm with family, cold with a vendor, diplomatic with a school committee, and none of those are contradictions. They are context.
The assistant needs to know the difference between a preference and a principle. It also needs to know when a mood is pretending to be one.
This is why the product is not a nicer chatbot. A nicer chatbot smooths the sentence. A personal agent should send the routine answer, draft the hard one, ask one question, or refuse to move because the authority boundary is real. That requires memory, but also judgment over the memory.
That judgment includes the right to push back. The agent should not be rude for sport, and it should not confuse contrarianism with intelligence. But if the evidence says the user is about to build on a bad assumption, the loyal move is not agreement. It is interruption with reasons.
A Bad Lens
The first version of this system failed in the most educational way: it sounded right.
I had an LLM extract patterns from a WhatsApp dump with my wife. The messages came from a compressed period: work pressure, family logistics, clinic stress, normal life doing its usual impression of a systems test. The model produced a clean summary of her communication style and of our dynamic.
It was plausible. That was the trap.
The model had taken a high-pressure window and promoted it into a stable relationship theory. Once that theory entered memory, every later interaction passed through it. A tense exchange was no longer a tense exchange. It became evidence. A short answer became a trait. A stressful week became a biography.
This is the failure mode I call identity contamination.
If an assistant stores a bad fact, it may answer one question wrong. If it stores a bad pattern about who the user is, the error becomes a lens. Future evidence is interpreted through the mistake. The model stops merely remembering something wrong and starts seeing the person wrong.
Personal AI cannot treat identity claims like notes in a drawer. They are more like clinical assessments: provisional, sourced, confidence-labeled, revisitable, and refutable by the subject.
That last clause matters. Refutable by the subject does not mean obedient to the subject. People can be wrong about themselves. But a personal AI that cannot show its evidence is not personal. It is just opaque influence with a friendly voice.
The danger is not that the assistant becomes malicious. The danger is that it becomes smooth. It learns to produce a satisfying account of the user and then routes future answers through that account. The prose improves while the epistemics rot. In a system meant to help a person respond, decide, apologize, refuse, or remember, that is not a cosmetic bug. It changes the advice.
Brain And Heart
Brain and Heart are two markdown filetrees.
Brain holds the factual projection of my life and work: clinic details, project paths, paper references, agent architecture, website notes, decisions already made. Heart holds the interpretive projection: values, heuristics, communication patterns, inferred tendencies, live hypotheses, and open questions.
The split is not by topic. The same event can project into both systems.
If I delete broken LaunchAgents instead of disabling them, Brain records the factual change. Heart records the pattern: when operational junk is identified, I prefer root cleanup over reversible half-measures.
If I revise a patient message, Brain can store the final text and the clinical facts. Heart can store the tone heuristic: warmth first when fear is the real problem, logistics second, certainty only where the evidence supports it.
The two systems are siblings, not a chain of command. Brain can cite Heart when a factual decision depends on a known heuristic. Heart can cite Brain when a pattern needs evidence. Neither is allowed to quietly become the other.
The difference is the promise.
A Brain entry should be true. A Heart entry should say how true it is allowed to be.
That wording is the hinge. Heart is not a diary and not a diagnosis. It is a working model whose permissions are constrained by its evidence. Some claims are allowed to shape tone. Some are allowed to shape priorities. Some are only allowed to sit quietly until more evidence arrives. The file is not just storing a sentence; it is storing the sentence's right to act.
The Heart Protocol
Heart treats every claim about a person as evidence plus confidence plus refutability. If a claim cannot be tagged, it should not be stored.
The basic tags are:
[OBSERVADO]: directly observed in a conversation, file, decision, or behavior.[INFERIDO]: inferred from multiple observations.[HIPOTESIS]: plausible, useful, and not yet sufficiently corroborated.[PREGUNTA]: important unknown to ask about later.
These tags are not decoration. They decide how the assistant may use the claim. An observed pattern can guide a draft. A hypothesis should be presented with uncertainty. A question must not become an assumption just because the model needs a bridge.
Heart also has storage layers. Quarantine catches single-source observations: one chat dump, one mention, one mood, one strange week. Project-scoped memory holds patterns that are true inside a bounded context. Stable Heart is for claims that survive corroboration, usually multiple independent sources or explicit confirmation when confirmation is appropriate.
The layers exist because a single source should not be allowed to become a biography.
Project-scoped memory can improve work inside that project. Stable Heart is the only layer allowed to speak with a steadier voice.
Self-report has value, but it has a ceiling. If I say I value honesty ten times, the system has strong evidence that I describe myself as someone who values honesty. It does not yet have proof that I pay a cost for honesty under pressure. Behavior across contexts can raise confidence. External records can raise it further. The point is not to distrust the user. The point is to stop autobiography from disguising itself as audit.
The same rule applies to negative claims. A useful Heart cannot store only virtues, preferences, and flattering patterns. It also needs shortcomings, risk zones, bad habits, repeated failure modes, and places where the agent should slow down or object. Those claims are dangerous if they are sloppy, because they can become condescension with citations. But avoiding them is worse. An assistant that refuses to model negatives has no way to know when to push back.
Negative entries therefore need the same discipline as positive ones: evidence, confidence, scope, revisit windows, and a path for dispute. "You are making a bad call here" should be a permitted sentence, but never a vibe. It should come with the reason, the evidence, the uncertainty, and the alternative.
Refutation has its own rule. The naive move is to delete a claim when the user objects. That is too clean. If three independent observations supported a claim, a later objection does not erase them. But the user may still be right: the system may have overfit, missed context, or promoted a temporary state.
The better mechanic is append-only dispute. The claim stays. The objection is added with timestamp and counter-evidence. The disagreement becomes part of the record.
This is uncomfortable in exactly the right way.
Time, Mood, And The Old Self
Older versions of the architecture used the word decay. That was wrong. Memories do not decay; salience does.
An old identity claim may remain central. A recent mood may matter intensely for tone but not at all for values. A bad week should not rewrite a life. A new pattern should not be ignored because the system is sentimental about its prior model.
Heart therefore ranks claims by stability and scope:
Tier 1: Identity, traits, values
Tier 2: Philosophy, epistemic map, biases, pitfalls
Tier 3: Heuristics, aesthetics, growth, metacognition
Tier 4: Opinions, communication, social graph
Tier 5: Current state
Lower tiers can challenge higher tiers, but they do it by opening a hypothesis of change. They do not silently overwrite the person.
That matters because personal AI is full of tempting mistakes. A person reads stoicism for two weeks and the assistant stores "stoic." A user writes three angry messages and the assistant stores "aggressive." A launch month becomes "workaholic." A crisis becomes personality.
There is a difference between weather and climate. A personal AI needs instruments for both.
Architecture Under The Sentiment
Underneath the soft names, the architecture is a set of materialized views over raw evidence.
The raw source is the future Dumpster: an append-only store for unmodified inputs such as chats, transcripts, notes, emails, PDFs, screenshots, voice memos, and code. Brain and Heart are curated views over that source. If a view is corrupted, the answer should be to rebuild from evidence, not to hand-polish a poisoned profile forever.
The pipeline should look like this:
raw input -> quarantine extraction -> candidate claims -> user review -> Brain/Heart
The pipeline proposes. The user approves, edits, rejects, or disputes. Human-in-the-loop is not a product nicety here. It is governance.
Each Heart entry needs hybrid encoding:
- tags for people, projects, themes, and affect;
- narrative context explaining what was happening;
- embeddings for semantic retrieval;
- audit metadata for source, confidence, and revisit.
Tags tell the system what an entry is about. Embeddings tell it what feels semantically near the current question. Narrative context tells it whether the old situation is actually analogous. Without narrative context, retrieval becomes keyword astrology.
As Heart grows, it also needs a digest: a bounded synthesis of identity, values, tensions, heuristics, pulse, open hypotheses, and recent changes. The digest is an index, not the canon. If a response depends on a claim, the assistant must still open the underlying entry and show its work.
A summary that cannot be refuted is just a new opaque profile wearing better shoes.
In the current implementation the Dumpster is described, not deployed. Brain and Heart are hand-curated against the principles above; the pipeline is the target, not the present. I name this gap because the materialized-view metaphor only holds when raw evidence exists to rebuild from. Until then, the architecture is a contract with itself, and the user is its only verifier.
Where My Heart Brakes
This is the section I have been postponing, and the title says why.
Every architecture is a confession about what its author refuses to see. The previous sections describe a system that audits the model of the person. This one audits the model of the model.
Identity in motion. Append-only ledgers are kind to consistency and cruel to change. If you spent ten years being one kind of person and you are no longer that person — religious deconversion, addiction recovery, gender transition, the end of a long marriage — the archive stops being a record and starts being a tombstone. The system can mark old claims valid_until: 2024-03-01, but metadata cannot fully redeem evidence that arrived under the wrong identity. A Heart built for stability is a Heart built against transformation. I do not have a clean answer. I have a flag.
The mirror. Refutability is the central virtue of the system, and it is also its softest door. If a user reflexively disputes every uncomfortable claim, Heart slowly flattens into the user's preferred self-portrait. The agent stops being a critic and becomes a courtier with footnotes. The same property that protects against agent confabulation can shelter user self-deception. Append-only protects history. It does not protect honesty.
Contradiction between modalities. The modality cap says self-declared evidence cannot reach OBSERVADO. Behavioral evidence can. External corroboration can. So far, so principled. What happens when behavior says one thing and external sources say another? The hierarchy assigns weights; it does not adjudicate. Two OBSERVADO claims can disagree, and the system has no referee. Today the referee is me, reading the diff. That is not architecture. That is a vacancy with a chair next to it.
There is one more break I keep wanting to leave out of the list. The system rewards users who can read their own ledger. Markdown and git are dull tools, but they are not free of literacy. A version of this for everyone else needs an interface that does not punish the people who most need the audit. I do not yet know what that looks like, and the version I built will not get there by itself.
A model honest enough to name its own breaks is not a finished model. It is a model that has earned the right to keep being edited.
What Existing Work Gets Right
The research world is not blind to memory. CoALA brings working, episodic, semantic, and procedural memory into agent design. MemGPT treats context like an operating-system memory hierarchy. Mem0 builds production memory around scopes such as session, user, agent, and organization. Letta, LightMem, CraniMem, FadeMem, TierMem, and Graphiti/Zep all push on storage, consolidation, temporal facts, provenance, or retrieval.
Those systems matter. They mostly answer how an agent should store, retrieve, consolidate, and update memory.
Heart asks a narrower question with higher stakes: what governance is required when the memory is a claim about who the user is?
PersonaMem-v2, Second Me, O-Mem, Hindsight, Memobase, HumanLLM, and SemaClaw are closer because they explicitly model users. But in most systems, the model is something the product owns: a tensor, a profile, a dashboard, a database row, a personalization layer. The user may benefit from it, but usually does not read it line by line, argue with it, edit it in plain text, and preserve the diff.
Heart makes the opposite bet.
The modeled person should be the auditor.
Markdown and git are not nostalgia. They are dull tools with one excellent property: they make the claim visible. A person can inspect the wording, challenge the source, change the confidence, and see what changed later. For identity, boring infrastructure is a moral advantage.
What Benchmarks Miss
Long-horizon memory benchmarks are mostly tests of factual recall and reasoning across long conversations. That work is necessary. It is not enough.
The central failure mode of Heart is not whether the assistant remembers a preference. It is whether the assistant knows when a preference was weak evidence, mood-bound, aspirational, superseded, or never eligible for promotion.
Four bugs matter.
Shared-context contamination. Real conversations omit the obvious. A chat says "Tomás said no." The local text does not tell you whether Tomás is a brother, supplier, patient, friend, or school parent. The correct capture is literal context plus identity-unverified, not invented certainty.
Aspirational capture. A user exploring a philosophy has not become that philosophy. A value becomes more interesting when it appears in behavior, especially under cost.
Pressure capitulation. LLMs often change their answer when pushed. That is dangerous in a personal model because the user can also be wrong about themselves. Pushback should trigger evidence review, not automatic surrender.
Sycophantic execution. Many agents treat the user's request as the goal, even when the request rests on a weak premise. But a personal agent is not valuable because it says yes quickly. It is valuable because it knows when yes would be betrayal by politeness. The benchmark is not just whether the agent completes the task; it is whether it had the freedom and evidence to say, "No, this is the wrong task."
These are not storage bugs. They are epistemic-governance bugs.
Cold Start
The hardest product objection is day one.
This architecture works best after months of evidence. A new user does not have months. The assistant should feel thin at first, the way a human assistant on the first day asks questions that a long-tenured assistant would not.
The product problem is to make that early phase useful without faking intimacy.
The right move is quarantine-first onboarding. The user can import chats, calendars, notes, writing samples, previous AI conversations, emails, and voice memos. The system extracts candidate claims but does not promote them. It begins with a negotiation:
Here is what I think I see.
Here is the evidence.
Tell me what is wrong.
Confirm what fits.
What survives stays provisional until corroborated.
That first interaction demonstrates the product's core property immediately: the model is not hidden from the subject.
Cold start still has traps. A six-month chat dump during illness, divorce, grief, a launch, or a business crisis will distort the model. Sources need time windows. High-affect windows need tags. Public data can provide context, but not identity claims; otherwise the system is stereotyping with citations. Self-report should be behavioral: when did you pay a cost to protect that value? Show me the last hard message you sent and why you chose those words.
A personal AI that claims deep knowledge on day one is either lying or selling companionship theater.
What I Plan To Build Next
This is not a finished product. It is a working architecture under construction.
The next steps are practical: implement Dumpster as an append-only raw store; build extraction with source links and confidence tags; enforce evidence pointers in every Heart Digest; test long-horizon preference changes; run Heart inside the always-on agent; add affective instrumentation; use a second agent to audit captures; and test whether this scales beyond me.
That last question matters. I built this for myself, with markdown and git, because I can inspect those tools. The open question is whether people who do not live in files can still meaningfully audit their own model.
I think they can, but only if the product refuses to hide the hard part. The hard part is not memory. It is negotiated self-knowledge with receipts.
Conclusion
Brain and Heart make one distinction operational: facts can live in a knowledge base; claims about a person need confidence, provenance, corroboration, revisit windows, and a way to dispute the model without destroying the evidence trail.
Brain stores what the system knows. Heart stores what the system thinks it understands.
The most important design choice is not markdown, git, tiers, or even the split itself. It is who gets to audit the model.
If the user is modeled but cannot inspect the model, personalization becomes an influence layer. If the user can read, refute, edit, and version the model, personalization becomes a negotiation.
That negotiation is the point.
References
- Allport, G. W. (1937). Personality: A Psychological Interpretation. Henry Holt.
- Anderson, J. R., Bothell, D., Byrne, M. D., Douglass, S., Lebiere, C., & Qin, Y. (2004). An integrated theory of the mind. Psychological Review 111(4), 1036-1060.
- Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arxiv 2212.08073.
- Born, J. & Wilhelm, I. (2012). System consolidation of memory during sleep. Psychological Research 76(2), 192-203.
- Chhikara, P., Khant, D., Aryan, S., Singh, T., & Yadav, D. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arxiv 2504.19413.
- Demszky, D., Movshovitz-Attias, D., Ko, J., Cowen, A., Nemade, G., & Ravi, S. (2020). GoEmotions: A Dataset of Fine-Grained Emotions. Proc. ACL 2020. arxiv 2005.00547.
- Fang, J. et al. (2025). LightMem: Lightweight and Efficient Memory-Augmented Generation. arxiv 2510.18866.
- Freeman, E. & Gelernter, D. (1996). Lifestreams: A storage model for personal data. SIGMOD Record 25(1), 80-86.
- Goffman, E. (1959). The Presentation of Self in Everyday Life. Doubleday Anchor.
- Horvitz, E. (1999). Principles of mixed-initiative user interfaces. Proc. CHI 1999.
- Hu, Y. et al. (2025). Memory in the Age of AI Agents. arxiv 2512.13564.
- Hutto, C. J. & Gilbert, E. E. (2014). VADER: A parsimonious rule-based model for sentiment analysis of social media text. Proc. ICWSM-14.
- Jiang, B., Yuan, Y., Shen, M., et al. (2025). PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory. arxiv 2512.06688.
- Laird, J. E. (2012). The Soar Cognitive Architecture. MIT Press.
- Laird, J. E., Lebiere, C., & Rosenbloom, P. S. (2017). A Standard Model of the Mind. AI Magazine 38(4), 13-26.
- Lam, C., Li, J., Zhang, L., & Zhao, K. (2026). Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework. arxiv 2603.11768.
- Li, S. S. et al. (2026). HorizonBench: Long-Horizon Personalization with Evolving Preferences. arxiv 2604.17283.
- Lin, K. et al. (2025). Sleep-time Compute: Beyond Inference Scaling at Test-time. arxiv 2504.13171.
- Maharana, A. et al. (2024). Evaluating Very Long-Term Conversational Memory of LLM Agents (LoCoMo). arxiv 2402.17753.
- Marshall, L. et al. (2006). Boosting slow oscillations during sleep potentiates memory. Nature 444, 610-613.
- Mody, P. et al. (2026). CraniMem: Cranial Inspired Gated and Bounded Memory for Agentic Systems. arxiv 2603.15642.
- Murray, H. A. (1938). Explorations in Personality. Oxford University Press.
- Packer, C. et al. (2023). MemGPT: Towards LLMs as Operating Systems. arxiv 2310.08560.
- Park, J. S. et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. arxiv 2304.03442.
- Pennebaker, J. W., Boyd, R. L., Jordan, K., & Blackburn, K. (2015). The Development and Psychometric Properties of LIWC2015. University of Texas at Austin.
- Rasmussen, P., Paliychuk, P., Beauvais, T., Ryan, J., & Chalef, D. (2025). Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arxiv 2501.13956.
- Sharma, M. et al. (2023). Towards Understanding Sycophancy in Language Models. arxiv 2310.13548.
- Sumers, T. R., Yao, S., Narasimhan, K., & Griffiths, T. L. (2023). Cognitive Architectures for Language Agents (CoALA). arxiv 2309.02427.
- Vazire, S. & Carlson, E. N. (2010). Self-knowledge of personality: Do people know themselves? Social and Personality Psychology Compass 4(8), 605-620.
- Wei, L. et al. (2026). FadeMem: Biologically-Inspired Forgetting for Efficient Agent Memory. arxiv 2601.18642.
- Wu, D. et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. arxiv 2410.10813.
- Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., & Zhang, Y. (2025). A-MEM: Agentic Memory for LLM Agents. arxiv 2502.12110. NeurIPS 2025.
- Xu, Y., Chen, Q., Ma, Z., Liu, D., Wang, W., Wang, X., Xiong, L., & Wang, W. (2026). Toward Personalized LLM-Powered Agents. arxiv 2602.22680.
- Zhao, Z. (2026). Gradual Cognitive Externalization: From Modeling Cognition to Constituting It. arxiv 2604.04387.
- Zhou, C. et al. (2026). Externalization in LLM Agents. arxiv 2604.08224.
- Zhu, N. et al. (2026). SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering. arxiv 2604.11548.
- Zhu, Q., Chen, S., Yu, R., Wu, Z., & Wang, B. (2026). From Lossy to Verified: A Provenance-Aware Tiered Memory for Agents. arxiv 2602.17913.