How AI Remembers: A Field Guide to Memory in LLM Apps, Agent Platforms, and Wrightery

How the leading conversational assistants and agent frameworks generate, keep, retrieve, and manage memory — and how Wrightery approaches the same problem for multi-user, multi-agent teams.
Who this is for. It's written in three layers so it works for everyone. Decision-makers get the "why it matters" and the comparison tables. Users get a plain-language picture of what these systems actually remember about you and how to control it. Engineers get the architecture — extraction pipelines, storage shapes, retrieval strategies, and the cost trade-offs behind each.
TL;DR
A chatbot without memory restarts from zero every session. Memory is what turns a clever demo into a product that compounds — it gets more useful the more you use it.
Designs still differ, but by 2026 they're converging on background extraction + consolidation: ChatGPT's "Dreaming" synthesizes memory offline, Gemini builds a personal-context profile, Mem0 runs a two-phase extract-and-reconcile pipeline. Claude keeps an agent-writes-its-own-files model; Zep builds a time-aware knowledge graph.
A useful cost insight: the leading assistants don't run "RAG" (vector search) every turn for their core memory. They inject the relevant slice directly. Vector search is the scale valve, not the default.
Wrightery composes all four memory modes onto one governed store — curated facts, verbatim recall, agent self-edit, and time-travel — as a multi-tenant, metered primitive inside a broader agent platform. Its design tracks where the 2026 frontier converged (ChatGPT's background "Dreaming," Claude's scoped managed-agent memory); the edge is integration + governance, not a unique capability.
1. Why memory matters
Talk to an assistant with no memory and you'll notice it immediately: it forgets your name, your preferences, the decision you made yesterday, and the mistake it made an hour ago. Every conversation starts cold. That's not just annoying — it caps how valuable the product can ever become.
Memory changes the economics:
For the business: memory is a compounding moat. A tool that remembers your context, your standards, and your history gets stickier and harder to switch away from every week. It's the difference between a stateless utility and a system of record.
For the user: less repetition, more relevance. The assistant behaves like a colleague who's been in the room the whole time, not a stranger you re-brief every morning.
For the engineer: memory is how you beat the context-window limit. You can't paste a year of history into every prompt — but you can distill it into a few durable facts and feed those back. Memory is context engineering made persistent.
The catch: memory done badly is worse than no memory. It can remember the wrong things, contradict itself, leak one user's private data into another user's session, or quietly balloon your costs. Good memory design is mostly about discipline — what to keep, how to reconcile it, and how to retrieve it cheaply and safely.
2. The mental model: memory has four jobs
Strip away the branding and every memory system — from ChatGPT to a research framework — is doing the same four things in a loop:
flowchart LR
C(["💬 Conversation"])
K["② Keep<br/>store + reconcile"]
S[("🧠 Memory<br/>store")]
M["④ Manage<br/>view · edit · forget · isolate"]
C ==>|"① Generate"| K
K ==> S
S ==>|"③ Retrieve"| C
M -.->|"governs"| S
classDef store stroke:#6366f1,stroke-width:2px;
classDef step stroke:#3b82f6,stroke-width:2px;
classDef gov stroke:#8b5cf6,stroke-width:2px;
class S store;
class K step;
class M gov;
① Generate — Turn a messy conversation into memory. Who decides what's worth remembering? Options: the user says "remember this," the system extracts facts automatically, or the agent writes its own notes.
② Keep — Store it, and reconcile it with what's already there. What structure? A flat list, a vector index, files, or a graph. How do you avoid duplicates and contradictions?
③ Retrieve — Bring the relevant memory back into a new conversation. How? Inject everything, search semantically, let the agent read on demand, or traverse a graph.
④ Manage — Keep humans and rules in control. Can the user see and delete it? Is one person's memory walled off from another's? Does stale memory expire?
We'll use these four verbs as the yardstick for every system below.
3. How the big assistants remember
ChatGPT (OpenAI) — from a curated list to background "dreaming"
ChatGPT splits memory in two: saved memories (an explicit, editable list of durable facts) and reference chat history (looser recall of past conversations, launched April 2025). Historically the saved list was small — reportedly ~8–12 facts, injected straight into the system prompt with no vector search — and written via an inline bio tool the model called mid-chat.
That changed in mid-2026. On June 4, 2026 OpenAI shipped "Dreaming V3", which replaced the hand-curated list with an asynchronous background synthesis process: during idle time it reads across years of conversations, consolidates them into structured "memory chains" with weighted relationships, and updates facts over time on its own (its example: "you're going to Singapore in July" rewrites itself to "you went to Singapore in July 2026" after the trip). At read time, memory relevant to your prompt is still injected as a hidden context block. In other words, ChatGPT moved from an inline-tool write to a background extract-and-consolidate pipeline — and picked up temporal awareness along the way. It stays auditable (a readable memory-summary page) and controllable (toggles + Temporary Chat).
Google Gemini — automatic recall of your past chats
Gemini leans toward automatic memory: it learns your preferences from past conversations without being asked and applies them to new ones ("you mentioned a comic book last week → it suggests a themed party around it"). It calls this Personal Context. Recall is automatic — ask about something you discussed before and it pulls from relevant past chats. It's on by default, with a settings toggle to turn it off, controls to delete conversations, and a Temporary Chat mode for private, unremembered sessions.
Claude (Anthropic) — a filesystem the agent operates
Claude's memory tool takes a different philosophy: instead of the platform distilling facts for the model, it gives the agent a directory of memory files it reads and writes with ordinary commands (view, create, str_replace, …). The agent decides what to jot down and, at the start of a task, lists its memory directory and reads what's relevant — "just-in-time" retrieval rather than always-on injection. Your application controls where the files live.
Then on April 23, 2026, Anthropic GA'd Memory for Claude Managed Agents for enterprise — and it added exactly the governance you'd expect from a platform: stores that are shared across multiple agents, per-user vs. org-wide scoping, read-only vs. read-write access scopes, and audit trails on every change. (Worth noting, because it closes much of the "governance gap" a platform like Wrightery used to own alone.)
Generate | Keep | Retrieve | Manage | |
|---|---|---|---|---|
ChatGPT | You say "remember," or it auto-extracts | Small curated fact list + implicit history index | Inject the list every turn; recall history implicitly | See / edit / delete list; toggle; Temporary Chat; auto-condense when full |
Gemini | Automatic from your chats | Personal Context derived from past conversations | Automatic recall of relevant past chats | On by default; toggle Personal Context; delete chats; Temporary Chat |
Claude | The agent writes notes via tools | Files in a directory (your storage) | Agent reads the directory on demand | Scoped permissions, audit logs, concurrent-safe |
The pattern to notice: the leading assistants keep their curated memory compact and inject the relevant slice at read time — rather than running a heavy vector search every turn. That's the cost insight the rest of this article builds on (even Dreaming, which synthesizes a lot in the background, still injects a selected subset).
4. How agent platforms remember
Consumer apps optimize for one user and zero setup. Agent frameworks optimize for scale, autonomy, and developer control — and they've explored richer storage shapes.
Mem0 — two-phase extract-and-reconcile
Mem0 is the reference design for "keep memory clean automatically." It runs a two-phase pipeline:
Extraction — an LLM reads a rolling window (the latest exchange, a running summary, the last few messages) and emits a concise set of candidate facts.
Update — for each candidate, it retrieves semantically similar existing memories and asks an LLM to choose one of four operations: ADD (new fact), UPDATE (refine/correct), DELETE (contradict), or NOOP (duplicate/irrelevant).
That reconciliation step is what stops memory from becoming an ever-growing, self-contradicting pile. Mem0 stores facts in a vector store, with an optional knowledge-graph layer for multi-hop questions.
Letta (formerly MemGPT) — memory as virtual memory
Letta treats the context window like RAM and memory like disk. A small core memory always sits in context; a larger archival memory lives outside and is paged in on demand. Crucially, the agent edits its own memory through function calls — it decides what's worth promoting or evicting. Maximum autonomy, at the cost of depending on the model's judgment.
Zep / Graphiti — a temporal knowledge graph
Zep argues that for real agents, time is a first-class dimension. Its Graphiti engine turns conversations into a temporal knowledge graph: entities become nodes, relationships become edges, and every edge carries two timestamps — when the fact became true (event time) and when the agent learned it (ingestion time) — plus a validity window and a confidence level. That lets the agent answer "what was true at the time" and cleanly supersede outdated facts, which flat stores struggle with. The trade-off is complexity and cost: graph builds are slower and more token-hungry than plain vector storage, for a modest accuracy gain on most benchmarks.
The field has borrowed a three-part taxonomy from cognitive science that's worth knowing:
Semantic memory — durable facts ("prefers ETFs," "benchmark is ACWI").
Episodic memory — what happened, and when ("on July 20 we shortlisted three REITs").
Procedural memory — how to do something ("the way this user likes the weekly report formatted").
5. Under the hood: how memory actually gets generated
We've said memory is "extracted" or "distilled" — but by what? This is the part most overviews skip, and it's where the real design choices live. The core question: who writes the memory, and with which model? There are two paradigms.
Paradigm A — the main model writes it inline (a tool call)
The same frontier model that's answering you decides, mid-conversation, to save a memory by calling a tool. There is no separate, cheaper model — the "prompt" is just a passage in the assistant's own system prompt telling it when to save.
Claude's memory tool and Letta/MemGPT are the clearest examples: the main agent calls memory commands (
create/str_replace) or functions (core_memory_append) itself.ChatGPT used to be here too — an inline
biotool (to=bio) the model called mid-chat. But as of Dreaming V3 (June 2026) it moved to Paradigm B — background synthesis (see below). Its history is the clearest sign of where the field is heading.
Trade-off: best judgment (a frontier model with the full conversation in view) and zero extra infrastructure — but it spends expensive frontier tokens on bookkeeping, adds work and latency to the chat turn, the model can simply forget to save, and writing memory inline opens a security surface: a prompt injection can plant false memories (demonstrated against ChatGPT).
Paradigm B — a separate, cheaper model runs an extraction pass
After the turn, a small, cheap model runs one or two tightly-scoped prompted calls. This is Mem0's design — and ours — and, as of 2026, where the frontier converged: ChatGPT's Dreaming V3 and Gemini's Personal Context both generate memory with background synthesis, not an inline tool. Mem0's reference implementation and its published benchmarks run on GPT-4o-mini for all memory operations; the newer default is gpt-5-mini. Two prompts do the work:
1. The extraction prompt — pull candidate facts from the recent window:
SYSTEM: You extract durable memories from a conversation. Return a JSON array of
short, self-contained facts worth remembering about the user and the task —
preferences, decisions, standing instructions, key facts. Ignore small talk and
one-off logistics. If nothing is worth saving, return [].
USER: <last exchange + rolling summary + last N messages>
→ ["Prefers ETFs over single stocks", "Max position size 5%", "Benchmark is ACWI"]
2. The update (consolidation) prompt — reconcile each candidate against what's already stored:
SYSTEM: You maintain a memory store. Given a NEW candidate fact and the most
SIMILAR existing memories, choose exactly one action:
ADD — a genuinely new fact
UPDATE — refines/corrects an existing one (give merged text + its id)
DELETE — contradicts an existing one (give its id)
NOOP — duplicate or irrelevant
Return JSON.
USER: candidate="Max position size 8%" existing=[{id:7,"Max position size 5%"}]
→ { "action": "UPDATE", "id": 7, "text": "Max position size 8%" }
Trade-off: cheap, fully async (never slows the chat turn), decoupled, and testable in isolation. The cost is a second call and a dependence on how good the window is.
So — "are they just using a cheaper LLM?" Yes, and on purpose.
For Paradigm B, the answer to your question is a flat yes: a cheap model with two specific prompts. And that's not a compromise — it's the right tool. Extraction and consolidation aren't open-ended reasoning; they're a narrow, well-specified task (pull salient facts, dedupe, resolve contradictions). Small models do that well, which is exactly why Mem0 benchmarks the whole system on gpt-4o-mini and still beats full-context baselines while cutting tokens and latency. Using a flagship model for it would be paying Ferrari prices to run errands. This is precisely why Wrightery routes memory through a dedicated cheap/fast tier (the catalog's default-for-memory model) rather than the flagship it uses for chat.
Is there a better way? A menu, by cost vs. quality
Approach | Cost | Quality | Who uses it |
|---|---|---|---|
No LLM (heuristics / embeddings only) | lowest | can't judge salience or rewrite | rarely used alone |
Cheap extractor model + prompts | low | strong for a narrow task | Mem0, Wrightery ← sweet spot |
Main model inline (tool call) | high (frontier tokens) | best judgment, but coupled to chat + injectable | Claude, Letta (ChatGPT until mid-2026) |
Fine-tuned small extractor | lowest at scale | best once you have data | an optimization on #2 |
Skip generation — store verbatim, retrieve later | none up front | perfect recall, no clean "profile" | transcript-recall systems |
There's no universal winner. If you optimize for scale and cost, a cheap extractor (Paradigm B) wins. If you optimize for judgment and don't mind the price, the main-model inline tool wins. If you need perfect recall, skip generation and lean on retrieval. Most mature systems end up hybrid — a cheap extractor for the volume, with a frontier model or a human reviewing the high-stakes memories.
The Wrightery memory-generation design
Here is that pipeline made concrete — at the design level. It is Paradigm B, run as a governed, async platform service so generation cost never lands on the user's chat turn.
flowchart LR
T["Conversation turn<br/>completes"] -. "enqueue · async" .-> Q(["Queue<br/>memory-jobs"])
Q --> W["memory-worker"]
W --> E["① Extract<br/>default-for-memory"]
E --> C["Candidate facts"]
C --> D["② Consolidate<br/>ADD·UPDATE·DELETE·NOOP"]
D --> S[("Memory store<br/>org × store × subject")]
S -. "similar items" .-> D
classDef store stroke:#6366f1,stroke-width:2px;
classDef async stroke:#f59e0b,stroke-width:2px;
class S store;
class Q,W,E,D async;
Trigger + debounce. A completed turn on an agent with a writable memory store enqueues a job and returns — nothing blocks the reply. Extraction is an LLM call, so it's debounced: run on conversation idle, every N turns, or on close — whichever first. A per-(conversation, store) cursor records the last processed message, so overlapping jobs are idempotent and never re-extract the same window.
Phase 1 — Extract (prompt memory.extract from the catalog). The store's own rubric (its instruction) steers what's worth keeping:
SYSTEM
You are {agent}'s memory extractor. From the recent conversation, extract
durable facts worth remembering for THIS user, following the store's rubric.
Rubric: {store.instruction}
e.g. "Remember risk tolerance, sector preferences, and standing
instructions; ignore day-to-day market chit-chat."
Rules:
- One self-contained fact per item (no dangling pronouns).
- Keep only durable/salient facts; skip greetings, acknowledgements,
one-off logistics.
- Return JSON: [{ "text": "...", "category":
"preference" | "fact" | "instruction" | "episodic" }]
- Nothing worth saving → [].
USER
{rolling window: last exchange + running summary + last N messages}
Phase 2 — Consolidate (prompt memory.consolidate). Candidates fan out to every writable store; for each store, each candidate is reconciled against the most similar existing items (retrieved from that store) so the memory never duplicates or contradicts itself:
SYSTEM
You maintain a user's memory store. For the NEW candidate and the SIMILAR
existing memories, choose ONE action and return JSON:
ADD { "action": "ADD" }
UPDATE { "action": "UPDATE", "id": "...", "text": "<merged fact>" }
DELETE { "action": "DELETE", "id": "..." }
NOOP { "action": "NOOP" }
USER
candidate: { "text": "Max position size 8%", "category": "preference" }
similar: [ { "id": "m_7", "text": "Max position size 5%" } ]
→ { "action": "UPDATE", "id": "m_7", "text": "Max position size 8%" }
Store + isolation. The resulting operations write MemoryItem rows scoped by organization × store × subject (the (agent, user) relationship by default; user / agent / org scopes for other granularities). Writes are soft — UPDATE/DELETE flip a status and keep an audit trail — so a bad extraction is recoverable and every memory is traceable to the conversation that produced it.
Why it's cheap. Both phases run on the catalog's default-for-memory tier (small + fast), routed through the shared LLM gateway — never a hardcoded provider SDK, so the model is swappable without code changes. A typical small per-user store needs no embeddings at all (Phase 2 compares against the whole set); embeddings only enter once a store grows large enough to need vector retrieval.
Cheap by default, upgradeable at every level. The model isn't fixed — it cascades most-specific-first: a store can pin its own model, an agent can pick one for its memory work, and an org admin can set the org-wide default, each falling back to the next down to the cheap platform default. (An org can surface this as a named mode — Economy / Balanced / Quality.) Every tier chooses from the same memory-tagged catalog models through the shared gateway. So a team that wants higher-fidelity memory just raises the tier; everyone else pays cheap-by-default.
Both prompts live in the platform's prompt catalog (memory.extract / memory.consolidate), so their wording is tunable by ops without a deploy. Room remains to swap in a fine-tuned small extractor (#4) or add a hybrid frontier/human review of high-stakes memories later. The full internal spec is in the platform's Memory design doc (see Sources).
6. The master comparison
Here's every system measured against the four jobs. Read it as a menu of choices, not a ranking — each design is right for someone's constraints.
System | Generate | Keep | Retrieve | Manage |
|---|---|---|---|---|
ChatGPT | User command + auto-extract | Small curated list (+ history index) | Inject list every turn | View/edit/delete; toggle; auto-condense |
Gemini | Automatic from past chats | Personal Context from history | Automatic recall of past chats | On by default; toggle; delete; temporary |
Claude | Agent self-writes (tools) | Files in a directory | Agent reads on demand | Scoped perms + audit logs; concurrent |
Mem0 | Two-phase LLM extraction | Vector store (+ optional graph); ADD/UPDATE/DELETE/NOOP | Vector search top-k | Programmatic; dedup keeps it lean |
Letta/MemGPT | Agent self-edits | Core (in-context) + archival (paged) | Agent pages in archival | Agent-managed; developer-configured |
Zep/Graphiti | Extract entities + relations | Temporal knowledge graph (bitemporal) | Graph traversal + semantic + keyword | Validity windows, supersession, time-travel |
Wrightery | Passive two-phase extraction (queue) + agent self-edit tools | Curated items + optional verbatim layer; ADD/UPDATE/DELETE/NOOP; bitemporal soft-delete | Tiered: inject → vector at scale → tool (+ transcript recall) | Org × store × subject isolation; reusable, auditable, metered |
Storage architectures, compared
The "Keep" column hides the deepest design fork. Four families:
Architecture | Used by | Strengths | Weaknesses |
|---|---|---|---|
Curated fact list (injected) | ChatGPT, Wrightery (default) | Cheap, transparent, always-on, easy to audit | Bounded size; loses un-distilled detail |
Vector store (RAG) | Mem0, Wrightery (at scale) | Scales to huge memory; semantic recall | Per-query cost; can miss or over-fetch |
Agent-managed files | Claude, Letta archival | Flexible, agent-structured, great for procedural notes | Depends on model judgment; harder to govern |
Temporal knowledge graph | Zep/Graphiti, Mem0ᵍ | Models change-over-time + relationships; "what was true when" | Complex, slower, more expensive to build |
7. How Wrightery remembers
Described here at the design level — this is the architecture Wrightery's memory is built around, borrowing the proven pieces from the systems above.
Wrightery treats memory as a first-class Context — the same primitive an agent already uses to read a database or a document library — with one difference: a Memory Context is written by the agent's own conversations and read back later. The design deliberately borrows the best-proven idea from each system above and adds the governance a multi-user platform requires.
① Generate — passive two-phase extraction, off the critical path. When a conversation turn completes, the platform queues a memory job and returns immediately — the user's reply is never slowed down by memory work. A background worker then runs Mem0's two phases: a memory model extracts candidate facts from a rolling window, then reconciles each one against existing memories with ADD / UPDATE / DELETE / NOOP. Chit-chat produces nothing; only durable, salient facts are kept.
② Keep — curated, distilled, auditable items. Each memory is a human-readable fact with a category (preference / fact / instruction / episodic), a link back to the conversation that produced it, and a soft lifecycle — updates and deletes flip a status rather than erasing, so the store stays auditable and recoverable.
③ Retrieve — cheap by default, RAG only at scale. This is the cost answer. Because a curated per-user store is small (dozens of facts, not thousands), Wrightery — like ChatGPT — injects the whole set directly into the prompt with no embedding and no vector search. Only when a store grows past a token budget does it switch to vector search for the top matches. A model can also explicitly search memory with a tool when it wants to dig deeper. Three tiers, cheapest first:
flowchart TB
Q["User turn needs memory"] --> D{"Store fits the<br/>inject budget?"}
D -- "yes · the common case" --> T1["<b>Tier 1</b> · inject the whole set<br/>no embedding, no search — free"]
D -- "no · large store" --> T2["<b>Tier 2</b> · vector search top-k<br/>the scale valve"]
T1 --> A["Grounded reply"]
T2 --> A
T3["<b>Tier 3</b> · memory_search tool<br/>on demand — zero cost unless called"] -. "any size" .-> A
classDef free stroke:#10b981,stroke-width:2px;
classDef paid stroke:#f59e0b,stroke-width:2px;
classDef opt stroke:#3b82f6,stroke-width:2px;
class T1 free;
class T2 paid;
class T3 opt;
④ Manage — what a single-user app never has to solve. Because one Wrightery agent serves a whole organization of users, memory is isolated on three dimensions at once — organization × memory-store × subject. A store's subject scope decides whose memory it holds — and the default is the unit other platforms miss: agent-user (what this agent knows about this user, from their conversations together — the relationship, not one global profile). A person who talks to two of your agents builds two independent memories. Also available: user (a cross-agent user profile, the ChatGPT model), agent (the agent's own, user-agnostic memory), or org (team knowledge). An agent commonly attaches two — its default agent↔user store and its own agent store. Every memory is browsable and editable, memory is a reusable object you can attach to many agents, and each attachment carries two simple switches — Read and Write — so a memory can be read-only, write-only, or both.
Four modes, one store
The baseline above is the semantic store — curated facts, injected cheaply. But the design composes four capabilities, each opt-in per store, so one Memory Context can be all four at once (the frontier labs now offer most of these individually — see the scorecard):
Mode | What it adds | Whose lead it matches | How |
|---|---|---|---|
1 · Semantic | Durable curated facts, always considered | ChatGPT saved memories |
|
2 · Episodic | Recall the exact words, even months later | ChatGPT reference-history, Gemini | An optional layer that chunks + embeds the raw messages, retrieved on demand |
3 · Active | The agent jots or fixes memory mid-task | Claude memory tool, Letta |
|
4 · Temporal | Answer "what was true as of date T" | Zep/Graphiti | Bitemporal |
All four sit behind the same organization × store × subject firewall, are auditable in the Record Page, and meter through the credits system. Composing them as one coherent primitive — rather than four bolt-ons — is the point; the frontier labs now offer most of these pieces individually (ChatGPT's Dreaming covers semantic + temporal; Claude's managed agents cover active + scoped stores). (These are design-stage, phased after the semantic baseline.)
8. Where Wrightery stands — an honest scorecard
No hand-waving. Here's where the design leads, ties, and trails.
Capability | Leaders | Wrightery |
|---|---|---|
Always-on curated memory, cheap to serve | ChatGPT | ✅ Matches (Tier-1 injection) |
Automatic, principled de-duplication | Mem0 | ✅ Stronger — explicit ADD/UPDATE/DELETE/NOOP + audit trail |
Cost discipline (no RAG until needed) | ChatGPT, Gemini | ✅ Matches, made explicit as tiers |
One shared agent, many users — memory walled per user and per org | consumer apps are single-user; frameworks leave org tenancy to you | ✅ org × store × subject, built in |
Shared / team memory | rare | ✅ Opt-in shared store |
Memory as a reusable object, attachable to many agents with access scopes | Claude (managed agents, Apr 2026) | ✅ Matches — reusable Context + read/write toggles |
Auditable, editable by end users | ChatGPT, Gemini, Claude | ✅ Matches |
Usage metering / billing (per-call credits) | none | ✅ By design — rides the platform credits system |
Raw recall of everything you ever said | ChatGPT, Gemini | ✅ By design — optional episodic layer indexes the raw messages for verbatim recall, per-user-isolated |
Agent writing its own notes mid-task | Claude, Letta | ✅ By design — |
Time-travel / "what was true when" | Zep/Graphiti, ChatGPT (Dreaming) | ✅ By design — bitemporal |
The honest verdict (mid-2026). When this comparison began, Wrightery's governed, multi-user, four-mode design looked like a superset. The frontier has since converged on the same ideas: ChatGPT's Dreaming V3 (June 2026) moved memory to background synthesis with idle-time consolidation and temporal self-updating facts; Claude's managed-agent memory (April 2026) added stores shared across agents, per-user vs. org scoping, read/write access, and audit trails. So Wrightery's design is now in line with where the best systems landed — not ahead of them on raw capability. Its durable edge is narrower and more honest: memory here is one primitive inside a broader agent platform — composed with Contexts, Capabilities, and Workspaces, isolated per tenant, and metered through the same credit system as every other action — rather than a standalone memory feature. If you want a drop-in memory layer, the frontier labs are excellent. If you're building governed, metered, multi-user agents end-to-end, the value is that memory is a native, consistent part of the same platform — and that its design tracks the proven frontier patterns instead of diverging from them.
9. Principles of good memory design
If you take nothing else away, take these — they hold across every system:
Distill, don't dump. Store meaningful facts, not raw transcript. (But keep verbatim recall as an escape hatch for the long tail.)
Reconcile on write. Never blindly append. Merge, correct, or drop — or your memory contradicts itself within a week.
Retrieve as cheaply as the size allows. Small memory → just inject it. Reach for vector search only when you must. RAG is a scale valve, not a starting point.
Make it transparent. Users should see, edit, and delete what's remembered about them. Opaque memory erodes trust fast.
Isolate ruthlessly. In any multi-user product, one person's memory leaking into another's session isn't a bug — it's a privacy incident.
Keep writes off the hot path. Extraction is expensive; do it in the background so conversations stay fast.
10. What to take away
If you're a decision-maker: memory is the feature that makes an AI product compound. Evaluate any vendor on all four jobs — especially Manage (isolation, auditability, control), because that's where consumer-grade tools fall short for real businesses.
If you're a user: these systems remember more than you might think. Learn where the memory settings are, check what's stored about you, and use "temporary" modes when you want a clean slate.
If you're an engineer: you rarely need a fancy store on day one. Start with a curated, injected fact list and a two-phase extractor; add vector search when a store outgrows the prompt, and a knowledge graph only when time and relationships genuinely matter. Match the machinery to the memory's size, not to the hype.
Sources
Consumer assistants
OpenAI — How "Reference saved memories" works · Memory and new controls for ChatGPT · the (historical)
biotool: TheBigPromptLibrary · Dreaming V3 (Jun 2026): TechTimes · Enterprise DNAGoogle — The Gemini app can now recall past chats · Temporary Chats & privacy controls
Anthropic — Memory tool docs · Memory for Claude Managed Agents · enterprise GA (Apr 2026): testingcatalog · SD Times
Agent platforms & research
Mem0 — Building Production-Ready AI Agents with Scalable Long-Term Memory (paper) · Long-Term Memory for AI Agents · source (extraction + update prompts)
Letta / MemGPT — Mem0 vs Letta
Zep / Graphiti — Zep: A Temporal Knowledge Graph Architecture for Agent Memory (paper) · What is a Temporal Knowledge Graph?
Overview — AI Agent Memory Architectures · counterpoint: Verbatim Chunks Beat Extracted Artifacts
Your next report, with a team behind it.
Wrightery is in private build. Join the waitlist and we'll bring you in as soon as a seat opens — then describe one deliverable you already produce, and watch your AI analysts research it, cross-check it and draft it for your review.
No credit card. We'll email you the moment your seat is ready.
Rather talk it through? Book a 30-minute meeting ↗