Skip to main content
Add long-term memory to Hermes Agent, a self-improving AI agent CLI by Nous Research. The standalone Mem0 plugin learns facts from conversations and recalls relevant memories for the current question. You can run Mem0 in three ways:
  • Platform mode (default): managed Mem0 Cloud. Add your API key and you are ready.
  • Self-hosted server mode: point the plugin at a Mem0 server you run yourself (the Docker-shipped server). The plugin only talks HTTP to your server.
  • OSS mode: run Mem0 in-process with your own LLM, embedder, and vector store. No Mem0 server required.

How It Works

Hermes runs a built-in memory system (file-based MEMORY.md and USER.md) alongside one external provider. When Mem0 is active, it works additively with the built-in system at two points in every conversation turn.

1. Current-turn recall (bounded wait)

When you send a message, Hermes searches your stored memories for the current question and waits up to 3 seconds for results. If they arrive in time, they are injected into the system prompt so the model can see them. If the backend is slower, Hermes skips the injection and the model can still call mem0_search itself after the bounded recall wait.

2. Background fact extraction (sync)

Once the model finishes, the plugin sends the user message and assistant response to Mem0 in a background thread for fact extraction. Each write includes the agent identifier and gateway channel.
Automatic capture truncates each message to 450 characters by default in every mode, preferring a sentence boundary. Adjust sync_max_chars for your model’s context limit. Capture is best effort: if the previous sync is still running after a five-second wait, the next turn is skipped. Use mem0_add to store specific text verbatim.

Agent Tools

When Mem0 is active, the model gets four tools it can call during a conversation:

Installation

Install Hermes Agent with memory-provider plugin support and Python 3.11 or later. Once the plugin directory is available on Mem0’s main branch, install it from the repository subdirectory:
Select mem0 in setup and choose one of the modes below. Start a fresh Hermes conversation after setup. Hermes installers with plugin dependency support install mem0ai>=2.0.10,<3 and httpx>=0.27,<1 from the plugin’s pyproject.toml. Older hosts such as Hermes v0.21.3 require those packages to be installed explicitly into the Hermes Python environment. The OSS setup flow installs additional provider packages as needed.
Hermes versions that still bundle Mem0 prefer the bundled provider. Use a Hermes release that has completed the standalone-provider migration; installing this plugin alone does not replace the bundled implementation. Existing users should keep their current configuration; see Migration for existing users.
Run the setup wizard in an interactive terminal. On Hermes hosts whose hermes memory setup --help lists only a provider argument, options such as --mode, --host, and --oss-llm are rejected by Hermes before the plugin runs. Use hermes memory setup mem0 or the manual configuration below. Redirected input cannot select the mode picker and falls back to Platform.

Platform Setup

Platform mode uses managed Mem0 Cloud and is the fastest way to start.
Choose Platform and paste your API key when prompted. The wizard writes settings to $HERMES_HOME/mem0.json and keeps the key in that profile’s .env. The default Hermes home is ~/.hermes; named profiles use their own home directory.
Get your API key from app.mem0.ai.

Option 2: Manual Configuration

Add your key to the active Hermes profile’s .env:
Set these values in the active profile’s mem0.json, choosing a stable user identity:
Remove any stale MEM0_HOST from the environment and profile .env, and remove an old inline api_key from mem0.json so the new .env key is used. The config command sets memory.provider: mem0 in that profile’s config.yaml. Restart Hermes and check hermes memory status.

Self-Hosted Server Setup

Run the Mem0 server (FastAPI + pgvector) from its Docker image and point the plugin at it. Unlike OSS mode, the plugin just talks HTTP to your server.

Interactive

Manual configuration

Select mem0 with hermes config set memory.provider mem0. Set these values in the active profile’s mem0.json:
Add the server key to that profile’s .env:
Remove an old inline api_key from mem0.json so the .env key is used. MEM0_HOST can also supply the server URL, but a non-empty host in mem0.json overrides it. Keep mode set to platform for the HTTP server backend. Then start a fresh Hermes session and call mem0_search — it connects to your server. The plugin authenticates with X-API-Key and uses the server’s /search and /memories routes. The API key is optional only for servers running with AUTH_DISABLED.
Setting host routes to the self-hosted server automatically. Don’t combine it with mode: oss — OSS takes precedence and ignores host.

OSS (Self-Hosted) Setup

OSS mode runs the Mem0 SDK in the Hermes process with your chosen LLM, embedder, and vector store. It does not use Mem0 Cloud or require a Mem0 API key. Data goes to the model services you configure; use local Ollama models and local storage for a fully local setup.

Interactive

The wizard uses the listed default OpenAI models and local Qdrant storage. For custom OpenAI-compatible endpoints, deployment names, or a Qdrant server, use manual configuration below.

Supported providers

Manual configuration

Add the model key to the active profile’s .env:
Set the following in that profile’s mem0.json. Use your existing storage path when migrating; for a new named profile, choose a path inside that profile’s home.
For an OpenAI-compatible service such as Azure’s /openai/v1 endpoint, add OPENAI_BASE_URL to the profile’s .env and set each model to its deployed name. Both the LLM and embedder use this endpoint unless their config.openai_base_url overrides it. The main Hermes chat model is configured separately; this JSON configures Mem0’s extraction and embedding models. For a Qdrant server, replace vector_store.config.path with url, for example "url": "http://localhost:6333". Manual setup does not install optional provider dependencies: install qdrant-client, psycopg2-binary, or ollama in the Hermes Python environment as needed for your selected providers. Start a fresh session and verify a memory write and search; hermes memory status reports configuration availability, not a full backend health check. Desktop sessions in the same process and profile share local Qdrant storage when their OSS settings match. Operations are serialized, and storage closes after the last session releases it. Conflicting settings are rejected without changing existing memories; close active sessions before changing models or credentials. For concurrent CLI and Desktop processes, use a Qdrant server or the self-hosted Mem0 HTTP API instead of sharing a local directory.

Switching Modes

Run hermes memory setup mem0 in an interactive terminal and choose the new mode, or edit the active profile’s mem0.json using the examples above. Switching backends does not transfer memories between them. Preserve existing OSS storage paths when editing configuration. When returning to Platform, set mode to platform, clear host, and remove any stale MEM0_HOST setting from your environment and profile .env.

Configuration

Settings live in $HERMES_HOME/mem0.json and are written by hermes memory setup. API keys normally live in that profile’s .env; distinct OpenAI LLM/embedder keys and database credentials are stored in the OSS configuration. Setup writes these files atomically with owner-only permissions. When editing these files manually, restrict both .env and mem0.json to their owner (chmod 600 on Unix). Keep configuration and secrets in the same active Hermes profile. MEM0_MODE, MEM0_HOST, MEM0_USER_ID, and MEM0_AGENT_ID supply environment defaults. Non-empty values in mem0.json take precedence. MEM0_API_KEY supplies the Cloud or server key unless api_key is set in the file.

Cross-channel memories

Hermes can run from the CLI and from gateways like Telegram, Slack, and Discord. The user_id setting controls how memories are scoped across them:
  • Set a user_id other than hermes-user and it applies to every gateway, so one person gets a single merged memory store no matter where they talk to the agent.
  • Leave it unset (or at the default hermes-user) and each gateway uses its own native ID when available, falling back to hermes-user.
Every write is tagged with metadata.channel (for example telegram or cli). Plugin searches filter by user identity across sessions; they do not restrict recall to the current channel or session.

Migration for Existing Users

Keep memory.provider: mem0, mem0.json, MEM0_* settings, user identity, and OSS database paths unchanged. Moving from the bundled provider to this standalone plugin does not require rerunning setup or moving stored memories. Automatic migration depends on Hermes rollout as well as this repository:
  1. Users need a Hermes build containing PR #114569.
  2. Hermes maintainers must approve a catalog entry named mem0 with repo: https://github.com/mem0ai/mem0, subdir: integrations/hermes-plugin-mem0, and a reviewed full commit SHA.
  3. The bundled Mem0 provider must be removed so the standalone provider can load.
With these in place, Hermes installs a missing configured provider during hermes update across profiles or at agent startup. Startup installation respects security.allow_lazy_installs. Offline or disabled installation needs manual action; merging the plugin directory alone does not complete automatic migration. CLI setup and status are supported. This plugin does not ship a Desktop configuration panel or provider-specific CLI commands.

Reliability

  • Circuit breaker: five consecutive backend failures pause calls for two minutes. The agent can continue without memory during that window. Expected update/delete errors such as a missing memory do not trip the breaker.
  • Bounded waits: recall waits up to three seconds. Capture runs in the background, but an overlapping turn may wait up to five seconds for the previous sync before being skipped.
  • Graceful shutdown: shutdown and Python process exit wait for active recall and capture workers before closing the backend. Backend network timeouts still apply. Self-hosted HTTP capture uses a 120-second read timeout and a 30-second connection timeout; other self-hosted HTTP operations use 30 seconds.
  • Best-effort capture: there is no durable queue. Forced termination, including Hermes’ 30-second exit watchdog, can interrupt pending writes even while graceful shutdown is waiting.
  • OSS data protection: an embedding dimension mismatch fails initialization without deleting the existing collection or table.

Troubleshooting

”Mem0 temporarily unavailable”

The circuit breaker tripped after five consecutive failures and resets after two minutes.
  • Platform mode: check your API key and internet connection.
  • Self-hosted server mode: check that the server is running and reachable at the configured host URL.
  • OSS mode: make sure your vector store (Qdrant or PGVector) is running and reachable.

OSS: vector store connection refused

OSS: Ollama not reachable

Memories not appearing

  • mem0_add stores text verbatim with no extraction. Ordinary conversation turns are extracted automatically by the background sync.
  • Search is semantic, so try a broader query.
  • Confirm user_id is the same across sessions (check $HERMES_HOME/mem0.json).
  • Check sync_max_chars: facts beyond the per-message limit are not sent for extraction.

OSS: embedding dimension mismatch

Restore the embedding model and dimensions that created the existing collection, or choose a new collection and migrate data explicitly. The plugin leaves the original collection intact when dimensions differ.

OpenClaw Integration

Add memory to OpenClaw agents with auto-recall and auto-capture

Mem0 Platform

Get your API key and explore the Mem0 dashboard
Using Mem0? Star us on GitHub to help more developers discover memory for AI apps.