AsterMind MCP
Cut the tokens you send to an LLM, on your own machine.
AsterMind MCP is a Model Context Protocol server that ranks and filters your retrieved context locally, so you send the model the few passages that actually answer the question instead of everything your retriever returned. No network calls, no API keys, and no data leaves the machine.
npx @astermind/astermind-mcpHow it saves tokens
It saves tokens one specific way: you retrieve broadly, then send the model only the passages that matter. It doesn't compress prompts or shrink a model's vocabulary. If your pipeline puts 8 to 12 retrieved chunks into every prompt, ranking them locally and sending 3 is a direct, repeatable cut to your input tokens.
Measured, not promised
On our published benchmark (six hand-built scenarios: a support knowledge base, API docs, an HR policy, an e-commerce FAQ, a DevOps runbook and a fintech help center), keeping the top 3 of 10 retrieved passages:
| Context tokens | 926 → 307 (66.8% fewer) |
| Top-ranked passage relevant | 6 of 6 scenarios |
| A passage that answers the question kept | 6 of 6 scenarios |
| All relevant passages kept | 55.6% |
The tradeoff is yours to choose
Keep fewer passages and you save more tokens, but risk dropping one you needed:
| Keep top K of 10 | Token saving | Answer kept | All relevant kept |
|---|---|---|---|
| 1 | 87.9% | 6 of 6 | 38.9% |
| 3 (default) | 66.8% | 6 of 6 | 55.6% |
| 5 | 45.9% | 6 of 6 | 75% |
| 6 | 36.7% | 6 of 6 | 75% |
For one-fact questions, keep the top 1 to 3. For questions that need several passages, keep 5 or 6.
This is a small benchmark (60 single-sentence passages, relevance labeled by us, tokens counted with an OpenAI tokenizer). Reproduce it yourself with npm run benchmark.
Set it up
You need Node.js 18 or later. There's nothing to install by hand: you give your AI app one command, and the first time the app starts the server, npx downloads it from npm (about 94 MB, once). After that it runs on your machine, and the server itself makes no network calls.
Every app below starts the same local (stdio) server: the command npx with the arguments -y @astermind/astermind-mcp. No GPU, no build step, no API key.
Claude Desktop (Mac and Windows)
- Open Settings → Developer → Edit Config. That opens
claude_desktop_config.json, and creates it if it isn't there yet. On a Mac it lives in~/Library/Application Support/Claude/, on Windows in%APPDATA%\Claude\. - Add AsterMind under
mcpServers. If the file already lists other servers, add the"astermind"entry beside them:
{
"mcpServers": {
"astermind": {
"command": "npx",
"args": ["-y", "@astermind/astermind-mcp"]
}
}
}- Quit Claude completely and open it again.
- To check, click + at the bottom left of the chat box, then Connectors → Manage connectors. You'll see
astermindwith its 10 tools.
Claude Code
Run this in the project where you want it:
claude mcp add astermind -- npx -y @astermind/astermind-mcpTo use it in every project, add --scope user after add. Check with claude mcp list, which shows it as connected. The very first check can fail while npx downloads the package; run it again.
Codex, and the ChatGPT desktop app
One command adds it:
codex mcp add astermind -- npx -y @astermind/astermind-mcpOr add it to ~/.codex/config.toml yourself (on Windows, %USERPROFILE%\.codex\config.toml). The longer start-up timeout gives the first download time to finish:
[mcp_servers.astermind]
command = "npx"
args = ["-y", "@astermind/astermind-mcp"]
startup_timeout_sec = 30Check with codex mcp list, or type /mcp in a Codex session. The Codex IDE extension and the ChatGPT desktop app, which runs Codex, read the same file.
OpenAI Agents SDK (Python)
Install it with pip install openai-agents and set OPENAI_API_KEY. The server runs on your machine, beside your agent:
import asyncio
from agents import Agent, Runner
from agents.mcp import MCPServerStdio
async def main() -> None:
async with MCPServerStdio(
name="astermind",
params={
"command": "npx",
"args": ["-y", "@astermind/astermind-mcp"],
},
client_session_timeout_seconds=30, # the default is 5 s; the first start downloads the package
) as server:
agent = Agent(
name="Assistant",
instructions="Use the AsterMind tools when they help.",
mcp_servers=[server],
)
result = await Runner.run(agent, "List the tools you can use.")
print(result.final_output)
asyncio.run(main())OpenAI Agents SDK (TypeScript)
Install it with npm install @openai/agents zod, then:
import { Agent, run, MCPServerStdio } from "@openai/agents";
async function main() {
const server = new MCPServerStdio({
name: "astermind",
command: "npx",
args: ["-y", "@astermind/astermind-mcp"],
});
await server.connect();
try {
const agent = new Agent({
name: "Assistant",
instructions: "Use the AsterMind tools when they help.",
mcpServers: [server],
});
const result = await run(agent, "List the tools you can use.");
console.log(result.finalOutput);
} finally {
await server.close();
}
}
main().catch(console.error);ChatGPT in the browser, and the MCP tool in OpenAI's Responses API, connect only to servers on the internet. They can't start a program on your computer, so use the ChatGPT desktop app, Codex or the Agents SDK above.
Cursor
Add this to ~/.cursor/mcp.json for every project, or to .cursor/mcp.json for one project, then restart Cursor. The server and its tools then appear in Cursor's MCP settings.
{
"mcpServers": {
"astermind": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@astermind/astermind-mcp"]
}
}
}VS Code (GitHub Copilot)
Add this to .vscode/mcp.json in your workspace, or run MCP: Open User Configuration from the Command Palette to use it everywhere. VS Code's top-level key is servers, not mcpServers.
{
"servers": {
"astermind": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@astermind/astermind-mcp"]
}
}
}Or add it to VS Code in one click. Check with MCP: List Servers.
LM Studio (local models)
In LM Studio 0.3.17 or later, open the Program tab in the right sidebar and choose Install → Edit mcp.json. Add this block and save. LM Studio loads it straight away and asks before each tool call.
{
"mcpServers": {
"astermind": {
"command": "npx",
"args": ["-y", "@astermind/astermind-mcp"]
}
}
}Cline
Open Cline's MCP servers settings, choose to edit the configuration, and add the same block as for LM Studio above. In the Cline CLI, cline mcp walks you through it.
Goose
In Goose Desktop, open Extensions → Add custom extension, choose Standard IO, and enter the command npx -y @astermind/astermind-mcp. In the Goose CLI, run goose configure, then Add Extension → Command-line Extension with the same command.
Any other MCP app
Add a local (stdio) server with the command npx and the arguments -y @astermind/astermind-mcp.
If it doesn't start
- "npx: command not found": install Node.js 18 or later, then restart your app.
- The first start times out: the first launch downloads the package, which can take longer than an app's start-up timeout. Start the app again, or give it longer:
MCP_TIMEOUT=60000 claudefor Claude Code,startup_timeout_secfor Codex, orclient_session_timeout_secondsfor the Agents SDK. - See its tools yourself: the MCP Inspector (it needs Node.js 22.19 or later) lists all 10:
npx -y @modelcontextprotocol/inspector npx @astermind/astermind-mcp
The 10 tools, rated honestly
| Tool | What it does | How good it is |
|---|---|---|
| rerank_documents | Scores and orders passages against a query | Strong: the core of the product |
| filter_context | Keeps the top passages | Strong |
| compress_context | Ranks, keeps the top passages, and returns one context block | Strong |
| count_tokens | Counts tokens (OpenAI tokenizer) | Exact for OpenAI models |
| estimate_savings | Before-and-after token counts for a filtering choice | Exact |
| semantic_search | Ranks a collection against a query | Good (keyword-based) |
| detect_language | Identifies the language of a text (6 European languages) | Good, with a confidence flag |
| classify_text | Labels text into categories you supply | Weak: a low-confidence hint, flagged |
| generate_embeddings | A character-level vector for a text | For near-duplicates only |
| compare_texts | Similarity between two texts | For near-duplicates only |
Known limitations (stated on purpose)
- Ranking is keyword-based. It matches shared terms, so a pure synonym can be missed: a search for "payment methods" won't strongly rank a passage that only says "Visa/Mastercard."
- Classification is a hint, not a decision.
- Embeddings are character-level: good for spotting near-duplicates, not deep meaning.
- A token saving only counts if the answer survives, and the benchmark measures both.
What it's built on
AsterMind MCP is built on AsterMind's open-source Community Edition library. Its ranking tools use keyword-based scoring; its classification and language detection use Extreme Learning Machines, a fast form of machine learning explained on the Brain page.
Two ways to use it
| AsterMind MCP (free) | AsterMind MCP Pro | |
|---|---|---|
| What it is | The MCP server, running on your machine | A hosted API at mcp.astermind.ai |
| Price | Free, MIT license | $19 a month |
| Where your text goes | Nowhere: it stays on your machine | To our hosted API on AWS, which processes it and doesn't store it |
| Tools | All 10 | All 10, and every call reports the exact tokens saved |
| Get it | npx @astermind/astermind-mcp | Subscribe to Pro |
The free server is available on npm.