In one sentence
ToolOps is to AI tools what a service mesh is to microservices — a framework-agnostic middleware SDK that upgrades any Python function with caching, resilience and observability, with zero changes to your business logic.
When you build AI agents, every external call — to an LLM, an API, a database — is a tool call. In production, those calls are expensive, unreliable and slow. ToolOps sits between your agent and its tools the way a service mesh sits between microservices, absorbing that complexity through a decorator interface deliberately chosen to be the thinnest possible integration surface.
1. Philosophy
The analogy that best describes ToolOps is the service mesh. Just as a service mesh — think Istio or Linkerd — sits between microservices to handle retries, timeouts and circuit breaking transparently, ToolOps sits between your AI agent and its tools. The application code knows nothing about the infrastructure layer beneath it.
ToolOps is to AI tools what a service mesh is to microservices.
This separation of concerns is intentional. Tool authors should focus on what a tool does, not on managing the distributed-system complexities of calling it reliably at scale. ToolOps absorbs that complexity through the decorator interface.
2. The production wall
Every agent developer hits the same wall when moving from demo to production. The symptoms are predictable: API bills that scale faster than usage, agents that crash on the third retry, request queues that bottleneck under any real concurrency load, and workflows that are completely opaque when they fail. ToolOps addresses each bottleneck at the infrastructure layer.
Problem → business impact → with ToolOps
- Redundant API calls. 10× cost spikes → 100 calls become 1 real call plus 99 cache hits.
- Similar queries phrased differently. Wasted LLM tokens → semantic match returns the same cached result.
- API instability. Agent crashes and retry loops → circuit breaker plus automatic retry.
- Concurrency bursts. Thundering herd → request coalescing collapses it to one real call.
- Sensitive data exposure. Tokens and PII in cache keys or logs → SHA-256 hashing plus automatic masking.
- Zero observability. Blind operations → structured JSON logs and OpenTelemetry traces.
Semantic caching
Near-duplicate calls resolve from cache — the agent retrying a paraphrased query stops costing money.
Circuit breakers
A failing dependency trips open and degrades gracefully instead of stalling the whole run.
Request coalescing
Concurrent calls to the same tool collapse into one upstream request, multicast to every caller.
Observability
Every call emits structured telemetry: hits, misses, retries, circuit-state transitions.
3. Installation
ToolOps is available on PyPI. Since v1.0.0 the package is fully batteries-included: a single install ships every cache backend — Memory, File, SQLite, Postgres, MySQL/MariaDB, Valkey/Redis, Semantic — plus the database drivers, the embedding and OpenAI integration libraries, and OpenTelemetry/Prometheus observability support. No optional extras are required.
# One command installs everything — backends, drivers, telemetry pip install toolops # Verify the installation toolops doctor
:: One command installs everything pip install toolops :: Using the launcher py -m pip install toolops
It is strongly recommended to install ToolOps within a virtual environment to avoid dependency conflicts.
# Create and activate (.venv) python -m venv .venv source .venv/bin/activate # Linux/macOS # .venv\Scripts\Activate.ps1 # Windows (PowerShell) # Install and verify pip install toolops toolops doctor
Legacy extras such as toolops[postgres],
toolops[semantic] or toolops[all] remain
available as empty compatibility aliases, so existing CI and deployment
scripts keep working — new installations should simply use
pip install --upgrade toolops.
Contributors need two more steps: a Docker environment with the full backend matrix, and the standardised Make targets that wrap the project's test, lint and format tooling.
# One-command setup with PostgreSQL
docker-compose up -d
docker-compose exec toolops make test
make test # Run full test suite with coverage make lint # Ruff + Black checks make format # Auto-format with Black make typecheck # mypy strict mode make coverage # HTML coverage report make clean # Remove all build artifacts
Since v1.0.0, pip install toolops is the complete
installation contract: the default dependency set includes the supported
production backends and integration libraries, so persistent caching and
distributed tracing work from day one. Contributors install from
requirements.txt, which adds the editable package and the
development toolchain.
4. Core architecture
The architecture has three layers: a decorator interface that sits on your tool functions, a composable middleware pipeline that orchestrates resilience and observability, and a pluggable backend system that handles storage and embedding. The three layers are entirely decoupled — you can swap backends or reorder middlewares without changing a single line of tool code.
A decorator, not a rewrite
- Async / await. Native, where
@lru_cachehas none. - Semantic cache. Vector embeddings, where
@lru_cacheonly matches exact keys. - Distributed cache. Postgres, SQLite, MySQL, Valkey/Redis, where
@lru_cacheis in-memory only. - Circuit breaker. Built in with backoff, where
@lru_cachehas none. - Request coalescing. Multicasts the result, where
@lru_cachelets a thundering herd through. - Stale-if-error fallback. Serves the last good value, where
@lru_cacheraises an exception. - Security. SHA-256 keys and automatic masking, native — where
@lru_cachehas none. - Observability. Structured OpenTelemetry and Prometheus telemetry, where
@lru_cachehas none. - AI-native. Native MCP and framework integrations, where
@lru_cacheis generic.
4.1 Decorators
ToolOps provides two decorators that map cleanly onto the two categories of tool operations. The distinction between read and write is a first-class concept: reads are idempotent and safe to cache and retry; writes are not — they get resilience patterns only, never automatic retries that could cause double submissions.
from toolops import readonly, cache_manager from toolops.cache import MemoryCache cache_manager.register("memory", MemoryCache(), is_default=True) @readonly( cache_backend="memory", cache_ttl=3600, retry_count=3, sensitive_params=["api_key", "auth_token"] ) async def get_market_data(ticker: str, api_key: str) -> dict: return await api.fetch(ticker, api_key=api_key) # Cached, retried, traced — api_key excluded from the cache key # and masked in logs. @sideeffect(circuit_breaker=True, timeout=5.0, retry_count=2) async def execute_trade(order: dict) -> bool: return await broker.submit(order) # No caching — protected by circuit breaker and timeout only.
4.2 Cache backends
Register backends once at application startup, then reference them by
name across all your decorators. Multiple backends can coexist — a fast
in-memory layer for hot data, a persistent SQL layer for audit trails, a
distributed Valkey/Redis layer for multi-process deployments, and a
semantic layer for NLP workloads. Since v1.0.0 every backend below is
installed by default, and all of them share identical lifecycle rules:
calling any operation on a closed backend raises a clear
RuntimeError instead of silently auto-reconnecting or
failing with an AttributeError — the same contract across
Memory, File, SQLite, Postgres, MySQL, Valkey/Redis, and Semantic.
from toolops import cache_manager from toolops.cache import ( MemoryCache, SQLiteCache, PostgresCache, ValkeyCache, MySQLCache ) cache_manager.register("memory", MemoryCache(), is_default=True) cache_manager.register("sqlite", SQLiteCache("toolops_cache.db")) cache_manager.register("db", PostgresCache("postgresql://user:pass@localhost:5432/mydb")) cache_manager.register("valkey", ValkeyCache(host="localhost", port=6379)) cache_manager.register("mysql", MySQLCache(dsn="mysql://root:secret@localhost:3306/myapp"))
Seven backends, one lifecycle
- MemoryCache. In-process — development, testing, single-process deployments.
- FileCache. Persistent, file system — lightweight local persistence without a database.
- SQLiteCache. Persistent, single file — serverless persistence via aiosqlite, with tag invalidation.
- PostgresCache. Persistent SQL — full audit trail, shareable across processes.
- MySQLCache. Persistent SQL — MySQL 8+ and MariaDB 10.5+ via aiomysql, DSN connection.
- ValkeyCache / RedisCache. Distributed, in-memory — full async connection pooling.
- SemanticCache. Vector embeddings — intent matching for NLP and RAG pipelines.
4.3 Middleware pipeline
The decorator you call is a thin wrapper: internally, ToolOps refactored
into a composable pipeline of independent middlewares, each handling one
concern, orchestrated in sequence by a ToolExecutor. A shared
ToolContext carries mutable state across the pipeline — and
the decorator behaviour is 100% preserved, so existing code using
@tool, @readonly, @sideeffect or
@stateful requires zero changes to benefit from it.
from toolops.middlewares import build_executor, DEFAULT_PIPELINE executor = build_executor(pipeline=DEFAULT_PIPELINE) # → [LoggingMiddleware, CacheMiddleware, CircuitBreakerMiddleware, # RetryMiddleware, CoalescingMiddleware, FallbackMiddleware]
LoggingMiddleware
Structured JSON logging for every tool call.
CacheMiddleware
Cache lookup, stale-if-error, cache write.
CircuitBreakerMiddleware
Circuit breaker protection.
RetryMiddleware
Retry loop with exponential backoff.
CoalescingMiddleware
Request coalescing — deduplication of concurrent calls.
FallbackMiddleware
Fallback execution on failure.
5. Resilience patterns
Beyond basic try/except blocks, ToolOps implements three deterministic patterns drawn from distributed-systems engineering. Together they ensure an agent never gets trapped in a failure loop, never exhausts its API budget on a degraded service, and never serves stale data when the upstream is healthy.
5.1 Circuit breaker
Stops all calls to a failing service after a configurable failure threshold. Once open, the circuit fails fast — returning immediately rather than waiting for a timeout — and enters a recovery window before attempting to re-establish the connection. This prevents a single failing tool from cascading into full agent failure.
@readonly( circuit_breaker=True, circuit_failure_threshold=5, # opens after 5 consecutive failures circuit_recovery_timeout=60 # retries after 60 seconds ) async def get_exchange_rates() -> dict: return await forex_api.fetch()
5.2 Stale-if-error
When an upstream service fails and no live data can be retrieved, ToolOps can automatically fall back to the last known good value from the cache — even past its normal TTL. The production equivalent of "serve something useful rather than crashing."
@readonly( cache_ttl=3600, stale_if_error=True, stale_ttl=86400 # serve stale data for up to 24h on failure ) async def get_exchange_rates() -> dict: return await forex_api.fetch()
5.3 Request coalescing
When multiple agent instances call the same tool simultaneously — a common pattern in multi-agent pipelines — ToolOps detects the in-flight request and holds subsequent callers until the first completes. The single real result is then multicast to all waiting callers.
50 → 1
Concurrent calls collapsed
98%
Reduction in credit consumption
0
Changes to the calling agents
In a benchmark with 50 concurrent agent calls to the same weather tool, request coalescing reduced upstream API calls from 50 to 1.
6. Semantic caching
Traditional caches operate on exact key equality. This works for deterministic systems, but agents are not deterministic systems — the same user intent surfaces in dozens of different phrasings. ToolOps uses vector embeddings to understand the meaning of a tool call, not just its literal arguments.
# Call 1 — cache miss, real API call query: "What is the status of invoice #442?" # Call 2 — semantic similarity 0.97 → cache hit query: "Check the current status for invoice 442" # Call 3 — semantic similarity 0.94 → cache hit query: "Invoice 442 — is it paid?"
from toolops.cache import SemanticCache, SentenceTransformerEmbedder embedder = SentenceTransformerEmbedder("all-MiniLM-L6-v2") semantic = SemanticCache(embedder=embedder, threshold=0.92) cache_manager.register("semantic", semantic) @readonly(cache_backend="semantic") async def ask_agent(query: str) -> str: return await llm.complete(query) # Reduces LLM latency by up to 90% on repeated intent patterns.
The similarity threshold — 0.92 above — is the primary tuning lever. Higher values require tighter semantic alignment before a hit is declared; lower values are more aggressive. The right value depends on how much variation is acceptable in your tool's input domain: a factual lookup tolerates a higher threshold than a creative-generation task.
Performance note
In v0.2.0, SemanticCache eviction was refactored from
O(n) list operations to O(1) collections.deque —
ensuring predictable latency and memory usage under high concurrency.
7. Observability
Debugging non-deterministic agent workflows requires instrumentation that goes deeper than application-level logging. ToolOps emits structured telemetry at every stage of the tool lifecycle — hits, misses, retries, circuit-state transitions — giving a complete audit trail without manual instrumentation.
from toolops import configure_opentelemetry, prometheus_metrics # Configure OpenTelemetry tracing — accepts any standard tracer instance configure_opentelemetry(tracer) # Expose Prometheus metrics as a raw text string metrics_string = prometheus_metrics()
Structured logs
Every cache hit, miss, failure, and retry is emitted as machine-readable JSON. Fields include tool name, backend, latency, cache key, and outcome — ready for ingestion by any log aggregator.
OpenTelemetry
Native OTEL traces and spans wrap every tool execution. Pass any standard tracer to configure_opentelemetry(tracer) and visualise the full call graph in Jaeger, Honeycomb, or Datadog.
Prometheus metrics
Real-time gauges for cache hit rate, circuit state (closed / open / half-open), and tool latency percentiles. Key metrics include toolops_cache_hits_total, toolops_tool_latency_seconds, and toolops_circuit_opens_total — ready to drive alerting rules and dashboards.
8. Ecosystem and MCP
ToolOps tools are plain Python functions. That design choice is not accidental — it means they work natively with every agent framework that accepts Python callables, with no adapter code and no framework-specific configuration.
Integrations, all available today
- LangChain / LangGraph. Built-in helper.
- CrewAI. Built-in helper.
- LlamaIndex. General compatibility.
- Model Context Protocol. Built-in helper.
- PydanticAI. General compatibility.
- AutoGPT and custom frameworks. Any Python callable.
The MCP integration deserves particular mention: a built-in adapter exposes any decorated tool as an MCP-compatible definition without writing a line of JSON Schema — so a resilient, production-grade tool is available to Claude Desktop, Cursor, or any MCP-compatible host instantly.
from toolops.integrations.mcp import MCPIntegration # get_weather is already decorated with @readonly definition = MCPIntegration.to_mcp_definition(get_weather) # → MCP-compatible tool definition, ready for Claude Desktop or Cursor.
9. CLI and operations
ToolOps ships with a command-line tool for inspecting and managing tool infrastructure in production — built for operators and CI pipelines, not just developers.
# List all available commands toolops --help # Check system health and backend readiness toolops doctor # View real-time cache statistics toolops stats --app my_app:setup_toolops # Print current metrics toolops metrics --app my_app:setup_toolops # Inspect a specific cache key toolops inspect-key --app my_app:setup_toolops # Clear a specific cache backend toolops clear postgres --app my_app:setup_toolops
toolops doctor is particularly useful in deployment
pipelines: it validates backend connectivity, checks embedding-model
availability, and reports circuit-breaker state — a readiness check you
can wire directly into your health endpoint.
10. Roadmap
v1.0.0 marks the stable, batteries-included release: a composable middleware pipeline, hardened security, a unified lifecycle across all cache backends, and a single-command install with no extras required. The following are planned for upcoming releases, ordered by expected delivery.
What's next
- Web dashboard. Real-time metrics, cost attribution and cache hit rates in a browser UI — no Prometheus or Grafana setup required.
- Budget control. Hard limits on tool-induced API costs per hour or per day, configurable per tool and per backend.
- Native MCP server. One-click deployment of ToolOps tools as a standalone MCP host — no Claude Desktop configuration required.
- Streaming middleware. Full support for streaming tool outputs in agent pipelines.
- New backends. ChromaDB and Pinecone support, extending the vector-native cache options. (MariaDB is already supported since v1.0.0 via MySQLCache.)
ToolOps is open source under Apache 2.0. Star the repository, open an issue, or submit a pull request — the project is built in the open, and roadmap priorities are shaped by real-world production use cases from the community.
1
Decorator to adopt
0
Framework lock-ins
7
Cache backends, batteries included
FAQ
Frequently asked questions
ToolOps is framework-agnostic. While LangChain or CrewAI have basic retry logic, ToolOps provides industrial-grade patterns like circuit breakers, request coalescing and semantic caching that work across any Python tool with zero migration cost.
ToolOps protects your system in three ways: circuit breakers stop the hammering, automatic retries handle transient blips, and stale-if-error fallback can serve the last known good value from the cache so your agent keeps moving.
If 50 agents call the same tool simultaneously during a cache miss, ToolOps executes the real API call once and multicasts the result to all 50 callers. This prevents overwhelming your upstream API rate limits.
No. Since v1.0.0, pip install toolops is
batteries-included: all standard cache backends, database drivers,
embedding libraries and telemetry support are installed by default.
Legacy extras remain as empty compatibility aliases so existing
scripts do not break.
Yes. ToolOps is designed as a foundation for Model Context Protocol servers and LangGraph stateful agents. It provides the industrial-grade infrastructure those frameworks lack natively.
v1.0.0 formalises ToolOps as a stable, batteries-included SDK. The
default install ships all cache backends and drivers — including
SQLiteCache, ValkeyCache/RedisCache and MySQLCache — plus a unified
backend lifecycle with clear closed-state errors and the
configure_opentelemetry / prometheus_metrics
observability exports. No application code changes are required.
Three layers: SHA-256 hashing of all cache keys ensures no
plaintext arguments — tokens, PII — ever appear in cache stores.
Automatic masking detects known sensitive keywords and replaces
their values with a redacted marker in structured logs. The
sensitive_params decorator argument lets you explicitly
exclude any parameter from cache key generation.
Running agents in production?
Try it on your flakiest tool first. Issues and war stories welcome.