ToolOps: The Service Mesh for AI Tools

In one sentence

ToolOps is to AI tools what a service mesh is to microservices — a framework-agnostic middleware SDK that upgrades any Python function with caching, resilience and observability, with zero changes to your business logic.

When you build AI agents, every external call — to an LLM, an API, a database — is a tool call. In production, those calls are expensive, unreliable and slow. ToolOps sits between your agent and its tools the way a service mesh sits between microservices, absorbing that complexity through a decorator interface deliberately chosen to be the thinnest possible integration surface.

1. Philosophy

The analogy that best describes ToolOps is the service mesh. Just as a service mesh — think Istio or Linkerd — sits between microservices to handle retries, timeouts and circuit breaking transparently, ToolOps sits between your AI agent and its tools. The application code knows nothing about the infrastructure layer beneath it.

ToolOps is to AI tools what a service mesh is to microservices.
The design premise

This separation of concerns is intentional. Tool authors should focus on what a tool does, not on managing the distributed-system complexities of calling it reliably at scale. ToolOps absorbs that complexity through the decorator interface.

2. The production wall

Every agent developer hits the same wall when moving from demo to production. The symptoms are predictable: API bills that scale faster than usage, agents that crash on the third retry, request queues that bottleneck under any real concurrency load, and workflows that are completely opaque when they fail. ToolOps addresses each bottleneck at the infrastructure layer.

Problem → business impact → with ToolOps

  • Redundant API calls. 10× cost spikes → 100 calls become 1 real call plus 99 cache hits.
  • Similar queries phrased differently. Wasted LLM tokens → semantic match returns the same cached result.
  • API instability. Agent crashes and retry loops → circuit breaker plus automatic retry.
  • Concurrency bursts. Thundering herd → request coalescing collapses it to one real call.
  • Sensitive data exposure. Tokens and PII in cache keys or logs → SHA-256 hashing plus automatic masking.
  • Zero observability. Blind operations → structured JSON logs and OpenTelemetry traces.

Semantic caching

Near-duplicate calls resolve from cache — the agent retrying a paraphrased query stops costing money.

Circuit breakers

A failing dependency trips open and degrades gracefully instead of stalling the whole run.

Request coalescing

Concurrent calls to the same tool collapse into one upstream request, multicast to every caller.

Observability

Every call emits structured telemetry: hits, misses, retries, circuit-state transitions.

3. Installation

ToolOps is available on PyPI. Since v1.0.0 the package is fully batteries-included: a single install ships every cache backend — Memory, File, SQLite, Postgres, MySQL/MariaDB, Valkey/Redis, Semantic — plus the database drivers, the embedding and OpenAI integration libraries, and OpenTelemetry/Prometheus observability support. No optional extras are required.

Install and verify Shell
# One command installs everything — backends, drivers, telemetry
pip install toolops

# Verify the installation
toolops doctor
Windows Shell
:: One command installs everything
pip install toolops

:: Using the launcher
py -m pip install toolops

It is strongly recommended to install ToolOps within a virtual environment to avoid dependency conflicts.

Virtual environment setup Shell
# Create and activate (.venv)
python -m venv .venv
source .venv/bin/activate  # Linux/macOS
# .venv\Scripts\Activate.ps1 # Windows (PowerShell)

# Install and verify
pip install toolops
toolops doctor

Legacy extras such as toolops[postgres], toolops[semantic] or toolops[all] remain available as empty compatibility aliases, so existing CI and deployment scripts keep working — new installations should simply use pip install --upgrade toolops.

Contributors need two more steps: a Docker environment with the full backend matrix, and the standardised Make targets that wrap the project's test, lint and format tooling.

Docker development environment Contributors
# One-command setup with PostgreSQL
docker-compose up -d
docker-compose exec toolops make test
Makefile — standardised commands Contributors
make test      # Run full test suite with coverage
make lint      # Ruff + Black checks
make format    # Auto-format with Black
make typecheck # mypy strict mode
make coverage  # HTML coverage report
make clean     # Remove all build artifacts

Since v1.0.0, pip install toolops is the complete installation contract: the default dependency set includes the supported production backends and integration libraries, so persistent caching and distributed tracing work from day one. Contributors install from requirements.txt, which adds the editable package and the development toolchain.

4. Core architecture

The architecture has three layers: a decorator interface that sits on your tool functions, a composable middleware pipeline that orchestrates resilience and observability, and a pluggable backend system that handles storage and embedding. The three layers are entirely decoupled — you can swap backends or reorder middlewares without changing a single line of tool code.

A decorator, not a rewrite

  • Async / await. Native, where @lru_cache has none.
  • Semantic cache. Vector embeddings, where @lru_cache only matches exact keys.
  • Distributed cache. Postgres, SQLite, MySQL, Valkey/Redis, where @lru_cache is in-memory only.
  • Circuit breaker. Built in with backoff, where @lru_cache has none.
  • Request coalescing. Multicasts the result, where @lru_cache lets a thundering herd through.
  • Stale-if-error fallback. Serves the last good value, where @lru_cache raises an exception.
  • Security. SHA-256 keys and automatic masking, native — where @lru_cache has none.
  • Observability. Structured OpenTelemetry and Prometheus telemetry, where @lru_cache has none.
  • AI-native. Native MCP and framework integrations, where @lru_cache is generic.

4.1 Decorators

ToolOps provides two decorators that map cleanly onto the two categories of tool operations. The distinction between read and write is a first-class concept: reads are idempotent and safe to cache and retry; writes are not — they get resilience patterns only, never automatic retries that could cause double submissions.

Decorators — read vs. write Python
from toolops import readonly, cache_manager
from toolops.cache import MemoryCache

cache_manager.register("memory", MemoryCache(), is_default=True)

@readonly(
    cache_backend="memory", cache_ttl=3600,
    retry_count=3, sensitive_params=["api_key", "auth_token"]
)
async def get_market_data(ticker: str, api_key: str) -> dict:
    return await api.fetch(ticker, api_key=api_key)
# Cached, retried, traced — api_key excluded from the cache key
# and masked in logs.

@sideeffect(circuit_breaker=True, timeout=5.0, retry_count=2)
async def execute_trade(order: dict) -> bool:
    return await broker.submit(order)
# No caching — protected by circuit breaker and timeout only.

4.2 Cache backends

Register backends once at application startup, then reference them by name across all your decorators. Multiple backends can coexist — a fast in-memory layer for hot data, a persistent SQL layer for audit trails, a distributed Valkey/Redis layer for multi-process deployments, and a semantic layer for NLP workloads. Since v1.0.0 every backend below is installed by default, and all of them share identical lifecycle rules: calling any operation on a closed backend raises a clear RuntimeError instead of silently auto-reconnecting or failing with an AttributeError — the same contract across Memory, File, SQLite, Postgres, MySQL, Valkey/Redis, and Semantic.

Registering backends Python
from toolops import cache_manager
from toolops.cache import (
    MemoryCache, SQLiteCache, PostgresCache, ValkeyCache, MySQLCache
)

cache_manager.register("memory", MemoryCache(), is_default=True)
cache_manager.register("sqlite", SQLiteCache("toolops_cache.db"))
cache_manager.register("db", PostgresCache("postgresql://user:pass@localhost:5432/mydb"))
cache_manager.register("valkey", ValkeyCache(host="localhost", port=6379))
cache_manager.register("mysql", MySQLCache(dsn="mysql://root:secret@localhost:3306/myapp"))

Seven backends, one lifecycle

  • MemoryCache. In-process — development, testing, single-process deployments.
  • FileCache. Persistent, file system — lightweight local persistence without a database.
  • SQLiteCache. Persistent, single file — serverless persistence via aiosqlite, with tag invalidation.
  • PostgresCache. Persistent SQL — full audit trail, shareable across processes.
  • MySQLCache. Persistent SQL — MySQL 8+ and MariaDB 10.5+ via aiomysql, DSN connection.
  • ValkeyCache / RedisCache. Distributed, in-memory — full async connection pooling.
  • SemanticCache. Vector embeddings — intent matching for NLP and RAG pipelines.

4.3 Middleware pipeline

The decorator you call is a thin wrapper: internally, ToolOps refactored into a composable pipeline of independent middlewares, each handling one concern, orchestrated in sequence by a ToolExecutor. A shared ToolContext carries mutable state across the pipeline — and the decorator behaviour is 100% preserved, so existing code using @tool, @readonly, @sideeffect or @stateful requires zero changes to benefit from it.

The default pipeline Python
from toolops.middlewares import build_executor, DEFAULT_PIPELINE

executor = build_executor(pipeline=DEFAULT_PIPELINE)
# → [LoggingMiddleware, CacheMiddleware, CircuitBreakerMiddleware,
#    RetryMiddleware, CoalescingMiddleware, FallbackMiddleware]

LoggingMiddleware

Structured JSON logging for every tool call.

CacheMiddleware

Cache lookup, stale-if-error, cache write.

CircuitBreakerMiddleware

Circuit breaker protection.

RetryMiddleware

Retry loop with exponential backoff.

CoalescingMiddleware

Request coalescing — deduplication of concurrent calls.

FallbackMiddleware

Fallback execution on failure.

5. Resilience patterns

Beyond basic try/except blocks, ToolOps implements three deterministic patterns drawn from distributed-systems engineering. Together they ensure an agent never gets trapped in a failure loop, never exhausts its API budget on a degraded service, and never serves stale data when the upstream is healthy.

5.1 Circuit breaker

Stops all calls to a failing service after a configurable failure threshold. Once open, the circuit fails fast — returning immediately rather than waiting for a timeout — and enters a recovery window before attempting to re-establish the connection. This prevents a single failing tool from cascading into full agent failure.

Circuit breaker Python
@readonly(
    circuit_breaker=True,
    circuit_failure_threshold=5,   # opens after 5 consecutive failures
    circuit_recovery_timeout=60    # retries after 60 seconds
)
async def get_exchange_rates() -> dict:
    return await forex_api.fetch()

5.2 Stale-if-error

When an upstream service fails and no live data can be retrieved, ToolOps can automatically fall back to the last known good value from the cache — even past its normal TTL. The production equivalent of "serve something useful rather than crashing."

Stale-if-error Python
@readonly(
    cache_ttl=3600,
    stale_if_error=True,
    stale_ttl=86400   # serve stale data for up to 24h on failure
)
async def get_exchange_rates() -> dict:
    return await forex_api.fetch()

5.3 Request coalescing

When multiple agent instances call the same tool simultaneously — a common pattern in multi-agent pipelines — ToolOps detects the in-flight request and holds subsequent callers until the first completes. The single real result is then multicast to all waiting callers.

50 → 1

Concurrent calls collapsed

98%

Reduction in credit consumption

0

Changes to the calling agents

In a benchmark with 50 concurrent agent calls to the same weather tool, request coalescing reduced upstream API calls from 50 to 1.

6. Semantic caching

Traditional caches operate on exact key equality. This works for deterministic systems, but agents are not deterministic systems — the same user intent surfaces in dozens of different phrasings. ToolOps uses vector embeddings to understand the meaning of a tool call, not just its literal arguments.

Same intent, different words Example
# Call 1 — cache miss, real API call
query: "What is the status of invoice #442?"

# Call 2 — semantic similarity 0.97 → cache hit
query: "Check the current status for invoice 442"

# Call 3 — semantic similarity 0.94 → cache hit
query: "Invoice 442 — is it paid?"
Configuring the semantic cache Python
from toolops.cache import SemanticCache, SentenceTransformerEmbedder

embedder = SentenceTransformerEmbedder("all-MiniLM-L6-v2")
semantic = SemanticCache(embedder=embedder, threshold=0.92)
cache_manager.register("semantic", semantic)

@readonly(cache_backend="semantic")
async def ask_agent(query: str) -> str:
    return await llm.complete(query)
# Reduces LLM latency by up to 90% on repeated intent patterns.

The similarity threshold — 0.92 above — is the primary tuning lever. Higher values require tighter semantic alignment before a hit is declared; lower values are more aggressive. The right value depends on how much variation is acceptable in your tool's input domain: a factual lookup tolerates a higher threshold than a creative-generation task.

Performance note

In v0.2.0, SemanticCache eviction was refactored from O(n) list operations to O(1) collections.deque — ensuring predictable latency and memory usage under high concurrency.

7. Observability

Debugging non-deterministic agent workflows requires instrumentation that goes deeper than application-level logging. ToolOps emits structured telemetry at every stage of the tool lifecycle — hits, misses, retries, circuit-state transitions — giving a complete audit trail without manual instrumentation.

Tracing and metrics Python
from toolops import configure_opentelemetry, prometheus_metrics

# Configure OpenTelemetry tracing — accepts any standard tracer instance
configure_opentelemetry(tracer)

# Expose Prometheus metrics as a raw text string
metrics_string = prometheus_metrics()

Structured logs

Every cache hit, miss, failure, and retry is emitted as machine-readable JSON. Fields include tool name, backend, latency, cache key, and outcome — ready for ingestion by any log aggregator.

OpenTelemetry

Native OTEL traces and spans wrap every tool execution. Pass any standard tracer to configure_opentelemetry(tracer) and visualise the full call graph in Jaeger, Honeycomb, or Datadog.

Prometheus metrics

Real-time gauges for cache hit rate, circuit state (closed / open / half-open), and tool latency percentiles. Key metrics include toolops_cache_hits_total, toolops_tool_latency_seconds, and toolops_circuit_opens_total — ready to drive alerting rules and dashboards.

8. Ecosystem and MCP

ToolOps tools are plain Python functions. That design choice is not accidental — it means they work natively with every agent framework that accepts Python callables, with no adapter code and no framework-specific configuration.

Integrations, all available today

  • LangChain / LangGraph. Built-in helper.
  • CrewAI. Built-in helper.
  • LlamaIndex. General compatibility.
  • Model Context Protocol. Built-in helper.
  • PydanticAI. General compatibility.
  • AutoGPT and custom frameworks. Any Python callable.

The MCP integration deserves particular mention: a built-in adapter exposes any decorated tool as an MCP-compatible definition without writing a line of JSON Schema — so a resilient, production-grade tool is available to Claude Desktop, Cursor, or any MCP-compatible host instantly.

Exposing a tool over MCP Python
from toolops.integrations.mcp import MCPIntegration

# get_weather is already decorated with @readonly
definition = MCPIntegration.to_mcp_definition(get_weather)
# → MCP-compatible tool definition, ready for Claude Desktop or Cursor.

9. CLI and operations

ToolOps ships with a command-line tool for inspecting and managing tool infrastructure in production — built for operators and CI pipelines, not just developers.

Operating a running app Shell
# List all available commands
toolops --help

# Check system health and backend readiness
toolops doctor

# View real-time cache statistics
toolops stats --app my_app:setup_toolops

# Print current metrics
toolops metrics --app my_app:setup_toolops

# Inspect a specific cache key
toolops inspect-key --app my_app:setup_toolops

# Clear a specific cache backend
toolops clear postgres --app my_app:setup_toolops

toolops doctor is particularly useful in deployment pipelines: it validates backend connectivity, checks embedding-model availability, and reports circuit-breaker state — a readiness check you can wire directly into your health endpoint.

10. Roadmap

v1.0.0 marks the stable, batteries-included release: a composable middleware pipeline, hardened security, a unified lifecycle across all cache backends, and a single-command install with no extras required. The following are planned for upcoming releases, ordered by expected delivery.

What's next

  • Web dashboard. Real-time metrics, cost attribution and cache hit rates in a browser UI — no Prometheus or Grafana setup required.
  • Budget control. Hard limits on tool-induced API costs per hour or per day, configurable per tool and per backend.
  • Native MCP server. One-click deployment of ToolOps tools as a standalone MCP host — no Claude Desktop configuration required.
  • Streaming middleware. Full support for streaming tool outputs in agent pipelines.
  • New backends. ChromaDB and Pinecone support, extending the vector-native cache options. (MariaDB is already supported since v1.0.0 via MySQLCache.)

ToolOps is open source under Apache 2.0. Star the repository, open an issue, or submit a pull request — the project is built in the open, and roadmap priorities are shaped by real-world production use cases from the community.

1

Decorator to adopt

0

Framework lock-ins

7

Cache backends, batteries included

FAQ

Frequently asked questions

ToolOps is framework-agnostic. While LangChain or CrewAI have basic retry logic, ToolOps provides industrial-grade patterns like circuit breakers, request coalescing and semantic caching that work across any Python tool with zero migration cost.

Running agents in production?

Try it on your flakiest tool first. Issues and war stories welcome.

Star on GitHub

Contact

Got a system to build?

If you're building a RAG pipeline, an AI agent workflow, or an MCP integration and want to compare notes, swap ideas, or just avoid the mistakes I've already made — reach out.

Or get the Playbook in your inbox — one email when a new lesson ships, and nothing else.

No noise, no spam. Unsubscribe anytime.