System Design

System Design for AI Systems: Caching, Rate Limits & State Persistence

Architectural blueprints for scaling agentic systems to production: semantic caching, backpressure handling, distributed state, and failure recovery.

Diagram for System Design for AI Systems: Caching, Rate Limits & State Persistence
On this page

Deploying an autonomous agent or multi-step LLM workflow in production presents challenges that traditional web services never face: unpredictable execution latency (from 2 seconds to 45 seconds), volatile API rate limits, non-deterministic cost profiles, and stateful multi-turn interactions.

Architectural Architecture: Decoupled Job Queues

Never execute long-running agent workflows synchronously within an HTTP request/response cycle. If the client disconnects or an upstream model hangs, resources are orphaned.

Instead, implement an asynchronous job queue:

text
[Client] ---> POST /api/tasks ---> [API Gateway] ---> [Redis / SQS]
                                                            |
                                                            v
[Client] <--- SSE / Polling <--- [State DB] <--- [Agent Worker Pool]

Semantic Caching with Vector Thresholds

For high-volume query workloads, identical or semantically identical prompts should bypass the LLM entirely:

python
import numpy as np
 
def cosine_similarity(v1, v2):
    return np.dot(v1, v2) / (np.linalg.norm(v1) * np.linalg.norm(v2))
 
def check_semantic_cache(query_vector, cache_records, threshold: float = 0.95):
    """Return cached answer if query is semantically indistinguishable."""
    for record in cache_records:
        sim = cosine_similarity(query_vector, record["vector"])
        if sim >= threshold:
            return record["cached_response"]
    return None

Implementing Token Bucket Rate Limiting

To avoid sudden 429 Too Many Requests errors from model providers, implement client-side token bucket limiters that buffer spikes and enforce smooth request cadence.

By treating AI agents as distributed, asynchronous systems rather than simple HTTP handlers, your systems maintain high availability even under extreme upstream latency and provider volatility.

Keep learning with Sri

More practical tutorials and experiments on the channel.

Watch on YouTube
Back to articles