- Exact-Match — returns cached responses only when the request is identical
- Semantic — returns cached responses when requests are semantically similar, even if worded differently
Prompt Caching vs Gateway Caching
LLM providers and the TrueFoundry AI Gateway offer caching at different layers. Understanding the difference helps you pick the right approach — or combine both for maximum savings.Cache Types
Exact-Match Cache
Exact-match caching stores responses keyed by a hash of the complete request — messages, model, and all parameters. A cached response is returned only when every part of the request matches exactly. Best for:- API calls with identical parameters
- Deterministic queries that always need the same response
- Development and testing environments
- Applications with predictable, repetitive queries
Semantic Cache
Semantic caching uses embeddings and cosine similarity to match requests that express the same intent, even when worded differently. For example, “How do I reset my password?” and “What’s the password reset process?” would match. The gateway extracts the last message, generates an embedding, and compares it against cached embeddings. All other request parameters (model, prior messages, temperature, etc.) are hashed and must still match exactly — only the last message is compared semantically. Best for:- Customer support chatbots
- FAQ systems where users phrase questions differently
- Conversational AI applications
- Any scenario where query variations express the same intent
Semantic cache is a superset of exact-match cache. Setting the cache type to
semantic will also return results for exact-match hits, so you don’t need to configure both.How to control the similarity required for a cache hit?
How to control the similarity required for a cache hit?
The
similarity_threshold parameter (0 – 1.0) controls how close two queries must be for the cached response to be returned. A higher value demands closer matches; a lower value allows broader matching.How is similarity computed?
How is similarity computed?
The gateway converts the last message of each request into a vector embedding using an embedding model. When a new request arrives, its embedding is compared against cached embeddings using cosine similarity — a score between 0 and 1.0 that measures how close two vectors are in meaning. If the score meets or exceeds the configured
similarity_threshold, the cached response is returned.All other request parameters (model, prior messages, temperature, etc.) are hashed separately and must match exactly — only the last message is compared semantically.Namespacing and Cache Isolation
The gateway isolates cache entries at two levels to prevent data leaking across boundaries.Level 1 — User / Virtual Account (automatic)
Every cache entry is scoped to the user or virtual account that created it. This happens automatically — you don’t need to configure anything. A cache entry created by User A is never visible to User B, even if they send the exact same request.Level 2 — Custom Namespace (optional)
Within a user’s or virtual account’s cache, you can further partition entries by providing anamespace string. Entries in one namespace are invisible to requests with a different namespace (or no namespace).
This is useful when a single virtual account serves multiple downstream end-users or application contexts and you want per-context cache isolation. For example, an application that serves many tenants through a single virtual account can set namespace to each tenant’s ID so that cached responses never cross tenant boundaries.
Configuration
Enable caching by adding thex-tfy-cache-config header to your requests.
Examples
Replace
{GATEWAY_BASE_URL} with the base URL of the TrueFoundry AI Gateway and your-truefoundry-api-key with your API key.Response Headers
The gateway returns headers indicating cache status on every response:How It Works
- Exact-Match
- Semantic
1
Hash the request
The gateway generates a hash of the complete request (messages, model, parameters).
2
Look up the cache
If a cached response exists for this hash, it is returned immediately.
3
Forward on miss
On a cache miss, the request is forwarded to the model provider. The response is cached before being returned.
Infrastructure Setup
For SaaS customers, caching infrastructure is fully managed — no additional setup is required. For On-Premise deployments, caching requires two infrastructure components:- Redis — cache store for both exact-match and semantic caching
- Embedding Model — required for semantic caching to generate vector embeddings (see Embedding Model)
Redis Setup
You can either enable the bundled Redis instance in thetfy-llm-gateway Helm chart, or connect your own Redis (or Redis-compatible store such as Valkey).
- Bundled Redis
- Bring Your Own Redis
Enable Redis directly in the
tfy-llm-gateway chart values:How do I connect my external Redis?
How do I connect my external Redis?
Enable an encrypted connection by setting
REDIS_TLS_ENABLED: true. TLS applies to both standalone and Sentinel connections. The certificate material (REDIS_TLS_CA_CERT, REDIS_TLS_CERT, REDIS_TLS_KEY) accepts either a path to a mounted PEM file or the inline PEM contents.- PEM material (
REDIS_TLS_CA_CERT,REDIS_TLS_CERT,REDIS_TLS_KEY) can be a path to a Kubernetes-mounted file or the inline PEM string. A value that is neither a readable file nor valid PEM fails loudly at startup. - Mutual TLS requires both
REDIS_TLS_CERTandREDIS_TLS_KEY— setting only one is a configuration error. - These TLS settings apply to both standalone and Sentinel connections.
How do I connect to a highly available Redis using Sentinel?
How do I connect to a highly available Redis using Sentinel?
For high-availability Redis deployments that use Redis Sentinel for automatic failover, point the gateway at your Sentinel nodes instead of a single Redis host. In Sentinel mode the gateway discovers the current master through the Sentinels and automatically follows failovers.Enable Sentinel by setting
REDIS_SENTINEL_ENABLED: true along with the Sentinel node list and master name:- Sentinel mode activates only when
REDIS_SENTINEL_ENABLEDistrueand bothREDIS_SENTINEL_NODESandREDIS_SENTINEL_MASTER_NAMEare set. If any of these is missing, the gateway falls back to standalone mode usingREDIS_HOST. - Data-node authentication (to the Redis master and replicas) uses
REDIS_USERNAME/REDIS_PASSWORD. TheREDIS_SENTINEL_USERNAME/REDIS_SENTINEL_PASSWORDvariables authenticate to the Sentinel processes only. - TLS applies to both the Sentinel and data-node connections — enable it with the same
REDIS_TLS_*variables used for standalone Redis.
Embedding Model
For embedding model, here is the configuration required based on your deployment mode:- TrueFoundry SaaS: No setup required, we use OpenAI’s
text-embedding-3-smallmodel by default. This model is not configurable. - Self Hosted (Control Plane + Gateway Plane): You can configure the model by going to “Settings” → “Semantic Cache” and select the embedding model. (should already be added as in integration)
- Hybrid: (TrueFoundry Control Plane + Self Hosted Gateway Plane): For this, you need add an environment variable in the values of tfy-llm-gateway helm chart.
You can add your embedding model (already registered on TrueFoundry) using the following change: