Visibility. Control. Governance. For every AI call.

Fastly AI Runtime Control (ARC)

Stop guessing your AI risk and spend. ARC is a unified control plane that sits between your applications and AI models, giving you centralized visibility, governance, and control over all your AI traffic. Track real-time spend, enforce budget and rate limits, automate multi-provider failover, and benefit from low-latency edge performance.

Benefits

Take control of all AI and agent traffic

Control internal AI usage and govern your AI-enabled applications. Fastly’s AI Runtime Control lets you safely manage multi-provider complexity with real-time visibility, cost and access controls, plus automated failover to guarantee continuous uptime.

Audit 100% of your AI activity

Satisfy legal, compliance, and security requirements. Verify how customer data is being handled and trace exactly who requested what.

Get unified multi-provider metering

ARC consolidates AI traffic into a single pane of glass with real-time, granular tracking of input and output token counts, request volumes, and calculated costs mapped to virtual keys and specific client sessions.

Quickly isolate high-spending hotspots

Use real-time dashboards and leaderboards to identify exactly which applications, teams, or developers are consuming the most tokens and dollars.

Turn provider outages into a non-event

Keep your customer-facing AI features running smoothly through rate limits or upstream provider outages. If a primary model goes down, ARC transparently retries and fails over to your backup models at the edge, ensuring a continuous user experience.

Enforce monthly budget limits

Transition from reactive spending to proactive financial guardrails by applying enforceable monthly dollar-budget limits directly to individual virtual keys.

Stop runaway agents from blowing your AI budget or quota

Rate limiting per virtual key protects your upstream provider quotas from being exhausted by runaway autonomous agent loops or external traffic spikes.

Protect against LLM-based threats with AI Firewall

AI Firewall is a security module within ARC that protects your AI-powered applications from prompt injection attacks with near-zero latency impact. It leverages Fastly’s SmartParse detection engine and custom guardrails to detect attacks like prompt injection and other LLM specific attacks. Deploy with a flip of a switch.

Explore AI Runtime Control

One control plane for all your AI traffic

Lock down access, cap spend, and stop attacks before they start. Fastly’s AI Runtime Control gives you one place to manage access, spend, and risk across every model and provider so you can run AI in production with real-time visibility and control.

Secure access with managed virtual keys

ARC replaces raw provider API keys with secure, rotatable virtual keys, and also supports connecting to providers via federated SSO pass-through for centralized identity and access control.

Per-key spend limits and alerts

Set a monthly dollar budget per virtual key, with your choice of hard-block or continue-serving once exhausted. Get 80% and 100% threshold alerts via Fastly's Notification Service.

AI Firewall

AI Firewall provides an optional layer of security utilizing Fastly’s proprietary SmartParse detection engine to protect AI-powered applications from prompt injection attacks with near-zero latency impact.

Multi-provider failover and retries

Virtual keys route to an ordered list of fallback targets. ARC automatically retries a failed request, then fails over to the next target if errors persist, transparent to the application or service making the request.

Rate limiting per virtual key

Independent requests-per-minute and tokens-per-minute caps per key protect provider quota and help protect your customer-facing AI products and features from attack vectors.

Detailed logging and dashboards

ARC captures full request and response logs, including prompts and completions, for debugging and analysis. Visualize month-to-date spend, request volumes, and token counts to instantly spot runaway usage and anomalies.

Frequently Asked Questions

What is Fastly’s AI Runtime Control?

ARC is a control plane that governs AI traffic across your organization tracking real-time spend, enforcing access controls, and enabling automatic failover between AI models. It gives teams a single view into token usage, costs, and requests across every AI provider they use.

How can I control or reduce AI spend across my organization?

ARC tracks input/output token counts, request volumes, and calculated costs in real time, mapped to virtual keys and client sessions. Real-time dashboards and leaderboards show exactly which apps, teams, or developers are driving the highest spend, so you can spot and fix hotspots quickly.

What happens if my AI provider (like OpenAI or Anthropic) has an outage?

ARC keeps AI features and workloads running by automatically retrying requests and transparently failing over to backup models at the edge when a primary provider goes down.

How do I prevent a runaway agent from blowing through my budget or quota?

ARC applies rate limiting per virtual key, which caps how much any single application, agent, or user can consume. This protects your budget and upstream provider quotas from being exhausted by runaway autonomous agent loops or unexpected traffic spikes.

What is Fastly’s AI Firewall?

AI Firewall protects AI applications from attacks that specifically target models, like prompt injection attempts where a natural-language input tries to trigger unintended actions. Fastly's AI Firewall is purpose-built for these AI-specific threats, unlike a traditional web application firewall.

How does Fastly's AI Firewall detect and block prompt injection attacks?

It uses a three-step process: requests are first inspected by Fastly's SmartParse detection engine for LLM-specific threats, guardrails are injected to sanitize responses, and outputs are validated for refusal responses and leak protection.

What's the difference between AI Runtime Control and AI Firewall? Do I need both?

ARC is the control plane that manages AI traffic, cost, and reliability across providers. AI Firewall is a security layer that deploys inside ARC adding threat detection and guardrails on top. Most organizations run AI Firewall as part of their ARC deployment for both governance and protection.

Looking for more?

Fastly joins the Agentic AI Foundation
Fastly has joined the Agentic AI Foundation (AAIF) to ensure that the protocols governing agentic interactions are built for the scale and speed of real-world production environments
MCP at the Edge
Agent loops make dozens of tool calls per task. Learn how stateless MCP at the edge turns each one into a sub-millisecond hop.
Building an actually secure MCP server with Fastly Compute
MCP can be a security minefield if you build it wrong. This article provides a step-by-step walkthrough for deploying a secure, scalable MCP server on Fastly.

Ready to get started?

Get in touch or create an account