Control

One API for every model

Point your existing SDK at the gateway and keep your code. The same endpoint reaches OpenAI, Anthropic, Azure AI Foundry, AWS Bedrock, Google Gemini, Mistral and self-hosted models, so you can change model or provider without a rewrite.

  • /v1/chat/completions, /v1/responses, /v1/embeddings and /v1/messages
  • Works with the official OpenAI and Anthropic SDKs, plus any HTTP client
  • Switch provider by changing a model name, not your application
  • Self-hosted and private models sit behind the same controls
Control

Virtual keys, teams and budgets

Issue a gateway key to each application, team or person. Each key carries its own allowed models, spend budget and rate limit, and can be revoked in a click without touching a provider account.

  • Gateway virtual keys with expiry and revocation
  • Teams and users, with per-key model permissions
  • Budgets with alerts, and requests-per-minute rate limits
  • Provider keys stay inside the gateway. Applications never hold them
Secure

PII protection, tuned for Australia

Sensitive data is detected before a prompt leaves your control. Choose what happens for each kind of data: watch it, mask it, swap it for a reversible token, or stop the request.

  • Australian identifiers: Medicare number, TFN, ABN, driver licence, passport
  • Common identifiers: names, emails, phone numbers, addresses and card numbers
  • Four modes: monitor, redact, tokenise (restored on response) or block
  • Per-policy control over which entity types are acted on

Same prompt, four modes (illustrative)

OriginalPatient Jane, Medicare 2123 45670 1, asks about a rebate.
RedactPatient Jane, Medicare [MEDICARE_NUMBER], asks about a rebate.
Tokenise (restored in the reply)Patient Jane, Medicare <MEDICARE_1>, asks about a rebate.
BlockRequest stopped by policy. The application receives a clear error. Monitor mode records it and lets it through.
Secure

Secrets and credential detection

Developers paste logs and config into chat. The gateway recognises common key formats, tokens and private key blocks and can block or mask them before they reach a third party.

  • Cloud, source-control and payment API key patterns
  • Private keys, bearer tokens and connection strings
  • Block or redact, with an audit record of what was caught
  • Applies to prompts and to model responses
Secure

Prompt-injection detection

Requests and retrieved content are checked for attempts to override your instructions or extract hidden prompts. Run in monitor mode first to see what would be caught, then enforce.

  • Instruction-override and system-prompt extraction patterns
  • Monitor first, enforce when you are confident
  • Events logged with the key, team and model involved
  • Layered defence: a control, not a guarantee against every attack
Secure

Response DLP scanning

Model output can contain sensitive data too, from your own context or from a connected tool. Responses are scanned with the same detectors and the same four modes.

  • Same detectors and modes as requests
  • Tokenised values restored on the way back to your application
  • Streaming-aware handling
  • Separate policy for requests and responses
Optimise

Safe prompt optimisation

The gateway normalises whitespace, compacts JSON and removes duplicated context. You see a before and after token count for each change. It never strips punctuation blindly, because that changes what a prompt means.

  • Whitespace, JSON and duplicate-context normalisation
  • Before and after token savings shown per request
  • Code, quoted text and structured data left intact
  • On or off per policy

Optimisation report (illustrative)

Before{ "items": [ { "id": 1, "name": "A" } ] }    (extra whitespace, repeated context block)
After{"items":[{"id":1,"name":"A"}]}   (compact JSON, duplicate block removed)
ReportedInput tokens before and after, per request, and the total saved.
Optimise

Response caching

Identical requests can be answered from cache at no model cost and with lower latency. Semantic caching, which matches similar rather than identical prompts, is off until you switch it on, because it suits some workloads and not others.

  • Exact-match response caching
  • Opt-in semantic caching with a similarity threshold
  • Cache keyed per policy so tenants never share answers
  • Hits and estimated savings reported in the console
Optimise

Smart model routing

Route simple work to faster, lower-cost models and keep demanding work on premium ones. Every routing decision includes a plain-language "why this model" explanation so you can trust it and tune it.

  • Rules by task type, size, team or key
  • "Why this model" explanation on every decision
  • Never routes outside the models a key is allowed to use
  • Easy to turn off for sensitive workloads

Why this model (illustrative)

DecisionRouted to a faster, lower-cost model.
ReasonShort classification task under the size threshold, allowed for this key, within the team budget.
Optimise

Retries and failover

Transient errors are retried and, when a provider or model is unavailable, traffic fails over to the alternatives you have approved.

  • Automatic retries with back-off
  • Provider and model failover chains
  • Failover respects data-residency and permission rules
  • Failover events visible in logs
Control

Spend tracking and alerts

Every request is attributed so finance and engineering can see where money is going. Set budgets, get alerts before limits are hit, and see estimated savings from caching, optimisation and routing.

  • Breakdowns by app, team, user, model and provider
  • Budgets, thresholds and alerts
  • Estimated savings clearly labelled as estimates
  • Export for chargeback and reporting
Control

Audit logs and OpenTelemetry

Who called which model, under which policy, and what the gateway did about it. Send traces and metrics to the observability stack you already run using OpenTelemetry.

  • Request, policy-action and admin-change records
  • Content logging is configurable, so you can keep prompts out of logs
  • OpenTelemetry export for traces and metrics
  • Retention set to your requirements
Control

Policy presets, every feature a switch

Seven presets give you a sensible starting point, and every individual feature is a simple on/off switch you can override per team or key.

  • Default, Healthcare, Finance, Government, Developer, High Security and AI Agent presets
  • Every feature is an on/off switch
  • Assign a preset to a team, key or environment
  • Preview in monitor mode before enforcing
DefaultHealthcareFinanceGovernmentDeveloperHigh SecurityAI Agent
Control

Microsoft Entra single sign-on

Administrators sign in to the console with Microsoft Entra, so access follows your joiner, mover and leaver process and your multi-factor policy.

  • Sign in with Microsoft Entra ID
  • Role-based access to the admin console
  • No separate admin passwords to manage
  • Applications authenticate with gateway keys

Put a firewall between your people and AI

Request access and we will help you set up your first policy, key and budget.