Skip to content

Now routing to self-hosted models

Ship AI features without the plumbing

Halcyon is the control plane between your product and every model you use. Route, cache, evaluate and observe from one typed API, and change providers without changing code.

38%
Lower cost per request
11ms
Routing overhead
1,200+
Teams building

Why Halcyon

The model is the easy part

Getting a good answer from a model takes an afternoon. Getting a good answer every time, at a price you can defend, from a provider you can replace, is the part that takes a year.

Halcyon is that year, already done. It sits between your product and the models, and it is the only thing your code talks to.

One API across every provider and your own weights

Quality measured continuously, not argued about

Costs that fall as the router learns your traffic

No prompt or completion stored unless you ask

Platform

Everything between the request and the answer

Six systems that most teams end up building badly. They are the product, not add-ons, and every plan includes all of them.

Adaptive routing

Send each request to the model that answers it best, scored on live quality and cost rather than a fixed preference order.

38%median cost reduction

Semantic caching

Cache on meaning, not on string equality, so paraphrased questions hit the same entry and never reach a model.

4.1xeffective hit rate

Continuous evaluation

Every deployment replays your graded set before it takes traffic, and rolls back on its own if quality drops.

2,400graded cases per run

Structured pipelines

Compose retrieval, tools and generation as a typed graph. Each node is independently versioned, cached and traced.

11msorchestration overhead

Managed retrieval

Chunking, embedding and reranking behind one endpoint, with incremental reindexing as your corpus changes.

92%recall @ 10

Policy enforcement

Redaction, jailbreak detection and per-tenant rate limits applied at the edge, before a prompt reaches a provider.

0prompts stored by default

How it fits

A day to adopt, not a quarter

Halcyon is a base URL and a key. Point your existing client at it and nothing else has to change on the first day.

Drop-in compatible

Speaks the API your SDK already speaks. Change the base URL and your calls keep working.

Versioned everything

Prompts, pipelines and routing policies are versioned artifacts you can diff, review and roll back.

Bring your own models

A self-hosted checkpoint is just another route, scored and cached like any hosted one.

Private by default

Nothing is retained unless you turn retention on. Redaction runs before a prompt leaves your VPC.

Fails to a good answer

A provider outage reroutes mid-flight. Your users see a slower response, not an error page.

One bill

Every provider, your own capacity and the platform on a single invoice, broken down per feature.

Questions

The things engineers ask first

No. Halcyon accepts the request shape your SDK already sends. Routing, caching and evaluation happen around your prompt, not to it. Prompt management exists if you want it, and is entirely optional.

Contact

Tell us what you are building

An engineer reads every message. If you are moving an existing workload across, mention the provider and we will send the migration notes for it.

Email

engineering@halcyon.dev

Docs

docs.halcyon.dev

Status

status.halcyon.dev

Office

Lisbon, remote-first