Skip to content

Pricing

Pay for what you serve

Per-token pricing with no seat minimum. Every plan includes the full platform; the tiers differ in throughput, retention and support.

Plans

Every plan is the whole platform

Tiers differ in throughput, retention and support. None of them withhold a feature.

Build

$0to start

For prototypes and the first production workload.

  • 1M tokens included each month
  • Adaptive routing across 6 models
  • Semantic cache, 7-day retention
  • Community support
Start building
Recommended

Scale

$0.40per 1M tokens

For teams serving real traffic with a quality bar.

  • Unlimited throughput
  • Continuous evaluation and auto-rollback
  • Managed retrieval, 90-day retention
  • Private models and BYO keys
  • Slack channel with an engineer
Start free trial

Enterprise

Customannual

For regulated workloads and dedicated capacity.

  • Everything in Scale
  • VPC or on-premise deployment
  • SOC 2 Type II, HIPAA, data residency
  • Reserved capacity and volume pricing
  • 99.99% uptime SLA
Talk to us

Questions

The things engineers ask first

No. Halcyon accepts the request shape your SDK already sends. Routing, caching and evaluation happen around your prompt, not to it. Prompt management exists if you want it, and is entirely optional.

Contact

Tell us what you are building

An engineer reads every message. If you are moving an existing workload across, mention the provider and we will send the migration notes for it.

Email

engineering@halcyon.dev

Docs

docs.halcyon.dev

Status

status.halcyon.dev

Office

Lisbon, remote-first