Skip to content

OpenAI ships a Decisions API that returns a fixed answer in about 150 milliseconds

· by Pondero Newsdesk

The short version

OpenAI's October 6 beta launch of the Decisions API returns one of a developer's predefined answers from text or images, a classification-only endpoint OpenAI says runs about 10 times faster than its Responses API.

OpenAI ships a Decisions API that returns a fixed answer in about 150 milliseconds

OpenAI's new Decisions API, released in beta on October 6, skips generated text entirely. A developer sends a question, a list of allowed answers, and a piece of text or image context, and the API returns one of those answers about 10 times faster than a standard call to its Responses API, per OpenAI's developer changelog.

What

The endpoint runs on gpt-6-luna through a dedicated POST /v1/decisions call and supports three question types, per OpenAI's Decisions guide: predicate returns the probability that a stated condition is true, choice picks one option from a fixed set, and score rates an input against ordered severity-style levels. OpenAI priced it at $0.10 per 1 million input tokens with no charge for output, cache-read, or cache-write tokens, and the guide lists support for Zero Data Retention and HIPAA compliance for eligible customers per the same guide. OpenAI first showed the feature as a limited preview at its DevDay event on September 29, then widened it to the public October 6 beta; the guide says general availability is expected "in the coming weeks." The specific 150-millisecond figure, measured against roughly 1.6 seconds for an equivalent standard Luna call, comes from OpenAI's own DevDay materials rather than this week's documentation per The New Stack, so it remains a vendor claim without independent benchmarking two days into the wider beta.

A cheaper gate for routing and moderation calls

Teams that currently burn a full chat completion to classify a support ticket, flag risky content, or decide which tool an agent should call next get a narrower, output-token-free option: Decisions bills only for input tokens and returns a probability or a labeled choice instead of prose the application still has to parse. OpenAI itself draws the line in its guide, pointing developers toward Structured Outputs when they need a generated object with their own schema and toward function calling when a model needs to request a tool call, reserving Decisions for the narrower predicate/choice/score cases. No maximum number of allowed answers appears in the published docs yet, so teams building high-cardinality routers should test their own limits before committing production traffic to the beta endpoint.

Context and reactions

OpenAI is not alone in building this category. Amazon shipped an open-weight rival, Strands Decider 2B, on October 1, positioning it against both OpenAI's hosted endpoint and TypeSafe's Jev, the model AWS credits with starting the current wave of dedicated "decision models" (Pondero's coverage of the Strands Decider launch has the AWS and TypeSafe detail). Where AWS shipped a downloadable checkpoint teams can run offline, OpenAI's Decisions API stays a hosted, metered call tied to its own infrastructure and pricing.

What to watch next

Watch for OpenAI to publish independent or third-party latency numbers once broader beta access lands, and for documentation on any cap on the number of allowed answers a single request can carry. The guide's "coming weeks" GA timeline is also worth tracking against OpenAI's actual release cadence.

Sources