Newsletter
Join the Community
Subscribe to our newsletter for the latest news and updates
Source: Cloudflare's Clef launch post. The official artwork identifies the release; it does not show a product interface or prove model performance.
Most agent stacks still ask a language model to make a yes-or-no call, then spend another step parsing the prose. The Cloudflare Clef decision model takes a narrower route: turn a fixed state and a fixed set of questions into typed probabilities. That could be useful for routing and triage, but only if the decision is well-defined enough to audit when it goes wrong.
Cloudflare Clef is a new 27B multimodal decision model designed to answer a fixed schema of questions rather than generate free-form chat. For a developer, that means sending a state—text, JSON, images, or video—and receiving probabilities for allowed choices, not a paragraph that needs another parser. Cloudflare announced Clef and the smaller Clef-flash on October 1, 2026, and its Workers AI documentation lists a public model ID, input schema, and supported question types. The launch post says the models are open source and available through Workers AI. MarkTechPost's report describes the weights as Apache 2.0 and deployable through the hosted service, but that does not replace testing on a team's own decisions. Cloudflare's decision-index chart labels its own results self-reported; no independent reproduction establishes comparative performance. The practical question is narrower: can your agent step be reduced to a fixed, reviewable decision?
Cloudflare's launch post introduces Clef and Clef-flash as decision models for classification and agentic workflows, alongside a reinforcement-learning fine-tuning offering. Its model documentation describes Clef as a 27B model that accepts text, JSON, images, or video and returns a probability for every permitted option.
That is a different contract from a chatbot. The documentation allows one to 64 typed questions in a request, and the returned answer stays keyed to the question ID. The public material does not prove that a typed answer is correct, reliable in every domain, or cheaper for a particular workload. MarkTechPost's independent report is useful confirmation of the release details, not an independent performance test.
Cloudflare's release is more than a new endpoint: it makes the decision-model idea available as both a hosted Workers AI model and a weight release. That matters for teams that want a deliberately small output contract in an agent workflow without treating natural-language output as an API.

Source: Cloudflare's Clef launch post. The chart is Cloudflare's self-reported evaluation and compares a specific Decision Index 0.2.1 setup; it is not an independent benchmark or a promise for a production workload.
The documented workflow starts with a state and a schema. The state can be a support ticket, a JSON payload, an image, or a video; the schema spells out the decisions the system is allowed to make. Instead of asking for an open-ended explanation, a caller asks questions such as whether a request belongs in a queue, which policy category applies, or how strongly an option fits a defined scale.
The answer is still model output, so it needs controls. A probability is not a permission to approve a refund, change a production system, or reject a person. The useful design move is to keep the model's choice, confidence, source state, and the human or deterministic rule that takes the next action in the same audit trail.
Start with a decision that already has a known answer or an inexpensive review step. For example, a team could route a bug report into a small, predeclared taxonomy and sample disagreements against a human label. The Workers AI docs make the contract concrete: callers provide typed questions and receive answers under the same IDs, so validation can happen before an agent invokes another tool.
Do not substitute a decision model for a policy. If the schema is vague, the labels are inconsistent, or a wrong answer causes irreversible harm, a tidier payload will not make the workflow safe. Keep the action behind a threshold, a reviewer, or a reversible rollout.
The interesting part is not that Clef can sound more certain than a chatbot; it is that it declines to be a chatbot at all. That makes it easier to inspect what the system was asked to decide. It does not make the model a general assistant, an autonomous operator, or proof that every agent stack should replace its LLM.
My take: Clef is worth a small pilot when the next action already depends on a finite, inspectable decision. It is a poor excuse to automate a messy policy just because the response now arrives as a probability.
What is Cloudflare Clef?
Cloudflare Clef is a 27B multimodal decision model. According to Cloudflare's documentation, it takes a state and typed questions, then returns probabilities for permitted answers. It is designed for bounded decisions such as classification or routing, not for open-ended conversational output.
Is Cloudflare Clef available now?
Cloudflare's October 1, 2026 launch post and Workers AI documentation list Clef and Clef-flash as public models. Before using either in a product, verify the current model availability, pricing, regional requirements, and account terms in the official documentation because those details can change.
How is Clef different from a chat model?
A chat model usually generates free-form tokens that an application must interpret. Clef's documented interface asks a fixed set of typed questions and returns structured answers and probabilities. That can simplify validation, but it does not show that the underlying classification is correct.
What should developers test before using Clef?
Use a reversible task with known labels, such as routing a limited set of support issues. Compare its answers with human review, define what confidence threshold triggers action, retain the input and output for audit, and stop the rollout when disagreement exposes a schema or data problem.
What is still unverified about Cloudflare Clef?
Cloudflare's release includes a self-reported Decision Index versus latency chart, but the public material does not establish independently reproduced production performance. Teams still need to test quality, latency, cost, reliability, and data-handling constraints against their own workload before treating its output as operationally trustworthy.
Discover practical AI products and emerging tools at AIToolHunt.