Typed Outputs, Token Diets, and Trusted Context
OpenAI split off a beta Decisions API on gpt-6-luna that returns typed answers it bills as ~10x faster than Responses, while New Relic turned AI Evaluation into a platform feature over OpenTelemetry traces. SAP's TechWolf buy and GoodData's governed serving plane show token efficiency and runtime governanc...
Highlights top of the board
- OpenAI shipped a beta Decisions API powered by
gpt-6-luna, turning text and images into typed answers it bills as ~10x faster than the Responses API - a purpose-built path for the classify/route/extract calls buried inside most agent graphs. (OpenAI) - OpenAI also collapsed its five API usage tiers into three - Build, Launch, Grow - with orgs auto-promoted as credit purchases cross thresholds, reshaping how teams reason about rate limits and budget headroom. (OpenAI)
- New Relic made AI Evaluation a platform capability inside its AI observability suite, scoring agent responses for quality, relevance, and hallucination on top of OpenTelemetry GenAI traces. (New Relic)
- SAP agreed to buy Belgian firm TechWolf, folding its "context graph for work" into Joule and SuccessFactors - and pitching it explicitly as a way to make HR-agent token usage more efficient. (The New Stack)
Key Signals ranked scan
-
OpenAI splits off a typed-decision path from Responses
The new Decisions API returns structured, typed outputs via
gpt-6-lunaat a claimed 10x latency edge over Responses, aimed at the high-volume routing, scoring, and extraction steps that dominate agent runtimes. For builders, it's a signal that the "one big chat endpoint" era is giving way to specialized primitives you compose - and a reason to benchmark your hot-path calls before the next cost review. (OpenAI) -
Observability vendors move from traces to judgments
New Relic's "Trusted Context for AI Operations" push promotes AI Evaluation to a first-class feature: semantic and qualitative scoring of non-deterministic outputs, auto-captured model/token/tool-call data for LangGraph, Strands, and AutoGen, all riding OpenTelemetry GenAI conventions. The bar for "production-ready agent" is shifting from does it run to is the answer good, and can you prove it. (New Relic)
-
Governed execution planes court the agent as a first-class query client
GoodData.AI's Agentic Serving Plane frames governance, metric consistency, and access enforcement as runtime concerns for high-concurrency agent traffic, exposed through an MCP server so agents operate analytics end-to-end inside guardrails. It's the data layer answering the same question governance teams are: what happens when the agent, not the human, asks next. (GoodData)
Market Heat topic map
- APItyped outputs
- Costtokens
- Evalobservability
- Govpolicy
- MCPserving
- Opsruntime
- Contextgraph
- Agentsharness
Why It Matters / What To Watch watch items
-
The quiet theme this week is token efficiency as a platform feature, not a tuning afterthought.
- Watch how a faster, typed Decisions API changes the economics of agent inner loops - and whether Responses stays the default once teams measure per-call cost. (OpenAI)
- Note SAP's framing of TechWolf's context graph as a way to cut HR-agent token spend: context engineering is now sold as a cost lever, not just a quality one. (The New Stack)
-
Evaluation and governance are converging on the same control surface.
- If you run LangGraph, AutoGen, or Strands, check whether your OTel GenAI spans already flow into evaluation scoring before you buy a second tool. (New Relic)
- For analytics-backed agents, treat the serving plane's access and metric-consistency controls as part of your governance story, not a BI detail. (GoodData)
- Context for the trend: Manus's late-September 2.0 release leaned on its Cascade harness to claim ~23% fewer tokens and ~32% lower cost per task - the same efficiency pitch, now showing up across runtimes. (TestingCatalog)
Quick Links sources
- OpenAI API Changelog - Decisions API, gpt-6-luna, and simplified usage tiers OpenAI
- New Relic Now (October 2026): Trusted Context for AI Operations and AI Evaluation New Relic
- Agentic Serving Plane: Governed Analytics for AI Agents GoodData.AI
- SAP Acquires TechWolf to Feed More Context Into Its HR Agents The New Stack
- Manus 2.0 Launches With Studio, Cloud Computer and Cue TestingCatalog