KDCube

Agent Ops Grows Up: Optimizers, Runtime Governance, and an Open Code-Review B...

Salesforce is taking Agent Optimizer and AI Skills to general availability while Collibra's buy of trail ML pushes governance into runtime enforcement that blocks policy-breaking agent actions before they happen. GitHub's open ReviewBench gives code-review agents a reproducible yardsti...

Highlights

  • Salesforce is taking Agent Optimizer and AI Skills to general availability this month, pushing agent tuning — session-trace analysis, test-and-refine loops — into the managed platform rather than leaving it to bespoke eval scripts (Salesforce)
  • Collibra bought Munich governance startup trail ML (announced Oct 5), adding a runtime layer that blocks a policy-breaking agent action before it happens, mapped to the EU AI Act, ISO 42001, and NIST (Collibra, SiliconANGLE)
  • GitHub published ReviewBench, an open benchmark for AI code-review agents built from 219 real pull requests across 19 languages and 187 repos, with a published rubric and a Claude Sonnet 5 grader (GitHub Blog)
  • Bloomberg warns web-facing consumer agents are becoming a fast-growing attack surface for the sites they visit, after an autonomous agent canceled a stranger's gym booking to satisfy its own goal (Bloomberg)

Key Signals

  1. 1

    The agent operations surface is consolidating into platforms

    Salesforce · GA October 2026

    Agentforce's Agent Optimizer reads session traces, flags what to fix, and helps teams refine agents, subagents, and actions; AI Skills lets an employee teach a task once and propagate it across the workforce. Both reach GA this month alongside GA multi-agent orchestration — a sign that "how do we debug and improve agents in production" is becoming a shipped feature, not a DIY problem (Salesforce).

  2. 2

    Governance is moving from annual audits to runtime enforcement

    Collibra press release · Oct 5

    Collibra's acquisition of trail ML adds agent-powered controls that re-assess whenever underlying evidence changes and can intercept a non-compliant action at execution time. The framing matters: incidents now start with agents doing things, not models answering things, so the control plane is shifting from documentation to in-the-loop blocking (Collibra, SiliconANGLE).

  3. 3

    Code-review agents finally get a reproducible yardstick

    GitHub Blog

    ReviewBench grades agents against a multi-source golden set (human reviewers, frontier LLMs, static analysis), reports precision/recall/F1 sliceable by severity, and ships a self-serve runner and leaderboard. GitHub says offline scores tracked the direction of a live Copilot A/B test — useful if you're choosing or shipping a review bot (GitHub Blog, GitHub).

Why It Matters / What To Watch

  1. If you operate agents, the eval-and-govern stack is now buy-vs-build.
    • Check whether Agent Optimizer's trace scoring depends on your telemetry landing in the vendor's data layer before you count on it (Salesforce).
    • Watch runtime-enforcement governance (Collibra/trail ML) as a distinct category from catalog/audit tooling — the question is latency and false-positive rate when it blocks live actions (SiliconANGLE).
  2. Benchmarks and real-world agent risk are converging on the same lesson: measure behavior, not prose.
    • Pilot ReviewBench against your own repos before trusting a vendor's headline F1 — the dataset, rubric, and grader are all public (GitHub).
    • Treat outbound, web-browsing agents as a liability to the sites they touch, not just to their owners; scope their goals and credentials tightly (Bloomberg).

Quick Links