AGP Picks
View all

New Open Source Benchmark Scores AI Agents on Their Ability to Learn and Perform Complex Actions

Brackett releases The Agent Effectiveness Index for evaluating AI agents based on their ability to learn complex operational tasks as demonstrated by knowledge workers

SAN FRANCISCO, Sept. 16, 2026 (GLOBE NEWSWIRE) -- There is now a way to measure how well an AI agent is able to learn and take action on the job it was built to do. The Agent Effectiveness Index (AEI), released today as a free and open-source benchmark, scores and ranks AI agents on their ability to understand complex, real-world processes, take proactive actions, and keep learning without drifting as processes change. It was built by Brackett, which has also launched its Connected Agentic Workforce platform.

Organizations are rushing to build AI agents, but they have no standard way to evaluate whether those agents will perform in production. As agents are increasingly tasked with inference and applying context, this includes whether they are able to make judgment calls, exceptions, and escalation patterns that make a process reliable, or if they're producing inconsistent results at a high cost. It's the equivalent of hiring a company's worth of people with no way to evaluate performance. The Agent Effectiveness Index addresses that gap.

“Most evaluations for AI agents focus on static knowledge, or what it knows. But this doesn't tell you whether or not that agent can complete the task it was created for, because real world tasks require things like judgement calls, exceptions, or knowing when to bring in a human. These are learned from experiences, not training data. Until now there has been no way to score that capability,” said Ehsan Azarnasab, co-founder and Chief Scientist of Brackett and formerly Principal Scientist on Microsoft’s GenAI Platform team.

AEI evaluates agent systems across three dimensions:

  • Business Understanding: whether an agent grasps how a specific company actually works, and grounds its answers in real evidence rather than plausible guesses.
  • Operational Execution: whether an agent produces correct results, handles exceptions, and stays inside the authority it was given.
  • Learning Persistence: whether teaching an agent genuinely changed its behavior, and whether that holds on new cases, after time passes, and when the rules change.

The Index publishes its first scores today, measuring learning and comprehension across three agent systems evaluated on the same demonstration: Brackett, OpenAI's Codex, and Anthropic's Claude. The full task set, scoring code, and methodology are available on Brackett's GitHub under the MIT License, with scoring for execution, transfer, and retention to follow as the Index expands toward a complete picture of agent effectiveness.

"The next era of work isn't humans versus agents, it's humans and agents becoming genuinely better together," said Jaideep Sarkar, Co-Founder and CEO of Brackett. "We built Brackett because we’ve seen a recurring gap in which companies try to automate tasks, but without a way to build compounding, connected intelligence that they actually own. It's also why we're opening the Agent Effectiveness Index to the world. You can't build a trustworthy agentic workforce without a way to measure it."

Also live today is Brackett’s Connected Agentic Workforce Platform. Using its Capture, Codify, and Compound methodology, Brackett turns simple conversations, with no code required, into agents that learn how to execute complex processes. Agents then progress to running tasks at scale with real consistency and control, and then to acting on the judgment they've learned over time. Brackett connects these agents to each other, to the systems they run in, and to the people who trained them, so every workflow makes the next one smarter across the business. The intelligence the organization builds along the way is retained as something it owns rather than something it rents from a model vendor.

The Agent Effectiveness Index is available now on Github, free and open source. To learn more about Brackett's Connect Agentic Workforce Platform for Enterprises, book a demo here.

About Brackett

Brackett is the Connected Agentic Workforce platform. It enables enterprises to capture how their people actually work, codify it into agents that learn real organizational judgment, and compound that intelligence into an asset the company owns. Founded by former technical leaders from Microsoft, Rubrik, and Amazon, Brackett is headquartered in San Francisco and is backed by Focal and Heavybit. Learn more at www.brackett.ai.


Media Contact
Jennifer Lankford
Lankford Communications
jennifer@lankfordpr.com

Primary Logo

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Sci-Tech North Carolina

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.