Skip to main content
Back to blog
5 min read

How to set AI goals that actually measure progress

Most AI initiatives fail the measurement test before they begin. The goals are either too vague to evaluate or too narrow to matter. Here is how to set ones that work.

By Zylver Editorial

Most AI initiative goals fall into one of two failure modes. The first is aspirational vagueness: “become an AI-first organization,” “leverage AI across the business,” “build AI capability.” These sound strategic but provide no mechanism for knowing whether you are making progress.

The second is shallow activity tracking: “deploy five AI tools,” “train 80% of employees,” “complete three pilots.” These are measurable but measure the wrong thing. Deploying tools is not the same as creating value. Completing training is not the same as changing behavior.

Both failure modes share the same root problem: the goal is set before the team has figured out what AI is actually supposed to do for the business.

The three types of AI metrics

It helps to distinguish between activity metrics, outcome metrics, and capability metrics. Most organizations track only the first.

Activity metrics measure what you did: tools deployed, people trained, pilots run, hours of AI usage. These are easy to track and provide no signal about value. An organization that runs ten failed pilots is ahead on activity metrics.

Outcome metrics measure what changed as a result: time saved, error rate reduced, revenue influenced, cost eliminated. These are harder to track because AI effects are often distributed, lagged, and entangled with other changes. But they are the only metrics that answer the actual question.

Capability metrics measure what you can now do that you could not do before: time-to-insight on a new category of question, scale of a workflow that was previously bottlenecked by headcount, accuracy on a task that required expert review. These are particularly useful early in an AI initiative when outcome metrics have not yet materialized.

A well-constructed AI goal references at least one outcome or capability metric. Activity metrics can appear as leading indicators, but they should never be the primary measure of success.

The lag problem

AI value often appears in different places than where the investment happened.

A team that deploys an AI tool to reduce document review time may find that the time saved goes into more document reviews, not fewer reviewers. A team that uses AI to accelerate customer onboarding may find the benefit shows up as reduced churn eighteen months later, not as headcount reduction.

This matters for goal-setting because it means the measurement window needs to be longer than most planning cycles and the measurement location needs to be broader than the deployment site.

A common mistake is measuring the wrong downstream variable. If you deploy AI in the sales development function and measure SDR productivity, you will miss the impact on win rate, sales cycle length, and revenue per rep that shows up later and elsewhere. Set goals that look downstream from where the work happens.

When activity metrics become dangerous

The problem with activity metrics is not that they are wrong. They are right about what they measure. The problem is that they become targets, and once they become targets, they stop being accurate measures.

If the AI goal is “deploy ten tools,” teams will deploy ten tools. Whether those tools are used, whether they create value, whether they are the right tools for the problem, becomes secondary to hitting the number.

Activity metrics are useful for tracking execution pace during an initiative, but they should be subordinated to outcome goals. The right structure is: “We will deploy these tools [activity] in order to achieve this outcome [result] by this date [horizon].”

The learning phase problem

Early in an AI initiative, the honest answer is often that you do not yet know what the right outcome metrics are, because you do not yet know what you are building.

This is a legitimate problem. Goals for a learning phase should be epistemic rather than performance-based: “We will understand the top three use cases where AI could have meaningful impact on this specific function” or “We will determine whether AI-assisted triage reduces escalation rate in our support workflow.”

These are not vague. They are specific questions with defined investigation plans and clear definitions of what “answered” looks like. A learning-phase goal should always specify how you will know when the learning is complete.

The failure mode is extending the learning phase indefinitely. Learning goals should have defined end conditions, after which you commit to either a performance goal or a decision to stop the investment.

What to do about attribution

One of the practical barriers to outcome metrics is attribution. When AI assists a process that also involves many humans and many other systems, how do you know how much of the outcome improvement is due to AI?

For small initiatives, randomized controlled experiments are sometimes feasible: run AI-assisted and non-AI-assisted versions in parallel and compare outcomes. For larger initiatives, this is usually not practical.

The realistic approach is to establish a baseline before deployment, measure the same metric after deployment with a reasonable lag, and be honest about what you can and cannot attribute. The point is not statistical rigor; it is organizational honesty about whether the initiative is working.

Common bad AI goals and better alternatives

“Become an AI-first company” becomes: “By end of year, 60% of our high-volume document workflows will have AI assistance with defined quality checks, and we will have baseline data on accuracy and time savings for each.”

“Deploy AI tools across the business” becomes: “AI assistance is deployed in three functions, each with a defined problem statement, a measurable target outcome, and a quarterly review of results.”

“Train employees on AI” becomes: “The customer success team is using AI drafting assistance daily, with quality review in place, and we have baseline data on draft acceptance rates and edit distance from first draft to sent.”

The pattern is: specific scope, defined mechanism, measurable outcome, honest data plan.

The quarterly review cadence

AI goals should be reviewed more frequently than annual planning cycles because the technology, the use cases, and the organizational capability are all evolving faster than most planning processes assume.

A quarterly review should answer three questions: Is the initiative producing the expected outcome? Is the outcome metric still the right one? Has anything changed about the technology or the business context that would alter the approach?

This is not a performance review. It is a calibration. The goal is to catch misaligned initiatives early, not to evaluate people. Organizations that treat quarterly AI reviews as accountability mechanisms rather than learning mechanisms get answers optimized for surviving the review rather than for honest assessment.

Zylver ships AI products: Forge, Signal, Agents, Flows, and Meter. View all products.

Get insights like this delivered monthly.

No spam. Unsubscribe anytime.