Back to Blog
Back to Customer Stories
Summary

We have spent 2 years discussing the first optimization hurdle in AI Search, winning the answer. The next frontier is coming into sight, winning the action that the AI takes.

I see this in my daily work. AI can recommend one company when asked for advice and prefer another when asked to complete a task. This is most obvious in situations where I am using cloud services to build internal tools at AirOps. Supabase, Elevenlabs, Cloudflare, Vercel are commonly reached for by agents when they are tasked with building apps. As we know with Claude Code, what starts with software development usually comes to other areas of knowledge work.  

We decided to dig in. We conducted a set of research projects around this. In the first one, we conducted 1,350 AirOps Research trials across 50 enterprise software companies and three agent environments. In 232 of 449 matched comparisons, changing the request from a general recommendation to a concrete action changed which company finished first.

GitHub's peer-only win rate fell from 95.2% to 36.5%. MongoDB's rose from 6.1% to 32.3%. Cloudflare remained a strong choice under both requests.

The same category produced a different winner once the job became progressively more concrete.

This is the shift marketers need to understand. Some AI interactions end with an answer: explain, compare, recommend. Others move toward action: choose a tool, find a usable setup path, request permission, and try to complete the work. Additionally, the work to improve Action queries will require close collaboration across teams, including product and engineering. We will cover these in future articles. 

Our study examined the decision between those two endpoints: which company an agent preferred as a request moved closer to implementation. The agents did not install the products or complete the work, so these results measure implementation preference, not successful execution yet. 

For marketers, the practical change is in the criteria an agent evaluates. A recommendation can be shaped by relevance, authority, and how clearly a brand is represented. A setup decision also asks whether the product looks usable in the agent's environment. Is the shape of the tool, connector or app suitable to the task at hand. The agent is looking into its known and available toolbox and is picking out the “sharpest tool” for the job. It’s a paradigm shift. 

This study tested that decision point. It did not establish which individual website or product features caused each choice. Those factors need to be tested separately.

We asked the same buying question three ways

In each test, we named one software company and a set of likely competitors. The agent then answered three versions of the same buying question:

  1. General recommendation: Which tool should the team choose?
  2. Ready to implement: Which tool offers the clearest path to a safe first use?
  3. Ready for an agent: Which tool has the clearest public setup path for an AI agent?

We ran every comparison across 50 companies in Claude Code, Codex, and Cursor, repeating each trial three times. That produced 1,350 decisions. Because every company and competitor was supplied in advance, this is a study of preference within a known competitive set. It does not measure organic discovery or market share this time. 

Figure 1. Did the agent choose the same company?

In 232 of 449 matched comparisons, the agent changed its first choice. The products and competitive set stayed constant; only the requested job changed.

Being named did not guarantee preference. Seventeen of the 50 companies never finished first, and 23 won zero or one of their 27 trials. Together, those findings show that inclusion and preference are separate. A company can be named, compared, and still lose when the agent evaluates a more practical route into the work.

A concrete request can break a default, or strengthen one

To identify the brands agents reached for by default, we excluded the company named in the prompt and counted only cases where a company appeared as an alternative. The rates below show how often each alternative finished first.

Figure 2. The same brands performed differently as the request became more practical

One example to illustrate this. GitHub won 120 of 126 peer comparisons under the general-recommendation prompt, then just 46 of 126 under the clear-setup prompt. Databricks also fell sharply. MongoDB and Datadog moved in the opposite direction, while Cloudflare remained strong under both conditions.

The setup prompt did not help every company or make every category more competitive. It reshuffled preference: some established defaults weakened, some strengthened, and others held their position.

AI preference already varies with the model, retrieval, geography, session history, memory, and user context. A request closer to implementation adds another source of variation: what the agent believes it can use in its current environment.

Different agent environments produced different winners

The requested job was only one source of variation. The agent environment also changed the result. Across the general-recommendation tests, Claude Code, Codex, and Cursor selected the same winner just 57.3% of the time. On the clear-setup comparison, agreement fell to 28.7%.

For marketers, the practical conclusion is straightforward: performance in one agent environment cannot be assumed to carry into another. Brands need to test the tasks and environments their customers are most likely to use.

What marketers should do now

Start with three changes to measurement.

  1. Separate recommendation prompts from implementation prompts. A blended visibility score can hide a strong position at one stage and a weak position at the next.
  2. Test the same commercially important job across multiple agent environments. Track first choice, rank, and how often the winner changes as the request becomes more practical.
  3. Investigate the public path behind the result. When a company loses on setup, review the documentation, requirements, permissions, examples, integrations, and success signals an agent can see. This study did not isolate those factors, so treat them as a diagnostic checklist rather than a proven ranking formula.

Mention share still matters, but it no longer describes the whole decision. Marketers should also know whether their brand remains preferred as a customer moves from asking about a product to trying to use one.

The next article will examine the harder question: what makes a product easier for an agent to choose and use?

Methodology note

The study includes 1,350 trials: 50 companies x three agent environments x three prompt scenarios x three repeats. The environments were Claude Code, Codex, and Cursor.

Every prompt named the company under evaluation and supplied likely competitors. Results therefore measure ranking and first-choice preference inside a seeded comparison. They do not measure organic discovery, market share, buyer behavior, or completed actions.

Agents inspected public evidence but did not authenticate, install a product, or complete an end-to-end task. Peer-only rates include only trials where a company appeared as an alternative; denominators are shown because peer exposure was not identical for every company.

Alex Halliday
Co-Founder & CEO

Win AI Search.

Increase brand visibility across AI search and Google with the only platform taking you from insights to action.

Book a Demo

Get the latest on AI content & marketing

New insights every week
Thank you for subscribing!
Oops! Something went wrong while submitting the form.
Part 1: How to use AI for content workflows - ship winning content with AI