Pipeline playbook

How to build new logo pipeline for MLOps and AI production engineering

A company hires its first AI engineer, ships a feature, and discovers three months later that it cannot say whether the model got worse, what the inference bill will be next month, or who is on call when the answers go wrong. All three of those discoveries were predictable from the job post. This is how to read it and arrive before the bill does.

Who actually signs

At a company between fifty and a thousand people, the signer is the VP of engineering or the CTO, and increasingly a head of AI if the company has created the title. The champion is the one or two person AI team, who know exactly what is missing and are too busy shipping features to build it. The person who becomes the champion later is the CFO, when the inference bill arrives.

The buying reason is nearly always the same: something is in production, it is a demo that became a product, and there is no way to tell whether it is working. The company has models and no evaluation. It has prompts and no versioning. It has a bill and no attribution.

The one sentence version

Your buyer is an engineering leader with an AI feature in production, one engineer who owns it, no evaluation harness, and a cloud bill that doubled last month for reasons nobody can name.

The triggers, and where each one is visible

  • First AI or ML engineer job posts. A company advertising for its first AI engineer has decided to build and has budgeted for one person. The launch that person ships will need everything this page describes, and the posting is public months before it.
  • AI feature launches. Product changelogs, press releases and app store updates announce them. A launch is a production system, and a production system at a company with a two person AI team has no evaluation.
  • Tool names in job posts. Postings that name orchestration frameworks, vector databases, experiment trackers or model serving tools are companies that have adopted the tools and are looking for someone to make them work together.
  • Funding rounds with AI in the narrative. The press release names the AI product the round will fund. The round also funds the inference bill, which is the number that brings the CFO into the conversation.
  • AI governance and policy pages appearing. A company publishing an AI policy or a responsible use page has been asked about it by a customer, and is about to be asked how it evaluates the system in production.
  • Public model failures. A chatbot that gave a customer a wrong answer that made the news is a company that has just learned what evaluation is for, and every peer that read the story has the same gap.
  • Enterprise AI procurement policies. Large buyers are publishing what they require of AI vendors, including evaluation and monitoring evidence. A vendor selling to those buyers has a document to produce.

The first hire and the launch are the two to build on. One tells you the company is about to have the problem. The other tells you it has it.

Qualify in sixty seconds

  • Is there AI in production, or only a demo? A feature in the product with users is in scope. A prototype in a notebook is a future prospect.
  • How big is the AI team? One to three people is the sweet spot. None means nothing to build on. Twenty means they may build the platform themselves and buy your team for the hard parts.
  • Is there a cost signal? Inference spend in a funding narrative, a CFO involved, a product with real usage. The bill is what makes the platform work fundable.
  • Is there a regulated or enterprise buyer? Healthcare, finance, EU customers or enterprise procurement. Any of those turns evaluation from good practice into a required document.

The angle that gets replies

Lead with the evaluation gap. Not AI strategy, not model selection, not the framework. The reader shipped something and cannot tell whether it is getting better or worse. That is the sentence that gets read.

Three openers you can adapt

  • On a first AI engineer hire"Saw the AI engineer posting. When that person ships the first feature, the thing that is usually missing on day one is an evaluation set, meaning a few hundred real examples with known good answers that every prompt or model change gets scored against. Without it, the first regression reaches customers before anyone knows. Building the set takes about two weeks and does not need the hire to be in seat. One page on how, attached to nothing."
  • On an AI feature launch"Congratulations on the launch. Two questions most teams cannot answer in month three: did the outputs get worse after the last prompt change, and what does each customer cost in inference. Both are answerable with an evaluation harness and per request cost attribution, which together are about a month of work. Happy to send what the harness looks like for a feature shaped like yours."
  • On a job post naming the orchestration framework and the vector database"Your posting names the framework and the vector store, which usually means a retrieval system is in production. The failure mode with those is retrieval quality drifting as the index grows, invisibly, until answers get vague. A retrieval evaluation on your own queries takes about a week and shows whether it is happening. No charge to scope it."

Each one names a specific thing that is missing, in the engineer's vocabulary, with a duration. None of them says the word strategy.

What not to send

  • "AI transformation" or "AI strategy." The reader has already decided and shipped. Strategy is what they did last year, and the word marks you as a workshop seller.
  • "We build custom LLMs." Almost nobody needs one, the reader knows it, and the claim suggests you sell what you can build rather than what they need.
  • The word agentic in the first sentence. It is a category name and it has been on every vendor's homepage for a year.
  • "We will fine tune your model." Maybe, after the evaluation shows fine tuning would help. Offered before, it is a solution looking for a problem.

The objection you will hit

We have an AI engineer. One. Shipping features, on call for the production system, and building the evaluation harness in whatever time is left, which is none. The platform work, meaning evaluation, monitoring, cost attribution and the data pipelines, is a different job from the feature work, and the company has budgeted for one person doing both. Say that the engineer is exactly who you would work with, and that the harness is what lets them ship faster.

The second is we use the cloud vendor's managed AI service. Which hosts the model. It does not score the outputs, does not detect drift, does not attribute cost to a customer, and does not build the pipeline that keeps the retrieval index current. The managed service is where the platform starts.

The third is we are not ready for MLOps. Evaluation is not a maturity stage, it is day one. A company that cannot tell whether its outputs are getting worse is not early, it is exposed. Say it gently and offer the two week evaluation set as the whole first step.

Deal shape

  • Production readiness assessment of the current AI feature: commonly $10K to $30K, and the engagement that opens most relationships.
  • Evaluation harness and monitoring build: $40K to $120K, delivered as working code in the client's repository.
  • Platform build, including cost attribution, pipelines, versioning and deployment: $100K to $400K.
  • Fractional AI platform lead retainer: $8K to $25K a month, and the most durable revenue in the niche.
  • Retrieval quality evaluation on an existing system: $15K to $40K, and a second door into companies that already shipped.
  • Signer: VP Engineering, CTO or head of AI. Champion: the AI engineer. Cycle: two to eight weeks, and days after a public failure or a surprising bill.

The readiness assessment is the funnel. It is small, it is about a system that already exists, and it produces a written list of what is missing, which is the build.

A cadence you can actually run

  • Weekly, pull AI and ML engineer job posts and flag the ones that are a company's first, and the ones naming specific tools.
  • Weekly, pull AI feature launches from changelogs, press and app store updates at companies with small engineering teams.
  • Weekly, pull funding rounds with AI in the narrative.
  • Monthly, pull new AI policy pages and enterprise AI procurement requirements.
  • Qualify against the four checks, with the production question first. One message per account, naming the missing thing. Twenty accounts a week is a full program.
  • Three touches over two weeks, then stop. The next launch or the next public failure in their category is a fresh reason to write.

The company posts the hire and then posts the launch. The consultancies that grow are the ones writing about the evaluation set in between.

The sending mechanics most people get wrong

Everything above is about who and what. This is about how, and it is where most outbound in this niche quietly dies. Seven rules. None of them are optional.

1.Three to five sentences. That is the whole email.

Your reader is on a phone between meetings. One observable fact about their company, one consequence they have not thought about, one specific thing you would do. Anything past five sentences is a memo, and memos get archived unread.

2.Lead with a technical differentiator that turns into a number.

The messages that work best name something concrete you do differently and translate it into time or money saved. In this niche the differentiator is the evaluation harness delivered as working code. A consultancy that can say how many production AI systems it has instrumented, and what fraction of the regressions those harnesses caught before customers did, has the only number an engineering leader who shipped without one will believe. The second is cost: state the average inference cost reduction your attribution work produced across the last dozen clients.

Most services firms do not have a technical differentiator, and pretending to have one reads as exactly that. The substitute is a verticalized case study: a company like theirs, what you did, what happened, in one sentence. For this niche the line is: a 120 person legal software company, one AI engineer, retrieval feature in production for four months with no evaluation, harness and cost attribution built in five weeks, a prompt regression caught in week six before release, inference cost per customer reduced by a stated fraction through routing. The five weeks and the caught regression are what the reader will check.

3.Ten to twenty emails a day per mailbox. Not a hundred.

Sender reputation is scored per mailbox and per sending domain. One inbox pushing a hundred cold emails a day looks like exactly what it is, and the penalty lands on the domain, which means it lands on your client correspondence too.

If the math says you need more volume, the answer is more mailboxes on more warmed sending domains, separate from the domain you invoice from. It is never more volume per mailbox. Twenty accounts a week at three touches is about twelve emails a day, one warmed mailbox. AI hiring and launches are continuous rather than seasonal, so a second mailbox is only needed if the target list widens.

4.Write ten versions of every step and test them.

Versions A through J, not A and B. Rotate subject lines and bodies. You learn which angle is actually working instead of guessing, and there is a second reason that matters more: identical bodies going out over and over is one of the patterns postmaster tools flag. Variation is a deliverability tool as much as a testing one.

Subject line seeds for this niche, each of which should become several variants: "before your AI engineer starts", "month three after the launch", "the retrieval index since launch". Lower case, no punctuation tricks, and nothing that would look odd in a reply from a colleague.

5.Stop at three.

Most replies arrive on the first and second email. The third is already thin. Every touch past that raises the odds the whole thread gets classified as spam, and that classification follows the mailbox to the next person you write to. The long cadence is over. Three touches, each with something new in it, then leave them alone for ninety days.

6.Know what good looks like.

A one percent reply rate with a quarter of those replies positive is a healthy trigger based program. Anyone quoting you double digit reply rates is counting out of office messages or selling a course.

7.LinkedIn Sales Navigator is not optional.

Every other data source tells you who held a title at some point. Sales Navigator tells you who holds it today, because the person maintains it themselves. That is the difference between a three percent bounce rate and a fifteen percent one, and bounces are scored against the mailbox the same way spam complaints are. Verify the name there before anything goes out.

It is also the cheapest trigger detector you will own. The job change filter surfaces people who arrived in a role in the last ninety days, which is the moment they have budget and no incumbent. The posted recently filter surfaces companies talking about the exact problem you solve. Account lists with headcount growth alerts tell you who is scaling before the press release does. For this niche the saved search is headcount 50 to 1,000 in software and technology enabled services, titles CTO, VP Engineering, Head of AI, ML Engineer and AI Engineer, with the job change alert on for the AI titles specifically and a keyword alert on the orchestration frameworks, vector databases and experiment tracking tools across job listings. The tool names in a posting are the trigger itself. Navigator confirms the person, and the company's changelog supplies the launch.

Use it for the research and the verification, not for the message. InMail reply rates are a fraction of email, and the person who replies to a thoughtful email is the same person who ignores a connection request with a pitch attached. Pull the work email from a data provider once Navigator has confirmed the person is real and current.

None of this is specific to your niche. All of it is specific to whether anyone ever reads the angle you spent an hour getting right.

If you would rather not run it yourself

That is what we do. ExpertLayer runs this exact loop for expert led firms: the weekly hire and launch pull, the tool name matching, the qualification with the production question first, the angle per account naming the missing thing, the sending across warmed mailboxes, and the reply reading. You take the conversations and build the harness.

The first step is free and it is the same research described above. Send us your website and we will come back with 10 companies that hit these triggers right now, with the hire or launch, the contact, and the opening line for each.

Questions from people running this

Everyone is selling AI consulting right now. How is a production focused practice different?+

By refusing to sell strategy. The market is full of workshops about what a company could do with AI. A practice that only engages with companies that have already shipped something, and only on the question of whether it works in production, is in a different conversation from the workshop sellers, and the buyer can tell from the first sentence.

Is the first AI engineer hire really the trigger, or the AI feature launch?+

The hire is earlier and the launch is louder. The hire means the company has decided to build and has one person to do it. The launch means the one person now has a production system, an on call rotation of one, and no evaluation harness. Both are public. Write to the hire about what the launch will need, and write to the launch about what it is missing.

The cloud vendors offer managed AI services. Does that remove the need?+

It removes the need to host the model. It does not evaluate the outputs, monitor drift, control the inference bill or build the data pipelines that feed the thing. Most companies discover that in month three when the bill arrives and nobody can say whether the answers got worse. Managed inference is the beginning of the platform, not the platform.

Should I mention the EU rules for AI systems?+

Only to companies selling into Europe, and then briefly, because the obligations are real and phased and most of the reader's competitors are not tracking them. For everyone else, lead with the evaluation gap, which is the thing that hurts this quarter.

Related

Start with 10 free targets.

Send us your website and the kind of customers you want. We come back with 10 bullseye prospects, why each one is relevant, and the outreach angle we would use. No charge, no obligation.

See how the review works