Playbook
How to Run an Outsourced SDR Pilot That Actually Tells You Something
A structured outsourced SDR pilot should run 90 days, define success in qualified opportunities, not raw meetings, agree on ICP criteria and copy approval before outreach begins, and include weekly reviews with full access to call recordings and sequence data. The goal is to answer one question: can this team book meetings with real prospects who match your ICP? If a vendor won’t agree to a defined pilot with those terms, that’s the answer.
Most pilots don’t fail because the vendor is bad. They fail because nobody defined what “working” meant before the first email went out, so thirty days in, everyone’s arguing about a number nobody agreed to measure.
Why most SDR pilots fail before they start
A pilot without a written definition of success isn’t a pilot. It’s a trial subscription you can cancel if you get a bad feeling.
The most common failure mode is judging too early, on the wrong metric. A vendor running a real launch model will tell you plainly that the first 10 days produce close to zero meetings by design: ICP definition, list building, messaging, infrastructure setup, and compliance checks all have to happen before outreach starts. Judging a pilot at day 30 measures your list and your setup, not your program. That’s not an excuse for slow results forever, it’s a description of what an honest ramp actually looks like.
The second most common failure is the opposite problem: a vendor who never wants the pilot to end. If every conversation about “how long should we give this” gets answered with “you have to commit to see real results,” that’s not patience, that’s a sales tactic. A real pilot has a defined endpoint and a defined decision at that endpoint, extend, scale, or walk away, agreed before day one, not negotiated after the fact based on how the relationship feels.
The elements of a pilot that produces real signal
A pilot design needs six things in place before outreach starts, or the results at the end won’t actually tell you anything:
- A defined length, with an honest reason for that length (more on this below)
- A written definition of a qualified meeting, agreed before the vendor starts prospecting, not negotiated after the results come in
- ICP and target-list approval, so you know exactly who’s being contacted on your behalf
- A fixed review cadence, weekly, against pipeline quality, not just activity volume
- Full access to the raw data, call recordings, sequence performance, disposition notes, not a vendor-curated summary
- A pre-agreed decision framework, so day 90 (or day 30, or day 60) produces a decision, not a debate
Skip any one of these and you’re not running a pilot. You’re running an unstructured trial and hoping the vibe tells you something useful.
How big should the pilot’s target list and outreach volume actually be?
Most pilot conversations focus entirely on timing and qualification standards, and skip a question that determines whether either of those even gets a fair test: is the list big enough to produce a real signal in 90 days?
One outbound-calling agency that runs pilots for a living, ColdCalls.ca, publishes its own benchmark for this: a minimum of 2,000 to 4,000 verified, enriched contacts for a two-person calling pod running roughly 800 to 1,000 dials a week over 90 days (ColdCalls.ca, “How to run a 90-day outbound pilot that produces real evidence,” updated May 13, 2026, coldcalls.ca). Treat that specific number the way this article treats the 90-day timing convention: it’s one vendor’s published figure, from a company with an obvious interest in you agreeing that a serious pilot needs serious scale, and, not coincidentally, in you buying a bigger engagement from them. But the underlying logic holds regardless of who’s saying it: a list too small to sustain 90 days of consistent outreach volume produces a pilot that looks inconclusive not because the vendor is bad, but because there was never enough raw material to generate a reliable signal.
Ask your vendor directly, before signing: how many verified contacts are in the starting list, what’s the expected weekly outreach volume across whatever channels they’re using, and does that math actually support 90 days of consistent activity without exhausting the list by day 45? A vendor who can’t answer with real numbers, rather than a general sense of “we have plenty of data,” is telling you the pilot wasn’t sized to succeed.
How long should the pilot actually run?
Here’s where a lot of well-meaning advice gets it wrong, including some early guidance in this exact content area: 6 to 8 weeks sounds reasonable, but it’s not long enough to produce a real verdict, and worth saying plainly rather than repeating a number that doesn’t hold up.
Every 4 to 8 week figure found in a review of how vendors and buyers actually describe pilot timing turns out to describe time to first meeting, not the length of the full pilot. A typical honest breakdown looks like this: days 1 to 10 are setup with close to zero output by design, weeks 2 to 4 produce the first handful of qualified meetings, days 30 to 60 are for testing and adjusting messaging and targeting based on what’s landing, and days 60 to 90 are where a repeatable weekly cadence and a trustworthy cost-per-meeting number actually emerge. Cutting the pilot off at week 6 or 8 catches you somewhere in the middle of that curve, exactly when the data is least reliable.
Ninety days is the more consistent standard, and it’s worth naming the tension in that advice directly: nearly every source recommending a 90-day pilot also sells outsourced SDR services with a multi-month minimum contract, so the recommendation is financially convenient for the person making it. That doesn’t make it wrong, but it means you shouldn’t take the number on faith just because it’s repeated often. The more useful cross-check is that at least one voice in this space with the opposite financial incentive, a company that sells an alternative to long agency retainers, converges on the same 90-day figure as the point where cancelling looks premature and staying past looks like denial.
A reasonable middle-ground rule: don’t judge before day 60, and don’t let a pilot without pipeline run past day 120 on hope alone.
How to define a “qualified meeting” before day one
This is the single most important thing to get right, and the thing almost every failed pilot skipped.
A qualified meeting definition needs three components in writing, agreed before outreach starts, not after the first report comes in:
Title and seniority
Specify the exact roles or levels that count. “Decision-maker” is not specific enough; “Director level or above, in a role with budget authority over this category” is.
A qualification standard beyond title
Some version of need, budget, and authority, defined in language you’d both recognize on a call, not just a checkbox.
A no-show and disposition standard
What happens when a scheduled meeting doesn’t happen, or happens but the person isn’t who was promised. Decide this before it happens, not while you’re arguing about whether it counts.
Without this, “qualified” means whatever keeps the invoices coming. With it, every meeting on the calendar is either a legitimate data point or a clear miss, and you’ll know which within the meeting itself, not weeks later when you’re trying to reconstruct what happened.
ICP criteria and list approval, the step most buyers skip
Most buyers assume the vendor will “figure out” the ICP from a kickoff call. That assumption is exactly how a pilot ends up producing meetings with the wrong companies, wrong seniority, or wrong use case, three months in, discovered only when an AE asks why they’re on a call with someone who can’t buy anything.
Before the first email goes out, you should see and approve the actual target list, or at minimum the exact filtering criteria: company size, industry, technology signals, title, and any disqualifying characteristics. This isn’t distrust, it’s the same discipline you’d apply to any campaign you were running yourself. If a vendor resists showing you the list before outreach starts, that’s worth noting as its own signal.
What to review each week, and what to watch for in the data
Weekly reviews are the one point where buyer advice and vendor best practice actually agree, this should happen every week, not monthly, and it should look at pipeline quality signals, not just how many touches happened.
Each week, look at: how many meetings were booked versus how many met the written qualification standard, the reply and positive-reply rate (a leading indicator before meetings show up at all), and specific call notes or recordings for at least a few meetings, not just a summary metric. If a vendor’s weekly update is a dashboard of activity counts with no qualitative detail, ask for the raw material directly. You’re not just tracking a number, you’re building a real read on whether the messaging and targeting are actually working.
Watch specifically for two failure patterns as the weeks go by: message and targeting change constantly without a clean read on any single version (a sign of chasing volume rather than testing a hypothesis), or the opposite, nothing changes at all despite weak early signal (a sign the vendor is running the pilot on autopilot rather than actually working the account).
A worked example: reading a mid-pilot data pull honestly
Most of the advice above is about what to set up before day one. Here’s what actually reading the data looks like partway through, because a pilot that isn’t reviewed at the midpoint is just a 90-day wait with extra paperwork.
ColdCalls.ca, the same agency cited above, published an account of exactly this kind of mid-pilot pivot: at the day-45 mark of a pilot it ran for a FinTech SaaS client, the agency reports finding a stark split in performance by company size, a 2% conversation-to-meeting rate for companies with 200 to 500 employees, versus 7% for companies with 501 to 1,000 employees (ColdCalls.ca, “How to run a 90-day outbound pilot that produces real evidence,” updated May 13, 2026). The agency’s account states it cut the smaller segment from the remaining call list, reallocated effort to the higher-converting one, and the pilot finished with a 7x pipeline return on spend as a result.
Treat the specific multiple as a single vendor’s self-reported result, not an independently verified benchmark; no third party has confirmed those numbers, and the agency has an obvious interest in publishing a strong outcome. What’s transferable regardless of the source is the mechanism, not the multiple: a written qualified-meeting standard and a mid-pilot data pull only produce value if someone actually acts on what the data says. If your day-45 or day-60 review shows one segment meaningfully outperforming another and nobody changes the targeting in response, you’ve collected data without using it, which is functionally the same as not collecting it at all.
What counts as real evidence: cost and pipeline benchmarks to ask for
A pilot can hit every process milestone in this article, written qualification standard, weekly reviews, full data access, and still leave you without a real answer if nobody defined what a good result looks like in dollar terms before day one.
Two numbers are worth asking any vendor to commit to before the pilot starts, so you have something concrete to hold the results against at day 90: a target cost per qualified meeting, and a target pipeline-to-spend multiple. ColdCalls.ca, publishing its own agency benchmarks (same source cited above), puts a healthy cost-per-qualified-meeting range at $1,500 to $2,500 for companies with an average contract value of $25,000 or more, and targets a 3x to 5x return in qualified pipeline generated relative to total pilot spend. Again, this is one vendor’s own published figure, not an independently audited industry standard, and it should be weighed with the same skepticism this article applies to any vendor-sourced number, including the 90-day recommendation discussed above. But the exercise of asking your own vendor for their equivalent numbers, in writing, before day one, is the useful part regardless of which specific benchmark they cite. A vendor unwilling to commit to a cost-per-meeting or pipeline-multiple target going in is a vendor who wants to grade their own pilot after the fact.
How to use pilot results to decide: extend, scale, or walk away
At the agreed decision point, whether that’s day 60, day 90, or somewhere in between based on your sales cycle, you should be answering one question with real data: is this producing qualified pipeline that a rep would actually want to have booked for them? If meetings are happening but not converting, that’s a different diagnosis, see diagnosing a conversion problem after the pilot for the specific root causes.
Walk away if, by day 90, the meeting-to-qualified-opportunity rate is low and hasn’t improved despite adjustments made along the way, or if the vendor has been unwilling to show you raw data throughout. Extend, without scaling, if the signal is mixed, some good meetings, some clear misses, and you can point to a specific, testable change (a different segment, a different message angle) worth trying before committing further. Scale if the qualified-meeting rate is consistent, the vendor has been transparent throughout, and you can see a repeatable weekly cadence rather than a lucky month.
One honest note worth naming here: real buyers don’t always wait for a clean 90-day endpoint before deciding. It’s a documented pattern that some buyers walk away as early as the first month on a gut read of ROI, well before the data is actually reliable enough to support that call. That’s understandable, sunk cost is real, and so is impatience, but it’s worth naming directly: a decision made at day 30 is a decision made on incomplete information. If you’re going to walk early, walk early with your eyes open about why, not because the pilot “felt” wrong before it had a real chance to produce a clean signal.
For the walk-away case specifically, it helps to know what the alternative actually costs. A fully loaded in-house SDR runs about $154,500 in year one once salary, benefits, tooling, recruiting, and turnover re-ramp are counted, and median annual SDR turnover industry-wide sits at 40% (The Bridge Group, “SDR Models, Motions & Metrics: 2025 Research Report,” February 6, 2025, bridgegroupinc.com; alleyoop.io/true-cost-of-an-sdr, 2026). A failed 90-day pilot is a real loss, but it’s a small one next to the cost of skipping the pilot entirely, hiring in-house on an unproven motion, and discovering the same targeting or messaging problem eight months and six figures into a ramp you can’t easily unwind.
How much of your own time a real pilot actually takes
A pilot isn’t a check you write and a calendar you wait on. It requires real hours from someone on your side, and buyers who go in expecting a fully hands-off engagement are usually the ones who disengage after kickoff and then wonder why the results are inconsistent.
One vendor’s own published onboarding framework puts a concrete number on this: roughly 5 to 10 hours from a sales leader in the first two weeks, covering the ICP workshop and messaging review, then a lighter but consistent cadence after that, around a 30-minute weekly check-in plus a monthly performance review through the rest of the pilot (Cold Call Me, “What to Expect in Your First 90 Days With an Outsourced SDR Team,” June 2, 2026, coldcallme.com). Treat the specific hours as one vendor’s estimate rather than a universal figure, since the honest range depends on how complex your ICP and product are. But the direction is consistent with everything else in this article: engagements where the buyer stays involved through the weekly reviews and the day-45 or day-60 data pull outperform the ones where the buyer signs the contract and checks back in at day 90.
Should pilot milestones be written into the contract?
Yes, and it’s worth treating this as a real negotiating point rather than boilerplate. Everything in the checklist below, the qualified-meeting definition, list approval, weekly review cadence, is only enforceable if it’s actually in the contract or SOW, not agreed verbally in a kickoff call and left to memory.
Specific milestones worth naming in writing: when the first qualified meeting should reasonably appear (by week three is a common marker in the buyer-side timelines vendors themselves publish), what the weekly reporting will include, and what happens if a milestone is missed, an extension, a remedy, or a defined off-ramp. A vendor who pushes back on putting their own stated timeline into the contract is telling you they don’t expect to be held to it.
A 90-day pilot checklist
Use this before you sign anything, whether the vendor on the table is Alleyoop or anyone else:
Before outreach starts:
- [ ] Written definition of a qualified meeting (title/seniority, qualification standard, no-show handling) agreed and in the contract or SOW
- [ ] Target list or exact targeting criteria reviewed and approved
- [ ] Messaging and copy reviewed before it goes out, not after
- [ ] Weekly review cadence scheduled in advance, with a named point of contact
- [ ] Agreement on what data you’ll have access to (call recordings, sequence performance, raw disposition notes)
- [ ] A defined pilot end date and a pre-agreed decision framework for what happens at that date
- [ ] List size and expected weekly outreach volume reviewed against the 90-day timeline (a list too small to sustain consistent volume for 90 days will produce an inconclusive pilot regardless of vendor quality)
- [ ] Key onboarding and reporting milestones written into the contract or SOW, not just agreed verbally at kickoff
Days 1 to 10:
- [ ] Confirm ICP, list, and messaging are locked before volume ramps
- [ ] Expect close to zero meetings; this is normal, not a red flag
Weeks 2 to 4:
- [ ] First qualified meetings should start appearing
- [ ] Review reply rate and positive-reply rate as early leading indicators
- [ ] Do not make a go/no-go call yet
Days 30 to 60:
- [ ] Review meeting-to-qualified-opportunity rate, not just meetings booked
- [ ] Test and adjust one variable at a time (segment or message), not everything at once
- [ ] Flag any pattern of vague or shifting definitions of “qualified”
Days 60 to 90:
- [ ] Assess whether a repeatable weekly cadence has emerged
- [ ] Calculate a real cost-per-qualified-meeting number and compare it against the target agreed at kickoff
- [ ] Calculate the pipeline-to-spend multiple generated and compare it against the target agreed at kickoff
- [ ] Make the extend, scale, or walk-away decision using the criteria above, not gut feel
Frequently asked questions.
How long should an outsourced SDR pilot run?
Ninety days is the most consistently supported window for a reliable go/no-go read. Shorter figures circulating online (4 to 8 weeks) almost always describe time to first meeting, not the full pilot length. A reasonable rule: don’t judge before day 60, and don’t let a pilot without any pipeline run past day 120.
What should I define before an outsourced SDR pilot starts?
A written definition of a qualified meeting (title, qualification standard, and no-show handling), an approved target list or targeting criteria, a weekly review cadence, and a pre-agreed decision framework for what happens at the pilot’s end date.
How do I know if my outsourced SDR pilot is actually working?
Track the meeting-to-qualified-opportunity rate, not raw meetings booked. Review call recordings or notes directly, not just a summary dashboard. A pilot producing a consistent, improving rate by day 60 to 90 is a good sign; one producing volume with no quality improvement is not.
Should I pay for a pilot per meeting or as a fixed fee?
Both models exist in the market; a flat monthly retainer plus a performance bonus is the more commonly reported structure, though some buyers negotiate a fixed scope with payment only for meetings that meet the written qualification standard. See how flat program pricing works for one example of this structure. Whichever model you use, the qualification standard matters more than the pricing structure itself.
What’s a reasonable reason to walk away from an SDR pilot early?
A vendor unwilling to define “qualified” in writing before outreach starts, unwilling to share raw call data, or unwilling to agree to a fixed pilot end date with a real decision point. Any one of those is a signal worth taking seriously before the pilot even begins.
How big should the target list be for an outsourced SDR pilot?
One agency’s published benchmark puts a minimum of 2,000 to 4,000 verified, enriched contacts for a two-person calling pod running roughly 800 to 1,000 dials a week over 90 days (ColdCalls.ca, 2026). Treat the specific figure as one vendor’s own number rather than an audited industry standard, but the underlying test holds regardless: ask your vendor whether their list math actually supports 90 days of consistent volume without running dry by day 45.
What cost-per-qualified-meeting and pipeline return should a pilot produce?
One agency’s published figures put a healthy range at $1,500 to $2,500 per qualified meeting for companies with an average contract value of $25,000 or more, and a target of 3x to 5x qualified pipeline generated relative to pilot spend (ColdCalls.ca, 2026). Treat these as one vendor’s own benchmark, not an independently audited standard, but ask your own vendor to commit to their equivalent numbers in writing before day one.
How much of my own time does running a pilot actually take?
Expect roughly 5 to 10 hours from a sales leader in the first two weeks for the ICP workshop and messaging review, then a lighter but consistent cadence after that, around a 30-minute weekly check-in plus a monthly performance review (based on one vendor’s published onboarding framework; Cold Call Me, 2026). Engagements where the buyer disengages after kickoff consistently underperform ones where the buyer stays involved through the weekly reviews.
Should pilot milestones be written into the contract?
Yes. A qualified-meeting definition, list approval, and weekly review cadence are only enforceable if they’re in the contract or SOW, not agreed verbally at kickoff. A vendor who resists putting their own stated timeline in writing is telling you they don’t expect to be held to it.
Done reading? Start measuring.
Twenty minutes, your numbers, and a straight answer on whether a program fits. If we’re the wrong fit, we’ll say so.
The assist is ours. The win is yours.