Same Sample

How We Test

By Cuong Tran · Published 2026-08-30 · Last updated 2026-08-31

Affiliate disclosure: Same Sample may earn a commission from some links on this page. The commercial relationship is stated next to each product, and products we have no approved affiliate relationship with carry ordinary links. Affiliate status and commission rates are not inputs to the scoring model.

Same Sample publishes the full test protocol before the test is run, so the procedure cannot be adjusted afterwards to suit a result. The first edition is a pilot of 40 records, small enough that every published figure carries a confidence interval and close pairs are reported as ties rather than ranked.

Why publish the method before the results?

A protocol written after the numbers are known can be shaped to fit them. Publishing it first fixes the sample, the metrics and the scoring rule in advance, so the results are constrained by the method rather than the other way round.

A vendor's own page will always describe that vendor favourably. A comparison assembled from several vendor pages inherits every one of those biases and adds nothing a reader can check. The only comparison worth reading is one where somebody ran the tools and said exactly how.

Stating the procedure in advance also makes this page falsifiable. Anyone who disagrees with a result can rebuild an equivalent sample from the source types listed below and check whether the same ordering comes out.

Is a vendor allowed to be benchmarked?

Every vendor's terms of service are read before a measured figure is published, and the answer is recorded per vendor with the clause, the source URL and the date checked. Apollo.io forbids disclosing benchmark results without prior written consent, so no measured Apollo.io figure appears here.

Apollo.io's terms, section 3(f)(vii), forbid disclosing the results of any platform or program benchmark test to third parties without Apollo's prior written consent. Same Sample treats that as binding. Apollo.io still appears in the comparison with its published pricing and features, because those are public facts, and the page states plainly why no measured figure sits beside them.

Snov.io, Hunter.io, Lusha and RocketReach carry no equivalent clause as of 31 August 2026, so measured figures for those four may be published. Each carries its own conditions instead: no scraping or automated extraction, and for Lusha and RocketReach no returned record may be republished. Those conditions shape how the test runs rather than whether it can be published.

The permission state is enforced by the site generator, not by memory. A vendor whose permission is anything other than allowed has every measured cell forced to TBA before the page is rendered, so a figure cannot reach the site by accident even if it is entered into the data file.

How do you know what the correct answer is?

The problem is inverted. Instead of verifying answers after the fact, the sample is built only from people whose work address a primary source has already published, so the answer key exists before the test begins.

Three common approaches were considered and rejected. Sending test mail and counting bounces gives the truest signal but means mailing hundreds of strangers who did not ask to be contacted, so it is out on ethical grounds alone. SMTP probing cannot show an address belongs to the intended person, and most corporate servers run catch-all, which returns a flattering result for addresses that do not exist. Treating agreement between tools as truth assumes the majority is right, when these tools share upstream data sources and therefore tend to be wrong together, and it cannot score a tool that beats the majority.

What remains is to select the sample from published ground truth. Four source types qualify: git commit metadata on public repositories, where developers configure git with their work address; corresponding-author lines on scientific papers, where the author supplied the address for contact; company press and newsroom pages; and conference speaker pages. Each record in the sample stores a full name, a company domain, the correct address, and the source URL that proves it.

What is in the sample, and what does a pilot of 40 support?

A frozen set of 40 records, balanced at the level of each individual dimension rather than across every combination of them. Forty records is a pilot: enough for one overall figure per tool with a stated confidence interval, not enough to name a winner inside any segment.

The constraint on size is the free tier of the tools being measured. Several grant only a few dozen lookups and one never refills. A larger sample is a paid exercise, and buying five subscriptions to produce a first result is not something Same Sample can claim to have done before it has.

Records are drawn across three dimensions: company size (under 50 people, 50 to 500, over 500), region (North America, Western Europe, elsewhere), and job function (engineering, sales and marketing, executive). Those three dimensions describe 27 combinations. Forty records is about 1.5 records per combination, so the balance holds along each dimension separately and not across the grid. A claim of the form "tool X is best for small European engineering teams" would rest on one or two observations, so Same Sample does not make it.

For reference, three dimensions of three levels each need roughly 270 records before every combination holds about ten observations. Segment conclusions start at that scale, not this one. The pilot exists to fix the protocol and produce a first honest overall figure, and the sample grows from there.

The core sample is frozen to a file and reused unchanged on every future run, because changing it destroys comparison across time. A rotating refresh sample is added each quarter alongside it so the benchmark does not slowly overfit one fixed set of people.

How is the test run?

All tools run inside one 48 hour window on identical input, first name plus last name plus company domain, with no extra hints, and each tool's raw export is kept unedited.

Contact data changes week to week, so a run spread across a month would measure the calendar rather than the tools. Every tool sees the same sample inside the same window.

Lookups run through each vendor's own interface or its documented API, never through scraping or any automated tool the vendor has not provided. Four of the five vendors covered here restrict automated or systematic extraction in their terms of service, so respecting that is a condition of running the test at all, not a courtesy.

The plan used for each tool is named on the comparison page next to its result, because a free tier and a paid tier of the same product do not behave alike. Where a vendor prices or behaves differently through its interface than through its API, the route used is stated. Lusha is a live example: its pricing page puts a revealed phone number at 10 credits while its API documentation says 5.

Raw exports are stored per tool and never hand-corrected. Scoring runs off those files, and no returned address is ever published. Lusha and RocketReach both forbid making returned records available to third parties, which is a further reason the sample and the raw results stay private.

What exactly is measured, and how is a result classified?

Every returned address falls into one of five classes rather than into right or wrong. An address that differs from the reference is not automatically an error, because one person can hold a primary mailbox, an alias and an address on an old domain, and all three can work.

The five classes are: exact reference match, where the address equals the one a primary source published; corroborated alternate, where the address differs but a second independent primary source confirms it belongs to that person at that company; unverifiable alternate, where the address differs and no source confirms or denies it; wrong or contradictory, where the address belongs to someone else or contradicts public evidence; and no result, where the tool returned nothing.

Usable rate counts exact reference matches plus corroborated alternates over the whole sample, and the ranking follows that figure. Unverifiable alternates are reported as their own number and never folded into the wrong column, because counting an unconfirmed address as an error would penalise a tool for something Same Sample cannot demonstrate.

Same Sample does not measure deliverability and does not use that word. Whether mail sent to an address arrives depends on the receiving server, and no part of this protocol tests that. An earlier version of this page said a mismatched address means mail does not arrive. That was an unsupported claim and it has been removed.

How is cost per usable result calculated?

Cost comes from the credits a run actually consumes, not from dividing a monthly plan price by the results in a 40 record sample. Running 40 lookups does not consume a monthly plan, so that division would overstate the true cost by a large multiple.

The unit cost of a credit is the plan price divided by the credits the plan includes. Cost per usable result is that unit cost times the credits the run consumed, divided by the usable results produced. Where a vendor charges a credit for a lookup that returns nothing, that is stated, because it changes the real cost materially.

A single figure hides how the answer moves with volume, so each edition also reports cost per usable result at 100, 1,000 and 10,000 lookups a month using each vendor's published overage rate. A tool that is cheapest at 100 lookups is often not cheapest at 10,000.

Tools also carry a scope label, because an email finder and a full sales intelligence platform are not the same purchase even when both return an address. Comparing headline prices without that label would mislead.

What are the limits of this benchmark?

Three limits are structural and stay on this page permanently. The sample is drawn from people whose address is already public, which is not a random slice of B2B contacts. Forty records gives wide confidence intervals. And no email is ever sent, so nothing here measures deliverability.

On selection bias: an address only enters the answer key if a primary source published it, so the sample leans towards people with a large public footprint. Developers who configure git with a work address, authors who print a contact line on a paper, and named press contacts are over-represented relative to a working prospect list. The honest description is therefore narrow: this pilot evaluates how well each tool recovers publicly documented B2B contact addresses from a defined sample, not how each performs on all B2B contacts. Each record stores its source type so a later edition can report whether a tool is stronger on one footprint type than another.

On uncertainty: every published rate carries a Wilson 95 percent confidence interval computed on the 40 observations. Wilson rather than the textbook normal approximation, because at this sample size and at rates near zero or one hundred the normal interval runs outside the possible range and covers badly. Where two tools ran on the same records they are compared with an exact paired test rather than by subtracting two independent rates, which is what lets 40 records say anything useful. Where that test does not separate two tools, the page reports a statistical tie rather than inventing a first and second place.

Records are dropped and replaced when the source URL stops resolving, or when the person visibly changed employer between sample construction and the run, since neither outcome says anything about the tool. The sample itself is never published. Same Sample writes about B2B contact data, and publishing a list of real people's addresses to prove a point about contact data would forfeit the standing to write about it. An anonymized dataset carrying result classes and stratum bands, with no names and no addresses, is published alongside each edition instead.

How is a benchmark versioned and re-run?

Every edition carries a version string such as email-finders-2026-q3, together with the protocol, dataset and scoring versions used. A re-run publishes a new edition and keeps the old one rather than overwriting it.

Each edition records the tools and plans tested, the run date, the sample design, and the limitations known at the time. The next edition shows what moved: which tools improved, which prices changed, what the method changed, and where the ranking changed and why.

A comparison site that silently replaces last quarter's numbers with this quarter's is unfalsifiable. Keeping both editions readable is what makes any claim here checkable.

How does this affect the rankings?

Rankings follow the measured results. Commission rates differ widely between the tools covered here and are not an input to the order, which is why every product on a page carries a referral link rather than only the top pick.

We earn a commission when a reader buys through a referral link. Because every tool on a comparison page carries one, revenue does not depend on which one a reader picks, only on the page being useful enough to act on.

Where a tool ranked highly pays us nothing, or pays less than the ones below it, that is stated on the page itself.

How are corrections handled?

A wrong number is corrected in place with a note saying what changed and when, rather than quietly edited.

Vendors change pricing and limits without notice. If a figure here no longer matches the vendor's own page, tell us and it will be re-checked.

Write to cuongtran@samesample.com.