Key takeaways
- We tested 10,000 addresses across 251 domains, every one of them a catch-all, using 500 real professionals (about 5,000 email permutations) and 5,000 fictional addresses, across 11 tools in August 2026.
- Seven of the fourteen tests left at least 94% of all 10,000 records unresolved. On a list made entirely of catch-all domains, most of the market returns almost nothing a sender can act on.
- Real contacts found ranged from 4% to 98.2% of the 500 real people, 20 people at the bottom against 491 at the top.
- Allegrow found the most real people: 491 of 500 (98.2%), ahead of ZeroBounce at 470 and BounceBan at 460.
- Million Verifier resolved far more of the list than its peers and called 640 of 5,000 invented addresses valid, a 12.8% false-positive rate.
Introduction
Every email verification vendor advertises accuracy above 99%. Those numbers get thrown around with almost nothing published behind them: no dataset, no method, no denominator, no way for a buyer to check any of it. They are not comparable to each other either, because each vendor defines accuracy in whichever way suits it.
On catch-all domains none of it measures the thing that decides whether verification was worth paying for, which is whether the tool can resolve a specific contact into a definitive, actionable status rather than handing it back undecided.
So we measured it. In August 2026 we assembled 10,000 addresses at 251 domains that are all catch-all, mixing 500 real professionals with 5,000 invented ones, and ran the identical list through 11 tools. We counted three things: how many of the 500 real people each tool found, how many of the 5,000 invented addresses it wrongly called valid, and how much of the full 10,000 it simply declined to make a decision on.
The headline result is that most tools are not cut out for the job. Seven of the fourteen tests left at least 94% of the list unresolved, returning "catch-all", "unknown", or "risky" on almost everything. Six of them found fewer than 45 of the 500 real people. That is not a marginal difference in accuracy; it is most of the market handing back a list you cannot use. If you are trying to work out which are the best email verification tools for a B2B list, the answer depends almost entirely on how your list behaves on domains like these.
It matters because catch-all is not a rare setup. Google runs one. So do Netflix, Stripe, Databricks, Nvidia and The New York Times, alongside 245 other domains in this test. Our census of the Fortune 500 found 47% of those companies behaving as catch-all and 69% catch-all, gateway-fronted, or both. These are the accounts B2B teams target most, and they are the accounts most verification tools are quietest about.
Study at a glance
The three metrics
- Real contacts found, out of 500. Counted per person: found if the tool returned a conclusive valid result on at least one of that person's permutations.
- False positives, out of 5,000. Measured only on the invented set, the one group where we know for certain no mailbox should exist.
- Unresolved, out of 10,000. Any status expressing uncertainty. A record counts as resolved only when the tool made a clear decision either way.
Email verification tools compared: how all 11 performed on catch-all domains
Every tool received the identical 10,000-record list. The table is ranked by real contacts found, and that ranking carries through the rest of this article.

Two groups, and almost nothing in between. Six tests found more than 400 of the 500 real people. Seven left at least 94% of all 10,000 records undecided, and six of those found fewer than 45 people. Allegrow found the most real people, 491 of 500, with one of the lowest false positive rates in the test.
One warning before you read the false-positive column, because it looks like a leaderboard and is not one. Debounce and Bouncify each recorded a single error in 5,000, and each found fewer than 30 of the 500 real people while declining to decide on roughly 95% of the list. A tool that answered "unknown" to all 10,000 records would post a perfect 0% and be worthless.
How to read these results
The three numbers are not a neutral trade-off where every tool has a defensible philosophy.
A verification tool exists to tell you whether you can email a specific person. There are two ways to fail at that. It can say yes when the answer is no, which is a false positive and shows up later as a bounce. Or it can refuse to answer, which shows up as nothing at all.
A tool that decides almost nothing will post a beautiful error rate, because it has barely given itself the chance to be wrong. Debounce made one mistake in 5,000 invented addresses, but it also only found 26 of 500 real people and left 94.8% of the list undecided. That single error is not evidence of precision, and reading it as accuracy gets the conclusion backwards. On this list, Debounce is not the most accurate tool tested, but is one of the least usable, because a list that comes back 95% unactionable.
The two failures also cost you differently, and only one of them is visible. A false positive costs sender reputation and you find out when the bounces land. An unresolved record costs you a person you could have reached. On a list of catch-all domains, those unresolved records sit at exactly the large companies a B2B team most wants to reach.
So the honest reading is this: the job is to resolve the list correctly, and a false-positive rate only means something measured against how much the tool was willing to decide. Read that way, the fourteen tests fall into three groups.
Hunter.io sits at the lower edge of the first group, resolving a fair percentage of the list but leaving 744 records undecided, which places it between the two clusters rather than squarely in either.
The practical upshot is a better set of questions for a vendor. An accuracy percentage on its own tells you almost nothing, because you cannot tell whether it was earned by being right or by refusing to answer. Ask instead: how many real people did you find? How many invented addresses did you call valid? And what share of the list did you leave undecided? A vendor who can answer all three is describing their performance.
How we ran the test
The whole point of a study like this is that someone else can check it or repeat it. So here is the full method, including the choices we made and the reasons behind them.
The dataset: 251 catch-all domains and 10,000 addresses
We started with the domains, not the people. Every one of the 251 domains in this test is a catch-all, checked before anything else went ahead. That is what makes this a catch-all benchmark rather than a general accuracy test: we did not take a mixed list of companies and let a few catch-alls fall in, we built the whole set out of them deliberately.
The 10,000 records split evenly. About 5,000 are name permutations for 500 real people, and 5,000 are fictional addresses at those same real domains.
These are not obscure companies. They are the accounts B2B teams actually chase:
- Big tech and consumer internet: Google, Netflix, Spotify, Uber, Lyft, Airbnb, Pinterest, Dropbox, Shopify, DoorDash, Grubhub, Roblox
- Developer and data infrastructure: GitLab, MongoDB, Databricks, Datadog, Figma, Vercel, Zapier, MariaDB, LaunchDarkly, Fivetran
- Fintech: Stripe, Coinbase, Chime, Betterment, Wealthfront, Ramp, Carta, Fidelity
- Industrials and enterprise: Honeywell, BP, Eaton, Air Liquide, Airgas, Avery Dennison, Trimble, Veolia, Kiewit, Nvidia
- Media and research: The New York Times, the Financial Times, Axios, Nielsen, Qualtrics
- Health and life sciences: Veeva, Zocdoc, Ginkgo Bioworks, 10x Genomics, AbCellera, Hinge Health
Two of the domains are worth a mention on their own: MXToolbox and RocketReach. Even companies that sell email tooling and contact data run catch-all setups.
How we identified the real people
The 500 real professionals came from public LinkedIn employee data at those 251 domains, so each one is a person who genuinely works there as far as public record shows.
For each person we generated around 10 candidate addresses using the standard business email formats: first.last@, flast@, f.last@, first@ and the other common patterns. That is the same permutation approach an email finder uses, and it matters for how you read the results. Finding a contact here is not a database lookup. It means picking the one live address out of roughly ten plausible guesses at a domain that accepts all of them.
How we built the control group
The 5,000 invented addresses use obvious fictional character names at the same real domains.
The obviousness is deliberate. A random string of letters might, by bad luck, match a real mailbox somewhere, and then we would be counting a correct answer as an error. A famous fictional character is very unlikely to work at Stripe, so when a tool calls that address valid we can say with reasonable confidence it got it wrong. Reasonable rather than total: at companies this large, an ordinary-sounding character email permutation will occasionally belong to a real employee, and the shorter permutation formats make that more likely. But even in these cases, we still counted them as errors.
How each metric was counted
Real contacts found is counted per person, not per address. A person counts as found if the tool returned a conclusive valid result on at least one of their permutations. Finding three working addresses for the same person still counts once, because the outcome that matters is whether you can reach them.
False positives are counted only on the 5,000 invented addresses. This is a deliberate limit and it is worth explaining, because it may look at first like a gap. We do not score false positives on the permutation set, for two reasons: we have no way of knowing which permutations are genuinely wrong, and a person can legitimately have several working addresses through aliases. Marking those as errors would punish a tool for being right. The invented set is the only group where we know for certain that no mailbox should exist, so it is the only clean test.
Unresolved covers any status that expresses uncertainty, whatever a vendor calls it. Tools use different words for this (catch-all, unknown, risky, accept-all), so we applied one rule across all of them: a record counts as resolved only when the tool made a clear decision, either safe to send or clearly not. Anything else is unresolved.
What this study cannot tell you
- It is a point-in-time result, run at the end of July into early August 2026. Vendors change how their tools behave.
- The real people were identified from public LinkedIn data, not from reply history, so a person recorded as real is real as far as public record goes.
- Real contacts found is measured per person, so a tool that finds several working addresses for one person gets credit once.
- False positives are measured only on the invented set, for the reason given above.
- Each tool was run on its documented settings at the time.
- There is no pricing or speed comparison here, because we did not capture that data. We are not going to estimate it as it constantly changes.
Tool-by-tool results, ranked by real contacts found
The sections below are ordered by how many of the 500 real people each tool found, from most to fewest. Where a tool offers more than one verification mode, each mode is reported separately so it can be judged on its own terms.
1. Allegrow: 491 of 500 real contacts found
On a list of 10,000 addresses at 251 catch-all domains, Allegrow found 491 of 500 real people (98.2%), called 11 of 5,000 invented addresses valid (0.22%), and left 197 of 10,000 records undecided (2.0%).
That is the highest number of real people found in the study, and the margin is the point. The next best result was 470, then 467 and 460, so Allegrow returned between 21 and 31 more reachable people per 500 contacts than the best tools tested. It did that while making 11 errors in 5,000, against 9 for the lowest recorded. Those 11 are worth explaining rather than just counting, and we come back to them at the end of this section.
The gap in real contacts found is easy to shrug at until you scale it. Run a list of one million real contacts at catch-all domains through each tool and Allegrow's 98.2% returns 982,000 reachable people. The next three, at 92% to 94%, return between 920,000 and 940,000. Same list, same day, and a gap of 42,000 to 62,000 professionals you either hold usable contact details for or you do not.
We left 197 records undecided, a 2.0% rate, where three ZeroBounce modes and BounceBan left very few. This is worth reading properly though: those tools cleared their last few records by returning a decision on addresses we flagged as uncertain, and they did it while finding 21 to 31 fewer real people. Given the choice between handing back 197 honest unknowns or 24 fewer contacts a team can actually email, we would rather hand you more valid reachable contacts.
A note on those 11 errors. The control group uses obvious fictional character names, which is what keeps the test clean, but these are 251 of the largest employers in the world. Michael Scott from The Office and Marjorie Simpson from The Simpsons are unmistakable as characters, but the permutations they generate can match names that real people have.
An address like mscott@ or msimpson@ throws away almost everything that identifies a person, leaving one initial and a common surname. At that point it does not matter whether a Michael Scott works there. A Maria Scott or a Mark Simpson resolves to the same address, and a tool returning valid is right about the mailbox even though our label calls the person fictional.
We counted all 11 as errors regardless. The same caveat applies to every error count in this study.
2. ZeroBounce: 470 of 500 real contacts found
ZeroBounce offers three verification modes, so we tested all three. The differences between them are one of the more useful findings in the study, particularly for anyone weighing up whether the paid catch-all add-on is worth it.
Standard is the strongest of the three on the numbers that matter. It found 467 of 500 real people and called 9 of 5,000 invented addresses valid, the joint-lowest error count in the study alongside BounceBan, and unlike the tools further down this list it earned that while resolving 99.1% of the records. Finding 467 on a catch-all list means the tool is doing more than reading the server's blanket “accept” SMTP signal.
The paid catch-all scoring add-on changed nothing that mattered. It took the 90 records standard had left undecided, scored each one, and resolved every one of them to invalid. Real contacts found stayed at 467. False positives stayed at 9. What the extra credits bought was a tidier spreadsheet, not another reachable person. A different list composition might produce cases where scoring recovers a contact. On these 10,000 records it recovered none.
Verify+ found 3 more people and made 16 more errors. It is an opt-in second phase that sends real test emails to the records that came back catch-all, then treats a bounce as proof the mailbox does not exist and the absence of one as proof that it does. That inference is the problem. Enterprise mail systems and secure email gateways are routinely configured to suppress non-delivery reports, so mail to an address that does not exist is silently discarded or rerouted rather than bounced. Nothing coming back is not evidence the person is real, and the false-positive count rising from 9 to 25 is what that looks like in the data.
Across all three modes, the ceiling is the finding. Standard found 467, scoring found 467, Verify+ found 470. Whatever a buyer enables, ZeroBounce lands within three people of where it started, because the modes differ in how they label the records they cannot resolve rather than in how many people they can find. 470 is 21 short of Allegrow, and no configuration closes it.
3. BounceBan: 460 of 500 real contacts found
BounceBan found 460 of 500 real people (92%), called 9 of 5,000 invented addresses valid (0.18%), and left a single record of 10,000 undecided.
That is a strong showing on two of the three measures at once. The 9 false positives match ZeroBounce for the joint-lowest count, and one unresolved record out of 10,000 is virtually complete resolution.
Where it falls short is the measure that decides how many contacts you can actually email. It sits 31 people behind Allegrow, roughly one in every sixteen real people on the list going unfound, and that is a wider gap than any of the three ZeroBounce modes managed. On a 100,000-contact enterprise list the same rate returns 92,000 reachable people against Allegrow's 98,200, a difference of 6,200 people who exist and are sitting at their desks.
Finding 460 of 500 while leaving a single record unresolved is a strong result, and it is not achievable by reading SMTP responses alone. Every permutation at a catch-all domain returns the same acceptance, so a tool that stops there cannot tell the live address from the nine dead ones. BounceBan is evidently doing something further, though the results do not tell us what.
4. Hunter.io: 427 of 500 real contacts found
Hunter.io found 427 of 500 real people (85.4%), called 17 of 5,000 invented addresses valid (0.3%), and left 744 of 10,000 records undecided (7.4%).
Hunter.io is the boundary case in this study. It sits well clear of the tools that decide on almost nothing, resolving 92.6% of the list and finding more than four in five of the real people, but it leaves noticeably more unresolved addresses than the tools above it, which cleared the list almost entirely. It also missed 73 of the 500 real people.
The pattern suggests a tool that tests beyond the server's blanket acceptance for most records, then stops short on a slice of them, rather than one that reads the catch-all response and gives up. That is what the results imply, not a claim about how the product works internally.
5. NeverBounce: 151 of 500 real contacts found
NeverBounce found 151 of 500 real people (30.2%), called 27 of 5,000 invented addresses valid (0.5%), and left 9,633 of 10,000 records undecided (96.3%).
That is the highest unresolved rate in the study. Fewer than four records in every hundred came back with a clear decision either way, which means a team running this list through NeverBounce would still have 96% of it to resolve by some other route, and 349 of the 500 real people unaccounted for.
The pattern defines the low-resolution group: reading the server's blanket acceptance, recognising it as a catch-all response, and returning that uncertainty rather than further testing whether the individual mailbox exists. Notably, the caution did not buy a clean error rate either, since it came back with 27 false positives.
6. Million Verifier: 147 of 500 real contacts found
Million Verifier found 147 of 500 real people (29.4%), called 640 of 5,000 invented addresses valid (12.8%), and left 6,771 of 10,000 records undecided (67.7%).
This is the clearest example in the study of the opposite failure to the cautious group. Million Verifier resolved substantially more of the list than its neighbours, 3,229 records against roughly 400 to 600 for the tools around it, and the extra resolution came with 640 invented addresses marked valid.
Be concrete about what a 12.8% false-positive rate means. Roughly one in eight of the invented addresses was waved through as sendable, and every one of those is a likely bounce waiting to happen. Bounce rates at that level damage sender reputation quickly, and the damage carries across every campaign sent from the same domain, not only the list that caused it. More than two-thirds of the list still came back undecided on top of that.
The pattern suggests a tool willing to return a positive answer on a catch-all domain without the evidence to support it, where the cautious group returns uncertainty instead. Again, that is what the numbers imply rather than a description of the product's internals.
7. MailerCheck: 44 of 500 real contacts found
MailerCheck found 44 of 500 real people (8.8%), called 49 of 5,000 invented addresses valid (1%), and left 9,436 of 10,000 records undecided (94.3%).
MailerCheck sits in the low-resolution group without the low error rate that usually accompanies it. Its 49 false positives are the second-highest count in the study, behind only Million Verifier, and they arrived alongside 94.3% of the list left undecided and 456 of the 500 real people unaccounted for. The tools directly around it recorded between 1 and 41 errors.
The pattern points to a tool that reads the catch-all response and returns uncertainty on nearly everything, as the rest of this group does, while the small share it did decide on carried a higher error rate than any other tool in the group: 49 against Emailable's 41, NeverBounce's 27, and three or fewer each from EmailListVerify, Debounce and Bouncify.
8. Emailable: 29 of 500 real contacts found
Emailable found 29 of 500 real people (5.8%), called 41 of 5,000 invented addresses valid (0.8%), and left 9,426 of 10,000 records undecided (94.2%).
Emailable is in the same position as MailerCheck: firmly in the low-resolution group, but with a higher error count than most of the tools in the group. Debounce and Bouncify each recorded one false positive at a similar resolution rate, where Emailable recorded 41. It left 471 of the 500 real people unaccounted for.
Put the two counts side by side, and Emailable called 41 invented addresses valid while finding 29 real people. Send to everything it marked valid and you are aiming at 70 accounts, of which 41 belong to nobody. Nearly six in ten. It said yes to more fictional characters than it found human beings, which is not a cleaned list so much as a list that now needs cleaning. Mail it as it stands and the bounces and dead engagement that follow degrade the same sender reputation you bought verification to protect.
The behaviour reads as returning uncertainty on the catch-all response for the overwhelming majority of records, which is the group pattern, while the small share it did decide on produced more errors than its neighbours. That is an inference from the results rather than a statement about how the tool works.
9. EmailListVerify: 29 of 500 real contacts found
EmailListVerify offers two scan modes, so we ran both. The deeper scan moves the numbers, though not enough to change the shape of the result.
EmailListVerify Normal Scan
The normal scan found 20 of 500 real people (4%), called 2 of 5,000 invented addresses valid (0%), and left 9,575 of 10,000 records undecided (95.7%).
That 20 is the lowest number of real people found in the study. It came with two errors in 5,000, which looks like precision until you notice it was recorded while answering almost nothing: 480 of the 500 real people went unaccounted for, and almost 96% of the list came back undecided.
The pattern is the group's: read the catch-all response, return uncertainty, do not further test the individual mailbox. That is what the numbers imply, not a description of the product's internals.
EmailListVerify deep scan
The deep scan found 29 of 500 real people (5.8%), called 3 of 5,000 invented addresses valid (0%), and left 9,401 of 10,000 records undecided (94%).
Set against the normal scan, the deeper mode found 9 more people, added 1 more false positive, and resolved 174 more records. It does more work and returns slightly more, but the shape does not change: both modes leave roughly 95% of a catch-all list undecided and find fewer than 30 of the 500 real people.
10. Debounce: 26 of 500 real contacts found
Debounce found 26 of 500 real people (5.2%), called 1 of 5,000 invented addresses valid (0%), and left 9,480 of 10,000 records undecided (94.8%).
That single error is the joint-lowest count in the study, tied with Bouncify, and it is the clearest illustration of why the false-positive column cannot be read alone. Judged on errors, Debounce looks flawless. Judged on whether it did the job, it found 26 of 500 people and handed back 94.8% of the list undecided. The error rate is not evidence of precision; it is what happens when a tool answers almost nothing. A list that comes back 95% undecided has not been verified.
11. Bouncify: 22 of 500 real contacts found
Bouncify found 22 of 500 real people (4.4%), called 1 of 5,000 invented addresses valid (0%), and left 9,561 of 10,000 records undecided (95.6%).
Bouncify and Debounce produced almost the same outcome: one false positive each, 22 and 26 people found, 95.6% and 94.8% left undecided. Two separate vendors landing this close together matters more than either result alone, because it suggests a characteristic of an approach rather than a quirk of one product. When a tool treats the catch-all response as the end of the inquiry, this is roughly what the numbers look like. That reading comes from the results themselves and is not a description of how either tool works internally.
What this means if you are buying an email verification tool?
The study points at a few things worth doing differently when you next evaluate a vendor.
- Ask for the unresolved rate alongside the accuracy figure: An accuracy claim on its own cannot be checked, because you cannot tell whether it was earned by deciding correctly or by declining to decide at all. Several of the tools here could advertise a very low error rate while finding fewer than 30 of 500 real people. The two numbers only mean something together.
- Test on your own catch-all and enterprise domains: A generic test list, heavy on consumer mailboxes that answer cleanly, will make almost any tool look competent. The domains where verification is hard are the ones your team actually targets, and that is where the differences show up. Every domain in this study was a catch-all for exactly that reason.
- Treat an unresolved result as a cost, not a neutral outcome: This is the habit most buyers need to change. An "unknown" feels like a safe non-answer, but it is a contact you paid to check and still cannot use. If a tool leaves 95% of your list undecided, you have not verified your list; you have paid to be told that verifying it is hard.
- Be wary of a near-perfect accuracy claim with no coverage figure attached. The cheapest way to a clean error rate is to answer almost nothing, and several tools here did exactly that. It is a design choice about which failure to avoid, and the buyer is the one who lives with it.
- Where Allegrow fits: Every tool in this study made a choice about which failure to avoid. Most chose not to be wrong, and paid for it by not answering. We chose to answer, and the bill for that shows up in our own numbers: 197 records we could not resolve and 11 invented addresses we called valid. What it bought, though, was 491 of 500 real people, more than any other tool tested.
Run this test yourself
You do not need our data to check any of this. The method is straightforward, and running it on your own list is more useful than any published benchmark, because it uses the domains you actually sell into.
- Choose your domains, and confirm they are catch-all. Pick companies from your own target list rather than a generic sample, then check which domains are catch-all and accept mail to an address that cannot exist. Any serious verification tool you already use should provide you with the domain configuration of your leads. If you skip this step you are testing general accuracy, not catch-all performance, and the results will look much better than they actually are.
- Identify real people you can confirm independently. You need ground truth, and the strongest form of it is a reply. If someone answered your email last month, that address works, whatever a tool says about it. After replies then comes your CRM, then public professional profiles. We used public profiles for this study because we cannot publish data about people who have emailed us, but that constraint does not apply to you. Your own reply history will give you a sharper test than ours.
- Generate the permutations. For each person, build the standard formats against their company domain: first.last@, flast@, f.last@, first@ and the rest. If you need help, we have an article mentioning the most common email formats for businesses. Around ten per person is enough.
- Build a control group of obvious fictional names. Use clearly invented people at those same real domains, and make them obvious. This is the only part of your list where you know for certain that no mailbox should exist, so it is the only clean way to catch a tool calling something valid when it is not.
- Run every tool on the identical list.
- Count the three numbers separately. How many of your real people did each tool find, counting a person as found if any one of their permutations came back conclusively valid. How many of your invented addresses did it call valid. And what share of the whole list did it leave undecided, using one consistent rule across vendors, since they all use different words for uncertainty.
The principle underneath all six steps is ground truth. Without knowing the right answers in advance, a comparison can only tell you which tool reports the highest percentage of something, and an arbitrarily high number is easy to claim.
Get the full dataset
The complete results file is available on request. It contains the full per-tool output across all 10,000 records, the list of all 251 catch-all domains, and the status each tool returned for every address, so you can check our counts or rerun the comparison against a tool we did not include.
The results are published for informational purposes and are designed to give you a framework for running your own assessment, on the domains and contacts that matter to you.
How to cite this study
If you are quoting these results, here is everything you need to attribute them correctly.
Study: Allegrow Catch-All Email Verification Benchmark
Publisher: Allegrow
Date: August 2026
Sample: 10,000 addresses at 251 catch-all domains, made up of around 5,000 name permutations for 500 real professionals and 5,000 invented addresses, tested across 11 tools in 14 tests
Metrics: real contacts found (out of 500), false positives (out of 5,000), unresolved (out of 10,000)
Canonical URL: https://www.allegrow.co/knowledge-base/best-email-verification-tools-benchmark
Suggested citation: Allegrow, "Catch-All Email Verification Benchmark," August 2026, https://www.allegrow.co/knowledge-base/best-email-verification-tools-benchmark
Conclusion
The measurement was simple. Take 10,000 addresses at 251 catch-all domains, hide 500 real people among 5,000 invented ones, and see which tools can tell the difference.
Most could not. Seven of the fourteen tests left at least 94% of the list undecided, and six of those found fewer than 45 of the 500 real people. That is not a small gap in performance. It is a split between tools that attempt the problem and tools that report back that the problem is hard.
The reason it goes unnoticed is that the market talks about accuracy, and accuracy alone cannot separate those two groups. A tool that answers almost nothing posts a beautiful error rate; Debounce and Bouncify each made a single mistake in 5,000 while leaving roughly 95% of the list unresolved. Coverage is the number that tells you whether the accuracy was earned.
Allegrow found 491 of the 500 real people, more than any other tool tested, with an error rate among the lowest recorded. The next best results were 470, 467 and 460, from ZeroBounce and BounceBan. Those are the only other tools that resolved the list rather than handing most of it back, so they are the fair comparison, and the lead over them is 21 to 31 people in every 500. On a list of 100,000 contacts at domains like these, that is 4,200 to 6,200 more people a team can actually email.
So the question worth carrying into a vendor conversation is not just: how accurate are you? It is whether the tool gives you coverage alongside accuracy on the contacts that are actually worth something to your business. If you sell to large companies, the catch-all servers that made this test difficult are commonplace among your prospects rather than an edge case.
If you want to see how this plays out on your own contacts rather than ours, start a 14-day free trial and run your catch-all and enterprise domains through it. The full dataset from this study is available on request if you would rather check our working first.
Frequently asked questions
What are the best email verification tools in 2026?
We tested 11 tools on 10,000 addresses across 251 catch-all domains. Ranked by how many of the 500 real people each one found: Allegrow 491, ZeroBounce 470 at best across its three modes, BounceBan 460, Hunter.io 427, then a steep drop to NeverBounce at 151 and Million Verifier at 147. The remaining five found fewer than 45 each and left 94% to 96% of records undecided.
How accurate are email verification tools on catch-all domains?
It varies enormously. In our August 2026 test of 11 tools, real contacts found ranged from 4% to 98.2% of 500 real people, meaning 20 people at the bottom and 491 at the top. Seven of the fourteen tests left at least 94% of all 10,000 records undecided, so on catch-all domains most tools resolve very little.
Why do email verification tools return "catch-all" or "unknown"?
Because a catch-all domain accepts mail addressed to anyone, so the standard SMTP check learns nothing from the server saying “accept”. It would say “accept” to a real employee and to an invented name alike. A tool that stops at that response has no basis for a decision and returns uncertainty instead.
Can any tool verify catch-all email addresses?
Some, out of the 11 tools tested, only 4 had a reasonable number of resolved valid contacts. Allegrow found 491 of 500 real people, ZeroBounce with Verify+ found 470, ZeroBounce standard 467, BounceBan 460, and Hunter.io 427.
Which email verification tool performed best on catch-all domains?
Allegrow, on the measure that determines how many people you can reach. It found 491 of 500 real contacts (98.2%), ahead of ZeroBounce with Verify+ at 470, ZeroBounce standard at 467 and BounceBan at 460, while calling 11 of 5,000 invented addresses valid (a level between the lowest error rates recorded) and resolving 98% of the full 10,000-record list.
What should you ask a verification vendor besides their accuracy rate?
Ask this: on catch-all domains, what share of addresses do you return a definitive valid or invalid for, rather than "catch-all" or "unknown"? An accuracy figure only covers the records a tool chose to decide on, so on its own it tells you almost nothing. In our test, tools making only one or two errors in 5,000 left 94% to 96% of records undecided. They were accurate. They were also useless.
Which domains were tested?
All 251 are catch-all domains, spanning big tech, developer infrastructure, fintech, industrials, media and life sciences. Recognisable names in the set include Google, Netflix, Stripe, Nvidia, Databricks, Honeywell, BP and The New York Times. The full list comes with the dataset, which is available on request.

.jpg)

.jpg)
.jpg)
.jpg)