Filter Databases vs Research Agents: One-Sentence ICP Test
A filter database can only answer questions someone already turned into a field. Most real ICPs contain one clause that was never a field, and it is usually the clause that predicts the sale.
Filter Databases vs Research Agents: One-Sentence ICP Test
Filter databases vs research agents is not a speed or cost comparison. A filter database answers only questions someone already turned into a field, like industry, headcount or installed technology. A research agent reads the open web to answer clauses that were never fields. Pick by whether your ICP is expressible.
- What is a research agent in sales prospecting?
- Why the filter database vs research agent choice is really about expressibility
- Three one-sentence ICPs worked through by hand
- Where filter databases win outright
- Firmographic filters limitations are not a claim about intent data
- What research agents cost you
- AI SDR vs research agent and why the confusion is expensive
- How to find companies matching a description in 2026
- Why this matters beyond one prospect list
- FAQ
Write your ICP as one sentence, then read it clause by clause. The clause that predicts whether someone buys is usually the one no database has a field for.
That is the entire comparison. Apollo publishes its filter list and it is long: job title, location, company size, industry, keywords, funding, job postings, news, technologies, headcount growth, revenue and AI research fields (Apollo Search Filters Glossary, accessed September 1, 2026). Long is not the same as complete.
Everything else people argue about, cost per record, run-to-run variance, whether a teammate can regenerate your list, follows from one prior question: can you state the criterion at all?
What is a research agent in sales prospecting?
A research agent is a system that investigates one account at a time by planning its own steps, reading sources and returning structured findings. Outreach defines its Research Agent as account-level research that extracts structured insights from internal or external sources and stores the results in account fields for personalization, planning, segmentation or automation (Outreach Research Agent Overview, accessed September 1, 2026).
The architectural difference is who decides the steps. Anthropic separates predictable workflows, which run predefined code paths, from agents whose language models dynamically direct their own processes and tool use, and positions agents for open-ended tasks where the number of steps cannot be hardcoded (Anthropic, Building Effective AI Agents, 2024).
That flexibility is bought with retrieval, not with a schema. AWS describes retrieval-augmented generation as first retrieving relevant information from an external source, then handing the user query plus that retrieved context to the language model (AWS, accessed September 1, 2026). Google's agentic version goes further: the agent reasons about the query, chooses among tools and retrievers, and combines structured and unstructured data before answering (Google Cloud Codelabs, accessed September 1, 2026).
A filter database looks things up. A research agent goes and reads. Both are legitimate, and they fail in completely different ways.
Why the filter database vs research agent choice is really about expressibility
Boolean filters vs AI prospecting gets sold as a speed and price argument, and that is the wrong axis. The first question is which criteria you can even state.
A filter query is a combination over a fixed vocabulary. Apollo applies AND logic across separate filters and OR logic among values within one filter, so adding a filter narrows the result set while adding alternatives inside a filter broadens it (Apollo, accessed September 1, 2026). If a term is not in the vocabulary, no arrangement of AND and OR reaches it.
Keyword search is the escape hatch, and its limits are published. Apollo's industry value is self-reported, while keyword matching can run over company descriptions, websites, social profiles and other free-text data, and Apollo warns that broad keywords can create noise or miss niche segments (Apollo, accessed September 1, 2026).
| Filter database | Research agent | |
|---|---|---|
| Query surface | The vendor's published field list | Whatever the agent can retrieve and read |
| Answer type | A row that matched or did not | A generated finding, ideally with evidence attached |
| Same query run twice | Same rows | Can differ between runs |
| Cost per record | Low, and known before you run it | Higher, and varies with how much reading each account needs |
| Speed | Seconds across thousands of rows | Minutes per account |
| Audit trail | The filter set is the definition of the segment | You re-read the cited evidence, account by account |
| Correct when | Every ICP clause is a field | A clause is observable but was never modelled |
Neither column is the winner. The row that decides your purchase is the last one.
Three one-sentence ICPs worked through by hand
These are illustrations of expressibility, not measurements. No benchmark was run for this guide and none is needed, because each clause below is checked against Apollo's published filter documentation, which anyone can open and repeat.
ICP 1: a pricing-model change
"US logistics software companies with 50 to 200 employees that moved their pricing page from self-serve to sales-assisted in the last two quarters."
| Clause | In the published filter list? | Where the answer lives |
|---|---|---|
| US | Yes, location | The database |
| Logistics software | Partly. Industry is self-reported, and keyword matching runs over descriptions and websites | The database, with noise |
| 50 to 200 employees | Yes, company size | The database |
| Moved pricing from self-serve to sales-assisted in the last two quarters | Not in the published filter list | The pricing page, its archived history, the changelog |
Three clauses of four are database work, and you should run them there first. The fourth clause is the one that says a company just decided humans should close its deals, which is the clause that predicts whether they buy a sales tool.
ICP 2: a stated public commitment
"UK manufacturers with 200 to 1,000 employees that have published a net-zero target with a dated deadline."
| Clause | In the published filter list? | Where the answer lives |
|---|---|---|
| UK | Yes, location | The database |
| Manufacturers | Partly, and self-reported industry is a known weak point | The database |
| 200 to 1,000 employees | Yes, company size | The database |
| Published a net-zero target with a dated deadline | Not in the published filter list | The sustainability report, a press release, an investor page |
The last clause is a fact the company published about itself, in prose, on purpose. It is completely observable and completely unmodelled, and answering it means opening a document and deciding whether "by 2040" counts as a dated deadline.
ICP 3: the one where the database nearly wins
"Companies running a headless commerce stack that are hiring their first data engineer."
| Clause | In the published filter list? | Where the answer lives |
|---|---|---|
| Running a headless commerce stack | Often yes. Apollo documents support for more than 1,500 technologies as searchable company attributes | The database, if that specific technology was modelled |
| Hiring a data engineer | Yes, Apollo documents a job postings filter | The database |
| Their first data engineer | Not in the published filter list. It is a claim about absence | The team page, the careers history, current staff titles |
Two clauses of three are database work, and the database will do them faster and cheaper than any agent. The word carrying all the predictive weight is "first", and "first" is a statement about what a company does not have yet. Absence is the hardest thing for a schema to hold, because nobody builds a field for what is not there.
ā Good: "US logistics software companies, 50 to 200 employees, that moved their pricing page from self-serve to sales-assisted in the last two quarters." Every clause names a fact someone could check.
ā Bad: "Fast-growing logistics companies that are ready to buy and have budget this quarter." Nothing here is checkable by a filter or an agent, because no clause names an observable fact.
If your sentence looks like the second one, the tool is not your problem. Go and define your ICP against customers you already closed, and read ideal customer profile examples that name checkable attributes.
Where filter databases win outright
If every clause in your ICP sentence is a field, use the database and skip the rest of this guide. An agent applied to a fully expressible ICP is a slower, more expensive way to get a worse version of the same list.
- Determinism. The same filter set returns the same rows. Because Apollo publishes its AND and OR semantics, you can predict what adding a filter does before you add it (Apollo, accessed September 1, 2026).
- Auditability. The filter set is the segment definition. Six months later you can re-run it and explain it to a board without re-reading anything.
- Reproducibility across a team. RevOps hands a saved search to a new rep and the new rep gets your list, not a similar one. This is the biggest practical advantage and it is rarely mentioned in vendor comparisons.
- Coverage is wider than critics assume. More than 1,500 technologies are searchable company attributes in Apollo, alongside job postings, news, headcount growth, funding and revenue (Apollo, accessed September 1, 2026). Plenty of criteria people assume are unmodellable are already fields.
- The architecture is built for cheapness. ZoomInfo's company-search API returns companies that meet specified search criteria with basic company information, and additional information is retrieved through a separate enrichment API (ZoomInfo, accessed September 1, 2026). Search and enrichment being separate operations is exactly why neither has to read anything.
Do not hire an agent to run a headcount filter. Run the filter, export the rows, and spend the agent budget on the one clause the filter could not reach.
There is also a moving boundary worth respecting: Apollo's published list now includes AI research fields alongside its conventional firmographics (Apollo, accessed September 1, 2026). The set of expressible questions grows every year, and it grows in the database's favour.
Firmographic filters limitations are not a claim about intent data
A missing field tells you what a database cannot express. It tells you nothing about whether a company wants to buy.
These are two different claims and they get merged constantly. Expressibility asks whether you can state a criterion. Propensity asks whether that criterion predicts a purchase now. A tool can be perfect at the first and silent on the second.
Every worked example above is an observable fact, not a buying signal. A pricing-page change, a published net-zero target and a first data-engineer hire are all things you can verify. Whether any of them predicts a sale is your hypothesis, and it gets tested by your reply rates, not by the tool that found it.
Apollo documents filters for news and job postings (Apollo, accessed September 1, 2026). Those are facts about a company, not evidence that anyone inside it wants to talk to you. The same caution applies to anything a research agent surfaces.
An unmodelled criterion is not automatically a valuable one. Most of them are just unmodelled. The trade only pays off when the missing clause is both hard to express and genuinely predictive, and you find that out by sending, not by shopping.
What research agents cost you
Research agents are slower, cost more per record, and do not return the same thing twice. Treat run-to-run variance as the headline cost, not a rounding error.
Anthropic's own framing explains why: agents dynamically direct their own processes and tool use, which is what makes them suitable for open-ended tasks where the number of steps cannot be hardcoded (Anthropic, 2024). A system that chooses its own steps does not choose the same steps every time.
- Reproducibility breaks first. You cannot hand a colleague a prompt and expect an identical list back, so segment definitions drift between people and between weeks.
- QA moves from per-filter to per-account. Checking a filter set takes a minute. Checking 300 generated findings does not.
- Retrieval bounds the answer. Since RAG retrieves first and generates second (AWS, accessed September 1, 2026), a sustainability report locked behind a form or a pricing page that never rendered produces a confident-looking nothing.
- Cost scales with reading, not with list size. Budget per account, and cap the accounts before you start.
Never let an agent's finding into a CRM field without the source link attached. Outreach's design stores extracted insights into account fields (Outreach, accessed September 1, 2026), which is the right pattern only when the evidence rides along with the value.
Every ICP clause you cannot filter on is a decision a product manager made years ago, for a different customer, about which facts were worth storing.
AI SDR vs research agent and why the confusion is expensive
An AI SDR sends. A research agent reads. Buying one when you needed the other is the most common expensive mistake in this category, and the search queries show people conflate them constantly.
Salesforce defines an AI SDR around top-of-funnel work: lead qualification, outreach, engagement, personalized messages, follow-ups and meeting scheduling (Salesforce, 2025). Outreach defines a research agent as account-level research producing structured insights stored in account fields (Outreach, accessed September 1, 2026).
| Research agent | AI SDR | |
|---|---|---|
| Job | Answer a question about an account | Contact prospects and handle replies |
| Output | Structured findings and evidence | Sent messages and booked meetings |
| Funnel stage | Before the list is finished | After the list exists |
| Who buys it | Whoever defines the segment | Whoever owns pipeline |
| Fails by | Returning an unsupported finding | Sending confident messages to the wrong list |
Do not buy an AI SDR to fix a targeting problem. It will send more messages to the same wrong list, faster, on the same domain you also send investor email from.
The reverse is equally true, and worth saying plainly. If your segment is already correct and you are simply not sending enough, a research agent will not help you, and an AI SDR tool for founders is the reasonable purchase.
How to find companies matching a description in 2026
Run the cheap step first, always. Filters shrink the universe by an order of magnitude before anything has to read a website, and that ordering is what makes AI research agent prospecting affordable at all.
- Write the ICP as one sentence. Not bullets. Bullets hide the clause you cannot express.
- Split it into clauses. One observable fact per clause. Delete any clause that names a feeling instead of a fact.
- Check each clause against the published filter list. Apollo's Search Filters Glossary is public. Open it and read it before you buy anything.
- Run every expressible clause in the database. This is the fast, deterministic, cheap part, and it should do most of the work.
- Count the surviving clauses. Zero means you are finished and you never needed an agent. One or two means you have a defined research question.
- Turn each surviving clause into a yes-or-no question with required evidence. "Does the pricing page require a demo request, and what is the URL that shows it?"
- Set the per-account budget and the account cap before you run. Agent cost scales with accounts and reading, not with rows.
| Your situation | What to use |
|---|---|
| Every clause is a field | Filter database, full stop |
| Every clause is a field but the industry value is unreliable | Filter database plus keyword search, and expect noise |
| One clause is observable on the web and not a field | Filters first to shrink the list, then an agent on the survivors |
| The clause is about absence, or about a change over time | Agent, with evidence required on every answer |
| You cannot name the observable fact at all | Neither. Fix the ICP first |
The mechanics of step 4 and step 6, including how to structure the export handoff, are covered in how to find companies matching your ICP.
Why this matters beyond one prospect list
The same expressibility problem appears anywhere a list is built from a description rather than a set of fields. Industry and headcount are fields, while "opened a second location in the last two quarters" is a sentence about behaviour, and that clause is usually the one deciding whether the message gets a reply. Causo is a research agent built for that job, and it carries the same tradeoffs described above: it reads more than a filter does, and it costs more per record than a filter does. If a plain industry-and-headcount filter already produces your list, use the filter.
FAQ
What is a research agent in sales prospecting? A system that investigates one account at a time, plans its own steps, reads sources and returns structured findings. Outreach defines its Research Agent as account-level research that extracts structured insights from internal or external sources and stores the results in account fields for personalization, planning, segmentation or automation (Outreach). It answers questions about accounts, and it does not send email.
Why can't I find my ICP with database filters? Because at least one clause in your ICP was never turned into a field. Apollo applies AND logic across separate filters and OR logic among values inside a single filter (Apollo), so a search is a combination of the vendor's published vocabulary and nothing else. If a criterion is not in that vocabulary, no combination of filters reaches it.
Are AI prospecting tools better than Apollo filters? Not generally, and often worse. Filters are faster, cheaper per record, deterministic and auditable, and Apollo's published set already covers a lot, including more than 1,500 searchable technologies (Apollo). An agent is better only for the specific clauses that are not fields, so the sensible pattern is filters first and an agent on the survivors.
How do you find companies that match a description? Split the description into clauses, run the clauses that exist as filters, then research only the ones that do not. Apollo warns that broad keyword matching can create noise or miss niche segments (Apollo), so keyword search is a weak substitute for a real field. For each remaining clause, define the exact evidence that would prove it before anyone starts reading.
What is the difference between an AI SDR and a research agent? An AI SDR works the top of the funnel: lead qualification, outreach, engagement, personalized messages, follow-ups and meeting scheduling (Salesforce). A research agent answers questions about accounts and returns structured findings with evidence (Outreach). They sit at different funnel stages and are usually bought by different people, which is why buying one to fix the other's problem is a common and expensive mistake.
Related on the hub
- Raising a seed round for an AI agent startup in 2026 ā for when the playbook turns into a raise.
- Buying a Lead List vs Building One: Cost per Verified Contact 2026 ā Related cold outreach guide.
- Apollo Alternatives 2026: 11 Best for Founder Sales ā Related cold outreach guide.
- Best B2B Lead Generation Tools 2026: 15 for Founders ā Related cold outreach guide.
Find your next customers with Causo.
Build a fit-ranked list of companies that match your ICP, draft hyper-specific outbound, and send from your own inbox, in one place.