Mika Pozo.
← index
In productionFlagship 02 — Lead gen & enrichment

My best-fit customers aren't in any database you can buy. An exercise in thinking outside of the box.

A from-scratch waterfall-enrichment pipeline that turns scattered public record into a scored, reachable, briefed lead — for the market the databases forgot. Running in production.

~8sources fusedper lead
0paid data vendorspublic record
65 / 35score tiers/130 ceiling
the 30-second version
01The problem

The ICP doesn't live on mainstream tools and databases.

My ideal customer isn't in any database you can buy. Sit with that for a second — it breaks the whole GTM playbook. Small and mid law firms in Brazil: local, WhatsApp-run, invisible to the firmographic backbone every enrichment tool resells.

I could scrape a list, of course — but a list of names is worth nothing if I can't trust a single field on it. My idea was simple: if waterfall enrichment is basically putting several data sources against each other to see what's actually true, it could be built from the ground up.

02The thesis, applied

Firmographics can't see fit. I stopped asking them to.

Here's the honest limit most enrichment decks skip: a list can tell you who exists. It cannot tell you who's a fit. In a market with signal — funding, hiring, tech stack — you paper over that gap with intent data. My market has no such gift: small Brazilian law firms are nearly invisible to the databases, and the ones with a website are the exception.

So I drew a hard line about what enrichment is allowed to conclude. It builds a picture — identity, rough revenue, practice area, real name, live channel — and writes a briefing good enough to open with a real question. It does not rule anyone in or out on fit, because it can't: a solo lawyer billing R$2k who wants to close more contracts might be perfect; the same R$2k split across partners isn't mine; R$40k is a coin toss until they talk. The conversation qualifies. Enrichment just earns the right to start it.

(There is stronger signal — actual intent and pain — but it lives where the ICP already communicates: strong, and far scarcer. I'm building for that now. Another story, another time.)

03The build
the sources

The turn was realizing I already had the data — I just hadn't been reading the right records. The raw material is public, if you know the law. In Brazil, whois for a .brdomain is open by statute, and the federal company registry — CNPJ, partners, share capital — is a matter of public record. So the pipeline reads what the vendors resell and the operators ignore: Jina for the site, registro.br for the domain's owner, Receita for partners and capital, Google for reviews, the page itself for a Meta pixel or GTM tag, Instagram for bio and reach. None of it fires blindly — no site means no pixel, no CNPJ means no Receita — and each source is there to fill or correct what the one before it missed. That's waterfall enrichment, except I own every stage.

And the sources overlap on purpose. Jina pulls an email, a phone, a name off the site; the whois pulls an email, a name, and the CNPJ; Receita takes that CNPJ and returns the partners, the capital, the practice area. The same fields surface two or three times from different places — and that overlap is the whole game: a CNPJ from the whois unlocks the record that confirms the rest, and a value that shows up in more than one source is a value I can trust. Then an LLM reads every source at once and reconciles them into one record — classifying the firm and resolving each field, preferring what more than one source agrees on.

waterfall enrichment — public record, cross-checkedfig.01 / pipeline
Jinasite
registro.brwhois owner
ReceitaWScnpj · partners
Googlereviews
pagemeta pixel / gtm
Instagrambio · reach
LLM reconcileprefer agreement
Reacheremail · smtp
WhatsAppreachable?
Twenty CRMperson · company · opp

The same field surfaces from two or three sources — and a value that shows up in more than one place is a value I can trust.

the name cascade

Here's the smallest decision on the page, and one of the ones I'm proudest of. The name I'm after isn't just any name — it's the decision-maker's. A cold WhatsApp addressed to the partner who actually runs the firm, by name, lands nothing like a generic intro to the institution: one reads as someone who already knows who's in charge; the other as a stranger who found a number and is hoping to be passed along. Get the name wrong, or leave it blank, and the lead knows in one second they're talking to software. So the name is the one field the system is never allowed to invent.

the name cascade — the one field the system can't inventresolution
1-- resolve the decision-maker's name, in order of trust
21 Receita partner list
32 registro.br registrant
43 site 'OAB' · 'Dr.' · 'Sócio Fundador'
54 Instagram bio
65 sheet raw name -- last resort only
7-- keep the name that appears in > 1 source
8-- no nameable decision-maker -> no contact
It walks the cascade in order of trust and keeps the name that shows up in more than one place. It would rather say nothing than open with 'Dear Contact.'
the score

With intent off the table, the next best question is cheaper: which of these firms even looks like it can afford to fix the problem? The score answers that, and nothing more. It's a plain additive model — every signal the enrichment found adds points. A professional site, not a drag-and-drop one. Ad spend already running. A Meta pixel on the page. Reviews in the range a mid-sized local practice tends to sit in. Real share capital and a sane number of partners at Receita. A high-ticket practice area — corporate, tax, real estate. Each is a proxy, and I treat it like one: no single signal means much, but stacked, they sort a firm into three tiers on two fixed thresholds. It's a read on capacity, not fit.

what the score is allowed to decide

Here's where it would have been easy to be clever, and wrong. The obvious move is to let the score gate: qualified firms get worked, the rest get dropped. I built the score and then refused to let it do that. It writes to a note on the CRM record — it colors the lead, it tells the human where to look first — and it decides nothing about who gets contacted. The reason is the discipline the whole system runs on: I hadn't validated this market, so a threshold I invented is just a guess wearing a number, and gating on it is the operator's mistake at smaller scale. Worse, the score reads capacity, and fit lives in the conversation — the firm that looks marginal on paper is exactly the solo lawyer who turns out to be perfect once he talks. So only facts gate, never the score. No WhatsApp, no contact — you can't work a channel that isn't there. No decision-maker I can name — I won't open blind. A firm too big to be my market — a different sale than the one I built this for. Everything else earns a conversation, and the conversation decides.

the briefing

I built this as a cold-call briefing: a profile by digital maturity — no site, site but no ads, already advertising — the pain most likely to be true, a hook, and three NEPQ questions pointed where it counts (cash tied up, time bled, the cost of standing still). Then I never made the calls — solo, there was no time. But it didn't go to waste; it became the seed of something better. It's what whoever takes the booked meeting walks in with, and it doesn't stay static: the WhatsApp BDR enriches it live, folding in what the lead actually says on the thread — the real pain, in their words. So by the time there's a demo, the briefing is two layers deep: the hypothesis enrichment guessed, and the signal the conversation confirmed. Enrichment writes the first draft of the meeting; the conversation finishes it.

the guardrails

The entry gates decide who gets in; a second layer verifies the channels before anything ships. Every email is checked at the SMTP level by Reacher — if it can't be confirmed deliverable, it's wiped, not saved, because an unverified address isn't a lead detail, it's a bounce waiting to poison your own sending domain. Every phone is checked against WhatsApp itself — a number without it is a dead end on the one channel that matters here. A record ships with a real, reachable channel or not at all — because everything downstream is only as healthy as the data it runs on.

04What broke / trade-offs

The honest failure mode isn't in the code — it's in the data. Public record is only as current as the last person who updated it. A domain's registrant can be the accountant who set it up, not the lawyer who runs the firm. A partner listed at Receita can have left years ago. The cross-check catches a lot of this — a name that shows up in one source and nowhere else gets treated with suspicion — but it's mitigation, not a cure. Some records still go out thin, or subtly wrong.

And I chose to let them. In a market this quiet, discarding every imperfect record means discarding the market. So the pipeline tolerates a noisier top of funnel than a bought list ever would — an incomplete picture, an occasional wrong field — because the conversation is the real validator, and a lead I can reach is worth more than a database entry I can't. The trade-off, stated plainly: I traded clean data for reachable data, and made the conversation clean up the difference.

05The numbers

This isn't the page with conversion metrics — those live in the conversation. What this build is measured by is cheaper, and for a GTM hire more telling:

  • No paid data vendor.The pipeline runs on public record and my own infrastructure — no Apollo, no ZoomInfo, no Clay seat. The list that wasn't for sale, built at the cost of the compute to build it.
  • ~8 sources fused per lead — site, domain registrant, federal company registry, reviews, ad-tech footprint, Instagram, email deliverability, WhatsApp presence — cross-checked into one record.
  • A transparent, additive score — no black box — summing to a documented 130-point ceiling, three tiers on two fixed cutoffs (65 / 35).
⟨match-rate / coverage⟩ — how often the pipeline turns a raw row into a usable record. Lives in the CRM; pulled before publish.

The point isn't “look how much data.” It's that the expensive part of enrichment was never the tool — it was knowing where the data already lived.

provenance

sources & score — derived from the workflowmatch-rate / coverage — lives in the CRM (Twenty), pulled before publish

Enrichment writes the first draft of the meeting; the conversation finishes it.