Finding AI-Forward Companies

How do you target AI-native startups for outbound?

You target AI-native startups by scoring the AI stack they run, not the AI language on their homepage. Run each domain through sniff_domain, filter on an ai_readiness verdict of shipping, and rank by the icp_signal. Bassethound returns the model providers, SDKs, orchestration, and vector stores in one correlated dossier, keyless, so an agent qualifies a list one self-contained call at a time.

Best move: score the AI stack each domain runs, then filter your list down to the ones shipping it.

Why it works: an AI-native company loads model SDKs, calls provider endpoints, and wires up a vector store. Those signatures live in the deployed front end. Marketing language is not evidence.

Key takeaways

  • AI-native is a stack fact, not a marketing claim. The deployed front end loads model SDKs and calls provider endpoints, and those signatures survive minification.
  • Bassethound’s AI-readiness layer scores six signal groups (providers, SDKs, orchestration, vector stores, observability, and the declarative layer of llms.txt, MCP, and .well-known files) into a number out of 1.0 and a verdict of shipping, experimenting, or none.
  • Firmographic “AI company” lists rely on self-reported tags and funding labels, which lag the deployed build by a quarter or more.
  • The icp_signal (hot, warm, cold) fuses the AI verdict with tech, infra, and firmographics, so you rank on fit instead of one axis.
  • The call is keyless and stateless, so an outbound agent qualifies each domain in one self-contained pass and enriches a list without storing anything.

What signals mark a startup as AI-native?

Six signal groups, read from the deployed stack. Model providers: calls to Anthropic, OpenAI, Gemini, Cohere, or Mistral endpoints. AI SDKs: the Anthropic SDK, OpenAI SDK, or Vercel AI SDK in the loaded scripts. Orchestration: LangChain, LlamaIndex, Haystack, CrewAI. Vector stores: Pinecone, Weaviate, Qdrant, Chroma, Milvus, each with a distinct client and API hostname. LLM observability: Helicone, LangSmith, Langfuse. Then the declarative layer: an llms.txt file (classified as the standard or a technical variant), an MCP endpoint, and .well-known AI files.

sniff_domain counts how many of those groups it hits, divides by six, and caps at 1.0. That number is the ai_readiness score. It also returns a verdict: shipping when the deployed evidence is strong, experimenting when signals are thin, none when nothing surfaces. A startup running a vector store behind an SDK behind a provider call did not experiment. It shipped. Treat the score as a floor, not a ceiling, because a static crawl reads what the front end exposes and no more.

How do you turn detection into a ranked list?

Start with domains, not names. Feed each domain to sniff_domain on the standard profile, which adds the AI stack and firmographics to the deterministic tech and infra layers. Filter on a verdict of shipping to drop the companies that only talk about AI. Then rank the survivors by icp_signal, which comes back hot, warm, or cold.

Because the endpoint is a remote MCP server, an agent host connects once and iterates. No keys, and no five separate tools to stitch together. The agent walks your domain list, pulls one correlated dossier per domain, and writes back the AI verdict, the primary stack, the security grade, and the ICP signal in a single row. For a longer list, run the fast profile first to drop the dead domains, then re-run the shipping candidates on standard or deep for firmographic contacts and the Wayback first-seen date. That last field tells you how long the company has run its AI stack, which separates a two-week experiment from a real deployment. You hand a rep or a sequencer a scored list, not raw domains.

Why not buy an “AI companies” list from Apollo or ZoomInfo?

Those lists are real and useful for firmographics: headcount, funding, titles, email. Apollo, ZoomInfo, and Clearbit maintain large contact corpora, and that corpus is their moat. What they do not have is the deployed stack. Their “AI company” tag comes from self-reported categories, press mentions, or a funding label, and every one of those lags the build. A company can raise an AI seed round and run nothing but a chat widget. Another can ship a production RAG pipeline and never touch its ZoomInfo category.

You want the second company, and the firmographic list cannot rank it. Bassethound reads the trail the application leaves: the SDK it loads, the vector store it calls, the provider endpoint it hits. That is present-tense evidence, not a label someone typed months ago. So the honest split: buy contact data from a firmographics vendor, get the qualification signal from the stack, then join them on the domain. Bassethound’s firmographics layer returns company name, vertical, contacts, and social profiles in the same dossier, so for many domains you get both in one call and skip the join.

What does the ICP signal add on top of the AI score?

The AI score answers whether a company ships AI. The icp_signal answers whether it fits you. It fuses the five layers. A hot signal means the verdict is shipping and the rest of the dossier lines up: a modern tech stack, real infra, a vertical you sell into. Warm means partial fit, a company experimenting or shipping in an adjacent space. Cold means the domain does not match, whatever the homepage says.

AI-readiness alone over-selects. Every well-funded startup is adding AI right now. If your product sells to teams already running a vector store and an orchestration layer, the score plus the specific detections confirm they hit that bar. If you sell the on-ramp to teams with none of it, invert the filter and target a verdict of experimenting or none with strong firmographics. Fusion lets you target the combination, not a single axis. An agent pulling five separate MCP tools cannot cheaply correlate them. One dossier already did.

Where does the AI-readiness signal miss?

Be honest about the gaps. A filter you trust without checking costs you deals. A static crawl reads what the front end exposes. A company that runs its entire AI stack server-side, behind its own API, with no client SDK and no provider call in the browser, can look quiet from the outside. Bassethound renders JavaScript when the signals call for it, which recovers client-only initialization, but a fully proxied backend stays dark. The dossier reports that gap rather than guessing past it.

So a verdict of none is not proof of absence. It means no signal surfaced on this crawl. For high-value domains, read the specific detections, not just the score, and treat a clean firmographic match with a quiet stack as a maybe, not a no. The inverse is the unusual case. A shipping verdict is hard to fake, since loading a vector store client and calling a provider endpoint costs real engineering and buys nothing when staged. False positives stay rare. False negatives are the risk you manage, so weight your list accordingly and reserve the deep profile for the domains that matter.

Bassethound perspective

Every firmographics vendor will tell you they know which companies are AI companies. They know which companies say they are. That is a labeling exercise, and labels lag the code by a quarter or more. The unclaimed axis is the deployed stack, and few competitors read it in depth. The AI-readiness checks you find elsewhere stop at the homepage and an llms.txt fetch. That rewards the company with the best copy, which is backward for outbound. You want the company that shipped, not the company that blogged.

We think the right unit of targeting is present-tense evidence: the SDK loaded, the vector store called, the provider hit. Read that across a list and you rank by what teams built, not what they announced. Do it keyless, in one MCP call per domain, and the qualification runs inside your outbound agent instead of a batch job you babysit. The moat is not the detection of any one layer. It is the fusion of five, correlated, in a single stateless request. A vendor selling you a contact list cannot hand you that.

Sources

Frequently asked questions

What separates an AI-native startup from a company that just mentions AI?

The deployed stack. An AI-native company loads model SDKs, calls provider endpoints, and wires up a vector store, and those signatures survive minification. A company that mentions AI has copy and nothing behind it. Bassethound scores the stack, not the story.

Can you target on the AI-readiness score alone?

You can, but it over-selects, because most funded startups are adding AI. Pair the score with the icp_signal, which fuses the AI verdict with tech, infra, and firmographics, so you rank on fit instead of one axis.

Does this run without an API key?

Yes. The free tier is keyless and rate-limited, so an outbound agent can qualify domains with no bring-your-own-key setup. OAuth unlocks higher depth and volume on the paid tier.

How do you qualify a long list of domains?

Run the fast profile first to drop the dead domains, then re-run the shipping candidates on standard or deep for contacts and the Wayback first-seen date. The endpoint is stateless, so each domain is one self-contained call.

What if a startup runs its AI fully server-side?

Then it can look quiet from the outside, and a verdict of none means no signal surfaced, not proof of absence. Bassethound renders JavaScript when the signals call for it and reports the gap. For high-value domains, read the specific detections, not just the score.

Sniff a domain.

Run sniff_domain on any site and read its five-layer dossier in one call.

Sniff a domain