
525
Customers
All inbound, product-led
120%
NRR
Per cohort, from a $25 entry
~50M
Requests / month
70M companies indexed
60%
Gross margin
Up from 50% at batch start
Thesis
AI has split data into three markets: pre-training corpora for the labs, post-training data for tuning models, and runtime data for agents doing work. This memo is about the third. Software is being rebuilt around agents, and an agent is only as useful as the system of record it works from. The skills are easy to replicate. Keeping the data underneath them complete and current is not.
The runtime market is young, and nearly everything sold today, by Context.dev or anyone else, is public open web data. The next layer is industry data, starting with public sources like permits, filings, and registries, then private datasets licensed for agent use. Context.dev's biggest deal is an early version of exactly that: Sunrun pulls 10M building permits a month through the API.[27] Underneath it the core business runs on 525 inbound customers with 120% net revenue retention.[27]
The bet: runtime data becomes core agent infrastructure. Every agent harness needs a system of record kept complete and current from public and private sources alike, and Context.dev is building that layer end to end.
- 01
Efficient, product-led revenue. 525 customers, 120% NRR per cohort, all inbound; ~45–50M requests/month; gross margin up from 50% to 60% during the batch, with profitability a month out excluding YC credits.[27]
- 02
Early enterprise pull without a sales motion. Sunrun signed in three days and added an enrichment order on the same PO; Klarna, Mintlify, Rho, and Passionfroot are also customers.[1][27]
- 03
One API covers the whole data workflow. Scrape, search, extraction, entity data, and monitoring share one credit system, backed by an entity index covering 70M companies. No competitor sells that combination today.
- 04
The founder has done this before. Yahia has two exits behind him and built this breadth with a team of four. Scrape, search, and entity data are each somebody's whole company.[3][7][20][26]
Problem
The web was built for humans to browse, not agents to consume.
Any AI product that needs live internet data ends up running its own scraping stack: browsers, proxies, crawlers, parsers, retry queues, caches, plus a few data vendors on the side.[1][2] The stack breaks whenever a target site changes.
Teams pay for this twice. Engineers spend their time on data plumbing instead of the product, and the agents fail anyway when the data they get is stale or badly parsed.
Why Now
Every SaaS is becoming an agent harness, and agents finally work.
The transition is early, and the plumbing just started working.
Software companies are rebuilding their products around agents, and most are only now shipping their first agent features. Every one of them needs live data behind its system of record and inside each agent session. That ties demand to the whole AI adoption curve rather than any single product cycle.
Browser and computer use also crossed a threshold this year. Agents can now operate real websites reliably, which multiplies how much live data they consume. It also strengthens the case for a structured API: when an agent can get the answer in one call, nobody wants to pay for a thirty-step browser session.
Meanwhile the open web is getting harder to read. Starting Sept 15, 2026, Cloudflare will block AI crawlers that mix search and agent traffic, and its pay-per-crawl program is shifting toward pay-per-use.[22][23] Harder access makes a managed data API more valuable and more expensive to operate, so this cuts for Context.dev and against it at the same time.
Product & Traction
One API in, structured context out.
Scrape, crawl, search, extract, entity data, monitor: one credit system.
Developers send the API URLs, domains, sitemaps, company names, or files, and get back clean Markdown, screenshots, entity data (brand, company, people, news), or JSON in whatever schema they define. Batches handle bulk jobs, and Monitors push updates over webhooks when a source changes. Context.dev runs the browsers, proxies, rendering, and retries underneath.[1]
Pricing is credit-based, starting at $25 a month with a 500-request free tier. The company started as Brand.dev, a logo and brand-data API, and grew into scraping, crawling, search, and extraction before taking the new name.
Traction (founder call, Aug 27, 2026)[27]
Traction- 525 customers, 120% NRR per cohort; all inbound and product-led
- ~45–50M API requests/month; 70M companies indexed; 900M people being indexed
- Gross margin 60%, up from 50% at batch start; profitable on Hetzner pre-YC; a month from profitability ex-YC credits
- Sunrun: closed in 3 days, enrichment add-on on the same PO, pilot converts mid-Sept 2026 · Klarna: in staging
Market
Nobody has sized this market yet, so size it from what buyers already spend.
Every company building an AI product that needs live web data already pays for it somehow: engineers maintaining a scraping stack, plus separate bills for proxies, browsers, and data vendors. Context.dev's contract replaces spending that already exists, which is why Sunrun could sign within three days.[27]
On the growth side, 525 customers generate 45–50M requests a month and expand at 120% NRR from a $25 entry point, so revenue rides two curves: how many products have agents in them, and how much data each agent session pulls.[27] Both are early, and the second is the one browser-capable agents are now bending upward. The move into industry data raises the ceiling again, because permits, filings, and licensed datasets sell as data products rather than metered requests.
Market map & competitive landscape
One product spanning at least four separately funded categories.
Each piece of the product has its own set of competitors. The map covers where each category ends and Context.dev begins, and who is moving across that line.
Founder
Risks
What we're watching
References
- [1]Context.dev: Live web data API for AI agents — YC Launch
- [2]Context.dev — Y Combinator company profile
- [3]Firecrawl announces $14.5M Series A (GlobeNewswire, Aug 19, 2025)
- [4]Firecrawl homepage — GitHub stars, SDK downloads, Semantic Index
- [5]Firecrawl GitHub — "The context API" tagline
- [6]Firecrawl vs Context.dev — Firecrawl comparison page
- [7]Exa raises $250M Series C at $2.2B — Exa blog (May 2026)
- [8]Parallel Web Systems hits $2B valuation (TechCrunch, Apr 29, 2026)
- [9]Parag Agrawal's Parallel raises $100M (Newcomer)
- [10]Training Data (Sequoia) — Parag Agrawal episode (Aug 2026)
- [11]The a16z Show — Building Search for AI Agents with Exa CEO Will Bryk (Jun 4, 2026)
- [12]Latent Space — Beating Google at Search with Neural PageRank, with Will Bryk of Exa (Jan 10, 2025)
- [13]Introducing the Perplexity Search API (Sep 2025)
- [14]Clay raises $100M Series C at $3.1B (Businesswire, Aug 5, 2025)
- [15]Apollo.io bags $100M at $1.6B (TechCrunch, Aug 29, 2023)
- [16]Browserbase raises $40M Series B (Jun 2025)
- [17]Browser Use raises $17M (TechCrunch, Mar 23, 2025)
- [18]Tavily raises $20M Series A (Calcalist, Aug 2025)
- [19]Bright Data reaches $300M ARR (Asymmetrix)
- [20]People Data Labs raises $45M Series B (GlobeNewswire, Nov 2021)
- [21]Crustdata closes $6M seed (Nov 2025)
- [22]Cloudflare's new policy pushes AI companies to pay for content (TechCrunch, Jul 1, 2026)
- [23]Cloudflare blocks mixed-use AI crawlers from Sept 15, 2026 (Technology.org, Jul 3, 2026)
- [24]New benchmark ranks search APIs for AI agents (The Decoder)
- [25]Firecrawl pricing complaints roundup (eesel — competitor-published, directional)
- [26]Stock Alarm acquisition — Acquire.com podcast, episode 92
- [27]Orange Collective — Founder call with Yahia Bakour, Aug 27, 2026 (internal)
- [28]ZoomInfo Form 10-Q, Q2 2025 (SEC)
- [29]Clay hits $100M ARR (arr.club, Dec 2025)

