Orange Collective
Context.dev

Context.dev

We give AI agents realtime web context.

Context.dev · YC S26 launch video

525

Customers

All inbound, product-led

120%

NRR

Per cohort, from a $25 entry

~50M

Requests / month

70M companies indexed

60%

Gross margin

Up from 50% at batch start

Thesis

AI has split data into three markets: pre-training corpora for the labs, post-training data for tuning models, and runtime data for agents doing work. This memo is about the third. Software is being rebuilt around agents, and an agent is only as useful as the system of record it works from. The skills are easy to replicate. Keeping the data underneath them complete and current is not.

The runtime market is young, and nearly everything sold today, by Context.dev or anyone else, is public open web data. The next layer is industry data, starting with public sources like permits, filings, and registries, then private datasets licensed for agent use. Context.dev's biggest deal is an early version of exactly that: Sunrun pulls 10M building permits a month through the API.[27] Underneath it the core business runs on 525 inbound customers with 120% net revenue retention.[27]

The bet: runtime data becomes core agent infrastructure. Every agent harness needs a system of record kept complete and current from public and private sources alike, and Context.dev is building that layer end to end.

  1. 01

    Efficient, product-led revenue. 525 customers, 120% NRR per cohort, all inbound; ~45–50M requests/month; gross margin up from 50% to 60% during the batch, with profitability a month out excluding YC credits.[27]

  2. 02

    Early enterprise pull without a sales motion. Sunrun signed in three days and added an enrichment order on the same PO; Klarna, Mintlify, Rho, and Passionfroot are also customers.[1][27]

  3. 03

    One API covers the whole data workflow. Scrape, search, extraction, entity data, and monitoring share one credit system, backed by an entity index covering 70M companies. No competitor sells that combination today.

  4. 04

    The founder has done this before. Yahia has two exits behind him and built this breadth with a team of four. Scrape, search, and entity data are each somebody's whole company.[3][7][20][26]

Problem

The web was built for humans to browse, not agents to consume.

Any AI product that needs live internet data ends up running its own scraping stack: browsers, proxies, crawlers, parsers, retry queues, caches, plus a few data vendors on the side.[1][2] The stack breaks whenever a target site changes.

Teams pay for this twice. Engineers spend their time on data plumbing instead of the product, and the agents fail anyway when the data they get is stale or badly parsed.

Why Now

Every SaaS is becoming an agent harness, and agents finally work.

The transition is early, and the plumbing just started working.

Software companies are rebuilding their products around agents, and most are only now shipping their first agent features. Every one of them needs live data behind its system of record and inside each agent session. That ties demand to the whole AI adoption curve rather than any single product cycle.

Browser and computer use also crossed a threshold this year. Agents can now operate real websites reliably, which multiplies how much live data they consume. It also strengthens the case for a structured API: when an agent can get the answer in one call, nobody wants to pay for a thirty-step browser session.

Meanwhile the open web is getting harder to read. Starting Sept 15, 2026, Cloudflare will block AI crawlers that mix search and agent traffic, and its pay-per-crawl program is shifting toward pay-per-use.[22][23] Harder access makes a managed data API more valuable and more expensive to operate, so this cuts for Context.dev and against it at the same time.

Product & Traction

One API in, structured context out.

Scrape, crawl, search, extract, entity data, monitor: one credit system.

Developers send the API URLs, domains, sitemaps, company names, or files, and get back clean Markdown, screenshots, entity data (brand, company, people, news), or JSON in whatever schema they define. Batches handle bulk jobs, and Monitors push updates over webhooks when a source changes. Context.dev runs the browsers, proxies, rendering, and retries underneath.[1]

Pricing is credit-based, starting at $25 a month with a 500-request free tier. The company started as Brand.dev, a logo and brand-data API, and grew into scraping, crawling, search, and extraction before taking the new name.

Traction (founder call, Aug 27, 2026)[27]

Traction
  • 525 customers, 120% NRR per cohort; all inbound and product-led
  • ~45–50M API requests/month; 70M companies indexed; 900M people being indexed
  • Gross margin 60%, up from 50% at batch start; profitable on Hetzner pre-YC; a month from profitability ex-YC credits
  • Sunrun: closed in 3 days, enrichment add-on on the same PO, pilot converts mid-Sept 2026 · Klarna: in staging

Market

Nobody has sized this market yet, so size it from what buyers already spend.

Every company building an AI product that needs live web data already pays for it somehow: engineers maintaining a scraping stack, plus separate bills for proxies, browsers, and data vendors. Context.dev's contract replaces spending that already exists, which is why Sunrun could sign within three days.[27]

On the growth side, 525 customers generate 45–50M requests a month and expand at 120% NRR from a $25 entry point, so revenue rides two curves: how many products have agents in them, and how much data each agent session pulls.[27] Both are early, and the second is the one browser-capable agents are now bending upward. The move into industry data raises the ceiling again, because permits, filings, and licensed datasets sell as data products rather than metered requests.

Market map & competitive landscape

One product spanning at least four separately funded categories.

Each piece of the product has its own set of competitors. The map covers where each category ends and Context.dev begins, and who is moving across that line.

Scrape/crawl infrastructure

Direct threat

Firecrawl, Diffbot. The sharpest threat: Firecrawl sells nearly the same product with open-source distribution Context.dev lacks (150K+ GitHub stars, 2.5M+ weekly SDK downloads), is adding a proprietary "Semantic Index," retagged on GitHub as "The context API," and publishes a dedicated "vs Context.dev" comparison page. Boundary: no entity/enrichment dataset. Pricing-opacity complaints about Firecrawl recur in third-party roundups.[3][4][5][6][25]

Agent-native search & index

Capital overhang

Exa ($335M+ raised, $2.2B val, search rebuilt as a neural index), Parallel ($230M, $2B, supplies Google Cloud; customers Clay, Harvey, Notion), Perplexity Search API (hundreds-of-billions-page index at $5/1K requests), Tavily. Whoever owns the agent's search box can add scraping as a feature. Boundary today: none sells site-targeted crawling, custom-schema extraction, monitors, or entity enrichment as first-class products.[7][8][9][10][11][12][13][18]

Entity/enrichment data APIs

Contested, fragmented

People Data Labs (3B records, batch-era) and Crustdata (YC F24, $6M seed, same "live data for agents" pitch, ~$4M revenue). The original Brand.dev product lives here, the least crowded category Context.dev touches and where a data asset could form. Boundary: Context.dev bundles enrichment with on-demand scraping in one credit system.[20][21]

GTM apps & infrastructure beneath

Indirect / adjacent

Clay ($3.1B) and Apollo.io ($1.6B) absorb enrichment budget from the application side: the buyer wants filled spreadsheets, not endpoints. But they sell to sales teams, not developers, and Clay's marketplace can also be a channel. Beneath: Browserbase and Browser Use sell sessions, not data. Exception: Bright Data, $300M+ ARR, moving up-stack via Web MCP; enterprise-priced, PE-owned, not a PLG product.[14][15][16][17][19]

Assessment

This is a crowded, heavily funded map: Exa has raised $335M+, Parallel $230M, Clay $204M+, Apollo around $250M, Browserbase $67.5M, and Firecrawl $16.2M, while Perplexity and Bright Data operate at another order of magnitude entirely. Crowded does not mean winner-take-all. The categories next door already support several large companies each: in enrichment, ZoomInfo built a business with about $1.2B in annual revenue on enterprise seats,[28] Clay reached $100M ARR on usage-based workflows,[29] and Apollo grew between them on SMB volume;[15] in scraping, Bright Data's $300M+ ARR coexists with Firecrawl's open-source funnel.[19][4] Runtime data spans an even wider range of buyers, from a solo developer on a $25 plan to Sunrun's operations team,[27] and no vendor holds both ends today. Context.dev's claim is the segment that wants the whole workflow from one metered API, plus data deals like Sunrun that no single endpoint was built to win. The endpoints themselves are converging anyway: a new benchmark scores the top search APIs within two points of each other (Parallel 75, Exa 74, Firecrawl 73).[24] The investment turns on Context.dev holding that segment as the bigger players broaden. In a market with several winners, it does not need to beat them all.

Founder

Yahia Bakour

Yahia Bakour

Repeat FounderExited

Founder

Solo founder, repeat founder with an exit. Co-founded Stock Alarm (2019–2023) as CTO: a bootstrapped stock/crypto alerting product that reached 225K+ users and was acquired by Center Mark Capital. Founded Essense.io (2023), an AI feedback-analysis tool, acqui-hired by Optic the same year. Amazon SDE2/tech lead on the retail checkout platform (2020–2022). Principal engineer, then AI engineering manager at Sunrun (2023–2026); left a ~$1M/year role to start this company. Team of 4 (Aug 2026).

Fit

He has built and sold data products twice[26], the company grew out of a problem he was solving for himself (brand data first, then everything around it), and he has taken 300+ customer calls during the batch.[27]

Risks

Risk

Well-funded competitors can converge on the bundle. Firecrawl is adding an index, Exa and Parallel could add scraping, and Bright Data is moving up the stack. The orchestration layer that would make the bundle hard to copy (request routing, self-maintaining extraction, learning from outcomes) is still on the roadmap.

Mitigation

Speed and focus. Matching the full stack would cost each competitor years of roadmap outside its core product, while the entity index gets deeper with every request served. Sunrun and Klarna are accounts that scrape or search alone would not have won, and once the API is wired into a customer's daily operations it competes as infrastructure, not as one scraper against another.

Risk

Margins face outside pressure just as the company scales. 60% gross margin is already below typical API-software margins; Cloudflare's crawler mandate and pay-per-use pricing raise costs and legal exposure for exactly this business; and 2–4s latency suggests the costs are structural, not temporary.

Mitigation

Margins already improved ten points during the batch, and rising access costs hit the well-funded competitors too. The catch is that they can absorb the pressure far longer than a company a month from break-even.

Risk

The industry-data shift is still a thesis, not a product line. Today's revenue is public open web work, where differentiation is thinnest and price pressure arrives first. The Sunrun and Klarna deals are early signs of vertical data products, not yet a dataset business, and both vertical incumbents and the bigger horizontal players can chase the same turn.

Mitigation

The entity index and Monitors are the beginnings of proprietary data products, and the index grows as a side effect of serving paid requests. The thing to watch is whether index-backed products grow faster than pass-through scraping.

What we're watching

  • Customer reference calls (Mintlify, Sunrun, one mid-market); the Sunrun pilot's contracted conversion in mid-Sept 2026 is the first checkable milestone.
  • Klarna staging → production. A signed European enterprise on a 6+ month compliance path converting would be the strongest single validation of the enterprise motion.
  • The orchestration layer (routing, self-healing extraction, managed datasets) shipping, versus remaining a bundle of primitives. It is the thesis's moat and does not exist today.
  • Benchmark presence and latency: appearing in the Artificial Analysis search-API benchmarks and getting below the reported 2–4s per request.
  • NRR and margin holding through Sept 15, 2026 as Cloudflare's crawler rules bite, and the first hires beyond the team of 4 (key-person exposure).

References

  1. [1]Context.dev: Live web data API for AI agents — YC Launch
  2. [2]Context.dev — Y Combinator company profile
  3. [3]Firecrawl announces $14.5M Series A (GlobeNewswire, Aug 19, 2025)
  4. [4]Firecrawl homepage — GitHub stars, SDK downloads, Semantic Index
  5. [5]Firecrawl GitHub — "The context API" tagline
  6. [6]Firecrawl vs Context.dev — Firecrawl comparison page
  7. [7]Exa raises $250M Series C at $2.2B — Exa blog (May 2026)
  8. [8]Parallel Web Systems hits $2B valuation (TechCrunch, Apr 29, 2026)
  9. [9]Parag Agrawal's Parallel raises $100M (Newcomer)
  10. [10]Training Data (Sequoia) — Parag Agrawal episode (Aug 2026)
  11. [11]The a16z Show — Building Search for AI Agents with Exa CEO Will Bryk (Jun 4, 2026)
  12. [12]Latent Space — Beating Google at Search with Neural PageRank, with Will Bryk of Exa (Jan 10, 2025)
  13. [13]Introducing the Perplexity Search API (Sep 2025)
  14. [14]Clay raises $100M Series C at $3.1B (Businesswire, Aug 5, 2025)
  15. [15]Apollo.io bags $100M at $1.6B (TechCrunch, Aug 29, 2023)
  16. [16]Browserbase raises $40M Series B (Jun 2025)
  17. [17]Browser Use raises $17M (TechCrunch, Mar 23, 2025)
  18. [18]Tavily raises $20M Series A (Calcalist, Aug 2025)
  19. [19]Bright Data reaches $300M ARR (Asymmetrix)
  20. [20]People Data Labs raises $45M Series B (GlobeNewswire, Nov 2021)
  21. [21]Crustdata closes $6M seed (Nov 2025)
  22. [22]Cloudflare's new policy pushes AI companies to pay for content (TechCrunch, Jul 1, 2026)
  23. [23]Cloudflare blocks mixed-use AI crawlers from Sept 15, 2026 (Technology.org, Jul 3, 2026)
  24. [24]New benchmark ranks search APIs for AI agents (The Decoder)
  25. [25]Firecrawl pricing complaints roundup (eesel — competitor-published, directional)
  26. [26]Stock Alarm acquisition — Acquire.com podcast, episode 92
  27. [27]Orange Collective — Founder call with Yahia Bakour, Aug 27, 2026 (internal)
  28. [28]ZoomInfo Form 10-Q, Q2 2025 (SEC)
  29. [29]Clay hits $100M ARR (arr.club, Dec 2025)