--- title: "Clean Architecture for Agents That Need Browser + Scraping APIs" category: "guides" url: https://keirolabs.cloud/blog/clean-architecture-agents-browser-scraping --- Route every URL to the cheapest layer that can read it. A search API discovers pages, an extraction API reads them, and a browser API handles login walls, JavaScript apps and infinite scroll. The router in front of those three contracts is the whole architecture. Everything else is configuration. This post lays out the stack layer by layer, gives you a routing rule you can implement in an afternoon, and shows the TypeScript that implements it. It also shows the arithmetic, because the architecture argument is mostly an arithmetic argument: reading 1,000 static pages through an extraction API bills $2.40 per 1,000 on Keiro's Startup plan, and the same 1,000 pages through a browser session bill $3.33 to $6.80 at published browser-hour rates, before proxies. Pick the wrong layer for a page and you pay two to three times more for the same text. All vendor prices below were checked on September 23, 2026 against the vendors' own pages: Browserbase, Browserless and Steel for the browser layer, Firecrawl for extraction, and Keirolabs for search and extraction. Where a number is my arithmetic rather than a published rate, the text says so. The one-line version: buy discovery by the query, reading by the page, and presence by the hour. Never let one contract do all three jobs. ## Clean Architecture for Agents That Need Browser + Scraping APIs Together Clean architecture gets invoked loosely, so here is the version that matters for agents: your agent's core logic should depend on interfaces you own, and every vendor should sit behind an adapter. An agent that imports one vendor's SDK all over its codebase is an agent you will rewrite the quarter that vendor reprices, degrades, or loses a legal fight. In this market those events are seasonal, not hypothetical. Prices in the search API category moved at least four times across major vendors in 2026. The stack has three layers, and each one answers a different question about a page. 1. **Search layer (discovery).** A query goes in, ranked URLs come out with snippets and metadata. This layer answers "where does the answer live?" It never reads pages. Keirolabs /search/lite lists at $0.25 per 1,000 queries, Serper at $1.00, Brave at $5.00, Exa at $7.00 for 10 results. 2. **Extraction layer (reading).** A URL goes in, clean markdown comes out. This layer answers "what does the page say?" It never searches. Keiro's /search/content and /extract bill 3 credits a request. Firecrawl bills 1 credit a page on basic scrape. 3. **Browser layer (presence).** A URL and a session go in, a rendered page comes out, along with screenshots, clicks and scroll state. This layer handles what the first two cannot: JavaScript-only apps, login walls, infinite scroll feeds, and sites that challenge datacenter traffic. Browserbase, Browserless and Steel sell it by the browser hour at $0.08 to $0.12 published overage rates. The layers bill in different units, and that is the deep reason they belong in separate contracts. Discovery bills by the query, reading bills by the page, and presence bills by the hour. A single API that does all three has to average those units into one price, and averages are where vendors hide margin and buyers overpay. The exa-shaped version of that sentence: nobody averages a $0.25 query and a $0.10 hour into one fair number. Here is the contract surface. The whole agent depends on these three interfaces and nothing else. ```ts // ports.ts - the only web access the agent core is allowed to know about. export interface SearchHit { url: string; title: string; snippet: string; } export interface SearchPort { search(query: string, maxResults?: number): Promise; } export interface PageDoc { url: string; markdown: string; fetchedAt: string; // ISO 8601 via: "extract" | "browser"; } export interface ExtractPort { // Returns null on a 404, a dead domain, or an empty read. extract(url: string): Promise; } export interface BrowserPort { // sessionKey lets the adapter reuse cookies and storage per identity. read(url: string, opts: { sessionKey: string; timeoutMs?: number }): Promise; } ``` The routing rule is the part teams skip, and it is where the money is. Route per URL, not per agent: - **URL unknown.** Search first, then read what comes back. Never let an agent guess URLs. - **URL known, static content.** Extraction, at 3 credits a read. This is the default, not the exception. On a general web corpus the extraction layer should handle the large majority of reads; measure your own escalation rate and you will have the number for your corpus within a week. - **Known URL, hard page.** JavaScript app, login wall, scroll feed, or a domain that bot-walls datacenter IPs: straight to the browser layer. - **Escalation.** When extraction returns a thin body (markdown under a few hundred characters, missing main content, or a challenge page), escalate to the browser once. One escalation, never a loop. In code, the router is about forty lines. ```ts // router.ts - send every URL to the cheapest layer that can read it. type Hint = { needsLogin?: boolean; needsJs?: boolean }; export async function readPage( url: string, deps: { extract: ExtractPort; browser: BrowserPort }, hint: Hint = {}, ): Promise { // Hard signals win before any request is spent. if (hint.needsLogin) { return deps.browser.read(url, { sessionKey: "identity:default" }); } const jsOnly = hint.needsJs || /(app\.|dashboard\.|docs\.google\.com|notion\.site|\/spa\/)/i.test(url); if (jsOnly) { return deps.browser.read(url, { sessionKey: "pool:js" }); } // Default path: pay 3 credits, not browser minutes. const doc = await deps.extract.extract(url); const thin = !doc || doc.markdown.length < 500; if (!thin) return doc; // One escalation. The browser is the escalator, not the default. return deps.browser.read(url, { sessionKey: "pool:fallback" }); } ``` What the router buys, numerically. Reading 1,000 static pages through /search/content on Keiro's Startup plan costs $2.40 (3 credits a read, $0.0008 per credit at that plan). Sending the same 1,000 pages to a browser costs $3.33 per 1,000 on Steel's Launch rate ($0.10 an hour at the 2-minute-per-page rule of thumb) up to $6.80 per 1,000 on Browserless' Starter overage, plus $6 to $10 per GB of residential proxy where the target demands it. The router is the difference between a $2.40 line item and a $6.80 one, on every run, forever. There is a second-order effect that matters more: when extraction fails and the browser picks the page up, you pay both meters for that URL. An agent without a router escalates blindly and pays double on a large share of reads. An agent with one escalates on a thin-body signal and pays double only where the page actually demanded it. Track that escalation rate as a first-class metric; it is the heartbeat of this architecture. A rate in the low single digits on a general corpus is healthy. A rate climbing past 20 percent means your extraction vendor or your user-agent hygiene is the problem, not the architecture. The alternative worth naming is the all-in-one platform that sells search, scrape and browser automation on one credit balance, Firecrawl being the clearest example (search at 2 credits per 10 results, scrape at 1 credit a page, Interact at 2 credits a browser minute, all drawing on one balance). The honest pros and cons: - **All-in-one.** Pros: one bill, one key, one SDK, one dashboard. Cons: every layer inherits the weakest layer's price, and time-based features price badly against specialist vendors (Interact works out to about $0.60 an hour at Firecrawl's $5-per-1,000-credits pay-as-you-go rate, six times Steel's published browser hour). - **Three specialists.** Pros: best rate at each layer, and any single vendor can be repriced or replaced without touching the agent. Cons: three keys, three dashboards, three failure modes to monitor. Keiro covers the first two layers on one credit balance, which is the part of the all-in-one convenience that costs nothing, and leaves the browser layer to a specialist, which is the part that does. That split is not a pitch; it is what the unit economics look like from the inside. The rest of this post stays vendor-neutral about the shape: any search port, any extract port, any browser port. ## What Is the Difference Between a Search API and a Scraping API for LLMs? The two APIs answer different questions and fail in different ways, and LLM pipelines need both because they play different roles in grounding. A **search API** takes a query and returns ranked URLs. Its contract is relevance, and its input is underspecified: the same question can be phrased ten ways, and the layer's whole value is mapping all ten onto the right result set. Its unit is the query. Published rates run from $0.08 per 1,000 (Keiro /search/lite on the Startup plan) to $8.00 (Tavily basic), with Exa at $7.00, Brave at $5.00, Serper at $1.00, and Parallel at $1.00 to $5.00 in between. Search indexes are caches of the web: ranked, deduplicated, and stale by minutes to days depending on the vendor. A **scraping or extraction API** takes a URL and returns the page's content. Its contract is fidelity, and its input is exact. The same URL always resolves the same way until the page changes. Its unit is the page (1 to 3 credits) or, once rendering is involved, the minute. There is no relevance problem here; there is a fidelity problem: did you get the whole page, cleaned of boilerplate, in a shape a model can use? Three consequences fall out of that split, and each one changes how you build: - **Empty results and broken HTML are different failures.** A search returning zero hits means reformulate; retrying the same string is waste. A scrape returning a challenge page or a 500-byte shell means escalate to a different layer. Pipelines that treat both as "retry" burn budget without changing the outcome. - **Snippets are triage, pages are proof.** Snippets are short, pre-ranked, and written for a results page, not for a model. On our deep search run against SimpleQA (1,000 seeded questions, retrieval only, no reader model), the full-page read layer is worth 15.3 of the 95.3 percent score. A snippet-only pipeline leaves roughly 15 points on the table because it never reads the page it cites. - **The caches differ.** Search results cache well because ranking changes slowly for most queries. Page content caches extremely well because most pages barely change, which is why the reading layer is where a content-hash cache earns its keep. More on that in the RAG section below. Where teams get confused is vocabulary. In 2026 vendor language, a "scraping API" usually means a managed read with proxies and rendering included, a "SERP API" means scraped Google results packaged as JSON, and a "browser API" means a session you drive yourself. The LLM-facing question for all three is the same: what arrives in the model's context window, at what price, and how stale is it? A $1.00 SERP result with a 30-word snippet and a 3-credit full-page read are not substitutes. One gives you 30 words for a fraction of a cent; the other gives you the page. The pros and cons, judged for LLM work specifically: - **Search-only agent.** Cheap ($0.08 to $8.00 per 1,000 questions) and fast, but it quotes snippets, and snippets misquote. Citation quality suffers, and numeric claims arrive without their surrounding sentence, which is where most hallucinations in grounded pipelines actually start. - **Scrape-only agent.** Great fidelity, no discovery. It needs URLs from somewhere, and "somewhere" is always a search API, a sitemap, or a user. - **Both, behind a router.** Discovery by the query, grounding by the page. This is the pipeline every serious agent stack converges on, and the rest of this post assumes it. A snippet is a promise. A page read is proof. Build the pipeline so the model can tell the difference. ## Best Scraping APIs and SDK for AI Agent Workflows With the contracts fixed, choosing vendors becomes a table exercise. Here is the layer-to-tool map with published list rates as of September 23, 2026. | Layer | Job | Keiro option | Other options | Published rate | |---|---|---|---|---| | Search (discovery) | Query to ranked URLs | /search/lite, 0.1 credit a request on plans ($0.25/1k list) | Serper, Brave, Exa, Parallel, Tavily | $0.08 to $8.00 per 1,000 queries | | Extraction | URL to markdown | /search/content or /extract, 3 credits ($2.40/1k on Startup, $7.20/1k on Essential) | Firecrawl scrape, Tavily extract | $2.40 to $7.20 per 1,000 pages (Keiro plans), $5.00/1k Firecrawl pay-as-you-go | | Browser | Render, log in, scroll, survive bot walls | Not offered; pair with a specialist | Browserbase, Browserless, Steel | $0.08 to $0.12 per browser hour | | Answers | Query to cited answer | /answer, 5 credits | Tavily, Exa answers, Perplexity | 5 credits a call on Keiro | | Full pipeline | Question to proof-scored answer | /agentic deep search, 20 credits ($4.44/1k list) | Exa deep research $12 to $15/1k | $4.44 per 1,000 (Keiro list) | The SDK story is shorter than the vendor list suggests, and that is the point. What a good adapter hides: auth, retries, unit conversion, and response normalization. What it should not hide: the vendor's billing unit, because the moment your code treats credits as if they were requests, your cost model detaches from your invoice. For the browser layer, the practical SDK choice is Playwright (or Puppeteer) over CDP rather than any vendor's proprietary client. All three specialists support it: Steel exposes CDP endpoints so Playwright and Puppeteer code works without rewrites, Browserless documents Puppeteer and Playwright connections alongside its BrowserQL language, and Browserbase maintains Stagehand, its own agent SDK on top of Playwright. Write against the CDP shape and the browser port stays swappable. Here are two adapters that implement the ports from the first section. Keiro for search and extraction: ```ts // adapters-keiro.ts const BASE = "https://api.keirolabs.cloud"; const KEY = process.env.KEIRO_API_KEY; async function keiro(path: string, body: unknown) { const res = await fetch(`${BASE}${path}`, { method: "POST", headers: { "x-api-key": KEY!, "content-type": "application/json" }, body: JSON.stringify(body), }); if (!res.ok) throw new Error(`keiro ${path} -> ${res.status}`); return res.json(); } export const keiroSearch: SearchPort = { async search(query, maxResults = 10) { // 0.1 credit a request on monthly plans. $0.24/1k on Essential. const r = await keiro("/api/v2/search/lite", { q: query, max_results: maxResults }); return (r.results ?? []).map((h: any) => ({ url: h.url, title: h.title ?? "", snippet: h.snippet ?? h.description ?? "", })); }, }; export const keiroExtract: ExtractPort = { async extract(url) { // 3 credits. Same balance as search, so one plan covers both layers. const r = await keiro("/api/v2/extract", { url, formats: ["markdown"] }); if (!r?.markdown) return null; return { url, markdown: r.markdown, fetchedAt: new Date().toISOString(), via: "extract", }; }, }; ``` And Steel for the browser port, with the two details that decide your browser bill: release the session in a finally block, and reuse session contexts per identity instead of paying for a fresh one every call. ```ts // adapters-steel.ts - any session browser (Steel, Browserbase, Browserless) // implements the same BrowserPort. This one uses Steel's REST + CDP surface. import { chromium } from "playwright-core"; export const steelBrowser: BrowserPort = { async read(url, { sessionKey, timeoutMs = 45_000 }) { const res = await fetch("https://api.steel.dev/v1/sessions", { method: "POST", headers: { "steel-api-key": process.env.STEEL_API_KEY!, "content-type": "application/json", }, body: JSON.stringify({ sessionContextId: sessionKey, // reuse cookies/storage for this identity blockResources: true, // images and fonts are not content timeoutMs: 60_000, // hard cap; Launch caps sessions at 15 min }), }); if (!res.ok) return null; const session = await res.json(); try { const browser = await chromium.connectOverCDP(session.cdpUrl); const context = browser.contexts()[0]; const page = await context.newPage(); await page.goto(url, { waitUntil: "domcontentloaded", timeout: timeoutMs }); await page.waitForTimeout(1_000); // let hydration settle; tune per target const markdown = await page.evaluate(() => document.body.innerText); return { url, markdown, fetchedAt: new Date().toISOString(), via: "browser", }; } finally { // Release unconditionally. An unreleased session bills by the hour. await fetch(`https://api.steel.dev/v1/sessions/${session.id}/release`, { method: "POST", headers: { "steel-api-key": process.env.STEEL_API_KEY! }, }); } }, }; ``` Two implementation notes on that adapter. First, `blockResources: true` is not a nicety; cutting images and fonts cuts proxy bandwidth, and proxies bill by the gigabyte ($6 to $10 per GB on Steel's plans, 6 units per MB on Browserless for residential). Second, `innerText` is a stand-in; production adapters convert the DOM to markdown with a library like Turndown and strip nav/footer boilerplate before caching, because the extraction quality bar for RAG is "what a reader would see," not "what the DOM contains." Vendor picks per layer, briefly and with the cons stated: Keiro for search plus extraction on one balance (cheapest published rates at both, catch: one-time credit packs bill lite searches at 0.5 credit instead of 0.1, so the monthly plan is the cheap path). Firecrawl if your pipeline wants search and full-page markdown from one vendor and can live with per-page billing that charges even for 403 and 404 responses. Browserbase if you want the most mature session platform and accept $0.12 an hour past the included 100. Browserless if unit-based pricing and SOC 2 Type II compliance matter more than raw price. Steel if you want the cheapest browser hour ($0.08 to $0.10) and an open-source core you can self-host if the relationship ends. ## ScrapingBee Alternatives for Autonomous AI Browser Workflows Teams usually arrive at the browser layer through a rendering API: the scraper started returning JavaScript shells, someone turned on a JS-rendering option, and the pages came back rendered. ScrapingBee built a large business on that motion, and it remains a reasonable tool for one-shot public page renders. The trouble starts when the workflow becomes autonomous. Autonomous means multi-step, stateful, and often logged in, and per-render credit models misprice all three. A render priced as one unit assumes one request, one response, no memory. An agent that logs into a portal, navigates three screens, and exports a report is none of those things, and pricing it per render is like pricing a phone call per syllable. What replaced it is the session API: you rent a browser by the hour (or by the 30-second unit), drive it over CDP with Playwright or Puppeteer, and manage sessions as first-class objects. Three vendors define that category in 2026, and all three published rates below are from their own pages as of September 23, 2026. **Browserbase.** Developer plan at $20 a month with 25 concurrent browsers and 100 browser hours included, $0.12 per browser-hour overage; Startup at $99 with 100 concurrent and 500 hours, $0.10 overage; a free tier with 3 concurrent browsers and 1 hour for evaluation. Their own rule of thumb: a typical web scrape finishes in under two minutes, so 100 hours covers roughly 3,000 page-level tasks. Watch the tiering: stealth mode is Basic on Developer and Startup with Advanced reserved for Enterprise, and data retention is 7 days on the lower plans. **Browserless.** Bills in Units, where a Unit is up to 30 seconds of browser time per connection. Prototyping at $35 a month (20,000 units, $0.0020 overage per unit, 15-minute session cap), Starter at $200 (180,000 units, $0.0017, 30-minute cap), Scale at $500 (500,000 units, $0.0015, 60-minute cap), and a free tier of 1,000 units with no card. Persisted sessions and replays run 7 to 90 days by plan, residential proxies bill 6 units per MB, and the platform holds SOC 2 Type II, which matters when the browser layer touches authenticated customer data. **Steel.** The cheapest published browser hour: $0.10 on Launch (billed per minute, rounded up, 10 concurrent sessions, 15-minute session cap) and $0.08 on Scale at $250 a month (100 concurrent, 1-hour cap), with $30 in one-time free credits that cover about 300 browser hours. Proxies run $6 to $10 a GB, CAPTCHA solves $1 to $3 per 1,000 on paid tiers, and the /scrape, /screenshot and /pdf browser tools bill $5 per 1,000. The core is open source, so the exit to self-hosting is real rather than marketing. Session and auth handling is where autonomous workflows live or die, and it is the part the rendering APIs never made you think about. Four rules that hold up in production: - **Identity is a resource, not a parameter.** Each tenant or persona gets a persistent session context (Steel session contexts, Browserbase contexts, Browserless persisted sessions at 7 to 90 days by plan). Cookies, storage and login state live there. The agent never sees credentials; it hands the browser a sealed profile. - **Refresh before expiry, not after 401.** A login replay mid-task burns minutes and CAPTCHA solves. If a portal's sessions die at 30 minutes, re-authenticate at 25. - **Cap wall clock and act on idle.** Every live session burns $0.08 to $0.10 an hour whether it works or waits. Set an inactivity timeout (Steel exposes one directly) and release the session the moment you have your artifacts. An agent that thinks for 40 seconds between clicks is paying for that thinking at browser-hour rates. - **Checkpoint and resume.** Long flows should write partial results to storage and finish in a fresh session rather than hold one open. Steel's Launch cap is 15 minutes and Scale's is 1 hour; Browserless caps sessions at 2 to 60 minutes by plan; design for the cap, not against it. The session-reuse pattern is small enough to inline: ```ts // sessions.ts - warm sessions per identity; idle ones die on a timer. const live = new Map(); export async function withSession( key: string, fn: (sessionId: string) => Promise, ): Promise { const s = live.get(key); if (s && Date.now() - s.lastUsed < 4 * 60_000) { s.lastUsed = Date.now(); // warm: cookies valid, no login replay return fn(s.id); } const id = await createSession({ contextId: key, idleTimeoutMs: 120_000 }); live.set(key, { id, lastUsed: Date.now() }); return fn(id); } ``` Do the idle math once and the timer writes itself. A 10-step flow with 40 seconds of model thinking between steps is about 7 minutes of wall clock, of which under 2 minutes is actual page work. At $0.10 an hour that is cheap per run and brutal at scale: 10,000 such flows a month is roughly 1,100 browser hours, or about $90 to $133 at the published overage rates, most of it spent watching a model think. Release on idle and the same workload lands nearer the working minutes. Idle browser hours bill at the same rate as working ones; that sentence is the whole browser cost discipline. Pros and cons, stated plainly: - **Browserbase.** Pros: the mature platform, Stagehand, observability with session recording, 3 to 250-plus concurrency by plan. Cons: $0.12 an hour is not the cheapest overage, advanced stealth is a sales conversation, and 7-day retention on lower plans complicates the "what did the agent do last Tuesday" debug. - **Browserless.** Pros: predictable unit pricing, SOC 2 Type II, self-host path on enterprise, 1,000 free units without a card. Cons: 30-second units reward short sessions and punish long ones; the 2-minute free-session cap makes evaluation fiddly. - **Steel.** Pros: cheapest published hour, per-minute billing, open-source core, honest operational guidance (release sessions, set idle timeouts). Cons: smaller ecosystem than Browserbase, 15-minute Launch session cap forces checkpointing on long flows, stealth bundled only from Scale. When should you stay on a per-render API? When the job is one-shot, stateless, and public: a product page, a news article, a search-result page. One render, one credit, done. The moment a workflow holds state across steps or logs in anywhere, move it to a session API and price it by the hour. ## AI Agent Web Scraping API Cost Comparison Here is the whole argument in one table: what each layer costs per 1,000 units of work at published rates, with the browser rows converted using the 2-minute-per-page rule of thumb (Browserbase's own figure; it is arithmetic, not a vendor claim, and it is the first number to replace with your own measurements). | Layer | Vendor and endpoint | Published rate | Per 1,000 units of work | |---|---|---|---| | Search | Keiro /search/lite (list) | $0.25 per 1,000 | $0.25 | | Search | Keiro /search/lite on plans | 0.1 credit: $0.24 Essential, $0.13 Pro, $0.08 Startup | $0.08 to $0.24 | | Search | Serper | $1.00 per 1,000 | $1.00 | | Search | Brave Search API | $5.00 per 1,000 | $5.00 | | Extraction | Keiro /search/content | 3 credits | $2.40 (Startup) to $7.20 (Essential) | | Extraction | Firecrawl scrape (basic) | 1 credit a page | $5.00 pay-as-you-go; $3.20 on Hobby annual | | Extraction | Firecrawl scrape (JSON format) | 5 credits a page (1 + 4 surcharge) | $25.00 pay-as-you-go | | Browser | Steel Launch | $0.10 per browser hour | $3.33 (at 2 min a page) | | Browser | Browserbase Developer overage | $0.12 per browser hour | $4.00 | | Browser | Browserless Starter overage | $0.0017 per unit, 4 units at 2 min | $6.80 | | Full pipeline | Keiro deep search (/agentic) | 20 credits, one flat meter | $4.44 list | Drawn, cheapest to most expensive per 1,000 page reads (browser rows at the 2-minute rule): COST PER 1,000 PAGE READS, SEPT 2026 (USD) Keiro content, Startup $2.40 Firecrawl Hobby $3.20 Steel Launch $3.33 Browserbase Dev $4.00 Firecrawl PAYG $5.00 Browserless Starter $6.80 Keiro content, Ess. $7.20 The chart compares reading paths, which is the comparison that matters, because reading dominates agent bills. A worked 1-million-pages-a-month scenario, using Keiro's Startup plan credit price ($0.0008 per credit), Steel Launch rates, and 1 MB per page for proxy sizing: 920,000 static reads at 3 credits is $2,208; 80,000 browser reads at 2 minutes is about 2,667 browser hours, $267 at $0.10; proxy bandwidth on those 80,000 pages is roughly 80 GB, $480 at $6 a GB; 50,000 lite searches at 0.1 credit is $4. The month prices out near $2,960, and 75 percent of it is the reading layer. Introduce a 90 percent content-cache hit rate and extraction falls to about $221 and the whole month lands near $940. At scale, the cache line item is the architecture. Where bills actually spike, in the order teams hit them: - **Idle sessions.** Billed wall clock at $0.08 to $0.12 an hour. The fix is inactivity timeouts and release-on-artifact, both of which cost zero. - **Retries.** A browser retry doubles the minutes; an extraction retry doubles the credits. Firecrawl's no-result scrapes are free, but a 403 or 404 response still returns to you and costs the credit (their policy, effective September 4, 2026). Keiro states that retries within a single request do not count as additional searches, which keeps in-request retry loops free. - **Structured extraction surcharges.** Firecrawl's JSON, Question and Highlight formats add 4 credits a page on top of the base credit. That is a 5x jump per page, and it is the single most common "why did this invoice triple" line item in the category. - **Proxy bandwidth.** $6 to $10 a GB on Steel, 6 units per MB residential on Browserless. Blocking images and fonts is a cost feature. - **CAPTCHA solves.** $1 to $3 per 1,000 on Steel's paid tiers. Cheap per solve, expensive when a retry loop triggers a thousand of them. - **Search API overage.** Browserbase includes 1,000 search calls on paid plans and bills $7 per 1,000 after, which is 28 times Keiro's lite rate and the clearest reminder that search belongs to the search layer. One more row deserves its own paragraph. Keiro's deep search bills 20 credits a request and lists at $4.44 per 1,000, and the notable thing is the meter: one flat price covering live-web resolution, full-page reads, proof scoring and re-ranking. Every other row in the table bills one stage and hands the rest to you. When you model an agent's cost, model the pipeline, not the request; the flat meter exists because vendors know most buyers underestimate the stages. ## Cheapest Reliable Scraping API for AI Agents Cheapest sticker and cheapest successful read are different titles, and the second one is the only one your finance team should care about. A $0.25-per-1,000 search that returns empty results on a third of hard queries is more expensive than a $1.00 alternative that hits. A $3.20-per-1,000 scraper that bills the credit on a 403 is not $3.20 on hostile domains; it is $3.20 plus every failed attempt, plus the browser escalation you needed anyway. Price the pipeline in successful documents and the ranking of vendors starts to stabilize. The failure modes differ by layer, so the responses have to differ too. This table is the operational core of the post: | Failure | Layer | Signal | First response | Never | |---|---|---|---|---| | Rate limit | search | HTTP 429 | backoff with jitter; move bulk to a batch path | a tight retry loop | | Empty result set | search | 200 with zero hits | reformulate the query, switch index or layer | resending the same string | | Dead URL | extraction | 404 or NXDOMAIN | log, mark, move on | retrying a known-dead URL | | Bot wall | extraction | 403 with a challenge body | one escalation to the browser layer | hammering the same endpoint | | JS shell | extraction | under ~500 chars of markdown, no main text | one escalation to the browser | parsing the empty document | | Thin or wrong page | browser | rendered text mismatch vs expected selector | widen selector, checkpoint, retry once fresh | blind retries in the same session | | Session cap hit | browser | Steel 15 min on Launch, 1 h on Scale | checkpoint artifacts, resume in a fresh session | one giant never-ending session | | Idle burn | browser | wall clock vs work ratio | inactivity timeout, release on artifact | holding a session across LLM turns | Retry policy per layer, compressed to rules: - **Search.** Retry only on 429 and 5xx, with jittered exponential backoff, and cap at three attempts. A 200 with zero hits is not a retry case; it is a reformulation case. Reformulate (drop a term, add a domain hint, switch from keyword to semantic) before spending another query. - **Extraction.** One escalation to the browser on a thin body, then stop. Add a per-domain circuit breaker: after five consecutive failures, park the domain for an hour and route its URLs straight to the browser layer or a human queue. Retries are cheap until they are metered, and they are always metered somewhere. - **Browser.** No blind retries, ever. Every browser retry re-pays wall clock, proxy bandwidth, and possibly CAPTCHA solves. Checkpoint after each successful step, and on failure resume from the checkpoint in a fresh session rather than replaying the whole flow. Firecrawl's pricing page is unusually explicit about this class of problem (error pages still cost a credit; no-result scrapes do not), and the same logic applies to every vendor's meter whether they publish it or not. Observability is what makes the reliability affordable, because you cannot fix what you cannot attribute. Log per request: the layer used, the URL, the status, the latency, the meter spent (credits or browser minutes), the content hash of what came back, and the escalation path. Retention windows decide whether that log is usable: Keiro keeps logs 7 days on Essential, 30 on Pro and 90 on Startup; Browserbase keeps 7 days on Free and Developer; Browserless gives 24 hours to 7 days of log search by plan. If your debugging loop is "a customer reports a failure next week, go replay it," then 7-day retention is a real cost of the cheap plan, not a footnote. Four metrics cover the whole stack, and all four are cheap to compute from that log: - **Escalation rate.** The share of reads that end in the browser layer. Low single digits on a general corpus is healthy; a climb means the extraction adapter or the URL classifier is rotting. - **Cache hit rate.** The share of reads served without a metered request. This number is your extraction discount. - **Cost per successful document.** Total metered spend divided by documents that actually reached the model. It is the only cost number that survives contact with reality, because it absorbs retries, escalations and dead URLs automatically. - **p95 per layer.** Search and extraction should sit far apart here (a browser page takes minutes; an extraction read takes seconds). If p95 on extraction approaches p95 on the browser layer, your extraction vendor is degrading, and the router will hide it from you until the invoice arrives. The cheapest reliable stack, in one sentence: extraction defaults, one escalation, circuit breakers per domain, sessions that die when idle, and a cost-per-successful-document metric on the wall. ## Best Web Scraping APIs for RAG and AI Agent Training Data RAG ingestion and training-data collection are the two workloads where extraction quality is the product rather than a step. What both need from the reading layer is the same list: markdown that keeps heading structure, boilerplate stripped, tables preserved (financial and pricing pages live in their tables), and provenance stored with every document: source URL, fetch time, and a content hash. The hash is the hinge of the whole caching design, so get it right early. The cache design that holds up: - **Two keys.** A normalized URL key (lowercase host, sorted query params, utm and tracking params stripped, trailing slash and fragment dropped) for lookup, and a content hash of the markdown for change detection. The URL key answers "have I seen this page?", the hash answers "is what I stored still what the page says?" - **TTL by vertical, not globally.** Prices move in minutes, news in hours, documentation in days, regulatory filings in years. A single global TTL either wastes metered reads on stable pages or serves stale prices to a shopping agent. Store TTL with the document, per source category. - **Serve stale, revalidate in the background.** A stale-while-revalidate pattern keeps the agent responsive and refreshes the cache as a background job, so the next run of the same pipeline is the cheap one. - **Use the origin when it offers it.** ETags and If-Modified-Since turn a full read into a free 304 where the site supports conditional requests. Fewer metered credits for the same freshness. - **Cache the read, never the answer.** The same page serves a hundred questions; an answer serves one. Answer caches only help identical inputs, which agents rarely produce. The arithmetic makes the case better than the design does. On Keiro, /search/content bills 3 credits a read: $7.20 per 1,000 on Essential, $4.00 on Pro, $2.40 on Startup. A corpus of 10,000 page reads a month at a 90 percent hit rate bills 1,000 metered reads: $7.20 on Essential instead of $72, $0.72 per 1,000 effective instead of $7.20. Push the hit rate to 95 percent with per-vertical TTLs and the effective rate halves again. The cheapest extraction API is the one you stop calling. Deduplication is the other quiet saver, and it matters more for training-data work than for RAG. The same article syndicated across a dozen domains should be fetched, hashed and stored once; body-hash matching (normalize whitespace, strip boilerplate, hash the main text) collapses the copies before you pay for anything but the first. Canonical URL resolution catches the rest. On a crawl-heavy corpus, dedup routinely removes a meaningful fraction of would-be purchases before they happen; measure it on your corpus rather than trusting ours. For training data at volume, throughput economics take over from latency: Firecrawl's crawl endpoint bills the same 1 credit a page for bulk jobs, Keiro's /search/content works for fresh corpus slices on one balance with search, and the batch path matters (Keiro's batch endpoint bills 1 credit a query with an unlimited allowance on the Startup plan). Whatever you run, respect robots.txt and site terms; compliance is not a vendor feature you can buy retroactively, and several of the vendors in this post publish guidance or compliance modes precisely because their enterprise customers ask. One quality bar worth automating: a small harness that samples reads weekly and checks heading preservation, link retention, table integrity, and boilerplate contamination against a hand-labeled set of your own pages. Vendors change parsers without notice. The harness is an afternoon of work and it catches the regression before your embeddings do. ## Scraping API for Real Time AI Research Agents Real-time research agents have one extra constraint: the answer has to be fresh enough to be true. That constraint reshapes the cache policy, not the architecture. The stack stays search, read, cite; what changes is when the cache is allowed to answer. Set freshness by vertical, and make it explicit in code rather than a global setting: - **Prices, availability, scores.** Minutes. Bypass the cache entirely for user-facing "right now" questions; a cached price is a wrong price. - **News and developments.** Minutes to hours. Search indexes differ here (this is where freshness-focused indexes earn their premium), so record each vendor's freshness behavior rather than assuming it. - **Docs, evergreen reference, profiles.** A day to a week. This is the bulk of research traffic and it caches almost perfectly. - **Filings, archives, historical data.** Months. Fetch once, cache forever, revalidate on hash drift. Three staleness signals are cheap to log and worth logging on every read: the content hash drift between fetches, any Last-Modified or sitemap lastmod the origin exposes, and the cache age at serve time. After a month of traffic, those three fields tell you the real TTL per vertical, which is always different from the one you guessed. In the hot path, keep the browser out of it. A real-time research loop should be search, extraction, answer; the browser tier joins only when an escalation fires, and it joins with a short timeout because the user is waiting. Anything that must run long (a portal login, a slow dashboard) belongs in a background flow whose artifacts land in the same cache the hot path reads from. When per-question quality matters more than per-question cost, the same architecture exists as a managed tier: hand the whole loop to a deep research endpoint and let it run the layers for you. Keiro's deep search bills 20 credits a request and lists at $4.44 per 1,000 requests, with one flat meter covering live-web resolution, full-page reads, proof scoring and re-ranking. On OpenAI's SimpleQA (1,000 seeded questions, seed 42, strict grading where near-misses score as wrong), the pipeline returned the gold answer for 95.3 percent of questions (953 of 1,000) with retrieval only and no reader model. The full-page read layer is worth 15.3 of those 95.3 points, which is the same snippet-versus-page gap from earlier in this post showing up as benchmark points. Published third-party runs collected in the same leaderboard put Tavily at 93.3 percent (GPT-4.1 answering, official classifier, self-published), Perplexity at 85.9 and Brave at 76.1. On the FinanceBench financial split, Keiro scores 78 percent, which is the vertical where thin snippets and shallow ranking fall apart hardest. The build-versus-buy line for the research loop, stated as a decision rule rather than a slogan: build the router when you need control of the meters (custom session logic, your own cache, per-tenant identity), buy the flat meter when you need answers more than you need architecture. The two are not rivals; the managed tier is what the router escalates to when a question's value justifies 20 credits instead of 3.1. Latency is the last honest unknown in real-time work. Most vendors in this space do not publish a p50, and the ones who do publish it under their own conditions. Measure your own: per-layer p50 and p95 from the request log, on your queries, against your target domains. The only defensible latency claim is the one your log produced this week. ## Recommended Scraping API for Startup AI Agents The picks, by situation, with the cons attached: - **The default startup stack.** Keiro for discovery and reading on one credit balance (1,250 free credits a month, no card, which is 12,500 lite searches or about 400 content reads a month to prototype with), plus Steel or Browserbase for the browser tail. Pros: cheapest published rates at both Keiro layers, one key, one invoice line. Cons: no browser API from Keiro, so pair a specialist; the free tier caps at 30 requests a minute, which is fine for development and not for production. - **Tightest possible budget.** Keiro's free tier plus Browserless' Prototyping plan at $35 a month. The browser tier stays under 20,000 units, and everything else runs on free credits. - **Compliance-sensitive workloads.** Browserless (SOC 2 Type II, GDPR and HIPAA-adjacent attestations) or Steel's self-hosted core when the data cannot leave your infrastructure at all. - **One vendor for everything.** Firecrawl, if the convenience of one balance across search, scrape and Interact outweighs paying about $0.60 a browser hour at pay-as-you-go rates, six times Steel's published hour. Now the full pipeline walkthrough, with numbers you can reuse as a template. Scenario: a four-person startup ships a company-research agent that answers 2,000 questions a month about 400 tracked companies. Assume 2.2 pages read per question after dedup, 90 percent of pages static, and a 2-minute browser page on the hard tail. | Step | Layer | Volume | Meter | Cost | |---|---|---|---|---| | Discover sources | /search/lite | 2,000 queries | 0.1 credit each = 200 credits | $0.16 | | Read static pages | /search/content | 3,960 pages | 3 credits each = 11,880 credits | $9.50 | | Read hard tail | Steel Launch | 440 pages, 2 min each | about 14.7 browser hours | $1.47 | | Proxy bandwidth | Steel residential | about 0.44 GB | $6 to $10 a GB | $2.64 to $4.40 | | Cited answers | /answer | 500 answers | 5 credits each = 2,500 credits | $2.00 | | Total | | | about 14,580 credits | about $16 | At Startup-plan credit prices the whole month reads about $16 of metered usage, comfortably inside the $100 plan whose 125,000 credits also cover every other workload the team runs, with the browser tail under $4 of it. The same pipeline runs on the free tier's 1,250 credits at a tenth of the volume, which is the honest evaluation path: prototype the router on free credits, then pay for the plan the usage curve demands. The meter reads about $16; the plan minimum is the biggest line item, and the router is the reason it stays that way. The boring takeaway. Three contracts, one router, one cache: search to find, extraction to read, a browser for the pages that fight back, and a content hash in front of the reading layer so the second run of anything costs a tenth of the first. Buy discovery by the query, reading by the page, and presence by the hour. Track escalation rate, cache hit rate, and cost per successful document, and the architecture will tell you which vendor to reprice next quarter. Every price in this post was checked on September 23, 2026, and every vendor here reprices quarterly, including the one writing the post. Check them again before you sign anything.