--- title: "Best Search APIs for RAG" category: "comparisons" url: https://keirolabs.cloud/blog/best-search-apis-for-rag --- ## The 60-second version A RAG pipeline is a retrieval pipeline. The search API you pick decides what your chunks look like, what your embedding bill looks like, and what your invoice looks like at 100,000 queries a month. We priced the eight APIs RAG builders actually shortlist, per 1,000 queries, with page text included, as of September 23, 2026: - **Keiro** (our product): /search/lite is $0.25 per 1,000 at list for ranked results, and /search/content is 3 credits for full page markdown, $7.20 per 1,000 on the $30 Essential plan, $4.00 on Pro, $2.40 on Startup. On monthly plans lite bills 0.1 credit per request: $0.24 per 1,000 on Essential, $0.08 per 1,000 on Startup. Free tier: 1,250 credits a month, no card, which is 12,500 lite searches. Deep search is $4.44 per 1,000 and retrieved 95.3% on SimpleQA retrieval-only. - **Tavily**: $8.00 per 1,000 basic search, $16.00 advanced, 1,000 free credits a month, no card. Basic returns snippets. The cleaned page content that RAG actually wants is the 2-credit advanced tier ([Tavily credits doc](https://docs.tavily.com/documentation/api-credits)). - **Exa**: $7.00 per 1,000 searches with token-efficient page contents included. Results above 10 add $1 per 1,000 each, so a 30-result search bills $27 per 1,000 ([Exa pricing docs](https://docs.exa.ai/reference/pricing)). - **You.com**: Web Search API at a flat $5.00 per 1,000 calls for up to 100 results, each with full page content included. Contents API is $1.00 per 1,000 pages. $100 free credit, no card ([you.com](https://you.com/resources/lower-search-api-cost)). - **Brave Search API**: $5.00 per 1,000 flat against a 30-billion-page independent index, $5 in free monthly credits, card required. Up to 5 snippets per result, no page text, and storing results (which is what a vector store does) needs a plan that explicitly grants storage rights ([brave.com/search/api](https://brave.com/search/api/)). - **Serper**: $1.00 per 1,000 Google queries at the entry pack, 2,500 free trial queries, 300 QPS. Raw Google SERP JSON, snippets only, credits expire in 6 months. - **Google Programmable Search (Custom Search JSON API)**: 100 free queries a day, then $5.00 per 1,000. Closed to new customers, scheduled for discontinuation on January 1, 2027 ([Google docs](https://developers.google.com/custom-search/v1/overview)). - **Bing**: retired. Microsoft shut the Bing Search APIs down on August 11, 2025 and points RAG builders at Grounding with Bing Search inside Azure AI Agents ([Microsoft Lifecycle](https://learn.microsoft.com/en-us/lifecycle/announcements/bing-search-api-retirement)). Read the output column before the price column. Snippets at $1.00 per 1,000 that force you to build a fetcher, a cleaner, and a parser are not cheaper than markdown at $4.00 per 1,000 that skips all three. ## Best Search API for RAG The direct answer, then the receipts. For most RAG pipelines in 2026, the short list is three APIs. Keiro's /search/content returns full page markdown with optional inline embeddings in one call, at $2.40 to $7.20 per 1,000 queries depending on plan. You.com's Web Search API returns up to 100 results per call with full page content included, at a flat $5.00 per 1,000. Exa returns token-efficient page contents inside its $7.00 per 1,000 base rate, and it is still the best index for meaning-shaped queries. Tavily belongs on the short list on quality and drops off it on price: the advanced tier that returns cleaned content costs $16.00 per 1,000, double its own basic rate. Brave at $5.00 per 1,000 is the index play, but it hands you snippets, not pages, so a RAG pipeline pays a second meter to read the URLs. Serper at $1.00 per 1,000 is the snippet budget option with the same catch, plus an unlicensed dependency on Google's results. | API | $/1k, page text included | Index | Output for RAG | The catch | |---|---|---|---|---| | Keiro /search/content | $7.20 Essential, $4.00 Pro, $2.40 Startup (3 credits) | Own index | Clean markdown, optional inline embeddings, chunked to a size you set (100-2,000 tokens) | maxResults caps at 5 pages per call | | Keiro /search/lite | $0.25 list; $0.24-$0.08 on plans | Own index | Ranked results, snippets, published dates | Snippets only; pair with /extract (3 credits) for known URLs | | Tavily | $16.00 advanced ($8.00 basic, snippets) | Hybrid crawl + cache + rerank | Cleaned page content, chunks | Highest per-unit price in the field; Nebius acquired it in February 2026 | | Exa | $7.00 (token-efficient contents in base) | Own neural index + live crawl | Results + contents + optional summaries ($1/1k pages) | Results 11+ bill $1/1k each; 30 results = $27/1k | | You.com Web Search | $5.00 flat, up to 100 results | Own systems | Full page content on every result | Research API tiers run $12-$450+/1k | | Brave | $5.00 (snippets only) | Own index, 30B+ pages, 100M updates/day | Up to 5 snippets per result | No page text; storage rights need a special plan; card required | | Serper | $1.00 + your own fetcher | Scraped Google SERP | Raw Google JSON, 1-2 s | No content layer; 2-credit rule past 10 results; platform risk | | Google CSE | $5.00, snippets | Google (existing customers) | URLs + snippets, 100 free/day | Closed to new signups; discontinued January 1, 2027 | | Bing | N/A | N/A | N/A | Retired August 11, 2025 | That table is the whole market. Everything below is the fine print, the math, and the failure modes. ## What Is the Best Search API for RAG Pipelines? A RAG retrieval call has to do four jobs: find the right pages, return their text, return text clean enough to chunk, and do it at a rate limit your traffic survives. Pricing pages are built to make the first job look like the only one. **Job 1: retrieval quality.** This is what benchmarks measure. Keiro's deep search retrieved the gold SimpleQA answer for 953 of 1,000 seeded questions (95.3%), judged with a strict retrieval-only check, no reader model. Tavily self-published 93.3 on the full 4,326-question set with a GPT-4.1 reader. You.com's own harness put its Research API at 92.09. Exa scored 91.9. The full board, with the caveat that harnesses differ, is in the accuracy section below. **Job 2: text, not links.** A snippet is 30 to 60 tokens. A chunk is 300 to 500. If your API returns snippets, your pipeline fetches every URL itself, and that fetcher is where RAG projects go to die: bot blocking, boilerplate, encoding bugs, and a second bill for bandwidth and parsing. **Job 3: clean output.** The difference between "page text" and "chunkable markdown" is nav bars, cookie banners, and repeated headers. Some APIs clean for you. Some hand you raw HTML and wish you luck. **Job 4: rate limits that match your ingest.** Keiro's tiers run 30 requests/minute on the free tier, 60 on Essential, 300 on Pro, 1,000 on Startup. Serper advertises up to 300 QPS. Exa runs 10 QPS on the free tier. If you backfill a corpus of 50,000 documents, do that arithmetic before you commit. > The two-meter trap: a search bill that looks like $1.00 per 1,000 and an extraction bill hiding behind it. Ask what one retrieval with page text costs, all in. One all-in number per vendor, then: Keiro /search/content $2.40-$7.20 per 1,000, You.com $5.00, Exa $7.00 (contents in base), Brave $5.00 plus a separate extractor at roughly $1.00 per 1,000 pages, Tavily advanced $16.00, Serper $1.00 plus your infrastructure and your block rate. ## Best Search API for AI Content Generation RAG Pipeline This query cluster is about generation pipelines: content generation, summarization, and writing tools that ground every claim. The requirements are stricter than chat RAG, because a wrong paragraph gets published. Generation pipelines need three things from search. Fresh pages, because a writing tool that cites a 2023 price in a 2026 article loses the reader. Verbatim text, because a paraphrased snippet cannot be quoted. And URLs worth citing, because the citation is the product. Per vendor, on those terms: - **Keiro**: /search/fast returns published_date per result, so you can filter to the last 30 days before you spend extraction credits. /search/content returns full markdown. /answer (5 credits) returns a synthesized answer with citations when a writing tool wants one sentence instead of five pages. - **Tavily**: the advanced tier returns cleaned content and it is good content. The topic-depth feature suits research-style generation. At $16.00 per 1,000 it needs your content budget to be real. - **Exa**: the neural index is the best tool for "pages like this one" and meaning-shaped queries, which is a generation workflow more often than a fact lookup. Contents are included in the $7.00 base. - **You.com**: full page content on every result at $5.00 per 1,000, and the Research API (from $12.00 per 1,000 at the low end) returns multi-source synthesis. You.com self-publishes 95.17% on SimpleQA for Web Search with Highlights. That is their own number on their own harness. - **Brave**: 100 million page updates a day feed the index, so freshness is real. The output is up to 5 snippets. For a generation pipeline that needs to quote, add an extraction step. - **Serper**: freshest possible Google index, and news verticals. Text comes back as snippets and positions, so the pipeline fetches pages itself. - **Google CSE**: same freshness argument, same snippet limitation, and a January 1, 2027 discontinuation date. Do not start a new pipeline here. The generation-specific failure mode is silent staleness. A snippet that looked right at index time can be months old. Ask every vendor how they expose page date, and treat "no date available" as a red flag in a writing pipeline. ## Which Search APIs Give You the Cleanest Results for RAG Pipelines? Clean means: markdown in, boilerplate out, one result object per page, and no client-side surgery. This is where the vendors differ most, and it is the difference the pricing pages bury. | API | What lands in your chunker | |---|---| | Keiro /search/content | Full page markdown, chunkable as-is; optional pre-split chunks with embeddings (chunkSize 100-2,000 tokens, default 500) | | You.com Web Search | Full page content per result, up to 100 results per call | | Tavily advanced | Cleaned page content, agent-tuned | | Exa | Token-efficient page contents in the search response; summaries $1/1k pages | | Firecrawl Search | Fresh-crawled full markdown per result (2 credits per 10 results on their credit system) | | Brave | Up to 5 snippets, ~30-60 tokens each | | Serper | Google snippets and positions | | Google CSE | Title, link, snippet | The honest ranking for clean output: Keiro and You.com return full page text by design. Tavily's advanced tier is close behind. Exa's contents are deliberately token-efficient, which is good for context windows and sometimes lossy for verbatim quoting. Firecrawl is the freshest but bills per crawl. Brave, Serper, and Google CSE return snippets, which means your chunker never sees a page. > A chunker fed snippets produces one thin chunk per URL. A chunker fed markdown produces 3 to 15 chunks per page with real headings. Same query, ten times the retrieval surface. One more cleanliness angle: JSON shape. Every API here returns JSON. The differences are field names, nesting, and whether scores and dates are included. Keiro returns score and published_date on /search/fast, and the content endpoint returns title, url, full_text, chunks, and score per result. Serper returns Google's own field vocabulary. Standardize inside your pipeline early, because swapping APIs later is a one-file change if your ingestion layer owns the shape. ## Compare Web Search APIs for RAG Systems Pros and cons, vendor by vendor, for RAG work specifically. Prices were checked on the vendor pages linked in each entry. **Tavily** ([tavily.com/pricing](https://www.tavily.com/pricing)) Pros: 1,000 free credits monthly with no card; cleaned content in one call; the API shape most "search for agents" products copied; a real reranking pipeline. Cons: $8.00 per 1,000 basic and $16.00 advanced, the highest agent-ready rate in the field; snippets on the cheap tier; acquired by Nebius in February 2026, which is a factor when you wire a core primitive to a company mid-absorption. **Exa** ([docs.exa.ai/reference/pricing](https://docs.exa.ai/reference/pricing)) Pros: the best neural and similarity retrieval in the category; contents included in the $7.00 base; $20 signup credit plus $10 monthly; pay-as-you-go with no subscription. Cons: keyword-shaped RAG overpays; result depth is expensive at $1 per 1,000 per extra result; contents are token-efficient, which trades some verbatim text for context-window economy. **Brave Search API** ([brave.com/search/api](https://brave.com/search/api/)) Pros: 30+ billion pages, about 100 million updates a day, the largest independent index; $5.00 per 1,000 flat; ZDR available; no dependency on Google. Cons: snippets only, up to 5 per result; storing results in a vector store requires a plan that grants storage rights; $5 monthly credits require a card on file; free tier was eliminated in February 2026. **Serper** ([serper.dev](https://serper.dev); pack detail via [ColdIQ](https://coldiq.com/blog/serper-pricing)) Pros: $1.00 per 1,000 at the entry pack, down to $0.30 at 12.5M; 2,500 free trial queries; up to 300 QPS; raw Google JSON in 1-2 seconds. Cons: no content layer at all; 2 credits for 11-100 results; credits expire in 6 months; the product runs on unlicensed Google scraping, which is the same exposure that put SerpAPI in court. **You.com** ([you.com/pricing](https://you.com/pricing)) Pros: flat $5.00 per 1,000 Web Search calls with full page content on up to 100 results; Contents API $1.00 per 1,000 pages; $100 free credit, no card; SOC 2 certified, ZDR available. Cons: Research API tiers climb from $12 to $450+ per 1,000 (third-party verified July 2026), so pick endpoints carefully; smaller ecosystem than the category leaders. **Google Programmable Search** ([developers.google.com/custom-search/v1/overview](https://developers.google.com/custom-search/v1/overview)) Pros: Google ranking; 100 free queries a day for existing customers; $5.00 per 1,000 after. Cons: closed to new customers; the docs state discontinuation on January 1, 2027; snippet-shaped results; a Programmable Search Engine config to maintain. Building here in 2026 is building on a demolition date. **Bing** ([Microsoft Lifecycle](https://learn.microsoft.com/en-us/lifecycle/announcements/bing-search-api-retirement)) Pros: none for new work; the API retired August 11, 2025. Cons: gone. The successor, Grounding with Bing Search, lives inside Azure AI Agents, is priced through Azure, and returns grounded responses rather than a result list you can chunk. **Keiro** ([keirolabs.cloud/pricing](https://keirolabs.cloud/pricing)) Pros: cheapest agent-ready result in the field ($0.25 per 1,000 lite at list, $0.08-$0.24 on plans); /search/content returns markdown plus optional inline embeddings at 3 credits; one credit balance across search, extract, answers, and deep research; 1,250 free credits monthly with no card. Cons: an own index can miss pages a Google-scale crawler has seen, /search/content caps at 5 pages per call, and one-time credit packs bill lite at 0.5 credit ($2.50-$3.33 per 1,000), so monthly plans are the cheap path. We sell it; check us against the others on your own queries. ## Search API for RAG Agent Cost Comparison Sticker prices first, then the invoice arithmetic that actually predicts your bill. | API | Rate | Unit | What 1 credit buys | Source | |---|---|---|---|---| | Keiro /search/lite | $0.25/1k list; 0.1 credit on plans | Own index | 1 lite search | [keirolabs.cloud/pricing](https://keirolabs.cloud/pricing) | | Keiro /search/content | 3 credits ($7.20 Essential / $4.00 Pro / $2.40 Startup) | Own index | Up to 5 pages, full markdown | [keirolabs.cloud/pricing](https://keirolabs.cloud/pricing) | | Keiro deep search | $4.44/1k | Full pipeline | Live-web resolution, page reads, proof scoring | [Deep search page](https://keirolabs.cloud/Deep-search) | | Tavily | $0.008/credit; basic 1 credit, advanced 2 | Hybrid | 1 basic search = $8.00/1k | [Tavily credits doc](https://docs.tavily.com/documentation/api-credits) | | Exa | $7.00/1k base (10 results) | Own index | +$1/1k per result above 10 | [Exa docs](https://docs.exa.ai/reference/pricing) | | You.com | $5.00/1k Web Search; $1.00/1k pages Contents | Own systems | 1 call, up to 100 results with content | [you.com](https://you.com/resources/lower-search-api-cost) | | Brave | $5.00/1k | Own index | 1 query, up to 5 snippets | [brave.com/search/api](https://brave.com/search/api/) | | Serper | $1.00/1k (packs to $0.30) | Google SERP | 1 query, 10 results | [serper.dev](https://serper.dev), [ColdIQ](https://coldiq.com/blog/serper-pricing) | | Google CSE | $5.00/1k after 100 free/day | Google | 1 query | [Google docs](https://developers.google.com/custom-search/v1/overview) | Now the same pipeline at three volumes. Assumptions: 10,000 / 100,000 / 1,000,000 queries a month, page text included, list or entry-plan rates. | API | 10k queries | 100k queries | 1M queries | |---|---|---|---| | Keiro /search/content, Startup | $24 | $240 | $2,400 | | Keiro /search/content, Pro | $40 | $400 | $4,000 | | You.com Web Search | $50 | $500 | $5,000 | | Keiro /search/content, Essential | $72 | $720 | $7,200 | | Exa (10 results, contents) | $70 | $700 | $7,000 | | Brave + extractor | ~$60 | ~$600 | ~$6,000 | | Tavily advanced | $160 | $1,600 | $16,000 | | Serper + DIY fetching | $10 + infra | $100 + infra | $1,000 + infra | | Keiro deep search | $44.40 | $444 | $4,440 | Three observations the table earns. First, the spread between cheapest and most expensive page-text pipeline is 6.7x at every volume, and it never narrows, because nobody gives enterprise discounts big enough to close a 6x unit-rate gap. Second, the DIY column is not free; at 1M queries the fetcher, proxy pool, parser, and block-rate losses usually cost more than the $4,000 you saved on the meter. Third, Tavily advanced at $16,000 per million is the price of the same cleaned-content product Keiro sells at $2,400 on Startup, which is the entire argument of this post in one row. > Cost per query is a contract with your future traffic. Pick rates you can afford at 10x your current volume, because RAG pipelines grow in steps, not curves. ## Cheapest Search API for 1M+ RAG Queries At a million queries a month, plan structure matters more than sticker price. Keiro on Startup: $100 a month buys 125,000 credits. Lite search bills 0.1 credit on plans, so that is 1,250,000 lite searches, $0.08 per 1,000, $80 per million. Content search bills 3 credits, so 1M content queries need 3M credits, or 24 Startup plans, $2,400, $2.40 per 1,000. The batch endpoint (1 credit per query) and unlimited batch on Startup are built for exactly this shape of backfill. Serper at 1M: the $3,750 pack buys 12.5M queries at $0.30 per 1,000, so $300 per million, if your workload is snippet-shaped and you accept the 6-month expiry and the scraped-Google exposure. You.com at 1M: 1,000,000 Web Search calls at $5.00 is $5,000, with content included per result. On the Contents API the meter flips: $1.00 per 1,000 pages, so a pipeline that already holds URLs pays $1,000 per million pages. Tavily at 1M advanced: 2M credits at the best plan rate ($0.005 on the $500 tier) is $10,000. At pay-as-you-go it is $16,000. Exa at 1M with 30 results each: $27 per 1,000, $27,000. Depth is where Exa's bill compounds; the $7.00 headline assumes 10 results. Brave at 1M: $5,000 in search, plus whatever extraction costs, plus the storage-rights plan if the results land in a vector store. The 1M+ lesson: the API with the lowest snippet rate is not the API with the lowest RAG bill. Compute your cost per retrieved page, not per query. On that metric, at 1M queries, Keiro content on Startup ($2.40) beats everything published, You.com ($5.00) is the best big-name bundle, and snippet APIs only win if your pipeline was already going to fetch pages for other reasons. ## Search API for RAG Free Free tiers decide which APIs a team actually evals. Here is the field, with the card requirement called out, because a card-gated free tier is a demo with paperwork. | API | Free allowance | What it buys | Card? | |---|---|---|---| | Keiro | 1,250 credits/month, recurring | 12,500 /search/lite or 416 /search/content calls per month | No | | Tavily | 1,000 credits/month | ~1,000 basic searches | No | | Exa | $20 signup + $10/month credits | ~2,800 searches to start | No | | You.com | $100 credit, one-time | 20,000 Web Search calls at $5/1k | No | | Brave | $5 monthly credits | ~1,000 requests | Yes | | Serper | 2,500 queries | One-time evaluation batch | No | | Google CSE | 100 queries/day | Existing customers only | No | | Bing | None | Retired | N/A | Keiro's free tier is the largest recurring allowance in the group: 12,500 lite searches every month, indefinitely, at 30 requests/minute, no card. You.com's $100 credit is the biggest one-time budget, and it prices out to 20,000 web searches with full page content, which is enough to run a real eval suite. Tavily's 1,000 credits are the easiest to reason about: one credit, one basic search. Brave's $5 renews but sits behind a card on signup, an anti-fraud measure by their own description. Two practical notes. Free-tier rate limits shape evals more than volumes do; 30 requests/minute means a 10,000-query eval suite takes about 5.5 hours on Keiro's free tier. And free credits on pay-as-you-go systems (Exa, You.com) are dollar balances, so an eval that calls deep research endpoints can drain them in an afternoon. Set a budget alarm before the eval, not after. ## Best Real Time Search API for RAG Pipelines Freshness is a pipeline requirement with a number attached: how old is the newest page you can retrieve, and how do you find out? - **Brave** publishes the strongest recency posture in the group: an index of 30+ billion pages refreshed by over 100 million page updates a day, plus the Web Discovery Project's opt-in human signal. Their API page positions it for exactly this use. - **Serper** queries Google live, which makes its news and "last 24 hours" verticals effectively real-time. You are paying $1.00 per 1,000 for Google's crawl schedule, not your own. - **Firecrawl Search** (outside the eight but relevant here) crawls pages at request time, so content can be minutes old. That freshness costs 2 credits per 10 results. - **Keiro** runs its own index with published_date on results, and the content endpoint has a noCache flag for pipelines that must bypass the cache. For news-shaped RAG, /search/fast is the always-fresh meter at 1 credit. - **Tavily** runs a hybrid of crawl, cache, and rerank; its topic and time filters work, but the vendor does not publish an index-refresh figure. - **Exa** offers livecrawl options on contents, so you can pay to pull a page at query time when the index copy is stale. - **Google CSE** would have been the freshness leader once. It is now a service with a documented end date. The practical pattern for real-time RAG is a two-tier setup: a cheap indexed API for the standing corpus, and a live-crawl path for the queries where staleness is detectable. Price the second tier per 1,000 pages, because it will be 5-20% of traffic and 80% of the freshness complaints. ## Most Accurate Search API for RAG Accuracy numbers, with harnesses disclosed, because a leaderboard without methodology is marketing. | Provider | SimpleQA | Harness | List price/1k | |---|---|---|---| | Keiro deep search | 95.3% (953/1,000, seed 42) | Own harness, retrieval-only judged check, no reader model | $4.44 | | Firecrawl | 94.7% | GPT-5.4 agent, official grader | $10.00 | | Tavily | 93.3% | Self-published, GPT-4.1 reader, full 4,326-question set | $8.00 | | you.com | 92.09% | Their open-source harness, GPT-5.4 synthesis + judge | $5.00 | | Exa | 91.9% | GPT-5.4 agent | $7.00 | | Parallel | 91.0% | GPT-5.4 agent | $1.00 | | Perplexity | 85.92% | GPT-4.1, official classifier | $5.00 | | Google (SERP via Serper) | 82.15% | GPT-4.1 answers | $5.00 | | Brave | 76.05% | GPT-4.1, official classifier | $5.00 | Source: the [Keirolabs deep search report](https://keirolabs.cloud/Deep-search), which publishes the script and scores 953 of 1,000 seeded SimpleQA questions with a deliberately strict standard: no credit for answers the page lets you compute but never states, and misses taken on contested golds. Read the harness column before the score column. Keiro's 95.3 is retrieval-only, which means the raw pipeline surfaces the gold answer with no LLM reader to rescue it. Tavily's 93.3 is a self-published run with a reader model on top of retrieval. You.com's 92.09 is on their own open harness, and You.com separately claims 95.17% for Web Search with Highlights on their own eval. Different questions, different graders, one honest conclusion: the top of this market clusters in the low-to-mid 90s, and the differences inside that band are smaller than the price differences underneath them. For RAG builders the operational lesson is narrower. SimpleQA measures short factual recall. Your retrieval benchmark should measure your domain: 100 of your own questions, your gold answers, and a judged check for whether the gold text appears in what the API returned. Run it on two APIs, not eight, and let the winner take production. ## RAG Chunk Size 500 Tokens Best Practice The cluster of queries around chunk sizing ("rag chunk size 500 tokens", "rag chunk size 300-500 tokens") is really a query about input quality, so here is the version that starts at the API. Chunk size 300-500 tokens with 10-20% overlap is the working default for text retrieval, and 500 happens to be the default chunkSize on Keiro's content endpoint. The reasoning: 300 tokens holds a coherent section of most prose; 500 gives the embedding model enough context to survive a heading boundary; overlap of 50-100 tokens keeps sentences that straddle a cut from falling between two chunks. What the search API hands you decides whether chunking works at all: - A snippet (30-60 tokens) is smaller than one chunk. You get no boundaries to respect and no context to embed. Snippet-fed RAG retrieves one weak chunk per result, which is why snippet pipelines over-fetch at answer time. - A clean markdown page (1,000-4,000 tokens) splits into 2-8 chunks along heading boundaries. Headings are the chunk boundaries you want, because models embed section text better than arbitrary windows. - Boilerplate (nav, footer, repeated headers) pollutes every chunk it touches and wastes embedding budget. Clean markdown in, clean chunks out. Here is a worked ingestion step in TypeScript. It calls Keiro's /search/content (POST /api/v2/search/content, 3 credits, about 3 seconds), splits returned markdown on headings, chunks to roughly 500 tokens with a 75-token overlap, and stores one row per chunk with its source URL. The same function works on any vendor that returns markdown (Exa contents, You.com page content, Firecrawl pages); only the response destructuring changes. ```typescript type ContentResult = { title: string; url: string; full_text: string; // markdown score: number; }; // Rough token estimate: 4 chars per token for English prose. const tokens = (s: string) => Math.ceil(s.length / 4); function chunkMarkdown(md: string, target = 500, overlap = 75): string[] { // Split on markdown headings; fall back to paragraphs. const sections = md.split(/\n(?=#{1,3}\s)/g).flatMap(s => s.length / 4 > target * 1.6 ? s.split(/\n{2,}/) : [s] ); const chunks: string[] = []; let buf: string[] = []; let size = 0; for (const section of sections) { const t = tokens(section); if (size + t > target && buf.length) { chunks.push(buf.join("\n\n")); // Keep the tail as overlap so sentences cut in half survive. const tail = chunks[chunks.length - 1].split(/\s+/).slice(-overlap); buf = [tail.join(" "), section]; size = overlap + t; } else { buf.push(section); size += t; } } if (buf.length) chunks.push(buf.join("\n\n")); return chunks.filter(c => tokens(c) >= 40); // drop stubs } async function ingestRAG(query: string, apiKey: string) { const res = await fetch("https://api.keirolabs.cloud/api/v2/search/content", { method: "POST", headers: { "X-API-Key": apiKey, "Content-Type": "application/json" }, body: JSON.stringify({ query, maxResults: 3, embeddings: { enabled: false } }), }); if (!res.ok) throw new Error(`search/content failed: ${res.status}`); const { results } = (await res.json()) as { results: ContentResult[] }; return results.flatMap(r => chunkMarkdown(r.full_text).map(text => ({ text, source_url: r.url, source_title: r.title, score: r.score, })) ); // Upsert rows into your vector store here, one row per chunk, // every chunk traceable to a URL. } ``` Count what the pipeline did not do: no HTML fetch, no readability pass, no nav stripping, no block-rate retries. At 3 credits per call, a 100,000-query month ingests about 300,000 pages (3 per call) for $400 on Pro. Set embeddings.enabled to true and the endpoint returns pre-split chunks with vectors inline (dimensions 384-1,024), which deletes the embedder from your own stack; you pay for that convenience in chunk control, since the vendor's splitter decides the boundaries. > Snippets give a RAG pipeline 10 facts. Markdown gives it 40 chunks with provenance. The second pipeline is the one that survives contact with hard questions. ## Best Web Scraping APIs for RAG and AI Agent Training Data RAG and training-data collection get answered by different tools, and the queries in this cluster mix them, so here is the split. Query-time RAG wants a search API: a query goes in, ranked page text comes out, the corpus updates itself. Keiro's /search/content, You.com's Web Search, Tavily advanced, and Firecrawl Search all fit this shape. The meter is per query. Corpus building and training data want a scraping pipeline: a URL list goes in, gigabytes of markdown come out, on your schedule, stored indefinitely. Firecrawl's crawl endpoint and Keiro's /extract (3 credits, full markdown per URL) fit here. The meter is per page, and throughput and politeness rules matter more than ranking. Two legal-adjacent lines matter in 2026. First, Brave's API terms: storing results, in part or whole, for training or tuning, requires a plan that explicitly grants storage rights; the general terms do not cover it. If your RAG cache is a vector store, and it is, read that clause before you build on Brave. Second, scraped SERPs: Google sued SerpAPI over exactly this practice, and Google's own Custom Search JSON API is closed to new customers with a discontinuation date of January 1, 2027. Training-data pipelines built on Google's rankings carry an unpriced dependency on Google's mood. Own-index APIs (Keiro, Brave, Exa, You.com) crawl and serve their own indexes, which is the boring, legally calm position. The trade is tail coverage: a page their crawler has not seen does not exist, so corpus-building jobs that need specific URLs belong on an extraction endpoint pointed at your URL list, not on search. ## Best Search API for RAG in United States The US cluster in this query set is large, so a direct answer about what "US" changes: billing, latency, legal exposure, and coverage of US-first content. Billing is USD everywhere in this post; nothing in the group is region-gated on price. Latency from US regions is single-hop for every API here; the field's published latencies run from Brave's "under 130 ms added" LLM-context claim to the 1-3 second band where content-bearing endpoints live (Keiro's content endpoint runs about 3 seconds). Test from your own region; no landing page latency number survives contact with your VPC. The US-specific news is that two of the three US first-party options are gone or going. Microsoft retired the Bing Search APIs on August 11, 2025, with Grounding with Bing Search in Azure AI Agents as the replacement, and Google's Custom Search JSON API is closed to new customers with discontinuation set for January 1, 2027. A US RAG stack built on either is a migration project on a timer. On compliance: You.com is SOC 2 certified with ZDR available, Brave offers ZDR and SOC 2 posture, and Keiro's plans add SSO/SAML at the Startup tier. If "US-based, audited, and purgeable" is the requirement, shortlist on those three documents rather than on price. And one US-market observation from the query data itself: the same questions arrive in Hinglish and other mixed-language queries, which means your RAG evaluation set should not be English-only either. Every API here claims multilingual coverage; only your own queries can verify it. ## Which Search API Is Best for RAG The picks, by workload, with the numbers attached. **Cheapest production RAG with page text:** Keiro /search/content on Startup, $2.40 per 1,000 queries, $2,400 per million, markdown plus optional inline embeddings at 3 credits. On Essential it is $7.20; the Pro tier at $4.00 is the middle of the market's price range for a better output format than most of it offers. **Cheapest big-name bundle:** You.com Web Search, $5.00 per 1,000 flat, full page content on up to 100 results per call, $100 free credit. The strongest straightforward competitor to everything on this list, and the first one to benchmark against Keiro on your own queries. **Meaning-shaped queries:** Exa at $7.00 per 1,000. If your queries are "pages like this" or conceptual rather than keyword-shaped, Exa's neural index returns results keyword APIs cannot, and contents are included in the base. **Regulated or provenance-sensitive products:** Brave at $5.00 per 1,000 for the 30-billion-page independent index, plus a separate extraction step, plus a storage-rights plan if results persist in your vector store. Boring, auditable, and about $6.00 per 1,000 all-in. **Snippet-shaped workloads on a budget:** Serper at $1.00 per 1,000, 300 QPS, freshest Google index in the group. Know what you are building on: unlicensed SERP data, 6-month credit expiry, and a fetcher you own. **Highest measured retrieval accuracy:** Keiro deep search, 95.3% on the SimpleQA retrieval-only run at $4.44 per 1,000, one flat meter for the whole resolve-read-verify pipeline. For agentic RAG where wrong grounding is the failure that matters, it is the number to beat. **The one to avoid for new builds:** Google Programmable Search (closed, ends January 1, 2027) and Bing (retired). Both appear in older tutorials, and both are dead ends with dates attached. ## The boring takeaway A RAG pipeline is a chunking pipeline, and a chunking pipeline is only as good as the text the search API hands it. Snippets make cheap chunks and expensive pipelines. Full page markdown makes slightly dearer queries and a much cheaper system, because the fetcher, cleaner, and parser you never build are the expensive parts. The numbers, one last time: Keiro /search/content $2.40-$7.20 per 1,000 with markdown and optional embeddings; You.com $5.00 with content on every result; Exa $7.00 with contents in the base; Brave $5.00 in snippets plus an extractor; Tavily $8.00 in snippets and $16.00 with content; Serper $1.00 in snippets plus your own stack; deep search at $4.44 with a 95.3% retrieval score. Google CSE is closing and Bing is closed. Run 100 of your own queries against two of these, judge whether the gold text came back, and price the winner at your real volume. Every number in this post links to a vendor page and was checked on September 23, 2026. Prices in this category move quarterly, so check them again before you sign anything, including with us.