A heterogeneous link corpus is a tempting input for browser automation. Dobar.info had 571 unique HTTP/HTTPS URLs available in the local corpus, spanning articles, documentation, GitHub, PDFs, news, Reddit, YouTube, X, and research pages. The Cloudflare Browser Run pilot tested whether direct /markdown extraction could be a useful ingestion primitive.
The result is useful precisely because it is limited:
20 requests
1 HTTP 200 success
19 HTTP 429 failuresThis is not a successful corpus-ingestion result.
What the pilot tested
The pilot used a deterministic selection of 20 URLs from Read-only corpus manifest: links.txt“, with concurrency two and no retries. It used direct /markdown only. /crawl was not used, and the report records no /scrape run.
The endpoint roles should therefore be kept separate in the architecture discussion:
/markdownwas the observed single-page normalization path;/crawlwould represent a crawl-oriented operation, but no result was collected in this pilot;/scrapewas not exercised, so its behavior for this corpus is unknown from local evidence.
For 571 unrelated URLs, /crawl is not a natural input shape without additional grouping and policy. A crawl usually assumes a site or navigable scope; this corpus is a set of independent targets. That is an architectural inference, not a measured /crawl failure.
Observed result
Request 001, the Docker article, returned HTTP 200 and 19,928 Markdown bytes. The output contained YAML frontmatter with a title and meta description, headings, links, and article body text. It was classified as good, while completeness remained uncertain because the pilot did not compare it with an independent source.
Requests 002 through 020 returned HTTP 429 with Cloudflare error code 2001 and the message Rate limit exceeded. No retries were attempted, and no additional requests were submitted after the bounded batch. The report classifies these as transient quota/rate-limit failures, not proof that each target site blocked Browser Run.
Metadata is weaker than the body
The successful direct /markdown response did not provide a structured canonical, author, or separate metadata object. The visible YAML frontmatter was observed inside the Markdown output, but the report explicitly marks consistency across domains as unknown because 19 requests were rate-limited.
That means an ingestion layer would need its own metadata normalization and provenance rules. It cannot assume that a title-like frontmatter field is equivalent to a verified canonical URL or author field.
The Docker article test
The Docker article is the only completed content-quality example. Its body was substantial and useful enough to demonstrate that /markdown can produce normalized Markdown for at least one article/blog page. It does not demonstrate complete extraction for the corpus, deterministic repeated output, or successful extraction for documentation, GitHub, PDFs, news, Reddit, YouTube, or X.
The report records no dedicated JavaScript-wait variation, no repeat request, and no fallback implementation. These are unknowns, not failures that have been resolved.
CAPTCHA, bot protection, and content classes
The pilot did not reach content-level CAPTCHA or bot-protection classification. It stopped at the observed rate limit. Therefore CAPTCHA/bot behavior remains unknown, as do login requirements, YouTube transcript availability, Reddit rendering, X media behavior, PDF extraction, paywall handling, and GitHub extraction quality.
The same caution applies to the browser timing: approximately 1,132 ms was observed for the one successful request. It is a measurement of one request, not a corpus estimate.
Architecture decision: still open
The evidence supports one narrow conclusion: Browser Run /markdown is technically capable of producing useful normalized Markdown for at least one article page. It does not establish suitability for a heterogeneous 571-URL corpus. The immediate operational blocker observed by the pilot is quota/rate-limit handling, not extraction quality.
The report leaves Cloudflare-only architecture neither recommended nor rejected and identifies Cloudflare plus a fallback as a reasonable hypothesis. That fallback was not implemented or tested. A final ingestion decision would require a controlled plan for rate limits, retries, per-host policy, endpoint selection, metadata provenance, and unsupported content types.
Observed, inferred, unknown
Observed: 20 requests; one 200; nineteen 429; one 19,928-byte Markdown result; no /crawl; no retries; no fallback.
Inferred: /markdown can be useful for at least one article; /crawl is not a natural direct representation of 571 unrelated URLs; rate-limit handling is the first operational blocker.
Unknown: /scrape behavior, crawl quality, CAPTCHA behavior, PDF/GitHub/Reddit/YouTube/X support, metadata consistency, repeatability, JS-heavy pages, and quota-adjusted behavior.
Self-review
- Every numeric result comes from the pilot report or its result artifact.
- The article does not call the pilot successful corpus ingestion.
- Observed facts are separated from architectural inference and unknowns.
- No Cloudflare token, request credential, or unverified production claim is included.
- The final architecture decision is explicitly left open.
Sources
Cloudflare pilot artifact: REPORT.mdCloudflare pilot artifact: results-summary.jsonCloudflare pilot artifact: results/001.jsonCloudflare pilot artifact: results/001.mdRead-only corpus manifest:links.txt“ (read-only corpus input referenced by the report)