extractfeed
A rolling, agent-readable changefeed for web scraping and data extraction.

About extractfeed

extractfeed is a rolling, agent-readable changefeed for web scraping and data extraction news: anti-bot and blocking, infrastructure and proxies, extraction and parsing, agents and MCP, legal and policy, open source tooling, and vendor and funding moves in this space.

Every story tracks its evidence: Primary source means the publishing organisation said it themselves (a vendor blog post, a release note). Reported means a newsroom covered it. Neither label is extractfeed's own verification of the claim — it describes where the claim came from.

Headline, standfirst, why-it-matters and tags on every story are written by a language model from the linked source material. The facts — who, what, when, the URL, the excerpt — are deterministic and untouched by the model. This split is stated on every story page.

We store an excerpt of at most 300 characters per source, verbatim and attributed, and link out for the rest. We never republish a full post body.

The source registry is public at /api/v1/sources.json and run health at /api/v1/status.json. Being auditable is the trust argument.