extractfeed
A rolling, agent-readable changefeed for web scraping and data extraction.

Chronological agents.md

Apify

Apify benchmarks four web unblockers across 384 challenging URLs

Apify published a benchmark comparing four web unblockers against a deliberately difficult set of 384 URLs.

Apify outlines six methods for downloading Instagram images in 2026

Apify published a guide covering six techniques for downloading Instagram images, ranging from a simple browser trick to a two-Actor pipeline for full profile extraction.

Apify streamlines Actor creation with one-click Git repo setup

Apify now lets users select a Git provider during Actor setup, automatically creating a repository, pushing template code, and configuring builds on every push.

Apify launches Apartments.com scraper for rental market analysis

Apify released a tool to scrape Apartments.com listings at scale and integrate them with ChatGPT for interactive market reports.

Apify publishes comparison guide pitting Firecrawl against its own ecosystem

Apify's blog compares Firecrawl's unified AI-driven scraping with Apify's flexible ecosystem, outlining the strengths of each.

Apify compares its e-commerce scraping platform with Bright Data's site-specific APIs

Apify published a blog post comparing its full-stack scraping platform to Bright Data's site-specific APIs for e-commerce data extraction.

Builder spotlight: Goldmine automated outreach and won on Apify

Apify profiles a developer who built a LinkedIn scraper for his team and later became a top-rated Apify developer, winning an EMEA prize in the Apify $1 Million Challenge.

Apify compares Oxylabs and Bright Data for web scraping

Apify published a comparison of Oxylabs and Bright Data, focusing on their proxy services and recent web scraping investments.

Apify Builder Spotlight: Non-developer wins LATAM prize with niche Actor built via prompt and no code

Apify profiles Gui Stetelle, a salesperson who used a prompt, Apify docs, and Gemini to build a niche Actor and win the $2,000 LATAM prize in the Apify $1 Million Challenge.

Apify Builder Spotlight: John Cole wins $20,000 in Apify Challenge without a CS degree

Apify profiles John Cole, a self-taught developer who placed second in the Apify $1 Million Challenge with nearly 100 Actors and 15,000 users.

Apify blog explains how to scrape Clutch.co and interpret its paid rankings

Apify published a guide on collecting Clutch.co data despite its blocking measures and explains how paid placements skew agency rankings.

Apify sets September 30 deadline for rental-to-pay-per-event Actor migration

Apify is sunsetting its rental pricing model for Actors, requiring developers to migrate to a pay-per-event model by September 30.

Apify shows how to build an AI research assistant in Slack using its MCP server and n8n

Apify published a tutorial on integrating its MCP server with n8n to create a Slack-based AI research analyst that answers team questions with live web data.

Apify migrates MCP server to new stateless specification

Apify has migrated its MCP server to the July 2026 stateless spec revision, detailing the changes and implications for hosted servers and existing clients.

Apify blog shows how to build a Dun & Bradstreet data API without an enterprise contract

Apify published a guide on using its platform to create a custom API for Dun & Bradstreet company data, targeting users who want to bypass expensive enterprise agreements.

Bright Data Blog

Bright Data Blog discusses robot training data from public web video

Bright Data's blog explores the use of publicly available web video as a source of training data for robotic AI systems.

Bright Data integrates with Paperclip for AI-driven data extraction

Bright Data announces a partnership or integration with Paperclip, an AI tool, as detailed in a blog post on their site.

Bright Data Defends Its Network Against Misuse Claims

Bright Data publishes a blog post arguing that its proxy network cannot be used in the ways critics allege.

Bright Data explains Context as a Service for AI data retrieval

Bright Data published a blog post defining Context as a Service, a concept where external context is provided to AI models via structured data feeds.

Bright Data compares its Cursor integration with default coding agent

Bright Data published a blog post comparing its Cursor integration against the default coding agent for web data tasks.

browserbase/stagehand releases

Stagehand Python adds WebMCP tool support inside iframes

Stagehand Python's latest dev release enables discovery and invocation of WebMCP tools registered inside iframes by routing calls through the main CDP session and all adopted OOPIF sessions.

Stagehand SDK exposes Browserbase Search and Fetch across TypeScript, Python, and Go

Stagehand SDK version 4.1.0a0.dev1494 adds Browserbase Search and Fetch APIs to its TypeScript and Python facades, with equivalent Go support via the existing HTTP transport.

Stagehand Python 4.0.3a0.dev1483 adds stagehand_facade tool surface for eval benchmarking

A dev release of Stagehand Python introduces a facade tool surface that allows evals to benchmark the exact byte-identical interface shipped to Claude Code, Codex, and Pi integrations.

Browserless

Browserless publishes enterprise Docker deployment guide for self-hosting

Browserless released a guide detailing how to deploy its browser automation service in an enterprise Docker environment.

Browserless introduces Skill Bucket for agent-based browser automation

Browserless announces a new Skill Bucket feature that allows AI agents to access and use predefined browser automation skills.

Browserless blog post examines the pitfalls of persisting browser profiles

Browserless published a blog post discussing the challenges and drawbacks of persisting browser profiles in automated environments.

Browserless publishes guide on browser infrastructure for computer use agents

Browserless has published a blog post discussing the infrastructure requirements for running browser-based agents that interact with web interfaces on behalf of users.

Browserless publishes guide on bypassing Datadome anti-bot protection

Browserless released a blog post detailing techniques to circumvent Datadome's anti-bot system.

Browserless introduces the Browser Automation Protocol

Browserless has announced a new protocol for browser automation.

Browserless launches Agent for AI-driven browser automation

Browserless has introduced a new product called Browserless Agent, designed to enable AI agents to control browser sessions.

Cloudflare Ai Bots

Cloudflare launches BotBase for Operators with submission management and behavior modeling

Cloudflare has introduced a dedicated dashboard section where bot operators can submit, track, and edit their bot declarations, including a behavior model for content usage.

Cloudflare Wallets launch to give AI agents programmable payments and identity

Cloudflare announced Cloudflare Wallets, a programmable wallet that enables AI agents to make autonomous payments and verify identity on the web using the x402 protocol.

Cloudflare Bot Management

Cloudflare launches Adaptive Intelligence to shift bot economics

Cloudflare introduces Adaptive Intelligence, a new engine that autonomously learns from live traffic meta-signals to counter bot operators' economic advantages.

Cloudflare launches BotBase for Operators to streamline bot directory submissions

Cloudflare has introduced a new dashboard interface for bot operators to manage their submissions to the company's bot directory, including status tracking, editing, and a behavior model for declaring content usage.

Cloudflare launches Bot Preference Sync to align robots.txt with AI bot policies

Cloudflare introduces Bot Preference Sync, a feature that automatically synchronizes a site's robots.txt file with its AI bot management policies for search, agent, and training bots.

Cloudflare shifts bot mitigation from risk scoring to continuous trust evaluation

Cloudflare announced a new approach to bot mitigation that moves from point-in-time risk assessment to continuous trust evaluation, introducing systems BotBase and Precursor to assess good and bad behaviors from bots and agents.

Crawlbase

Crawlbase shows how to build a web scraping pipeline with Zapier using async callbacks

Crawlbase published a guide explaining how to split web retrieval from Zapier's execution window by dispatching async crawls with a callback and catching the finished page in a second Zap.

Crawlbase introduces Web Bot Auth as an identity layer for bots, coinciding with pay-per-crawl pricing

Crawlbase ships Web Bot Auth, a specification that treats bot identity as a verifiable signature rather than a self-declared claim, on the same day it launches pay-per-crawl pricing.

Crawlbase publishes technical guide on scaling headless browser fleets to 10,000 concurrent sessions

Crawlbase's blog post details the architecture and capacity planning required to run 10,000 concurrent Playwright sessions across roughly 100 nodes.

Crawlbase details the infrastructure behind 8,000 CAPTCHAs per second

Crawlbase published a blog post explaining the throughput math, Go-based control plane, and scaling challenges required to solve 8,000 CAPTCHAs per second.

Study finds only 34.5% of free proxies work, thousands tamper with traffic

Crawlbase reports on two peer-reviewed studies that tested 640,600 free proxies, finding that just over a third were functional and many altered traffic.

Vacation rental intelligence platform scales to 1 billion monthly crawl requests with Crawlbase Enterprise Crawler

Crawlbase published a case study detailing how a vacation rental intelligence platform processed 5.52 billion requests in six months with 99.96% success using its Enterprise Crawler.

Crawlbase argues AI agent failures are infrastructure failures, not code problems

Crawlbase publishes a blog post claiming that most AI agent failures stem from infrastructure issues like Markdown normalization, retrieval circuit breakers, and storage-backed memory.

DataDome

DataDome launches hosted MCP server for AI agent access to trend reports

DataDome introduced a hosted MCP server that allows AI agents to query its Trend Reports data using natural language.

DataDome report finds AI agent requests hit 7.9 billion, scalper bots outnumber shoppers 10 to 1

DataDome recorded 7.9 billion AI agent requests in early 2026 and published a report highlighting widespread spoofing and visibility gaps in bot detection.

DataDome report finds European publishers face disproportionate AI bot scraping

DataDome recorded 7.9 billion AI agent requests in early 2026 and reports widespread spoofing and visibility gaps leaving organizations exposed.

DataDome CEO and Forrester Discuss Bot and Agent Trust Management Landscape

DataDome's CEO and Forrester analysts discuss the findings of The Forrester Wave: Bot and Agent Trust Management Software, Q2 2026.

DataDome report finds 7.9B AI agent requests and widespread spoofing in early 2026

DataDome published a report detailing 7.9 billion AI agent requests recorded in early 2026, highlighting spoofing and visibility gaps that leave retail systems exposed.

DataDome records 7.9B AI agent requests, warns of widespread spoofing and visibility gaps

DataDome published a report documenting 7.9 billion AI agent requests in early 2026, highlighting spoofing and detection blind spots.

DataDome publishes guide on blocking AI bots and scrapers

DataDome released a guide covering methods to block AI bots, from robots.txt to advanced bot management.

DataDome updates Account Protect with clearer visibility and stronger signals

DataDome has released updates to its Account Protect product, adding deeper visibility, broader detection across password-reset flows, and API control over protection status.

DataDome publishes guide on AI agent access control, arguing traditional IAM falls short

DataDome released a blog post explaining how AI agents require dedicated identities, scoped permissions, and real-time oversight beyond traditional identity and access management.

DataDome report finds 7.9B AI agent requests and widespread spoofing in early 2026

DataDome published a report documenting 7.9 billion AI agent requests in early 2026, highlighting spoofing and visibility gaps that leave organizations exposed.

Firecrawl

Firecrawl releases official ChatGPT plugin

Firecrawl has launched an official ChatGPT plugin, enabling users to extract web data directly through conversational AI.

Firecrawl introduces Anydoc and PDF Inspector for document extraction

Firecrawl announced two new tools, Anydoc and PDF Inspector, aimed at improving document and PDF data extraction.

Firecrawl introduces Agentic OCR for structured data extraction from images

Firecrawl has released a new OCR feature that uses AI agents to extract structured data from images and documents.

Firecrawl launches Research Index for life sciences

Firecrawl announced the launch of a Research Index tailored for the life sciences domain.

Firecrawl releases Convex component for web scraping integration

Firecrawl announced a new Convex component that allows developers to integrate web scraping capabilities directly into their Convex applications.

Firecrawl announces integration with Eden AI

Firecrawl has published a blog post announcing an integration with Eden AI.

Firecrawl publishes case study on 11x's use of its web scraping API for prospect research

Firecrawl's blog details how sales automation startup 11x uses its web scraping API to automate prospect research.

Firecrawl launches Developer Index for web scraping performance metrics

Firecrawl announced the launch of a Developer Index to provide benchmarks and performance data for web scraping tools.

Hacker News

Jactl brings virtual threads to Java 8 without Project Loom

The Jactl scripting language introduces virtual thread support for Java 8, bypassing the need for Project Loom.

Google's GOTO link changes dramatically increase scraping difficulty

A study published on Growtika claims that Google has made web scraping 10 times harder in less than a year through changes to its GOTO link structure.

Google Details AI-Assisted Rewrites of C/C++ Dependencies to Rust for Memory Safety

Google published a blog post on its bug hunters portal describing how it uses AI to assist in rewriting C and C++ dependencies into Rust at scale.

Aaron Swartz prosecution contrasted with Meta's scraping impunity

A blog post argues that Aaron Swartz was harshly prosecuted for scraping academic articles while Meta scrapes data at scale with little legal consequence.

PageSieve browser extension launches for Firefox-based web scraping

A new open-source browser extension called PageSieve enables users to scrape data from websites directly within Firefox.

PyScrappy brings self-healing selectors and an MCP server to web scraping

A new open-source Python library, PyScrappy, offers self-healing CSS/XPath selectors and an MCP server for resilient web scraping.

Anthropic confirms watermarking concerns in statement

Anthropic issued a statement acknowledging all concerns about watermarking.

LuaCAD brings parametric CAD scripting to Lua with operator overloading

A new tool called LuaCAD lets users model solids in Lua instead of the OpenSCAD language, offering a CLI and desktop app with preview and editor.

Kameleo blog post demonstrates watching browser fingerprinting calls via Chromium's Blink layer

A blog post on Kameleo's site describes a technique for observing browser fingerprinting calls by instrumenting Chromium's Blink rendering layer.

Draco: A self-hostable, single-binary Firecrawl alternative built in Rust

A developer released Draco, an open-source, single-binary web scraping tool in Rust designed as a self-hostable alternative to paid APIs like Firecrawl and Browserbase.

microsoft/playwright releases

Playwright 1.63.0 adds test locks for safe concurrent access to shared resources

Playwright 1.63.0 introduces named test locks that prevent concurrent execution of tests sharing the same lock name across files, workers, and projects.

Oxylabs

Oxylabs partners with Bellingcat via Project4beta to support open source investigations

Oxylabs announces a partnership with the investigative journalism collective Bellingcat through its Project4beta initiative.

Oxylabs announces collaboration with Trackcorona

Oxylabs has partnered with Trackcorona, a platform for tracking COVID-19 data.

Oxylabs partners with University of Michigan

Oxylabs announced a partnership with the University of Michigan, as detailed on their blog.

Oxylabs publishes AI predictions for 2026

Oxylabs has released a blog post outlining its predictions for AI developments in 2026.

Oxylabs receives investment, signaling continued growth in proxy and data extraction infrastructure

Oxylabs announces it has received an investment, though details on the investor and amount are not disclosed in the excerpt.

puppeteer/puppeteer releases

Puppeteer 25.10.0 adds video-stream screen recording and Firefox 155.0 support

Puppeteer released version 25.10.0 of puppeteer-core, introducing a video-stream-based screen recording feature via page.record() and rolling to Firefox 155.0.

Scrapfly

Scrapfly ranks top JavaScript web scraping libraries for 2026

Scrapfly published a guide ranking eight JavaScript and Node.js libraries for web scraping, covering HTTP clients, parsers, headless browsers, and the Crawlee framework.

Scrapfly reviews nine Scrapy extensions and middlewares for 2026

Scrapfly tested nine Scrapy extensions and middlewares against Scrapy 2.18, covering rendering, TLS fingerprints, proxies, extraction, shared queues, and deployment.

Scrapfly publishes 2026 diagnostic guide for blocked Scrapy spiders

Scrapfly released a walkthrough covering nine mechanisms that can block a Scrapy spider, from IP reputation to CAPTCHA, with evidence and mitigation steps for each.

Scrapfly publishes 2026 guide to e-commerce scraping tools

Scrapfly's blog post surveys nine e-commerce scraping tools for developers, covering retrieval, extraction, browser automation, and discovery layers.

Scrapfly ranks six open-source YouTube scrapers by job, flags two failures

Scrapfly published a blog post comparing six open-source YouTube scrapers, including a GitHub snapshot from August 11, 2026, and noting two projects that failed in their tests.

Scrapfly publishes guide to scraping Target.com via Redsky API and bypassing PerimeterX

Scrapfly released a blog post detailing how to extract product and pricing data from Target.com using the internal Redsky API while handling store-keyed prices and PerimeterX anti-bot defenses.

Scrapfly publishes tutorial on scraping Skyscanner flight prices with Python

Scrapfly released a blog post showing how to extract flight data from Skyscanner by constructing deep-link URLs and capturing itinerary JSON from the rendered page.

Scrapfly publishes guide on scraping Airbnb listings and prices

Scrapfly released a blog post detailing how to scrape Airbnb search results, listing details, prices, reviews, and availability using Python and its own scraping platform.

Scrapfly ranks five open-source LinkedIn scrapers on GitHub by auth model and ban risk

Scrapfly published a blog post evaluating five open-source LinkedIn scraping repositories on GitHub, ranking them by authentication approach, maintenance status, and real-world blocking risk as of August 2026.

Scrapfly publishes guide on scraping Lowe's product data and bypassing Akamai

Scrapfly released a blog post detailing how to scrape Lowe's product, price, search, and store location data using embedded page state and their maintained Python scraper.

Scrapfly Publishes Guide to Scraping DigiKey Data Past Cloudflare

Scrapfly's blog post details how to scrape DigiKey pricing, stock, and parametric specs while navigating its Cloudflare challenge, and compares this approach to using the official API v4.

Scrapfly ranks six open-source Instagram scrapers with notes on auth and ban risk

Scrapfly published a comparison of six open-source Instagram scrapers for 2026, covering auth models, ban risk, and maintenance status.

Scrapfly compares five MCP servers for web scraping and browser automation

Scrapfly published a blog post comparing five MCP servers by their capabilities in protected-site scraping, browser control, debugging, static fetching, and cross-browser automation.

Scrapfly compares six modern command-line tools as alternatives to cURL and Wget

A blog post from Scrapfly evaluates HTTPie, aria2, and other tools that address specific limitations of cURL and Wget, including a managed fetch tool for blocked requests.

Scrapfly compares 8 Python HTTP clients for web scraping in 2026

Scrapfly published a blog post evaluating eight Python HTTP clients on async support, HTTP/2, HTTP/3, TLS impersonation, and maintenance, with runnable examples.

Scrapfly ranks 7 lead scraping tools for 2026

Scrapfly published a ranked list of seven lead scraping tools covering no-code extensions and production APIs, with honest assessments of each tool's limitations.

Scrapfly publishes diagnostic guide for browser fingerprint testing tools

Scrapfly released a layer-by-layer guide covering fingerprint and bot detection tools, explaining what detectable results mean and how to fix each leak.

Scrapfly publishes tutorial on scraping Google Jobs with Python

Scrapfly released a blog post showing how to scrape Google Jobs listings using Python and its own scraping platform.

Scrapfly publishes guide on scraping Google Play app reviews and metadata with Python

Scrapfly released a tutorial showing how to extract full Google Play app reviews, ratings, and metadata using Python, bypassing the typical few-hundred-review limit of free libraries.

Scrapfly ranks four open-source proxy scrapers still viable in 2026

A blog post from Scrapfly filters the crowded open-source proxy tool landscape down to four actively maintained scrapers and checkers worth using this year.

Scrapfly ranks 7 AI browser agents for production scraping in 2026

Scrapfly published a ranked guide to the best AI browser agents for automation and scraping, evaluating them on production stability and anti-blocking capability rather than demo performance.

Scrapfly publishes guide on scraping Marriott hotel data through Akamai defenses

Scrapfly released a tutorial covering how to extract Marriott hotel prices and availability using Python, including bypassing Akamai bot protection.

Scrapfly publishes guide to scraping Kayak flight data with its SDK

Scrapfly released a blog post walking through the process of scraping Kayak flight search results using its own SDK, covering JavaScript rendering and parsing internal poll JSON.

Scrapfly publishes guide on scraping RS-Online for product data

Scrapfly released a tutorial on extracting pricing, stock, specifications, and datasheet links from RS-Online's North American listings and product pages.

Scrapfly compares Browser Use and Playwright for web scraping

Scrapfly published a blog post comparing Browser Use and Playwright, covering architectural differences, speed and cost tradeoffs, silent failure risks, and a hybrid approach for production scraping.

Scrapfly publishes guide on bypassing AWS WAF Bot Control for web scraping

Scrapfly released a blog post detailing how AWS WAF Bot Control detects scrapers across five layers and how to bypass it using their Scrapfly ASP product.

Scrapfly rounds up top open-source Facebook Marketplace scrapers on GitHub for 2026

Scrapfly published a blog post listing the five best open-source Facebook Marketplace scrapers on GitHub as of 2026, along with repos to avoid.

ScrapingBee

ScrapingBee argues for cascade approach to e-commerce scraping, calling LLM last

A ScrapingBee blog post advocates for a cascade strategy that uses local extraction methods before resorting to a language model for zero-shot e-commerce scraping.

ScrapingBee explains how to decode Google's new /goto redirect URLs

ScrapingBee published a guide on handling Google's /goto redirect URLs and the opaque CAES tokens they carry in search results.

ScrapingBee compares eight top Python web scraping tools for 2026

ScrapingBee published a guide comparing eight Python web scraping tools, covering parsers, browser automation, crawling, proxies, and full-stack scraping services.

ScrapingBee explores how local deep research agents break on web retrieval

ScrapingBee published a blog post examining the failure points in retrieval stages of open-source deep research systems like Local Deep Research, gpt-researcher, and LangChain agents.

ScrapingBee explains CAPTCHA solvers and when to avoid them

ScrapingBee published an article defining CAPTCHA solvers, how they automate challenge responses, and when developers should skip them in scraping workflows.

ScrapingBee explains how CodeWhale gives AI agents live web access

ScrapingBee publishes an article detailing how CodeWhale enables AI agents to access live web data by integrating web scraping tools, MCP servers, or framework tools.

ScrapingBee ranks top proxies for Amazon scraping in 2026

ScrapingBee published a guide evaluating residential, ISP, and mobile proxies for Amazon scraping, noting that Amazon's anti-bot updates can render previously effective proxies obsolete.

ScrapingBee publishes guide on scraping website text for LLM training

ScrapingBee released a tutorial covering how to extract all text from a website for use in LLM training pipelines.

ScrapingBee publishes guide to AI agent web scraping with wigolo and MCP

ScrapingBee released a guide covering how to equip AI agents with web scraping capabilities using wigolo and the Model Context Protocol.

ScrapingBee publishes guide on MCP servers for web scraping, emphasizing control over data

ScrapingBee explains how a Model Context Protocol (MCP) server can improve scraping agent performance by keeping context lean and avoiding page bloat.

ScrapingBee explores how 'Agent Skills' can make AI coding agents more reliable

ScrapingBee published a blog post examining the Agent Skills repository by addyosmani, which aims to improve the reliability of AI coding agents by teaching them structured workflows.

ScrapingBee publishes guide on building Python flight scrapers in 2026

ScrapingBee released a blog post outlining four methods for building a Python flight scraper, recommending a managed scraping API to bypass anti-bot measures.

ScrapingBee publishes tutorial on building a job aggregator scraping pipeline

ScrapingBee released a blog post walking through the full pipeline for a job aggregator, from sourcing and normalizing to deduplication and freshness management.

ScrapingBee launches Agentic Employee Search API for natural-language profile discovery

ScrapingBee has released an Agentic Employee Search API that accepts natural-language queries to find relevant employee profiles, eliminating the need for manual filtering or complex search parameters.

ScrapingBee publishes tutorial on scraping Algolia via its hidden API

ScrapingBee released a guide explaining how to scrape Algolia search results by targeting the search endpoint directly instead of parsing HTML.

ScrapingBee publishes tutorial on building an AutoGPT agent with its API

ScrapingBee released a blog post showing how to use its web scraping API as a data source within AutoGPT's visual Agent Builder.

ScrapingBee publishes guide on scraping Uber Eats data with legal and technical caveats

ScrapingBee released a tutorial covering how to extract public menu items and prices from Uber Eats, noting the platform's JavaScript rendering and terms of service restrictions.

ScrapingBee ranks top sneaker proxies for 2026 drops

ScrapingBee published a guide ranking ISP, residential, and datacenter proxies for sneaker bots, focusing on checkout success during limited drops.

ScrapingBee publishes tutorial on using Playwright in Ruby

ScrapingBee released a guide covering installation and usage of the community-maintained playwright-ruby-client gem for web scraping and testing in Ruby.

ScrapingBee adds geolocation support to classic proxies, covering 40+ countries

ScrapingBee announced that its classic proxy offering now allows users to select a target country from over 40 options via the request builder or API parameter.

ScrapingBee covers Ruflo, a multi-agent orchestration layer for Claude Code and Codex

A blog post on ScrapingBee introduces Ruflo, an open-source meta-harness that coordinates swarms of AI agents by wrapping around Claude Code or Codex.

ScrapingBee publishes blog post on /last30days AI agent skill for multi-platform research

ScrapingBee's blog introduces /last30days, an AI agent skill designed to research topics across multiple internet platforms while navigating access restrictions.

ScrapingBee publishes guide on scraping TCGplayer, highlighting ToS restrictions and JavaScript challenges

ScrapingBee's blog post explains how to scrape TCGplayer while noting the site's Terms of Service discourage scraping and that JavaScript rendering complicates simple requests.

ScrapingBee publishes guide on using CloudProxy for web scraping

ScrapingBee's blog post covers setting up CloudProxy, its limits, and its managed API for lightweight scraping tasks.

ScrapingBee publishes guide on bypassing CAPTCHA with Selenium in Ruby

ScrapingBee released a tutorial explaining how to avoid triggering CAPTCHAs when using Selenium in Ruby by reducing automation signals and using better IPs.

Scrapingdog

Scrapingdog launches MCP Server for AI-powered web scraping

Scrapingdog has introduced an MCP Server that integrates its web scraping and data extraction capabilities directly into AI-powered applications.

SerpApi

SerpApi resolves Google /goto URL redirect rollout

SerpApi announced it has fixed an issue where Google changed result links to use google.com/goto redirects, restoring direct destination URLs in its API output.

SerpApi publishes guide for scraping Zillow listings

SerpApi released a tutorial showing how to scrape Zillow real estate listings using its API with multiple programming languages.

SerpApi countersues Reddit over API access restrictions

SerpApi has filed counterclaims against Reddit in response to Reddit's lawsuit, alleging broken promises of an open internet and unfair API pricing.

SerpApi publishes roundup of best web scraping tools for 2026

SerpApi has published a guide covering popular web scraping tools, from open-source frameworks like Scrapy and Crawlee to commercial scraping platforms and search APIs.

Google introduces /goto redirect URLs in Search, SerpApi works on resolution

Google is rolling out new /goto redirect URLs across Search, and SerpApi is actively resolving affected links as the implementation evolves.

SerpApi publishes guide on scraping Walmart product reviews with its dedicated API

SerpApi released a tutorial showing how to use its Walmart Product Reviews API to extract ratings, review text, feedback counts, and reviewer details in structured JSON and export to CSV.

SerpApi explains HTTP 429 errors and rate-limit best practices

SerpApi published a guide covering the causes of HTTP 429 errors, how to read Retry-After headers, and strategies for retrying requests without escalating blocking.

SerpApi moves to dismiss Google's amended complaint in scraping lawsuit

SerpApi has filed a motion to dismiss Google's amended lawsuit after the court previously dismissed the original complaint in July.

SerpApi launches Markdown output for all search APIs

SerpApi now offers a Markdown output option for all 100+ of its search APIs, delivering the same data as JSON with roughly half the tokens and no additional configuration.

SerpApi surpasses 1.5 million activated user accounts

SerpApi announced it has reached 1.5 million activated user accounts, marking a significant growth milestone for the search engine results API provider.

SerpApi publishes case study on lead generation company scaling outreach with Google Maps API

SerpApi released a case study detailing how a lead generation company used its Google Maps API to scale lead discovery.

Zenrows

Zenrows blog walks through scraping 2026 FIFA World Cup data across three access patterns

Zenrows published a tutorial demonstrating how to scrape 2026 FIFA World Cup data from a JSON endpoint behind JavaScript, an API with per-session JWT, and server-rendered HTML using Python and its own scraping API.

Zenrows publishes practical guide on web data for LLM fine-tuning

Zenrows outlines how to source clean web data for LLM fine-tuning, warning that a single bad extraction can measurably degrade a small seed set.

Zenrows integrates with smolagents to give AI agents production-grade web access

A tutorial shows how to swap smolagents' plain-requests VisitWebpageTool for a Zenrows-backed tool that bypasses bot checks.

Zenrows MCP brings live scraping to Cursor's AI editor

Zenrows released an MCP integration that lets Cursor users scrape JavaScript-rendered and bot-protected sites directly from the editor.

Zyte

Zyte warns AI agents face a hostile web where scraped pages become attack surfaces

Zyte published a blog post arguing that the real web is becoming hostile for AI agents, with scraped pages posing as potential attack surfaces.

Zyte publishes guide on running Playwright at scale with its CDP browser

Zyte released a blog post explaining how to connect Playwright to its CDP browser for scalable web scraping.

Zyte study finds 40% of top sites block AI crawlers via robots.txt

A Zyte analysis of the world's leading websites reveals that nearly four in ten now use robots.txt to restrict AI bots.

Zyte tests Claude Fable 5.1 and GLM-5.3-Flash in a live extraction benchmark

Zyte published a personal benchmark comparing Claude Fable 5.1 and GLM-5.3-Flash on real extraction tasks, revealing that the GLM model matched a model the author had previously encountered.

Zyte analysis finds 75% of top sites use robots.txt, but few name specific crawlers

Zyte published a study showing that three-quarters of the world's top websites publish a robots.txt file, yet most do not name individual crawlers.

Zyte adds CDP support for browser automation on its infrastructure

Zyte has introduced Chrome DevTools Protocol (CDP) support, allowing users to run browser automation scripts on Zyte's managed infrastructure.

Zyte argues EU AI scraping guidelines rely on outdated robots.txt standard

Zyte publishes a blog post arguing that Europe's new generative AI scraping guidelines, which lean on the robots.txt protocol, will harm users and entrench monopolies.

Zyte pitches WebFetch as a drop-in replacement for coding agents' built-in fetch tool

Zyte published a blog post arguing that its WebFetch CLI tool outperforms the default webfetch tool in coding agents for research and coding workflows.

Zyte report reveals retailers as second most aggressive sector in blocking AI crawlers

Zyte published a blog post detailing how major retail marketplaces deploy anti-bot technology to block AI crawlers and automated data extraction.

Zyte research argues web scraping faces pricing barriers, not outright blocking

Zyte's State of Web Access report, discussed in an interview with the researcher, finds that new economic barriers are making web scraping more difficult rather than technical blocks.

Zyte report finds fashion websites among the most heavily defended against scraping

A Zyte blog post examines the anti-bot and blocking measures used by fashion e-commerce sites, finding they are some of the most aggressively protected on the web.

Zyte blog post examines how AI and web scraping turn scattered personal data into security risks

Domagoj Marić explores the intersection of AI, web scraping, and OSINT to show how fragmented personal data is assembled into profiles, scams, and security threats at Extract Summit.

Zyte Audit Reveals Industry-Specific Bot Access Policies

Zyte published a large-scale audit of web access controls showing how different industries enforce different policies toward bots.

Zyte blog profiles case study of $70 AI-coded app replacing $5,000 platform

Zyte published a blog post detailing how developer Fran Muñoz used AI coding and specification-driven development to build a production app that replaced a costly platform.

Chrome's new navigator.cpuPerformance API opens a fresh fingerprinting vector

Zyte reports that Chrome 152 will expose a navigator.cpuPerformance property, giving sites a new way to fingerprint browsers.

Zyte publishes largest ever audit of web access control mechanisms

Zyte released a comprehensive audit of how websites regulate programmatic visits, revealing the current state of web access barriers.

Zyte publishes tutorial on building custom fetch tools for AI agents with Claude Agent SDK

Zyte released a tutorial showing how to create a custom fetch tool using the Claude Agent SDK to help AI agents extract structured data from the web.

Zyte releases scrapy-spidey-sense, a preflight CLI for Scrapy projects

Zyte has open-sourced a command-line tool that performs static analysis on Scrapy projects before a crawl begins, scoring production-readiness and linking findings to fixes.

Zyte blog post explores rendering JavaScript pages with Playwright and Scrapy

Zyte published a guide on using Playwright to render dynamic content within a Scrapy workflow.

Zyte launches 'Modern Scrapy for experienced developers' tutorial series

Zyte published the first part of a new blog series aimed at experienced developers building production-ready Scrapy projects.