Yozh Scraper - Playwright web scraper with MCP
Yozh Scraper is an open-source, self-hosted web scraping engine that renders URLs inside a real Playwright browser environment to return extracted fields, raw HTML, or screenshots. Designed for developers and AI agents, it features native Model Context Protocol (MCP) support for direct integration with AI tools, alongside built-in proxy routing, anti-detection hardening, and pre-built scraping profiles for popular online platforms.
- Real Browser Rendering: Executes JavaScript-heavy pages via Playwright with configurable wait states and device profile selection.
- Structured Data Extraction: Extracts data fields directly into JSON objects using CSS or XPath rules, bypassing manual raw HTML parsing.
- Built-in Proxy Integration: Routes requests through residential, mobile LTE, or datacenter proxy pools with geographic targeting capabilities.
- Stealth Hardening: Applies anti-detection patches by default to minimize bot mitigation triggers and blocks during automated crawling tasks.
- Pre-Built Site Presets: Includes specialized extraction profiles for major platforms like Amazon, Google, and eBay with optional self-healing parsing.
- Authenticated Sessions: Manages persistent sessions and cookies to handle logged-in target sites through programmatic login execution scripts.
- Native MCP Endpoint: Exposes a Model Context Protocol interface directly at an HTTP endpoint for query-based automation workflows.
Built for backend developers, data engineers, and AI builders requiring controllable, self-hosted infrastructure for automated data collection. The project is distributed under the MIT license and runs on local or cloud environments via Docker Compose.