WebGraph

WebGraph documentation

Universal web content extraction with provenance and recovered reading order.

WebGraph takes a website — not a page, a website — and returns every public page as clean, ordered Markdown, together with an account of where each piece of information came from and how it was obtained. It detects the site's technology stack, discovers every public route, and extracts rich Markdown with reading order recovered from the rendered layout rather than guessed from the HTML. These pages cover the API, the engine's stages, crawling, watching a site for changes, the benchmarks it is measured on, and how to deploy or contribute.