Skip to content

Ultra-fast, self-healing, undetectable web scraping.

Rust does the fetching, parsing and crawling. You write Python, with the selectors you already know.

pip install netweir

Open source under the AGPL-3.0. Wheels for Linux, macOS and Windows, Python 3.10 and up.

Ten lines, nothing to compile.

Fetch a page, ask it questions, get strings back. If you've used Scrapy, ::text, ::attr(),get() and getall() mean what you think they mean.

Already have the HTML? netweir.parse(html) skips the request and gives you something you query the same way.

The wheels bundle BoringSSL and the parser, so there's nothing else to install. Free-threaded Python 3.14 is included.

quick_look.py
import netweir

page = netweir.get("https://books.toscrape.com/")

page.css("h1::text").get()  # "All products"
page.css(".price_color::text").getall()  # ["£51.77", "£53.74", ...]
page.css("article h3 a::attr(href)").getall()  # ["/catalogue/...", ...]

for book in page.css("article.product_pod"):
    print(book.css("h3 a::attr(title)").get(), book.css(".price_color::text").get())

Fast, and you don't have to take our word for it.

Pages are parsed by lexbor, the same C parser selectolax uses. The difference is what happens after: the whole query runs in Rust and Python gets a finished list of strings, not an object for every element along the way. The GIL is released while it works, so threads parse in parallel.

Parse a 1 MB page and pull out every price, title and link

  1. netweir5.0 ms
  2. selectolax5.3 ms
  3. BeautifulSoup (lxml)196 ms
  4. lxml + cssselect347 ms
  5. parsel353 ms
bench/parse.py, 1 MB page. Lower is better. Log scale.

Crawl 10,000 pages across 100 local sites

  1. netweir1,770
  2. Scrapy 2.19890
  3. httpx + selectolax220
  4. Scrapling 0.4180
bench/crawl.py, pages a second. Higher is better.

Measured on an Apple M-series laptop, Python 3.14 for parsing and 3.13 for the crawl. Both parsing benchmarks run in CI, and the build fails if any library beats netweir. The crawl is server-bound: with 100 requests in flight and 50 ms per answer, 2,000 pages a second is the most anything could do. All three tables, every caveat, and the commands to run them.

Three things go wrong on a long crawl. It handles each.

The site blocks you.

Every response is checked against the block pages of Cloudflare, Akamai, DataDome, HUMAN, Kasada, Imperva and AWS WAF, so a challenge that comes back as a 403 isn't mistaken for a page. A blocked request is tried again with a new session and the next of your proxies, the site is slowed down, and a 429's Retry-After is honoured. A site that keeps blocking is paused instead of hammered.

class Shop(netweir.Spider):
    settings = netweir.Settings(browser="on_block", proxies=[...])

With browser="on_block", a request still blocked after all that gets one go in Chrome, which waits for the challenge to pass and hands its cookies to the HTTP client. The rest of the site doesn't need Chrome.

The crawl dies.

Give it somewhere to keep its state and run it again after a crash, a reboot or a kill -9:

netweir crawl books.py -o books.jsonl -s checkpoint=crawls/books

It picks up where it stopped. Pages already done aren't fetched again, and no item ends up in books.jsonl twice: the tests kill a crawl at random moments until it finishes and check exactly that.

The site is redesigned.

Name a selector and netweir remembers what its element looked like.

price = page.css(".price_color::text", track="price")
price.relocated  # True when the selector missed and the element was found by similarity
price.score      # 1.0 for a match, the similarity when relocated

When a new build renames price_color, wraps it in another div or swaps h1 for h2, the selector stops matching, and netweir finds the element most like the one it remembers. You get a warning, not an empty column. Add repair = netweir.repair.llm() and a language model proposes a replacement selector for each one that broke; netweir checks every proposal on the page and never applies one itself.

On the wire, it's Chrome.

Most sites don't block scrapers by reading their code. They block them by the first few hundred bytes of the connection: which TLS ciphers and extensions arrive, what the HTTP/2 settings frame says, which headers come first. A Python HTTP library gives itself away before it has asked for anything.

netweir sends what Chrome 154 sends. Not something close to it: the same ClientHello, the same HTTP/2 settings and priorities, the same headers in the same order and the same capitalisation over HTTP/1.1. Cookies go where Chrome puts them, and redirects are followed hop by hop the way Chrome follows them.

Firefox 156 and Safari 27 are there too, as profile="firefox" and profile="safari", each with the habits of the real browser, down to Safari sending two cookies with the same path newest first.

JA4 fingerprintt13d1517h2_8daaf6152771_cb7bf5808d99What tls.peet.ws reports for Chrome 154. And for netweir.
How it's checkedChrome was recorded doing four navigations against a local server: a page that sets a cookie, a revisit, a redirect to another site and one within it. cargo test -p netweir-core has netweir do the same four and fails if any request differs from Chrome's.
One known differenceChrome sends a priority header only over HTTP/2, and netweir can't tell a server lacks HTTP/2 until it has answered once. So the first request to an https server that only speaks HTTP/1.1 carries that header; every request after it doesn't.

Bring your selectors with you.

Whatever you scrape with now, your selectors should work here unchanged. XPath is all of XPath 1.0 plus the extras parsel users lean on, and the same page answers Beautiful Soup's methods too.

XPath, as in parsel
page.xpath("//article//h3/a/@title").getall()
page.xpath("//p[has-class('price_color')]/text()").get()
page.xpath("//article//a[re:test(@href, '_\\d+/index\\.html$')]/@href").getall()
page.xpath("//li[@class=$cls]", cls="next").get()  # variables, as in parsel
page.css(".price_color::text").re(r"[\d.]+")  # ["51.77", "53.74", ...]
Beautiful Soup's find family
import re

page.find("ul", class_="pager").find("a")["href"]
page.find_all("a", title=True, limit=10)
page.find_all(string=re.compile("£"))
page.find("h3").find_parent("li").find_next_sibling("li")

Two things work slightly differently, on purpose. get() on an element gives its HTML the way Chrome's outerHTML writes it, and attribute values are plain strings, so node["class"] is "star-rating Three", not a list. Filters still match one class out of several.

A spider is a class with a parse method.

It yields what it found and the links worth following. Run it from the shell, and the file extension picks the format:

netweir crawl books.py -o books.jsonl

That run fetched all 50 pages and wrote 1,000 books in 18 seconds, most of it spent waiting on purpose. Out of the box netweir reads each site's robots.txt and stays out of what it disallows, skips pages whose owners reserve text and data mining rights, and spaces its requests to each site by how quickly the site answers, backing off when it gets a 429. A page linked three different ways, or with utm_ tags on the end, is fetched once.

books.py
import netweir


class Books(netweir.Spider):
    start_urls = ["https://books.toscrape.com/"]

    async def parse(self, page):
        for book in page.css("article.product_pod"):
            yield {
                "title": book.css("h3 a::attr(title)").get(),
                "price": book.css(".price_color::text").get(),
                "url": page.urljoin(book.css("h3 a::attr(href)").get()),
            }
        if next_page := page.css("li.next a::attr(href)").get():
            yield page.follow(next_page)

Or describe the item, and let Rust crawl.

Most spiders are the same three moves: follow the pagination, open each product, pull the same fields out of every one. You can say that instead of writing it.

Run against the real site, that crawled all 1,050 pages and wrote 1,000 books. Every page was parsed, searched and turned into an item in Rust, several at a time on separate threads. Python saw finished items and nothing else. If a price won't turn into a float, you get None and one warning naming the field, rather than a crash on page 600.

Callbacks whose own Python is the slow part can spread over processes: workers=4 ran a heavy spider 3.1 times faster, with the same output.

books.py, declarative
class Book(netweir.Item):
    title = netweir.css("h1::text")
    price = netweir.css(".price_color::text", re=r"[\d.]+", into=float)
    upc = netweir.xpath("//th[.='UPC']/following-sibling::td/text()")
    stock = netweir.css(".availability::text", re=r"\d+", into=int)


class Books(netweir.Spider):
    start_urls = ["https://books.toscrape.com/"]
    rules = [
        netweir.Follow("li.next a"),
        netweir.Follow("article.product_pod h3 a", extract=Book),
    ]

When a page needs a real browser.

Some pages are empty until their JavaScript runs. For those, netweir drives Chrome, and you get the rendered page back as the same node you'd get from a fetch.

There's no sleep in that, and there doesn't need to be. A click waits until its button exists, is visible, has stopped moving, is enabled and isn't covered by a cookie banner. If it never gets there, the error tells you which of those it was.

A driven Chrome normally gives itself away: navigator.webdriver is true, the user agent says HeadlessChrome, and the usual drivers switch on DevTools features that page scripts can notice. netweir doesn't do any of that, and a test checks all three. What it doesn't hide yet is in the docs.

netweir install chrome fetches the Chrome it's tuned for.

quotes.py
async with netweir.browser() as browser:
    page = await browser.new_page()
    await page.goto("https://quotes.toscrape.com/js/")
    await page.click("li.next a")
    root = await page.parse()
    quotes = root.css("span.text::text").getall()

Free software. A commercial licence if you need one.

netweir is AGPL-3.0. If that doesn't work for your company, a commercial licence is available. Issues, fixes and new browser profiles are all welcome either way.

Start with one line.

pip install netweir
Getting started, in about ten minutes

Please use it the way you'd want your own site scraped. Crawls obey robots.txt and TDMRep unless you switch that off, and netweir tells you when you have.