Skip to content

Every speed claim netweir makes is a script in bench/ that you can run. The numbers below are copied from the README, where they were measured on an Apple M-series laptop. Your machine will give different numbers; the shape should hold. Two of the three run in CI, and the build fails if any library beats netweir.

Parsing and CSS extraction

Pulls every price, title and link out of a generated shop page.

  1. netweir5.0 ms
  2. selectolax5.3 ms
  3. BeautifulSoup (lxml)196 ms
  4. lxml + cssselect347 ms
  5. parsel353 ms
Parsing and CSS extraction, 1 MB page. Lower is better. Log scale.
Parsing and CSS extraction, every column
Library1 MB page10 MB page
netweir5.0 ms61 ms
selectolax5.3 ms64 ms
BeautifulSoup (lxml)196 ms2.3 s
lxml + cssselect347 ms82 s
parsel353 ms82 s

Measured on an Apple M-series laptop with Python 3.14.

The lead over selectolax is small, because both parse with lexbor, and it comes from doing the extraction in Rust. On other machines the two can swap places by a few percent; on CI’s Linux runners they do.

Run it yourself, from a checkout of the repository:

bench/parse.py
uv run --group bench python bench/parse.py

XPath and find_all

Parses the same shop page and asks six XPath questions, or six Beautiful Soup ones.

  1. netweir, XPath9.9 ms
  2. lxml, XPath809 ms
  3. parsel, XPath814 ms
  4. netweir, find_all9.4 ms
  5. Beautiful Soup, find_all170 ms
XPath and find_all, 1 MB page. Lower is better. Log scale.
XPath and find_all, every column
Library1 MB page3 MB page
netweir, XPath9.9 ms33 ms
lxml, XPath809 ms8.8 s
parsel, XPath814 ms8.8 s
netweir, find_all9.4 ms29 ms
Beautiful Soup, find_all170 ms515 ms

Look at lxml across the row: three times the page took eleven times as long. netweir builds a flat index of the document on the first query, so //x is a scan over a few arrays, and its time grows with the page and no faster.

Both this benchmark and the parsing one run in CI, and the build fails if any library beats netweir.

Run it yourself, from a checkout of the repository:

bench/query.py
uv run --group bench python bench/query.py

A whole crawl

One local server plays 100 sites of 100 pages, each answer 50 ms late, and every crawler fetches all 10,000 pages and pulls a title and price from each, with the same limits (100 requests in flight, 8 per site).

  1. netweir1,770
  2. Scrapy 2.19890
  3. httpx + selectolax220
  4. Scrapling 0.4180
A whole crawl, Pages a second. Higher is better.
A whole crawl, every column
LibraryPages a secondCPU per pagePeak memory
netweir1,7700.22 ms68 MB
Scrapy 2.198901.1 ms125 MB
httpx + selectolax2203.2 ms190 MB
Scrapling 0.41800.8 ms76 MB

With 100 requests in flight and 50 ms per answer, 2,000 pages a second is the most any crawler could do here, so netweir is waiting on the server, not on itself.

Scrapling is held back by its HTTP session’s default of 10 connections, which it doesn’t let you change; that is how it ships.

Same laptop, Python 3.13. Run it once the others are installed.

Run it yourself, from a checkout of the repository:

bench/crawl.py
uv run python bench/crawl.py

Setting up to run them

The benchmarks live in the netweir repository and run against a development build. You need Rust, a C compiler, CMake and uv:

git clone --recurse-submodules https://github.com/netweir/netweir
cd netweir
uv sync --group dev
uv run maturin develop --uv

The parsing and query benchmarks install the libraries they compare against with --group bench. The crawl benchmark expects Scrapy, httpx, selectolax and Scrapling to be installed already.