Skip to content

Guides · Sessions and traffic

Cut proxy bandwidth in a headless browser

A headless browser is honest about what a page costs: it fetches everything a real visitor would, including the images, the fonts, the advertising and the analytics you have no use for. Blocking the request types you do not need is a handful of lines, and it is usually the difference between megabytes and hundreds of kilobytes per page.

Published Updated

Short answers

Why does a rendered page cost so much more than a fetch?

Because a fetch downloads one document and a browser downloads the document plus everything it references: images, fonts, stylesheets, scripts, video, analytics and advertising tags. A plain HTML fetch is commonly tens of kilobytes; the same page rendered is commonly one to three megabytes.

What can I safely block?

Images, media and fonts almost always, stylesheets usually, and any third-party host that is not the site you are reading. Scripts are the one to think about: on a server-rendered page they are free to drop, and on a client-rendered one they are the page.

How do I block requests in Playwright?

One route handler on the context: context.route('**/*', ...), check route.request().resourceType(), and call route.abort() for the types you do not want. Setting it on the context rather than the page means every new tab inherits it.

And in Puppeteer?

page.setRequestInterception(true), then a request listener that calls req.abort() or req.continue(). Once interception is on, every request must be answered by your handler, or the page hangs waiting for it.

How do I measure what I saved?

Count the bytes on the wire with a CDP session listening for Network.loadingFinished, run the same page set with and without the filter, and compare. Then confirm against gb_used from the API over a longer run, since that is the number you are billed on.

When should I not use a browser at all?

When the data is in the HTML, or in a JSON endpoint the page itself calls. Capture that endpoint once in the browser, then call it directly for the rest of the run. That is a tenfold saving that no amount of interception can match.

What a page pulls in, and what you actually need

Extraction almost always needs the document and the XHR or fetch responses behind it. Everything else exists for a human to look at. Blocking is not an optimisation trick: it is declining to download things you were never going to read.

Resource types, by how safe they are to abort
TypeBlock it?What you lose
imageAlmost alwaysScreenshots, and any check that reads an image
mediaAlways, for scrapingVideo and audio you were not watching
fontAlmost alwaysPixel-accurate screenshots and layout comparisons
stylesheetUsuallyVisibility checks, element positions, anything visual
scriptOnly on server-rendered pagesClient-rendered content, and the XHR calls that fetch data
xhr and fetchNoThe data itself, most of the time
documentNoThe page

A second filter earns as much as the first: allow the site’s own hosts and drop the rest. Advertising, consent platforms, session replay and tag managers are frequently more traffic than the article they sit around.

Playwright: one route handler

First establish a passing journey with the Playwright proxy setup guide. Verify the route using a page in the configured browser context, then add filtering and rerun the same assertions so the saving does not hide a broken test.

Set the route on the context, not the page, so every tab inherits it. Credentials go in the proxy object as separate fields; Chromium strips them from a URL.

Block by type and by hostnode 22 · playwright
import { chromium } from 'playwright';

const BLOCK_TYPES = new Set(['image', 'media', 'font', 'stylesheet']);
const ALLOW_HOSTS = [/(^|\.)example\.com$/];

const browser = await chromium.launch();
const context = await browser.newContext({
  proxy: {
    server: 'http://gw.portproof.org:7000',
    username: 'USERNAME-peer-us-rot-ondemand',
    password: process.env.PROXY_PASSWORD,
  },
});

await context.route('**/*', (route) => {
  const request = route.request();
  const host = new URL(request.url()).hostname;
  const thirdParty = !ALLOW_HOSTS.some((re) => re.test(host));
  if (BLOCK_TYPES.has(request.resourceType()) || thirdParty) return route.abort();
  return route.continue();
});

const page = await context.newPage();
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
console.log((await page.content()).length, 'characters of HTML');
await browser.close();

Two details matter more than the filter itself. Wait for domcontentloaded rather than networkidle, because idle never arrives on a page whose analytics you just aborted. And reuse one context per session instead of launching a browser per page, so TLS handshakes are not paid twice.

Puppeteer: interception with the same rules

Start with the Puppeteer proxy setup guide to check the connection and authentication scope before adding resource filtering.

The following is a separate illustration for a trusted endpoint. page.authenticate() can supply its credentials to website HTTP authentication as well as the proxy, so use it only when every requested origin is trusted. Do not combine this authentication or interception setup with the scoped CDP handler in the setup guide; review and test one coordinated handler for a broader job.

With interception on, every request must be answered by the handler. A missing continue() is a page that hangs rather than an error.

The Puppeteer equivalentnode 22 · puppeteer
import puppeteer from 'puppeteer';

const BLOCK_TYPES = new Set(['image', 'media', 'font', 'stylesheet']);

const browser = await puppeteer.launch({
  args: ['--proxy-server=http://gw.portproof.org:7000', '--blink-settings=imagesEnabled=false'],
});
const page = await browser.newPage();
await page.authenticate({
  username: 'USERNAME-peer-us-rot-ondemand',
  password: process.env.PROXY_PASSWORD,
});

await page.setRequestInterception(true);
page.on('request', (request) => {
  if (BLOCK_TYPES.has(request.resourceType())) return request.abort();
  return request.continue();
});

await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
await browser.close();

What breaks, and how to notice

  • Client-rendered pages go blank when scripts are blocked. Check for the text you expect before trusting an empty result.
  • Consent walls sometimes live in a third-party script: blocking it can leave a page that never reveals its content, or can skip the wall entirely. Both happen; test the site rather than assuming.
  • Screenshots stop being representative once images, fonts and stylesheets are gone. Keep a small unfiltered run for anything visual, including ad verification work.
  • Infinite scroll and lazy loading usually survive image blocking, because the trigger is the placeholder, not the picture.
Add an assertion to the job itself: after goto, require a selector or a string that only appears on a correctly loaded page. A filter that silently breaks extraction costs more than the traffic it saved.

Measure the saving, do not estimate it

encodedDataLength is what arrived for each request. It is a comparison instrument: run the same pages with the filter on and off, and read the ratio.

Count the bytes on the wirenode 22 · playwright · cdp
const cdp = await context.newCDPSession(page);
await cdp.send('Network.enable');

let bytes = 0;
cdp.on('Network.loadingFinished', (event) => {
  bytes += event.encodedDataLength;
});

await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
console.log(Math.round(bytes / 1024), 'kB on the wire for this page');

Then confirm against the bill. Read gb_used from GET /v1/traffic before and after a longer run and take the difference; usage refreshes every five minutes, so leave a window either side. What the meter includes is explained in what counts as proxy traffic. Once the changed workload still returns the required data, recalculate your bandwidth estimate using the measured sample.

From kilobytes to GB and euros

The arithmetic is worth doing once. A hundred thousand pages at 2.2 MB each is about 220 GB. The same run at 260 kB is about 26 GB. That is the same job, the same pages and the same data, at a different order of magnitude on the invoice.

Sizing an order from page counts, page weight and retries is worked through in how many GB do you need.

Price per GB by order size, EUR
Order sizePer GB
1 to 4 GBEUR 4.50
5 to 24 GBEUR 3.90−13 %
25 to 49 GBEUR 3.50−22 %
50 to 99 GBEUR 3.20−29 %
100 to 249 GBEUR 2.90−36 %
250 GB and moreEUR 2.50−44 %

GB never expire. What you buy stays on your balance until you use it; a new purchase adds to the same balance. The trial is 0.5 GB for EUR 2.90, once per customer.

The ladder applies to both pools and every rotation mode, so a saving here is a saving on whatever you are running.

When the right answer is no browser

Open the site once with the network panel on and look at what it calls. If the data arrives as JSON, call that endpoint directly with an HTTP client for the rest of the run: no rendering, no assets, a fraction of the bytes, and far less to go wrong.

Keep the browser for what needs a browser: rendered advertising, geo-dependent layouts, QA journeys. Which pool to point it at is in mobile 4G/5G against residential, and the per-GB ladder is on pricing, with the connection format on the setup docs.

What is not allowed

What is not allowed: blocking assets to save your own traffic is fine. Using a browser farm to hammer a site, ignore its stated limits, evade a block or create accounts is not, and nothing on the declined list becomes acceptable because it is cheaper. See the acceptable-use policy.

Cut proxy bandwidth in a headless browser · Portproof