Short answers
Why does a rendered page cost so much more than a fetch?
Because a fetch downloads one document and a browser downloads the document plus everything it references: images, fonts, stylesheets, scripts, video, analytics and advertising tags. A plain HTML fetch is commonly tens of kilobytes; the same page rendered is commonly one to three megabytes.
What can I safely block?
Images, media and fonts almost always, stylesheets usually, and any third-party host that is not the site you are reading. Scripts are the one to think about: on a server-rendered page they are free to drop, and on a client-rendered one they are the page.
How do I block requests in Playwright?
One route handler on the context: context.route('**/*', ...), check route.request().resourceType(), and call route.abort() for the types you do not want. Setting it on the context rather than the page means every new tab inherits it.
And in Puppeteer?
page.setRequestInterception(true), then a request listener that calls req.abort() or req.continue(). Once interception is on, every request must be answered by your handler, or the page hangs waiting for it.
How do I measure what I saved?
Count the bytes on the wire with a CDP session listening for Network.loadingFinished, run the same page set with and without the filter, and compare. Then confirm against gb_used from the API over a longer run, since that is the number you are billed on.
When should I not use a browser at all?
When the data is in the HTML, or in a JSON endpoint the page itself calls. Capture that endpoint once in the browser, then call it directly for the rest of the run. That is a tenfold saving that no amount of interception can match.
What a page pulls in, and what you actually need
Extraction almost always needs the document and the XHR or fetch responses behind it. Everything else exists for a human to look at. Blocking is not an optimisation trick: it is declining to download things you were never going to read.
| Type | Block it? | What you lose |
|---|---|---|
image | Almost always | Screenshots, and any check that reads an image |
media | Always, for scraping | Video and audio you were not watching |
font | Almost always | Pixel-accurate screenshots and layout comparisons |
stylesheet | Usually | Visibility checks, element positions, anything visual |
script | Only on server-rendered pages | Client-rendered content, and the XHR calls that fetch data |
xhr and fetch | No | The data itself, most of the time |
document | No | The page |
A second filter earns as much as the first: allow the site’s own hosts and drop the rest. Advertising, consent platforms, session replay and tag managers are frequently more traffic than the article they sit around.
Playwright: one route handler
First establish a passing journey with the Playwright proxy setup guide. Verify the route using a page in the configured browser context, then add filtering and rerun the same assertions so the saving does not hide a broken test.
Set the route on the context, not the page, so every tab inherits it. Credentials go in the proxy object as separate fields; Chromium strips them from a URL.
import { chromium } from 'playwright';
const BLOCK_TYPES = new Set(['image', 'media', 'font', 'stylesheet']);
const ALLOW_HOSTS = [/(^|\.)example\.com$/];
const browser = await chromium.launch();
const context = await browser.newContext({
proxy: {
server: 'http://gw.portproof.org:7000',
username: 'USERNAME-peer-us-rot-ondemand',
password: process.env.PROXY_PASSWORD,
},
});
await context.route('**/*', (route) => {
const request = route.request();
const host = new URL(request.url()).hostname;
const thirdParty = !ALLOW_HOSTS.some((re) => re.test(host));
if (BLOCK_TYPES.has(request.resourceType()) || thirdParty) return route.abort();
return route.continue();
});
const page = await context.newPage();
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
console.log((await page.content()).length, 'characters of HTML');
await browser.close();Two details matter more than the filter itself. Wait for domcontentloaded rather than networkidle, because idle never arrives on a page whose analytics you just aborted. And reuse one context per session instead of launching a browser per page, so TLS handshakes are not paid twice.
Puppeteer: interception with the same rules
Start with the Puppeteer proxy setup guide to check the connection and authentication scope before adding resource filtering.
page.authenticate() can supply its credentials to website HTTP authentication as well as the proxy, so use it only when every requested origin is trusted. Do not combine this authentication or interception setup with the scoped CDP handler in the setup guide; review and test one coordinated handler for a broader job.With interception on, every request must be answered by the handler. A missing continue() is a page that hangs rather than an error.
import puppeteer from 'puppeteer';
const BLOCK_TYPES = new Set(['image', 'media', 'font', 'stylesheet']);
const browser = await puppeteer.launch({
args: ['--proxy-server=http://gw.portproof.org:7000', '--blink-settings=imagesEnabled=false'],
});
const page = await browser.newPage();
await page.authenticate({
username: 'USERNAME-peer-us-rot-ondemand',
password: process.env.PROXY_PASSWORD,
});
await page.setRequestInterception(true);
page.on('request', (request) => {
if (BLOCK_TYPES.has(request.resourceType())) return request.abort();
return request.continue();
});
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
await browser.close();What breaks, and how to notice
- Client-rendered pages go blank when scripts are blocked. Check for the text you expect before trusting an empty result.
- Consent walls sometimes live in a third-party script: blocking it can leave a page that never reveals its content, or can skip the wall entirely. Both happen; test the site rather than assuming.
- Screenshots stop being representative once images, fonts and stylesheets are gone. Keep a small unfiltered run for anything visual, including ad verification work.
- Infinite scroll and lazy loading usually survive image blocking, because the trigger is the placeholder, not the picture.
goto, require a selector or a string that only appears on a correctly loaded page. A filter that silently breaks extraction costs more than the traffic it saved.Measure the saving, do not estimate it
encodedDataLength is what arrived for each request. It is a comparison instrument: run the same pages with the filter on and off, and read the ratio.
const cdp = await context.newCDPSession(page);
await cdp.send('Network.enable');
let bytes = 0;
cdp.on('Network.loadingFinished', (event) => {
bytes += event.encodedDataLength;
});
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
console.log(Math.round(bytes / 1024), 'kB on the wire for this page');Then confirm against the bill. Read gb_used from GET /v1/traffic before and after a longer run and take the difference; usage refreshes every five minutes, so leave a window either side. What the meter includes is explained in what counts as proxy traffic. Once the changed workload still returns the required data, recalculate your bandwidth estimate using the measured sample.
From kilobytes to GB and euros
The arithmetic is worth doing once. A hundred thousand pages at 2.2 MB each is about 220 GB. The same run at 260 kB is about 26 GB. That is the same job, the same pages and the same data, at a different order of magnitude on the invoice.
Sizing an order from page counts, page weight and retries is worked through in how many GB do you need.
| Order size | Per GB | Saving |
|---|---|---|
| 1 to 4 GB | EUR 4.50 | |
| 5 to 24 GB | EUR 3.90−13 % | |
| 25 to 49 GB | EUR 3.50−22 % | |
| 50 to 99 GB | EUR 3.20−29 % | |
| 100 to 249 GB | EUR 2.90−36 % | |
| 250 GB and more | EUR 2.50−44 % |
GB never expire. What you buy stays on your balance until you use it; a new purchase adds to the same balance. The trial is 0.5 GB for EUR 2.90, once per customer.
The ladder applies to both pools and every rotation mode, so a saving here is a saving on whatever you are running.
When the right answer is no browser
Open the site once with the network panel on and look at what it calls. If the data arrives as JSON, call that endpoint directly with an HTTP client for the rest of the run: no rendering, no assets, a fraction of the bytes, and far less to go wrong.
Keep the browser for what needs a browser: rendered advertising, geo-dependent layouts, QA journeys. Which pool to point it at is in mobile 4G/5G against residential, and the per-GB ladder is on pricing, with the connection format on the setup docs.
What is not allowed
What is not allowed: blocking assets to save your own traffic is fine. Using a browser farm to hammer a site, ignore its stated limits, evade a block or create accounts is not, and nothing on the declined list becomes acceptable because it is cheaper. See the acceptable-use policy.