Skip to content

Guides · Choose and compare

Choosing proxies for web scraping

Choose proxy traffic around the records your collector needs to produce. Check the country, client and session requirements, then measure usable results and transferred bytes before committing to a larger run.

Published

Short answers

What does Portproof provide for a collector?

Proxy traffic by the GB through shared mobile 4G/5G and residential pools. Your software handles fetching, rendering, parsing, scheduling and storing records. Use the connection for permitted price monitoring, market research or QA workflows.

Should I start with mobile or residential?

Choose the network type your task requires and a country currently available in that pool. Mobile supplies a carrier-network exit; residential uses opt-in peer devices. Compare a controlled sample when either fits. Neither pool guarantees a particular destination response.

How should I compare the cost of a run?

Measure all transferred traffic, including failed attempts, and count distinct usable records. Keep that result separate from requests sent or HTTP success responses. A cheap request that produces the wrong record does not complete the job.

Write the record you need before choosing a proxy

Start with a permitted source and a concrete output: a catalogue entry with its market and capture time, a comparable product offer, or a QA result from your own page. Prefer an authorised feed or API when it already supplies that output. Define missing fields and duplicate handling before running the collector.

A collection brief to keep in your own project
DecisionRecord it before the sample
Source and scopeAllowed destinations, required fields and collection schedule.
Market inputsCountry, storefront, language, delivery region and account state.
AcceptanceRequired field checks, freshness and a key for removing duplicates.
BoundariesPer-domain concurrency, task deadline, response-size limit and retry limit.

An HTTP 200 can contain an empty template or the wrong product variant. Validate content before counting a record. Preserve missing values rather than turning them into zero, and keep the collection timestamp separate from any date printed by the source.

Choose ordinary HTTP or a browser from the source

Inspect the permitted response first. If it contains the fields you need, an HTTP client can fetch them without loading a whole browser page. Scrapy’s dynamic-content guidance starts with locating the data source; a browser becomes useful when the job needs rendered content, interaction or a screenshot.

Use the Scrapy setup guide for a crawler or the HTTPX guide for an HTTP application. For rendered checks, use the Playwright guide. Test in the client that will run the job. An echo check confirms that request’s connection, not the collector’s parsing or every later browser request.

A browser may load scripts, images and background requests. Remove a resource only after checking that the accepted output stays correct. Keep certificate verification enabled, protect proxy credentials, and close responses and browser contexts when finished.

Match the proxy session to the observation

Independent observations can use per-connection rotation. Connection reuse means a new application request need not open a new proxy connection. A multi-step observation may need a sticky session alongside its own cookies and application state. Sticky keeps the selected device while available; its IP can still change or the device can disconnect.

Country routing, storefront settings and browser locale are separate inputs. Keep them consistent when comparing samples. Record requested settings without treating them as proof of the page’s actual market. The session guide explains the connection behaviour; price monitoring covers offer comparability.

Set limits that stop an unproductive run

Begin with low concurrency within the source’s permitted limits. Set a whole-task deadline, response-size ceiling and finite attempt count. For an initial diagnostic, disable retries so one failure remains one observation. Resolve proxy authentication, certificate or parser errors before repeating the sample.

A 429 response means too many requests and may include Retry-After. Honour the delay, reduce load, and stop repeated failures. Retry-After can specify seconds or an HTTP date. If the delay exceeds the remaining job deadline, defer the task instead of waiting indefinitely.

Retry only operations you know are safe to repeat; HTTP semantics distinguish idempotent requests from operations that may create another effect. Link retries to the original task and deduplicate saved output. A target refusal calls for review, not more addresses or more workers.

Measure traffic per task and per usable record

The following numbers are hypothetical arithmetic, not a Portproof benchmark. Suppose a sample plans one record per task. Its traffic total includes initial attempts, retries and failures, measured consistently across the whole sample.

Fictional sample: one intended record per task
Sample measureIllustrative value
Initial tasks100
Total attempts, including retries110
Distinct usable records90
Total transferred traffic60,000,000 bytes (60 decimal MB)
Traffic per initial task600,000 bytes
Traffic per usable recordAbout 666,667 bytes

At the same mix, 10,000 initial tasks would project 6 GB and about 9,000 usable records. That projection does not promise either outcome. If the requirement is 10,000 usable records, this schedule falls short: investigate missing results and resample before buying more traffic. If usable records are zero, the per-record figure is undefined.

Retries are already included here, so another retry multiplier would count them twice. Use the bandwidth calculator with consistent units, and choose headroom from observed variation. Compare client measurements with dashboard usage after an isolated run; the two counters need not measure exactly the same bytes.

Choose an order after the sample passes

Check current locations, then evaluate a small order using your acceptance conditions. The paid trial guide covers that decision. Both pools draw from one GB balance; the offer below shows current pricing and validity. Recheck the sample when the source, parser, browser resources or schedule changes.

Price per GB by order size, EUR
Order sizePer GB
1 to 4 GBEUR 4.50
5 to 24 GBEUR 3.90−13 %
25 to 49 GBEUR 3.50−22 %
50 to 99 GBEUR 3.20−29 %
100 to 249 GBEUR 2.90−36 %
250 GB and moreEUR 2.50−44 %

GB never expire. What you buy stays on your balance until you use it; a new purchase adds to the same balance. The trial is 0.5 GB for EUR 2.90, once per customer.

Choose the amount from your measured workflow. Proxy traffic does not include a hosted collector or extracted datasets.

What is not allowed

Collect only where you have permission. Follow the source’s terms, robots.txt and request limits. Stop when access is refused, and keep credentials and unnecessary personal data out of logs and exported records.

Choosing proxies for web scraping · Portproof