Skip to content

Guides · Connect and debug

Scrapy proxy setup with authentication

Set an authenticated proxy on a Scrapy request, verify the HTTPS destination and inspect the returned exit address. Start with this bounded echo check before adding proxy settings to a larger, authorised crawl.

Short answers

How do I set a proxy in Scrapy?

Set meta["proxy"] on the Request. The built-in HttpProxyMiddleware reads that value and prepares proxy authentication. Include percent-encoded credentials in the proxy URL when the gateway requires a username and password.

Does Scrapy use HTTP_PROXY and HTTPS_PROXY?

HttpProxyMiddleware can read environment proxy settings, but an explicit request proxy takes precedence. That explicit value also takes precedence over NO_PROXY exclusions.

Why does an HTTPS check use an HTTP proxy URL?

The URL describes the gateway connection. Scrapy opens an HTTP CONNECT tunnel, then makes the HTTPS request through it. Certificate verification must be enabled for the destination.

Before scaling a collector, read how to choose proxies for web scraping. It connects the client setup below to usable records, finite retries and a measured traffic budget.

Use the tested Scrapy release

Install Scrapy 2.19.0 in a project environment using Python 3.10 or newer. We checked the example with Python 3.12.14 and Twisted 26.4.0. Confirm the version from the same interpreter that will run your worker; a shell and a scheduled job may use different environments.

Install and check Scrapysh
python -m pip install "scrapy==2.19.0"
python -c "import scrapy; print(scrapy.__version__)"

Copy the complete username from the connection builder into your runtime’s PROXY_USER variable. Choose a currently available pool and country. Enter the proxy password at the hidden prompt. Your account password and API key are different credentials.

Run one HTTPS echo check

scrapy_check.pypython
import getpass
import ipaddress
import os
from urllib.parse import quote, urlsplit

import scrapy
from scrapy.crawler import CrawlerProcess


class ProxyCheck(scrapy.Spider):
    name = "proxy_check"
    ok = False

    async def start(self):
        yield scrapy.Request(
            "https://api.portproof.org/v1/echo-ip", callback=self.parse, errback=self.failed,
            meta={"proxy": self.proxy, "handle_httpstatus_all": True},
        )

    def parse(self, response):
        if response.status != 200:
            print("HTTP response after CONNECT:", response.status)
            return
        try:
            payload = response.json()
            if not isinstance(payload, dict) or not isinstance(payload.get("ip"), str):
                raise ValueError("Expected an IP address")
            address = ipaddress.ip_address(payload["ip"])
        except (ValueError, KeyError, TypeError):
            print("Check failed: invalid echo response")
            return
        self.ok = True
        print("HTTP", response.status, "exit", address)

    def failed(self, failure):
        print("No destination response:", failure.type.__name__)


if __name__ == "__main__":
    server = os.environ.get("PROXY_SERVER", "http://gw.portproof.org:7000")
    try:
        gateway = urlsplit(server)
        valid = (
            gateway.scheme == "http" and gateway.hostname
            and gateway.port is not None and gateway.username is None
            and gateway.password is None and gateway.path in ("", "/")
            and not gateway.query and not gateway.fragment
        )
    except ValueError:
        valid = False
    if not valid:
        raise SystemExit("Use an HTTP gateway URL without credentials or a path")
    user = os.environ.get("PROXY_USER", "")
    if not user or ":" in user:
        raise SystemExit("Set PROXY_USER to your complete proxy username")
    password = getpass.getpass("Proxy password: ")
    proxy = (
        "http://" + quote(user, safe="") + ":" + quote(password, safe="")
        + "@" + gateway.netloc
    )
    process = CrawlerProcess(settings={
        "TWISTED_REACTOR_ENABLED": True,
        "DOWNLOAD_HANDLERS": {
            "https": "scrapy.core.downloader.handlers.http11.HTTP11DownloadHandler",
        },
        "DOWNLOAD_VERIFY_CERTIFICATES": True,
        "CONCURRENT_REQUESTS": 1,
        "DOWNLOAD_TIMEOUT": 20,
        "DOWNLOAD_MAXSIZE": 65536,
        "RETRY_ENABLED": False,
        "REDIRECT_ENABLED": False,
        "METAREFRESH_ENABLED": False,
        "ROBOTSTXT_OBEY": False,  # This diagnostic schedules only the echo URL.
        "COOKIES_ENABLED": False,
        "LOG_ENABLED": False,
        "TELNETCONSOLE_ENABLED": False,
        "REMOTE_CONTROL_ENABLED": False,
    })
    crawler = process.create_crawler(ProxyCheck)
    process.crawl(crawler, proxy=proxy)
    process.start()
    raise SystemExit(0 if crawler.spider and crawler.spider.ok else 1)

Save the file and run python scrapy_check.py from a terminal. It schedules one GET, follows no links or redirects, and exits successfully only after validating an IP address in an HTTP 200 JSON response. The password prompt happens before downloading. For unattended work, replace that prompt with your existing secret-manager lookup.

The downloader has a twenty-second timeout and a 64 KiB response-body limit. These are example settings, not service performance promises. The body limit is not a cap on billable traffic: headers, tunnel overhead and buffering are separate. Give scheduled workers an overall deadline too, since startup and shutdown take additional time.

Encode credentials and keep them out of logs

quote(value, safe="") encodes punctuation such as @, /, # and % before the username and password enter the URL. Pass raw credentials into it once; encoding an already encoded password changes what the gateway receives. A colon is not valid inside a Basic-auth username, so this example rejects it.

The proxy middleware prepares Proxy-Authorization for the gateway. Do not substitute the destination’s Authorization header or Scrapy’s website-login settings. The encoded URL and Base64 header remain secrets. This diagnostic disables Scrapy logging and prints exception types only; protect headers, request metadata and persisted job files when integrating it into a crawler.

Check the tunnel and destination separately

The example selects the HTTP/1.1 download handler explicitly. It uses an HTTP gateway and an HTTPS destination; it does not configure SOCKS or TLS to the gateway. CONNECT carries destination TLS, but does not encrypt the initial proxy authentication on the HTTP connection to the gateway.

Scrapy 2.19.0 defaults to not verifying HTTPS certificates. The explicit DOWNLOAD_VERIFY_CERTIFICATES=True setting enables verification here; see Scrapy’s certificate guidance. For certificate failures, check the destination hostname, system clock and the worker’s approved trust store. Keep verification enabled while fixing the cause.

Make routing explicit on every request

An explicit meta["proxy"] wins over environment routing, including NO_PROXY. We tested conflicting environment proxies and NO_PROXY=*; the request still reached the selected fixture gateway. Check that this choice matches your worker’s intended network configuration. Newly created follow-up requests need their own proxy metadata; setting it on the first request is not a project-wide proxy setting.

Diagnose 407 before adding retries

No destination response: TunnelError
The gateway refused CONNECT. A gateway 407 caused this result locally, but so did 503: the printed type alone does not identify the status. Check the generated username, proxy password and traffic balance, then use the 407 diagnostic guide.
HTTP response after CONNECT: 403 or 429
An HTTP response arrived through the tunnel. Check the destination’s access rules and rate limits. Even 407 returned inside that HTTPS exchange is distinct from a gateway refusing CONNECT.
DownloadFailedError or DownloadTimeoutError
The request failed before a complete destination response. TLS rejection produced DownloadFailedError locally; it is not a certificate-specific diagnosis. Check trust, connectivity and timeout settings without printing raw exceptions that may include sensitive data. When the gateway refuses the connection or never answers, follow the connection-refused guide.

Throttle a permitted crawl deliberately

The echo-only diagnostic skips robots.txt so it does not schedule another URL. For a real crawl, enable ROBOTSTXT_OBEY, restrict allowed destinations and follow the site’s terms. Start with low per-domain concurrency and a download delay suited to its limits. AutoThrottle adapts delays using response latency; its target concurrency is not a strict ceiling or permission to increase load.

Middleware retries are disabled here. Add only bounded retries for permitted, repeatable reads, with deliberate handling of rate-limit responses. Count attempts and downloaded bytes. Connection reuse means a new Scrapy Request does not establish a new exit IP; choose gateway behaviour through the generated username and read the session guide.

What the local fixture established

On 1 October 2026 we ran this snippet through a local authenticated CONNECT proxy and HTTPS origin using synthetic credentials and certificates. Checks covered punctuation, environment overrides, refused tunnels, untrusted and mismatched certificates, destination errors, redirects, malformed responses, body limits and a stalled download. The origin received no proxy authentication header; no tunnel remained after process exit.

These checks validate the named client configuration, not live pool performance or exit-country accuracy. A successful echo describes that request only. Next, measure one permitted application workflow and use the bandwidth guide to plan its traffic. For a standalone HTTP client, see the HTTPX guide.

What is not allowed

Use proxies for authorised research, monitoring and QA in accordance with the destination’s terms and the acceptable-use policy.

Scrapy proxy setup with authentication · Portproof