Crawling at scale2025 – present
Keeping a bot fleet unblocked
Brand-protection & anti-piracy platform · outstaff
- Role
- Python engineer
- Team
- Crawling team
Problem
A few hundred bots scan marketplaces, social networks and piracy sites millions of times a day, and every site keeps adding anti-bot defences.
What I did
- Kept bots reliable against Cloudflare, Akamai, proof-of-work challenges, TLS fingerprinting, geo-blocking and captchas: diagnosis first, then the cheapest fix that passes.
- Fixed the shared framework, not single bots: a proof-of-work solver that unblocked a whole family of mirror sites, and a page-ownership contract for every Playwright bot.
- Shipped data-acquisition features: search by image, delivery-location-aware monitoring that never reports a false “closed”, signed-JWT auth for a mobile API.
Decision
Escalate per failure, not per site: classify the block first and climb only as far as the cheapest tier that works. Premium proxies stay the exception.
- repositories, incl. 5 shared libraries
- 46
- for ~18 country bots via a shared library
- 1 fix