You know know the drill. You check a few reviews, scan a buying guide, then watch prices for a week before you buy.
Small firms do the same thing at scale. They track rivals, spot promos fast, and keep their own pricing sharp. The hard part sits behind the scenes. Sites block bots, pages change layout, and your data turns stale.
This article shows a practical setup that holds up under real load. It stays simple enough for an SMB team, but it still scales.
Why price data breaks in the real world
Price pages look easy until you automate them. Retailers push changes often, and they serve different HTML to different users. You also hit rate limits fast when you check many SKUs.
Anti-bot tools watch for repeat IPs, odd headers, and tight request timing. You can trigger blocks even at low volume if your requests look fake. JavaScript can hide the real price behind an API call, too.
Speed adds another snag. Google research found 53 per cent of mobile visits end when pages take over three seconds to load. Slow pages drive users away, and they also slow your scraper.
Build a small, stable fetch layer
Start with a fetch layer that does one job well. It pulls pages, handles retries, and logs each request. Keep parsing and rules in a separate module so teams can change them fast.
Use a queue so you can cap your rate per site. Add jitter so your timing does not look like a bot. Store raw HTML for a short time so you can debug fast when a page shifts.
If your team tests requests by hand, keep it close to what code will send. A curl check helps you match headers and spot redirects. Byteful has a clear walkthrough for GET requests in this guide.
Start with a clean request
Set a real User-Agent and accept headers that match a normal browser. Follow redirects and keep cookies per target site. Log the final URL so you can spot geo or device splits.
Watch status codes in your logs. A run of 403, 429, or 503 tells you the site wants you to slow down or back off. Treat that as a signal, not a fight.
Pick the right proxy for the job
A proxy plan should match your task. For light price checks on a few stores, a small pool of datacentre IPs may work. They cost less and stay fast, but blocks can hit sooner.
For broad retail scans, rotate residential IPs with strict limits. Residential traffic looks like real shoppers, so block rates often fall. You still need to keep volume sane, or you burn your pool.
For local context, lock to an Australian exit when you can. Some sites change stock, shipping, and price by state. An AU IP also helps you see the same offers Aussie buyers see.
Mobile proxies suit edge cases. Some sites treat mobile traffic with lighter checks. They cost more, so save them for hard targets or small sample runs.
Keep the data tidy and useful
Price data fails when you mix like with unlike. Save the price, currency, unit size, and shipping cost as separate fields. A $99 item with $15 shipping does not match a $109 free ship offer.
Track time with the right level of detail. Use a timestamp in AEST or AEDT and store UTC too. That lets you line up runs with promos, email drops, or TV ads.
Add a confidence flag for each scrape. Mark a result low trust when the parser falls back to a weak rule. That keeps bad numbers out of your alerts.
Try a two-step check for high value items. Re-fetch the same URL from a second IP when you see a big drop. This cuts false alarms from A/B tests or blocked responses.
Play it straight: rules, risk, and respect
Price tracking can cross lines if you rush it. Read each site’s terms and check robots.txt. Treat “no scrape” rules as a risk marker even when they do not bind you in all cases.
Keep your rate low and avoid peak times. Do not hit checkout, account, or pay flows. Do not collect personal data unless you have a clear legal basis and a strong need.
Build a contact path for complaints. A simple email alias and fast response can save a lot of pain. It also helps you keep ties with key brands you may sell through.
When you should buy data instead
Scraping works best when you need custom rules or niche sites. It can cost more than you think when pages shift each week. A paid feed can win when you need broad cover and clean fields.
Use a hybrid plan when you can. Buy base catalog data, then scrape the few sites that drive your margin. That mirrors how Tech Guide mixes hands-on reviews with broad buying guides.
If you build your own stack, budget time for care and feed checks. The best systems act like products. They need owners, logs, and clear change control.

