Scraping Public Web Data to Track Competitor Movements
Competitors tell you a lot about what's working in your market — if you know how to read what they're doing. Their pricing pages, job listings, recently published content, and product changes all carry signals about where they're investing and where they see opportunities.
The problem is that manual competitive monitoring doesn't scale. You can't check twenty competitor sites every week across dozens of data points and still have time to act on what you find.
Python web scraping solves the collection problem. What you do with the data is a separate skill.
What Competitive Scraping Is Actually Useful For
Before writing a line of code, be clear about what question you're trying to answer. Scraping for its own sake produces data nobody uses. Scraping to answer a specific question produces insights that drive decisions.
The most useful categories of competitive intelligence you can gather programmatically:
Pricing and feature changes. If your competitors update their pricing pages or add new pricing tiers, you want to know quickly. Scraping key pages weekly and comparing the new version against a stored snapshot will catch these changes automatically.
Content publishing. Monitoring competitors' blog RSS feeds or tracking their newly indexed pages shows you which topics they're investing in. If a competitor publishes six articles about a topic you haven't touched, that's a signal worth investigating.
Job listings. Companies hiring for specific roles often telegraph strategy. A competitor posting three data engineering roles probably isn't about to launch a reporting feature by accident. LinkedIn, Greenhouse, and Workable are all scrapeable.
Product page changes. For e-commerce or SaaS, watching competitor product and feature pages for copy changes, new sections, or removed content can indicate A/B test results, strategic pivots, or market responses.
Python Tools You'll Actually Use
The standard scraping toolkit in Python is small:
requests handles HTTP requests. You send a GET request to a URL and get back the HTML content of the page.
BeautifulSoup parses HTML and lets you navigate it like a tree structure, extracting specific elements by tag, class, or ID.
Playwright or Selenium handles JavaScript-rendered pages — sites where the content you want only appears after scripts have run. These tools control an actual browser, so they can interact with dynamic pages the same way a real user would.
pandas is useful for storing and comparing data across scraping runs. You can keep a CSV or SQLite database of each week's scraping results and compare against the previous week to identify what changed.
A Basic Scraping Pattern
The simplest useful scraper fetches a page, extracts specific content, stores it, and compares it against the previous version.
In practice:
- Use
requests.get(url)to fetch the page - Parse with
BeautifulSoup(response.text, 'html.parser') - Use
soup.find()orsoup.select()with appropriate CSS selectors to extract what you need - Store the result with a timestamp in a JSON file or a small database
- On the next run, load the previous version and diff the two
The diff step is where the intelligence comes from. A pricing page that hasn't changed in 12 weeks and then suddenly shows a 20% price increase is information. A blog that published nothing for a month and then put out five articles in a week is information.
Handling JavaScript-rendered Content
Many modern sites load their content through JavaScript after the initial page request. If you fetch these pages with requests alone, you'll get the HTML skeleton but none of the dynamically loaded content.
Playwright is the cleanest solution for this. Install it with pip install playwright, run playwright install chromium, and then write code that launches a browser, navigates to the URL, waits for the content to load, and extracts the page source.
The tradeoff is that Playwright-based scraping is significantly slower and more resource-intensive than requests-based scraping. Use it only where you actually need JavaScript rendering; don't default to it for everything.
Avoiding Blocks
Most websites have rate limits, bot detection, or both. A scraper that hits the same site fifty times in five minutes will be identified and blocked.
Practical approaches to reduce detection:
- Add delays between requests (a few seconds is usually enough for politeness and to avoid blocks)
- Rotate user-agent strings so requests look like they come from different browsers
- Use residential proxies if you're scraping at significant volume and hitting blocks
- Respect
robots.txtand don't scrape pages that are explicitly disallowed
The robots.txt file is a standard that most scrapers should follow both legally and ethically. Pages that are blocked there usually have a reason to be.
Storing and Acting on the Data
Raw HTML is rarely the output that's useful. Build your scraper to extract structured data — specific fields from specific elements — and store it in a way that makes comparison easy.
A SQLite database is a good choice for small to medium competitive monitoring projects. It's serverless, requires no separate database server, and is fully queryable with Python's built-in sqlite3 library.
Set up a scheduled job (cron on Linux/Mac, Task Scheduler on Windows, or a simple cloud function) to run your scrapers regularly. Have it compare the new data against the stored version and send you a summary of what changed — either by email or by posting to a Slack channel.
The output should be a short digest: "Competitor X changed their pricing page. Competitor Y published 3 new articles about [topic]. Competitor Z added a new feature to their product page."
The Ethics and Legal Side
Web scraping occupies complicated legal territory that varies by jurisdiction and by the specific site being scraped. In general:
- Scraping publicly visible information is broadly legal in most jurisdictions, though there have been court cases challenging this
- Scraping behind a login or through circumventing access controls is much more likely to create legal problems
- Terms of service violations don't automatically create legal liability, but they can get your IP blocked and, in some cases, lead to more serious consequences
- Scraping at a scale that meaningfully impacts a server's performance is closer to a DDOS attack than data collection
Be reasonable about volume, respect robots.txt, don't scrape anything that requires bypassing authentication, and your competitive monitoring scraper will almost certainly stay in safe territory.