HomeServicesIndustriesPortfolioAbout Us
logo
menu-cross-icon
Home
Servicesarrow-down
Industriesarrow-down
Portfolio
About Us

Dark Mode

floating-arrow

Nightwatch — SERP Data Collection at Scale

Global search-results scraping that doesn't get blocked

Industry

SEO, Data Engineering, Web Scraping

Technologies

Golang

ClickHouse

PostgreSQL

Oracle Cloud

S3

Project description

Nightwatch can collect search-results data globally, at the volume its rank tracker demands, while staying resilient against the blocking that defeats naive scrapers. The distributed, auto-scaling proxy layer keeps throughput high and costs controlled, and the parsing engine produces clean, structured ranking data across many engines and result formats. Together they form the reliable data foundation the entire product is built on.

Challenge

A rank tracker is only as good as the search-results data behind it — and collecting that data at scale is genuinely hard. Search engines actively resist automated access, so reliably fetching results from the right country, language and device, for hundreds of thousands of keywords every day, requires a large, constantly-rotating pool of network infrastructure and sophisticated handling to avoid being blocked. On top of acquiring the pages, the raw results then have to be parsed into clean, structured ranking data across many different search engines and result formats. Nightwatch needed this entire pipeline to be fast, resilient and cost-efficient.

Solution

We built Nightwatch's end-to-end search-results collection pipeline in two parts. The first is a distributed proxy-and-orchestration layer that provisions and scales cloud instances across multiple regions on demand, distributes scraping jobs across them with priority handling so urgent work isn't starved, and uses anti-detection browser automation to fetch results reliably from the correct location and device. The second is a SERP parsing engine that turns the raw fetched pages into structured ranking data — positions, URLs, snippets and the many rich SERP features — across all the major search engines and their desktop and mobile variants. Results are persisted efficiently, with structured data in PostgreSQL, raw pages compressed into object storage, and analytics flowing into ClickHouse. The infrastructure uses cost-optimised, preemptible instances and exposes metrics for monitoring throughout.

Project Results

Nightwatch can collect search-results data globally, at the volume its rank tracker demands, while staying resilient against the blocking that defeats naive scrapers. The distributed, auto-scaling proxy layer keeps throughput high and costs controlled, and the parsing engine produces clean, structured ranking data across many engines and result formats. Together they form the reliable data foundation the entire product is built on.

Features and Benefits

  • Distributed Proxy Infrastructure: Auto-scales cloud instances across regions to fetch results at volume.
  • Anti-Blocking Collection: Browser automation fetches results reliably from the right location and device.
  • Priority Job Handling: Urgent and new work is processed without starving regular jobs.
  • Multi-Engine Parsing: Turns raw pages into structured rankings across major engines, desktop and mobile.
  • Efficient Storage: Structured data, compressed raw pages and analytics each stored in the right system.
  • Cost-Optimised & Monitored: Preemptible instances and metrics keep the pipeline efficient and observable.

Technologies Used

  • Go: Both the proxy orchestration layer and the SERP parsing engine.
  • Oracle Cloud (OCI): On-demand, multi-region instance provisioning for the proxy pool.
  • PostgreSQL: Structured ranking results.
  • Object storage (S3 / Backblaze B2): Compressed raw page storage.
  • ClickHouse: Analytics store for SERP data.

Development Process

We treated data collection as two distinct problems. For acquisition, we built an orchestration layer that provisions and scales a proxy fleet across cloud regions, balances scraping jobs across it with priority tiers, and uses anti-detection browser automation to fetch results that look like genuine traffic from the target location. For processing, we built a parsing engine that handles the quirks of each search engine and device variant, extracting positions and the full range of SERP features into a consistent structured form. We paired each storage need with the right system — relational for structured results, compressed object storage for raw pages, and columnar ClickHouse for analytics — and kept the whole pipeline cost-efficient with preemptible instances and instrumented with metrics for reliability.

banner-image

Look at our other projects!
Explore our work and get inspired for your next project.

Projects arrow_right

Have a question?
Let's start a conversation

Reach out to us through the contact form, email or phone. Our team is here to assist you!

Reach out to us through the contact form, email or phone. Our team is here to assist you!

business@altitudeit.org

+381 64 392 7915

Novosadskog sajma 3,

Novi Sad, Serbia

Have a question?
Let's start a conversation

Submit

arrow-right

business@altitudeit.org

+381 64 392 7915

Novosadskog sajma 3,

Novi Sad, Serbia

Altitude

Contact Information

/icons/email.png

business@altitudeit.org

/icons/phone.png

+381 64 392 7915

/icons/location.png

Novosadskog Sajma 3,

Novi Sad, Serbia

Follow Us

linkedininstagram

Services

Agile Software DevelopmentWeb Application DevelopmentAPI Development & IntegrationE-commerce DevelopmentMobile App DevelopmentCloud-Native Web DevelopmentCloud-Native DevOps PracticesUX/UI Design

Industries

FinTechE-CommerceEducationTravelSportsSystem Management

Copyright © 2026 AltitudeIT. All Rights Reserved.