SISTRIX GmbH Logo

SISTRIX GmbH

Bonn

SERP Parser for SEO Analytics at SISTRIX

Software Craftsman | DevOps

September 2016 - January 2018
1 year 5 months
Full-time
Product work
Bonn

Relevance

Why this case matters

The case shows how problem understanding, responsibility, decisions, and implementation came together in this context.

Impact

SERP parser, big data delivery, and platform operations for SEO analytics with 450M+ keywords, 200M seeds, SERP feature analysis, status dashboards, Docker Swarm, and CI/CD operations.

What carries forward today

AI-native systems craft depends on the same foundations: large data flows, reproducible pipelines, clear ownership, automated quality checks, and operationally reliable platforms.

SISTRIXSERP ParserYou Build It You Run ItDocker SwarmObservabilityStatus DashboardSEO Analytics

Defensible proof

450M+

Keywords worldwide

SERP parsing and Spark/Hadoop extraction for large SEO analytics across countries and devices.

200M

Crawler seeds

HTML crawler work for structured data extraction at large scale.

SERP Features

Parser product

Made organics, SEM, shopping, knowledge graph, maps, featured snippets, and other search result types inspectable.

Where this experience creates value

  • For teams with data-intensive SaaS products where pipelines, delivery, and operations belong together.
  • For organizations that need not just big data backend work, but reliable delivery and ownership.

Case context

Overview

SISTRIX needed SEO analytics pipelines that could be built, shipped, and operated by the same team. A central part was the SERP Parser: collect search results across countries and devices, classify them, and make parser status visible instead of treating the system as a hidden batch pipeline.

I worked with a "You Build It, You Run It" model across Docker/Docker Swarm, CI/CD, Spark/Hadoop extraction for 450M+ keywords worldwide, and Apache Mesos/Marathon platform work. The parser made SERP Features such as organics, SEM, shopping, knowledge graph, maps, featured snippets, and sitemaps inspectable across countries and devices.

Responsibility

Activities

  • SERP Parser: Made search results inspectable across countries, devices, and SERP Features
  • Parser status: Play Framework dashboard for weekly runs, success states, throughput, and calc nodes
  • Big Data Pipeline Development: HTML parser with XPath for millions of keywords per country, API integration
  • Data Extraction: HTML crawler with Spark/Hadoop for 200M seeds, structured data extraction
  • DevOps & CI/CD: "You Build It, You Run It" pipeline with Docker and automated deployments
  • Platform setup: Built container setup, dashboards, and observability close to parser operations on provided server access
  • PaaS Architecture: Apache Mesos/Marathon, AWS Route 53, scalable infrastructure
  • Quality Assurance: Automated acceptance tests for SaaS tools with Cucumber and Selenium
  • Monitoring & Operations: Status dashboard with Play Framework, operational transparency

Operating mode

Methodology

  • "You Build It, You Run It": End-to-end ownership for parser, pipelines, and operations
  • Operational visibility: Parser status, SERP Features, throughput, and failure states turned runtime behavior into product feedback
  • Big Data Processing: Spark/Hadoop and scalable data pipelines for repeatable search result analysis
  • CI/CD: Automated tests, continuous deployment, and deployment feedback kept delivery tied to runtime behavior

Technical context

Technology stack

The tools are not the point by themselves. What matters is which system layers had to work together.

10Areas
46Technologies

Databases & Storage

5
Google Guava CachememcachedMySQLArangoDBData Storage

Messaging & Event Streaming

1
Pub/Sub

DevOps

9
MavenDockerDocker SwarmGitApache MesosMarathonContainer OrchestrationStatus DashboardObservability

Tools

4
SBTCucumberSeleniumMockito

Practices

3
Test-Driven DevelopmentTestingQuality Assurance

Backend

9
JavaScalaPlay FrameworkAkkaGoogle GuiceREST APIsRxJavaGroovyjOOQ

Frontend

2
JavaScriptHTML Parsing

Data & AI

11
SparkHadoopBig Data ProcessingNutchTikaXPathSAXON-HESearch Result ParsingSERP ParserSERP Feature ExtractionData Extraction

CI/CD & Delivery Pipelines

1
CI/CD Pipeline

Architecture

1
PaaS Architecture

Next step

If you need similar responsibility, we can identify the next useful leverage point directly.

Send the situation, goal, and decision in front of you. I will respond personally with a clear view of the value I could add.