Back to projects

DataLocate

Project years 2025-2026

Turning public websites and Companies House records into cleaner business intelligence.

60 second pitch

What DataLocate does.

DataLocate combines a web crawler with a Companies House enrichment service. The crawler discovers public business websites, stores page snapshots, extracts useful signals, cleans noisy values and scores confidence.

The Companies House side adds official company records, sector codes, officers, filing information and searchable company lookup. Together, the systems turn messy public information into a more reliable business data product.

Role

Product strategy, crawler architecture, data pipeline design and Companies House enrichment

Sector

Business data, public web intelligence and company record enrichment

Stack

Laravel, Horizon, Redis queues, Spatie Crawler, MySQL, Scout, Meilisearch, S3 storage, AI prompts and Companies House data

Product shape

Business data only becomes useful when it is cleaned, connected and searchable.

DataLocate was shaped around repeatable data operations: discover a business, crawl the public pages, extract the useful signals, clean the values, score confidence and enrich the record with official company context. The result is a data product, not a pile of scraped pages.

Application showcase

Coming soon

Product screens are being prepared for this case study and will be added here shortly.

01

The problem

Useful business data was spread across websites, search results and official company records.

Business data is rarely clean when you first find it. A company might have a website, phone number, email address, social profiles, trading address, registration number, sector, directors and official Companies House record, but those details are usually scattered across different public sources.

DataLocate was built to turn that messy public information into something a business could actually search, trust and use. The challenge was not just collecting pages. It was finding the right signals, removing duplicates, cleaning values and connecting them to official company context.

02

Why it was built

The platform needed to behave like a data pipeline, not a one-off scraper.

The crawler discovers or imports websites, runs controlled scans, stores page snapshots and extracts useful business signals such as emails, phone numbers, social links, company numbers and page metadata. Those values are then normalised, deduplicated and scored before being stored in a searchable data bank.

The Companies House service adds the official layer around that data: company records, statuses, SIC codes, registered office details, officers, people with significant control, charges and filing history. Together, the two systems create a cleaner business profile than either source could provide alone.

03

How it works

The crawler finds the signal, then Companies House adds the official business context.

The crawl engine uses scheduled scans, proxy controls, user agents, page limits and response checks to gather public website content safely and repeatably. Each page is stored as a version, tagged by page type and made available for later analysis without needing to crawl the same site again.

The analyser then runs configurable formulas and AI prompts across the stored pages. Data points are cleaned into standard formats, hashed for deduplication and scored for confidence. The Companies House mirror makes the official company data searchable and available for enrichment through API and dashboard views.

04

Why it matters

The value is not more raw data. It is cleaner data that teams can use with confidence.

A raw crawl gives you pages. DataLocate turns those pages into structured business intelligence: contact details, categories, source trails, confidence scores and official company context.

That matters for teams who need to filter markets, enrich company lists, identify useful prospects, reduce duplicate records and make better use of public business information without relying on manual research.

Platform capability

Features

The platform combines crawler infrastructure, data cleaning, confidence scoring, official record enrichment and search so public business data can become usable.

  • Public website crawling

    Discover and scan business websites for useful public information.

  • Search result discovery

    Use search sources to find relevant domains and business pages.

  • Site imports and schedules

    Import large site lists and run scans on a repeatable schedule.

  • Proxy and user-agent controls

    Control crawl behaviour with device, country, proxy and browser settings.

  • Canonical page storage

    Deduplicate URLs and store page versions for later analysis.

  • Metadata and markdown extraction

    Capture titles, descriptions, page content and readable page versions.

  • Page type tagging

    Classify pages such as home, about, contact, services, careers and locations.

  • Formula-based extraction

    Run configurable patterns to find emails, phones, links and business identifiers.

  • AI prompt processing

    Use AI prompts where structured extraction needs more context.

  • Normalised data bank

    Clean noisy values before storing them as canonical records.

  • Confidence scoring

    Score extracted values so teams can judge quality and relevance.

  • Company number matching

    Connect extracted company identifiers to official company records.

  • Companies House mirror

    Import and search official company data locally for faster enrichment.

  • Officer and PSC records

    Store directors, officers and people with significant control.

  • SIC code enrichment

    Use official sector codes to categorise and segment companies.

  • Company status and filing data

    Include registered status, filing history, charges and health signals.

  • Meilisearch-backed lookup

    Make company names, numbers and sectors fast to search.

  • Queue-based processing

    Separate crawling, analysis, scoring, metadata and imports into background jobs.

The results

A cleaner route from public web data to usable company intelligence.

DataLocate gives the business a way to collect, clean, enrich and search company data through repeatable systems rather than manual research or one-off exports.

4

pipeline stages

Crawl, analyse, normalise and score before data becomes usable.

2

data sources

Public website signals connected with official Companies House records.

18

feature areas

Discovery, crawling, extraction, AI prompts, enrichment, search and background processing.

Next Case Study

Flexi-Orb renewable installer scheme operations portal

View project

Need this kind of data platform?

Let us turn scattered public data into a cleaner business intelligence pipeline.

Discuss a project