An offer from GESIS – Leibniz Institute for the Social Sciences

Technical Details

Learn more about how the monitor fetches, fingerprints, and saves the DSA Article 40(4) data catalogue pages of very large online platforms and search engines.

How our data catalogue monitor works

Each monitoring run works through the same five steps for every platform; the AI classification runs only where a change was detected.

  1. 01
    Fetch

    The monitor retrieves each platform's data catalogue page with the method that suits it: plain HTTP, JS-rendered headless browser, full Playwright automation, or PDF text extraction.

  2. 02
    Compare

    It computes a SHA-256 fingerprint of the page content and compares outgoing links against the previous snapshot. This detects even minor edits.

  3. 03
    Log

    It writes the result to an append-only audit log. Changed pages receive a full diff with line counts and link deltas; unchanged pages receive a brief status record.

  4. 04
    AI Classify

    An AI model reviews each detected change and classifies it as significant (worth a researcher's attention) or insignificant (no real change).

  5. 05
    Commit

    Every run is permanently saved: a tamper-evident git commit records what changed and when, together with the content fingerprint and the AI classification. Full page snapshots are kept alongside on the server filesystem.

Before anything reaches this public dashboard, a project team member reviews every detected change in an admin queue — confirming the AI's classification, correcting it, or flagging the capture itself as a bad fetch. This happens after the run above is committed, on the team's own schedule, not as part of the automated pipeline. The dashboard only ever shows a change's last reviewed state, never one still awaiting review.

← Back home