Technical Details
Learn more about how the monitor fetches, fingerprints, and saves the DSA Article 40(4) data catalogue pages of very large online platforms and search engines.
How our data catalogue monitor works
Each monitoring run works through the same five steps for every platform; the AI classification runs only where a change was detected.
-
01
Fetch
The monitor retrieves each platform's data catalogue page with the method that suits it: plain HTTP, JS-rendered headless browser, full Playwright automation, or PDF text extraction.
-
02
Compare
It computes a SHA-256 fingerprint of the page content and compares outgoing links against the previous snapshot. This detects even minor edits.
-
03
Log
It writes the result to an append-only audit log. Changed pages receive a full diff with line counts and link deltas; unchanged pages receive a brief status record.
-
04
AI Classify
An AI model reviews each detected change and classifies it as significant (worth a researcher's attention) or insignificant (no real change).
-
05
Commit
Every run is permanently saved: a tamper-evident git commit records what changed and when, together with the content fingerprint and the AI classification. Full page snapshots are kept alongside on the server filesystem.
Before anything reaches this public dashboard, a project team member reviews every detected change in an admin queue — confirming the AI's classification, correcting it, or flagging the capture itself as a bad fetch. This happens after the run above is committed, on the team's own schedule, not as part of the automated pipeline. The dashboard only ever shows a change's last reviewed state, never one still awaiting review.