Back to Blog
|5 min readCompliance Guide

The Data You Collect and Never Use

We mapped every field of personal data in Metriport, an open-source healthcare API, from the source code to where it actually travels. Sixty-four fields went nowhere.

By Scrutora · scrutora.com

Repository · metriport, publicFrameworks · DPDPA, HIPAAMethod · static, source only

Data minimisation is the rare compliance principle that everyone agrees with and almost no one can evidence. “Collect only what you need” is one line in a policy and a genuinely hard thing to prove against a live codebase. The schema that a service ships with is rarely the schema someone designed on purpose. Fields arrive for a feature, the feature ships and later gets cut, and the field stays in the model, quietly holding data that nothing reads.

A spreadsheet inventory will not catch that. The spreadsheet records what someone believed the system collected on the day they filled it in. The code records what it collects today. So we pointed Scrutora at a real one and let it derive the map from source.

Metriport is a good subject for this because it is open, it is a real healthcare API, and it handles data that regulation actually cares about. Everything below is checkable against the repository, field by field, down to the file and line.

What the Scan Settles First

Before a single flow is drawn, the question is which fields are personal data at all.

Scrutora begins by reading the codebase for every field name that looks like personal data and proposing a classification for each. That first pass is deliberately wide. Name-based detection catches real identifiers and also catches false positives: a config parameter called max_age, an entity property called name, a format string. So the map does not treat the first pass as truth. It asks for confirmation, and it shows its own uncertainty on screen.

Scrutora tracked data objects grid for Metriport, each field tagged and 33 marked as not personal data
The tracked-data view. Every field the scan believes is personal data, each with a proposed tag. The line at the top, 33 marked not personal data, hidden from the map, is the human confirmation step doing its work: matches that read as personal by name but are not, taken out before they can distort the flow map or the RoPA.

On the tags themselves

Scrutora scans against more than one framework, so the same field can carry a different tag depending on which regime applies to the codebase. A date of birth is plain personal data in one context and a protected health identifier in another. The tag follows the framework, not the field name, which is why confirmation matters and why the map is built from confirmed fields rather than raw matches.

Then It Follows Each Field

Collection, processing, and every sink the field reaches.

For each confirmed field, the scan traces where the value goes: into a database store, into application logs, out through an external transmit. The map is a live picture of that, the auto-detected flows merged with any destinations added by hand. In this run it resolves to 372 sinks across 211 entry points and 237 processing points.

Data flow ribbons from many personal-data fields into database store, external transmit and application log sinks
Fields to sinks. Ribbon width is field count multiplied by how far the value travels. The heavy paths land in the database store, the external transmit, and the application log. Hovering a field traces its own flows; clicking a sink drops to the file and line.

This is already more than a spreadsheet can hold, because it is derived rather than recalled. But the finding that stays with you is not on any of the loaded paths. It is the field that has no path at all.

Sixty-Four Fields That Go Nowhere

Declared in the code, never traced to a store, a log, or a recipient.

33
data elements, under 115 names
33
matches ruled not personal data
372
sinks traced across the code
64
fields declared, no flow traced

The map carries a sink of its own for these: Declared, no flow traced. Sixty-four fields are defined in the codebase and never reach a store, a log, or an external recipient that the code can demonstrate. They sit in the schema. Something, at some point, decided to collect them.

The Declared, no flow traced node showing 64 fields with faint ribbons and no downstream sink
The minimisation node. Sixty-four fields resolve here. The ribbons are faint because they carry field count but no distance travelled. Every one is a question: is this still needed, and if not, why is it being stored?
Static analysis reads what the code declares and where values move through it. It cannot see runtime configuration or a caller that supplies a field the source never references. So treat “no flow traced” as a strong prompt to check, not a proof of dead data. The value of the number is that it turns an open-ended audit into a finite list of 64 named things to look at.

That list is the most actionable minimisation artifact we know of. A privacy engineer cannot review a whole schema on instinct. They can review 64 named fields, and for each one answer a question that a RoPA spreadsheet never poses: do we still need this, and if not, why are we holding it?

Why Derive It From the Code

The usual way to build a Records of Processing inventory is to ask teams what their systems collect and write it down. That produces a document that is accurate on day one and drifts every day after, because the code keeps changing and the document does not. Deriving the map from source inverts the default. Instead of asking people to recall what they hold, you start from what the code demonstrably touches, and the fields that touch nothing fall out on their own.

  • The inventory is a product of the current code, not a memory of it, so it does not drift between audits.
  • Unused fields surface as a side effect of mapping, rather than needing a separate hunt.
  • Every entry is evidence: a field name, a file, a line, checkable by anyone with the repo.
  • The uncertain cases are labelled as uncertain instead of being presented as fact.
A screen recording of the same map: from the tracked-data view through to the 64 unrouted fields.

The Honest Limits

Two, stated plainly. The classification pass is wide by design and needs a human to confirm it; the 33 fields ruled out in this run are that step, not an afterthought. And the flow map is static, so it sees imports and assignments rather than runtime behaviour. Neither limit weakens the minimisation finding, because a field with no traced flow is exactly the field a human should look at next. It does mean the number is a starting point for review, not a verdict.

If you want to check any of this, the repository is public and the field names above are real. That is the point. A data map is only worth as much as its evidence, and evidence is checkable or it is nothing.

Vimalendukumar Dwivedi is Co-founder of Scrutora, a Code Compliance Platform.

Find your own 64

Scrutora traces every personal-data field in your codebase to wherever it actually lands, at rest or in motion, and hands you the evidence instead of another meeting.

Share this article