PretzelPort — home
industry

Healthcare & clinical research

Two studies ask the same clinical question and record it differently. Add data from routine care, which was never collected for research at all. We make them analyzable together.

We have built research data infrastructure for large clinical research networks, including AI-assisted harmonization that made clinical study data and real-world evidence analyzable together.

Deep stacks of tabbed and annotated paper files on a desk.

// two studies, one question, four vocabularies

the situation

The mapping exists. It lives in a spreadsheet, with one person.

Two studies record the same clinical concept under different variable names, in different units, against different terminologies, at different timepoints, with different eligibility definitions. Neither is wrong. Then real-world data joins the picture — collected for care, not for research, and structured accordingly.

To compare or pool any of it, someone builds the mapping by hand: this variable equals that one, this code maps to that concept, these units convert, this timepoint is close enough. It takes months. It is undocumented. It lives with one person, and the next question starts it over.

What we build: harmonization onto a common data model — CDISC-shaped study data and OMOP-mapped routine-care data brought into one queryable form — with AI doing the first pass at mapping variables, codes and units, and people reviewing rather than authoring. Every mapping is recorded, versioned and traceable to its source, so a pooled variable can be defended and not merely produced.

What it is worth: cross-study and study-plus-RWE questions answerable in weeks rather than quarters; the mapping becomes a reusable asset instead of a spreadsheet; and the provenance holds when a reviewer asks how a pooled variable was derived.

what we find

Problems we see

  • Clinical and research data fragmented across sites, systems and formats — hard to combine even inside one institution.
  • The same clinical concept recorded four ways, so every pooled analysis begins with months of mapping.
  • Harmonization done once, by hand, for one question — undocumented, and rebuilt from scratch for the next.
  • Clinicians and researchers dependent on someone else for every question worth asking, so the marginal question never gets asked.
  • Research networks that want to share insights without sharing raw data, and no infrastructure that lets them.
  • Generic AI tools that become a non-starter the moment the data protection officer looks at them.
how we help

Analysis where the data lives

Study data, registry data and routine-care data harmonized into one governed place — connected data foundations — with access researchers and clinicians can use directly rather than by requesting a report: answers on your data.

The architectural position matters more here than anywhere else on this site: analysis happens where the data lives, and insights are shared without raw data moving. That is a design choice made at the start, not a privacy feature added once someone objects. AI does the first pass at mapping and people review it — not the reverse.

both sides

The same problem, from either direction

We have worked on both sides. What differs is who owns the data and who has to approve the analysis.

Academic and hospital research

In German university hospitals this is the Datenintegrationszentrum and the Studienzentrale: research data infrastructure serving many studies at once, where the institution owns the data and governance is close at hand.

Sponsor-side study and RWE teams

Clinical data management, biometrics and real-world evidence groups, who own the study data and the pooling problem, and who have to defend a derived variable to somebody outside the organization.

the approval path

Built for the conversation with your data protection officer

The projects that fail are the ones that treated a required gatekeeper as an obstacle rather than a design input.

How data protection requirements are met in the architecture rather than in a policy document
RequirementHow it is built
Where analysis runs Inside your environment, on your data. Insights leave; raw records do not.
Purpose limitation Reflected in the architecture — what a given role can reach is enforced, not described in a policy document.
Access control Mirrors site, study and role boundaries, because those are the boundaries the approval is actually about.
The trail A documented record of what was accessed, by whom, and for which purpose — available before anyone asks for it.

Done this way it makes projects faster, not slower. Approval is where these things actually stall.

Which question can your researchers not ask today?

Tell us what stands between them and the answer — the mapping, the approvals, or both.

Get in touch