University project · Schema matching & data integration
Two hospitals, one patient, slightly different codes.
Hospitals store the same kinds of information in different databases, with different column names. To combine them you first need to know which column in one corresponds to which in the other. That is schema matching. A common way to do it is to look at the values: if two columns contain the same things, they probably mean the same thing.
Clinical data makes this harder. Diagnoses are stored as codes from huge medical vocabularies, and real-world coding errors are rarely random. Someone records "chronic kidney disease, stage 5" when it was stage 4: a neighbouring concept, not a typo. To an exact comparison those two are as different as "kidney disease" and "broken arm".
The project treats the vocabulary as what it is, a hierarchy. Two codes are similar if they are close in that tree. I built a way to generate realistic near-miss errors on OMOP-CDM data and a matcher that compares values by meaning instead of by spelling, then measured both against exact matching on the Valentine benchmark framework.
The code behind it →