informatik.hu-berlin.de

Schema Matching using Duplicates

Authors: 
Bilke, A.; Naumann, F.
Year: 
2005
Venue: 
ICDE, 2005

Most data integration applications require a matching
between the schemas of the respective data sets. We show
how the existence of duplicates within these data sets can be
exploited to automatically identify matching attributes. We
describe an algorithm that first discovers duplicates among
data sets with unaligned schemas and then uses these duplicates
to perform schema matching between schemas with
opaque column names.
Discovering duplicates among data sets with unaligned
schemas is more difficult than in the usual setting, because

Syndicate content