Graph Augmented Entity Resolution Using Representation Learning
Saksham Rai
Abstract
The process of entity resolution involves finding records that refer to the same real-world object even in the presence of missing values, variations in spelling, disagreements in schema, and duplicated identifiers. Although modern representation-learning systems are good at encoding pairs of records, many of them still make judgments about each pair separately and thus fail to take into account relational evidence such as shared addresses, co-authors, transactions, household connections, or agreement among the neighboring candidate records. This paper presents a graph-augmented approach in which a transformer-based record encoder provides the semantic features and a graph neural network spreads the contextual evidence before the match decisions are calibrated and the clusters are formed. The approach combines classical collective entity resolution with current representation learning methods and claims that the primary advantage of graph augmentation is not just greater expressiveness but the ability to enforce contextual consistency while at the same time preserving uncertainty. A practical architecture with its formal update equations, training objectives, an evaluation protocol with published baselines and testable hypotheses, and risk controls are given. The paper concludes that graph augmentation is most useful when the relations are informative and heterogeneous, but it must be constrained by blocking, edge confidence, cluster consistency, and bias-aware evaluation.
References
Back