Metadata-Driven Entity Matching for Real-Time Golden Record Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems require shutting down to integrate new data sets, leading to inefficiencies in managing and updating golden records, which are representations of real-world entities with multiple views and survivorship rules, in a big data context.
Innovation Solution
A unique architecture utilizing an n-layer model, including a multi-tenant platform with a configuration layer on top of the RELTIO™ platform, enabling efficient modeling of entities, relationships, and interactions, allowing seamless scaling and integration of data from various sources, with real-time matching and merging capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional systems are used to integrate new data sets, then data integration can be achieved, but system shutdown is required which reduces productivity and increases loss of time
Solution Approach 1:
The system transitions from a static schema to a dynamic schema where data models can be evolved and updated in real-time. The metadata configuration allows the system to adapt its structure dynamically without requiring shutdowns, enabling continuous integration of new data sets while maintaining operational productivity
Solution Approach 2:
The system performs preliminary actions by pre-configuring metadata schemas and data models before new data sets arrive. This allows the system to be prepared in advance with the necessary structure to accommodate new data, eliminating the need for disruptive shutdowns and enabling continuous integration operations
2Productivity
If golden records are changed in a datastore, then data updates can be performed, but the process becomes O(n) linear which reduces productivity in big data contexts
Solution Approach 1:
The system segments the golden record management into multiple independent data models and metadata configurations. Instead of managing a single monolithic schema, the system divides it into separable components that can be updated independently, reducing the complexity of data management and enabling faster updates without O(n) linear processes
Solution Approach 2:
The system changes parameters by using metadata configuration to define and modify data structures dynamically. Rather than changing physical data storage structures which are O(n) operations, the system changes metadata parameters that describe the data, enabling fast updates through configuration changes rather than physical data manipulation
3Adaptability or versatility
If an n-layer model with configuration layer is implemented, then real-time data integration and scalability are enabled, but device complexity increases
Solution Approach 1:
The system uses a nested architecture where the configuration layer is nested within the RELTIO platform layer. This nested structure allows the configuration layer to inherit and leverage existing platform capabilities while adding its own metadata configuration functionality, achieving real-time data integration and scalability without proportionally increasing overall system complexity
Solution Approach 2:
The configuration layer is designed as a universal component that can handle multiple data integration scenarios, data models, and metadata configurations through a single unified interface. This multi-functionality reduces the need for separate specialized systems, achieving adaptability and versatility while keeping the overall architecture manageable
Data Source
AI summary
Among other techniques, techniques for dynamic survivorship, cross-tenant matching, and lineage entity identifier (EID) promotion are described. A system utilizing these techniques can include an EID assignment engine, a progressive stitching engine, and a data item update engine. The progressive stitching engine can be at least conceptually characterized as comprising a data point onboarding subengine, a data point registration subengine, a data point matching subengine, and a data point merging subengine. A method utilizing these techniques can include assigning a data item EID to a data item, onboarding a data point, assigning a data point EID to the data point, matching the data point with the data item in a multitenant EID lineage-persistent relational database management system (RDBMS), merging the data point with the data item to create a merged data item, and changing the data item, triggering survivorship and lineage EID promotion rules.


