Entity Identifier Clustering via Context Scores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional master data management hubs face challenges in accurately managing customer data due to the complexity of customer interactions across multiple databases and applications, often resulting in incorrect or inconsistent data records, especially when dealing with dynamic customer information and data from various sources.

Innovation Solution

The implementation of entity identifier clustering based on context scores, which generates confidence scores for contact elements based on their context within different departments of an enterprise, allowing for the creation and updating of entity identifier clusters to provide accurate and consistent customer data records.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional master data management hubs are used to manage customer data, then data can be stored and accessed, but data accuracy and consistency deteriorate due to complexity of customer interactions across multiple databases and applications

Engineering Contradiction:
Improvedata accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments customer data into distinct entity identifier clusters, where each cluster represents a unique customer entity. This segmentation allows the system to manage complex customer interactions by breaking down the overall data management task into smaller, manageable clusters that can be processed independently, thereby maintaining data accuracy without requiring a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces entity identifier clusters as intermediary structures between raw customer data from multiple sources and the final unified customer view. These clusters act as mediators that consolidate and validate data from various databases and applications, resolving inconsistencies and improving data reliability without directly managing the complexity of all source systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If customer data is collected from multiple data sources to provide comprehensive customer information, then data completeness improves, but data consistency deteriorates due to duplicate and inconsistent records

Engineering Contradiction:
Improvedata completenessVSAvoiddata consistency
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent merges customer data from multiple data sources by consolidating contact elements into unified entity identifier clusters. The system combines information from various sources while applying matching algorithms to identify and merge records representing the same customer, thereby achieving data completeness without sacrificing consistency. Duplicate records are detected and merged, ensuring a single unified view of each customer.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms through confidence scores that evaluate the quality and consistency of matched customer records. The system continuously refines its matching algorithms based on feedback from data quality assessments, allowing it to improve data consistency over time while maintaining comprehensive information from multiple sources. The feedback loop enables the system to learn from inconsistencies and adjust its merging strategy accordingly.

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If traditional match rules are used to evaluate customer records, then matching process is simple, but matching accuracy deteriorates because match rules evaluate record attributes without assessing individual identifier reliability

Engineering Contradiction:
Improvematching process simplicityVSAvoidmatching accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes the parameters of the matching process by introducing confidence scores that assess the reliability of individual contact elements (email addresses, phone numbers, etc.). Instead of using simple binary match/no-match rules, the system evaluates each contact element's reliability and incorporates this into the overall matching accuracy. This parameter change enables more precise matching while maintaining a relatively straightforward process flow.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality assessment by evaluating the reliability of individual contact elements within customer records rather than treating all attributes uniformly. Each contact element (email, phone, address) is assessed independently for quality and reliability, allowing the system to weight different attributes differently in the matching process. This local quality approach improves matching accuracy by focusing on the most reliable identifiers while still maintaining a simple overall matching framework.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10733212B2Entity identifier clustering based on context scores
Publication Date: 2020.08.04 SALESFORCE INC
  • US10733212B2 patent drawing
  • US10733212B2 patent drawing
  • US10733212B2 patent drawing

AI summary

A system receives entity data and other entity data, including an identification element and another identification element submitted by an entity for identifying the entity, and a contact element and another contact element submitted by the entity for contacting the entity, from the entity via a department and another department of an enterprise. The system generates scores for each of the contact element the other contact element, the scores being based on the contexts associated with the departments of the enterprise and the contact elements. The system stores an entity identifier cluster including the entity data. The system stores another entity cluster including the entity data and the other entity data, if a match exists between the contact element and the other contact element. The system outputs data stored by any entity identifier cluster that includes query-identified data, the output data being based on the scores.