Entity Type Discovery in Master Data Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Master Data Management (MDM) systems are limited in their ability to link records into multiple related entity types, failing to support the notion of linking a record into multiple entities and thus cannot view the same record in multiple ways, which is necessary for various business applications.
Innovation Solution
A method and apparatus for discovering entity types by inputting records with associated attribute values, grouping them based on attribute values, domain ontologies, and dimension hierarchies, calculating an interestingness measure, and validating candidate entity types using correlation, query logs, and average group size to support multiple entity types and their relationships.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional MDM systems link each record into a single entity, then the system structure is simple and easy to manage, but the system cannot support multiple entity types and relationships between them
Solution Approach 1:
The patent segments the entity model into multiple independent entity types (e.g., Person, Household, Organization) that can coexist in the same system. Each entity type is treated as a separate segment with its own attributes and relationships, allowing the system to support multiple entity types without creating an unmanageable monolithic structure.
Solution Approach 2:
The patent introduces a new dimension to the traditional single-entity model by adding entity type classification. Instead of a flat one-to-one mapping between records and entities, the system creates a multi-dimensional structure where records can belong to multiple entity types simultaneously, enabling versatile entity relationships while maintaining structured organization.
2Adaptability or versatility
If the system discovers and validates multiple entity types with relationships, then the adaptability and versatility improve, but the computational complexity and processing time increase
Solution Approach 1:
The system performs preliminary actions by pre-defining domain ontologies and dimension hierarchies that guide entity type discovery. These pre-established frameworks provide templates and constraints that accelerate the entity type validation process, reducing the time required to analyze and validate multiple entity types while maintaining accuracy.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously refines entity type discoveries based on validation results. The feedback loop allows the system to learn from previous validations, adjusting its discovery process to reduce redundant computations and minimize processing time for subsequent entity type analyses.
Data Source
AI summary
Methods and arrangements for discovering entity types for a set of records. A set of records is input, with each record comprising attributes with associated attribute values. The records are grouped into candidate entity types in view of at least one of: the attribute values of the records, at least one domain ontology and at least one dimension hierarchy. An interestingness measure of each candidate entity type is calculated, via estimating interestingness based on at least one factor selected from the group consisting of: a correlation between attribute values of records, a number of attributes, a log of queries issued to a server, and an average group size for candidate entity types. At least one candidate entity type is validated based on the calculated interestingness measures. Other variants and embodiments are broadly contemplated herein.


