Synonymous Entity Merging via Clustering Algorithm

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data processing systems face inaccuracies and inefficiencies due to the lack of global identifiers in life sciences, leading to confusion and incorrect query results from multiple synonyms for the same entity across databases.

Innovation Solution

A clustering statistical algorithm is used to identify and merge synonymous terms from multiple structured sources by comparing paired terms from authoritative sources, generating a dataset with normalized versions of synonymous terms, thereby improving data aggregation and query efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple data sources with synonymous terms are aggregated, then data completeness and coverage are improved, but data accuracy and query reliability deteriorate due to confusion from multiple names for the same entity

Engineering Contradiction:
Improvedata coverageVSAvoidquery accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent merges synonymous entities from multiple data sources by identifying entities with the same semantic meaning but different names. It combines records for the same entity across different sources while eliminating duplicates, thereby maintaining data completeness while improving query accuracy through unified entity representation.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary entity resolution system that acts as a mediator between multiple data sources. This system uses entity identifiers and semantic analysis to bridge different data sources, matching synonymous terms and enabling accurate cross-source queries without direct integration of the original heterogeneous sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If synonymous terms from multiple sources are integrated without normalization, then data diversity and source coverage are improved, but data consistency and processing efficiency worsen

Engineering Contradiction:
Improvesource compatibilityVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent creates a universal entity representation system that can handle multiple data sources with different naming conventions. By establishing a standardized entity model that works across diverse sources, it enables consistent processing of heterogeneous data while maintaining compatibility with various source formats and structures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent transforms synonymous terms into normalized entity representations by changing the parameter of term representation. It converts various names and formats into a standardized form using entity identifiers and canonical names, thereby improving processing efficiency while preserving the ability to handle diverse source data.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If entity normalization is implemented across multiple data sources, then query accuracy and data consistency are improved, but system complexity and implementation difficulty increase

Engineering Contradiction:
Improveentity identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary entity resolution and normalization during data ingestion and integration phases. By pre-processing data to establish entity identifiers and canonical names before queries are executed, it reduces the complexity of query processing while maintaining high entity identification accuracy throughout the system.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the entity resolution process into distinct modules: entity identification, synonym matching, canonical name assignment, and identifier generation. This modular approach reduces overall system complexity by breaking down the complex normalization task into manageable, independently implementable components.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10671577B2Merging synonymous entities from multiple structured sources into a dataset
Publication Date: 2020.06.02 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10671577B2 patent drawing
  • US10671577B2 patent drawing
  • US10671577B2 patent drawing

AI summary

Merging synonymous entities from multiple structured sources into a dataset includes receiving a first set of paired terms from a first authoritative source for a domain and a second set of paired terms from a second authoritative source for the domain. The first set of paired terms is compared to the second set of paired terms with a similarity assessment based on a clustering statistical algorithm to identify paired terms from the first set of paired terms that share a synonymous term with one or more paired terms from the second set of paired terms. The paired terms associated with the synonymous term are merged and a dataset is generated that associates a normalized version of the synonymous term with any terms included in the merged paired terms.