Cluster-Based Lexicon for Cross-Domain Entity Relationship Deduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing cognitive question answering systems face difficulties in efficiently identifying and recognizing entity relationships across different industry domains due to conflicting named entity extraction results, as they lack contextualization and suitable methods for handling disparate industry domain dictionaries.

Innovation Solution

A system and method that utilize a cluster-based dictionary vocabulary lexicon with weighted or scored relationships, performing natural language processing and semantic analysis to identify and rank entity relationships across multiple knowledge databases, allowing for the construction of models that specify relationships between different industry domains with minimal human supervision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional named entity recognition processes are used in different industry domains, then entity extraction can be performed for each domain, but conflicting extraction results occur and entity relationships across domains cannot be effectively identified

Engineering Contradiction:
Improveentity extraction accuracyVSAvoidcross-domain entity relationship recognition
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal entity relationship model that works across multiple industry domains. The system extracts entities and relationships from different domains (e.g., healthcare, finance) using a unified approach, allowing the same model to handle diverse domains without domain-specific customization. This resolves the contradiction by making the system adaptable to multiple domains while maintaining consistent extraction accuracy through standardized relationship types and scoring mechanisms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary layer of relationship scoring and normalization that mediates between domain-specific entity extractions and cross-domain relationship identification. The relationship score calculation and threshold filtering act as intermediaries that harmonize conflicting extractions from different domains, enabling effective cross-domain entity relationship recognition while preserving domain-specific extraction accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If existing solutions attempt to identify entity relationships across different industry domain dictionaries, then comprehensive coverage can be achieved, but the process becomes extremely difficult and inefficient at practical levels

Engineering Contradiction:
Improvecross-domain entity relationship identificationVSAvoidrelationship identification efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the complex cross-domain entity relationship identification process into manageable components: entity extraction, relationship extraction, relationship scoring, and threshold filtering. This segmentation allows each component to be optimized independently and processed efficiently at scale, resolving the contradiction by making the comprehensive cross-domain process practical and productive through systematic breakdown of operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes parameters by introducing relationship scores and threshold values that quantify and filter relationships across domains. This parameter-based approach transforms the qualitative, difficult process of cross-domain relationship identification into a quantitative, efficient process where relationships are automatically scored and filtered, dramatically improving productivity while maintaining comprehensive cross-domain coverage.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If named entity extraction is performed without contextualization, then extraction speed can be maintained, but extraction results conflict across different industry domains

Engineering Contradiction:
Improveentity extraction speedVSAvoidentity extraction consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary contextualization by extracting and analyzing the context surrounding entities before final relationship determination. The system examines contextual information (surrounding text, document type, domain-specific patterns) in advance to disambiguate entities and relationships, ensuring consistent extraction across domains while maintaining speed through pre-computed contextual features and efficient processing pipelines.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10664505B2Method for deducing entity relationships across corpora using cluster based dictionary vocabulary lexicon
Publication Date: 2020.05.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10664505B2 patent drawing
  • US10664505B2 patent drawing
  • US10664505B2 patent drawing

AI summary

An approach is provided for identifying entity relationships based on word classifications extracted from business documents stored in a plurality of corpora. In the approach, performed by an information handling system, a plurality of cluster classifications are identified for the business documents so that entity information from the business documents can be classified or assigned to the cluster classifications, such as by performing natural language processing (NLP) analysis of the business documents. The approach applies semantic analysis to identify and score entity relationships between the entity information classified in the cluster classifications, and based on the scored entity relationships, cluster relationships between the cluster classifications are identified.