Cluster-Based Lexicon for Cross-Domain Entity Relationship Deduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cognitive question answering systems face difficulties in efficiently identifying and recognizing entity relationships across different industry domains due to conflicting named entity extraction results, as they lack contextualization and suitable methods for handling disparate industry domain dictionaries.
Innovation Solution
A system and method that utilize a cluster-based dictionary vocabulary lexicon with weighted or scored relationships, performing natural language processing and semantic analysis to identify and rank entity relationships across multiple knowledge databases, allowing for the construction of models that specify relationships between different industry domains with minimal human supervision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional named entity recognition processes are used in different industry domains, then entity extraction can be performed for each domain, but conflicting extraction results occur and entity relationships across domains cannot be effectively identified
Solution Approach 1:
The patent creates a universal entity relationship model that works across multiple industry domains. The system extracts entities and relationships from different domains (e.g., healthcare, finance) using a unified approach, allowing the same model to handle diverse domains without domain-specific customization. This resolves the contradiction by making the system adaptable to multiple domains while maintaining consistent extraction accuracy through standardized relationship types and scoring mechanisms.
Solution Approach 2:
The patent introduces an intermediary layer of relationship scoring and normalization that mediates between domain-specific entity extractions and cross-domain relationship identification. The relationship score calculation and threshold filtering act as intermediaries that harmonize conflicting extractions from different domains, enabling effective cross-domain entity relationship recognition while preserving domain-specific extraction accuracy.
2Adaptability or versatility
If existing solutions attempt to identify entity relationships across different industry domain dictionaries, then comprehensive coverage can be achieved, but the process becomes extremely difficult and inefficient at practical levels
Solution Approach 1:
The patent segments the complex cross-domain entity relationship identification process into manageable components: entity extraction, relationship extraction, relationship scoring, and threshold filtering. This segmentation allows each component to be optimized independently and processed efficiently at scale, resolving the contradiction by making the comprehensive cross-domain process practical and productive through systematic breakdown of operations.
Solution Approach 2:
The patent changes parameters by introducing relationship scores and threshold values that quantify and filter relationships across domains. This parameter-based approach transforms the qualitative, difficult process of cross-domain relationship identification into a quantitative, efficient process where relationships are automatically scored and filtered, dramatically improving productivity while maintaining comprehensive cross-domain coverage.
3Productivity
If named entity extraction is performed without contextualization, then extraction speed can be maintained, but extraction results conflict across different industry domains
Solution Approach 1:
The patent performs preliminary contextualization by extracting and analyzing the context surrounding entities before final relationship determination. The system examines contextual information (surrounding text, document type, domain-specific patterns) in advance to disambiguate entities and relationships, ensuring consistent extraction across domains while maintaining speed through pre-computed contextual features and efficient processing pipelines.
Data Source
AI summary
An approach is provided for identifying entity relationships based on word classifications extracted from business documents stored in a plurality of corpora. In the approach, performed by an information handling system, a plurality of cluster classifications are identified for the business documents so that entity information from the business documents can be classified or assigned to the cluster classifications, such as by performing natural language processing (NLP) analysis of the business documents. The approach applies semantic analysis to identify and score entity relationships between the entity information classified in the cluster classifications, and based on the scored entity relationships, cluster relationships between the cluster classifications are identified.


