Cross-Sentence Disease-Factor Relation Extraction With AI Graphs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for extracting disease-related factors from document data suffer from low accuracy and reliability due to independent sentence-based natural language processing, inability to recognize entities across multiple sentences, and limited extraction of new data beyond pre-defined dictionaries.

Innovation Solution

A system utilizing an AI deep learning technology to derive entities and relations between them from document data, considering the context of entire texts, and outputting integrated data in a graph form, using a pre-trained neural network model to recognize and link disease, gene, and protein-related terms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If natural language processing is performed independently on each sentence to extract relations between entities, then the processing complexity is reduced, but the accuracy and reliability of disease-related factor extraction deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidextraction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple sentences into a unified processing context, allowing the system to track and recognize disease-related factors across sentence boundaries. This combining approach enables the extraction of relations that span multiple sentences, thereby improving accuracy without proportionally increasing processing complexity through incremental analysis.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system performs preliminary identification and tracking of disease-related factors as they appear in the text, maintaining their context and relationships throughout the document. This preliminary action allows for more accurate extraction later, as the system has already established the contextual framework across multiple sentences before performing detailed relation extraction.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If entity recognition is limited to pre-defined dictionary objects, then the system reliability is improved through consistent extraction, but the ability to recognize new entities deteriorates

Engineering Contradiction:
Improvesystem reliabilityVSAvoidnew entity recognition capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary layer between the pre-defined dictionary and the text being analyzed. This intermediary component matches text entities against dictionary definitions while allowing for flexible interpretation and recognition of new entities that conform to the dictionary's structural patterns, thereby maintaining reliability through dictionary guidance while enabling adaptability for new entity discovery.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically adjusts the matching parameters when comparing text entities against dictionary definitions. By modifying matching criteria based on contextual evidence and entity characteristics, the system can reliably recognize both known dictionary entities and new entities that share similar structural patterns, balancing consistency with innovation.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If the entire text context is considered for entity recognition, then the accuracy of relation extraction is improved, but the processing time increases

Engineering Contradiction:
Improverelation extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the text processing into distinct phases: initial scanning for disease-related factors, contextual relationship establishment, and detailed relation extraction. This segmentation allows the system to process the entire text context systematically, improving accuracy by considering all sentences while managing processing time through structured, incremental analysis rather than simultaneous full-text processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12417849B2Method for identifying association between disease-related factors from document data, and system constructed using same
Publication Date: 2025.09.16 STANDIGM
  • US12417849B2 patent drawing
  • US12417849B2 patent drawing
  • US12417849B2 patent drawing

AI summary

The present invention relates to a method capable of recognizing disease-related factors included in document data and extracting relations between the disease-related factors.