Cross-Sentence Disease-Factor Relation Extraction With AI Graphs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for extracting disease-related factors from document data suffer from low accuracy and reliability due to independent sentence-based natural language processing, inability to recognize entities across multiple sentences, and limited extraction of new data beyond pre-defined dictionaries.
Innovation Solution
A system utilizing an AI deep learning technology to derive entities and relations between them from document data, considering the context of entire texts, and outputting integrated data in a graph form, using a pre-trained neural network model to recognize and link disease, gene, and protein-related terms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If natural language processing is performed independently on each sentence to extract relations between entities, then the processing complexity is reduced, but the accuracy and reliability of disease-related factor extraction deteriorates
Solution Approach 1:
The patent merges multiple sentences into a unified processing context, allowing the system to track and recognize disease-related factors across sentence boundaries. This combining approach enables the extraction of relations that span multiple sentences, thereby improving accuracy without proportionally increasing processing complexity through incremental analysis.
Solution Approach 2:
The system performs preliminary identification and tracking of disease-related factors as they appear in the text, maintaining their context and relationships throughout the document. This preliminary action allows for more accurate extraction later, as the system has already established the contextual framework across multiple sentences before performing detailed relation extraction.
2Reliability
If entity recognition is limited to pre-defined dictionary objects, then the system reliability is improved through consistent extraction, but the ability to recognize new entities deteriorates
Solution Approach 1:
The patent introduces an intermediary layer between the pre-defined dictionary and the text being analyzed. This intermediary component matches text entities against dictionary definitions while allowing for flexible interpretation and recognition of new entities that conform to the dictionary's structural patterns, thereby maintaining reliability through dictionary guidance while enabling adaptability for new entity discovery.
Solution Approach 2:
The system dynamically adjusts the matching parameters when comparing text entities against dictionary definitions. By modifying matching criteria based on contextual evidence and entity characteristics, the system can reliably recognize both known dictionary entities and new entities that share similar structural patterns, balancing consistency with innovation.
3Measurement precision
If the entire text context is considered for entity recognition, then the accuracy of relation extraction is improved, but the processing time increases
Solution Approach 1:
The patent segments the text processing into distinct phases: initial scanning for disease-related factors, contextual relationship establishment, and detailed relation extraction. This segmentation allows the system to process the entire text context systematically, improving accuracy by considering all sentences while managing processing time through structured, incremental analysis rather than simultaneous full-text processing.
Data Source
AI summary
The present invention relates to a method capable of recognizing disease-related factors included in document data and extracting relations between the disease-related factors.


