Bio-entity Relationship Extraction Using NLP and Graph Algorithms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting bio-entity relationships from literature are inefficient, prone to errors, and costly due to manual annotation, and existing automated methods have limited accuracy and coverage, especially with large datasets.
Innovation Solution
A computer-implemented software application using natural language processing and graph theoretic algorithms to extract textual relationships by building a decision support tool with multiple levels of decision nodes, allowing for the classification of triplets as true or false based on probability values and handling synonyms, and storing extracted relationships in a structured format.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation is used to extract bio-entity relationships from literature, then accuracy of extraction can be maintained, but time consumption and cost increase significantly
Solution Approach 1:
The patent introduces an automated text mining system as an intermediary between the literature corpus and the final extracted relationships. This system uses natural language processing, pattern matching, and machine learning algorithms to automatically extract bio-entity relationships, serving as a mediator that reduces the need for time-consuming manual annotation while maintaining reasonable extraction accuracy
Solution Approach 2:
The patent creates structured representations (copies) of the unstructured literature text through standardized formats and schemas. By transforming the original text into structured data with defined relationships and entities, the system enables automated processing and analysis without requiring manual annotation of each relationship instance
2Productivity
If automated text mining methods are used to extract relationships from literature, then efficiency and productivity improve, but accuracy and precision decrease due to false positives and limited coverage
Solution Approach 1:
The patent segments the text mining process into multiple independent modules including entity recognition, relationship extraction, pattern matching, and validation stages. Each module focuses on specific aspects of relationship extraction and can be independently optimized, allowing the system to process large volumes of text efficiently while maintaining precision through staged filtering and validation
Solution Approach 2:
The patent employs machine learning models that automatically learn and adapt extraction parameters from training data. The system adjusts sensitivity thresholds, pattern weights, and classification parameters based on performance metrics, dynamically optimizing the balance between recall (coverage) and precision (accuracy) for different extraction scenarios
3Measurement precision
If manually specified rules are used for relationship extraction, then precision can be improved by filtering false positives, but coverage decreases due to limited rule sets
Solution Approach 1:
The patent develops a universal extraction framework that combines multiple extraction strategies (rule-based, pattern-based, and machine learning-based) into a single system. This multi-functional approach allows the system to handle diverse relationship types and linguistic patterns across different biological domains, achieving both high precision through rule filtering and broad coverage through adaptive pattern recognition
Solution Approach 2:
The patent implements a dynamic rule system where extraction rules and patterns are not fixed but can be automatically generated, updated, and refined based on training data and performance feedback. The system adapts its rule set to cover new relationship types and linguistic variations while maintaining precision through continuous validation and optimization
Data Source
AI summary
Automated, standardized and accurate extraction of relationships within text. Automatic extraction of such relationships/information allows the information to be stored in structured form so that it can be easily and accurately retrieved when needed. Such information can be used to build online search engines for highly specific and accurate information retrieval. The current invention discloses a novel approach to extract such information from raw text based on natural language processing (NLP) and graph theoretic algorithm. The novel method can be applied, for example, to extract protein-protein relationships in biomedical literature. The method can be easily extended to extract other biological relationships between biological terms such as proteins, genes, pathways, diseases and drugs. The method can also be applied to other information domains to extract other relationships.


