Bio-entity Relationship Extraction Using NLP and Graph Algorithms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting bio-entity relationships from literature are inefficient, prone to errors, and costly due to manual annotation, and existing automated methods have limited accuracy and coverage, especially with large datasets.

Innovation Solution

A computer-implemented software application using natural language processing and graph theoretic algorithms to extract textual relationships by building a decision support tool with multiple levels of decision nodes, allowing for the classification of triplets as true or false based on probability values and handling synonyms, and storing extracted relationships in a structured format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation is used to extract bio-entity relationships from literature, then accuracy of extraction can be maintained, but time consumption and cost increase significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces an automated text mining system as an intermediary between the literature corpus and the final extracted relationships. This system uses natural language processing, pattern matching, and machine learning algorithms to automatically extract bio-entity relationships, serving as a mediator that reduces the need for time-consuming manual annotation while maintaining reasonable extraction accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates structured representations (copies) of the unstructured literature text through standardized formats and schemas. By transforming the original text into structured data with defined relationships and entities, the system enables automated processing and analysis without requiring manual annotation of each relationship instance

Inventive Principle:
Principle #26Copying

2Productivity

If automated text mining methods are used to extract relationships from literature, then efficiency and productivity improve, but accuracy and precision decrease due to false positives and limited coverage

Engineering Contradiction:
Improveextraction efficiencyVSAvoidextraction precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the text mining process into multiple independent modules including entity recognition, relationship extraction, pattern matching, and validation stages. Each module focuses on specific aspects of relationship extraction and can be independently optimized, allowing the system to process large volumes of text efficiently while maintaining precision through staged filtering and validation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs machine learning models that automatically learn and adapt extraction parameters from training data. The system adjusts sensitivity thresholds, pattern weights, and classification parameters based on performance metrics, dynamically optimizing the balance between recall (coverage) and precision (accuracy) for different extraction scenarios

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manually specified rules are used for relationship extraction, then precision can be improved by filtering false positives, but coverage decreases due to limited rule sets

Engineering Contradiction:
Improveextraction precisionVSAvoidextraction coverage
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent develops a universal extraction framework that combines multiple extraction strategies (rule-based, pattern-based, and machine learning-based) into a single system. This multi-functional approach allows the system to handle diverse relationship types and linguistic patterns across different biological domains, achieving both high precision through rule filtering and broad coverage through adaptive pattern recognition

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a dynamic rule system where extraction rules and patterns are not fixed but can be automatically generated, updated, and refined based on training data and performance feedback. The system adapts its rule set to cover new relationship types and linguistic variations while maintaining precision through continuous validation and optimization

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8886522B2Automated extraction of bio-entity relationships from literature
Publication Date: 2014.11.11 FLORIDA STATE UNIV RES FOUND INC
  • US8886522B2 patent drawing
  • US8886522B2 patent drawing
  • US8886522B2 patent drawing

AI summary

Automated, standardized and accurate extraction of relationships within text. Automatic extraction of such relationships/information allows the information to be stored in structured form so that it can be easily and accurately retrieved when needed. Such information can be used to build online search engines for highly specific and accurate information retrieval. The current invention discloses a novel approach to extract such information from raw text based on natural language processing (NLP) and graph theoretic algorithm. The novel method can be applied, for example, to extract protein-protein relationships in biomedical literature. The method can be easily extended to extract other biological relationships between biological terms such as proteins, genes, pathways, diseases and drugs. The method can also be applied to other information domains to extract other relationships.