Protein Interaction Analysis via Text Mining and Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in determining biological pathways involved in protein-protein interactions and identifying FDA-approved drugs that target these pathways, especially for diseases associated with genetic variants.
Innovation Solution
A processor-implemented method that receives information about a target protein, its variant, and the type of mutation associated with a disease. It generates search queries, gathers text descriptions, and uses neural networks to identify suggested drugs by analyzing protein-protein interactions and altered expression levels of other proteins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If text mining and neural network methods are used to identify protein-protein interactions, then the ability to discover biological pathways and drug targets is improved, but the complexity of the system increases
Solution Approach 1:
The patent introduces text descriptions from scientific literature as an intermediary medium to capture protein-protein interaction information. Instead of directly analyzing complex protein structures or conducting extensive experiments, the system mines text descriptions that already contain curated interaction data, thereby reducing the direct complexity of detection while improving the ability to identify interactions.
Solution Approach 2:
The patent replaces traditional experimental and manual analysis methods with computational and information-processing approaches. By using natural language processing, neural networks, and text mining algorithms, the system substitutes mechanical and experimental procedures with digital information processing, reducing physical complexity while enhancing interaction identification capabilities.
2Loss of information
If comprehensive text mining is performed to gather all relevant protein interaction descriptions, then the completeness of interaction data is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary filtering and selection of text descriptions based on relevance criteria before conducting detailed analysis. By pre-identifying and selecting only those text descriptions that are likely to contain relevant protein-protein interaction information, the system reduces the volume of data requiring intensive processing, thereby maintaining data completeness while reducing processing time and computational resources.
3Measurement precision
If the system analyzes expression levels of multiple proteins to identify therapeutic targets, then the accuracy of drug suggestions is improved, but the complexity of data analysis increases
Solution Approach 1:
The patent segments the complex task of drug target identification into distinct analytical components: (1) identifying protein-protein interactions from text, (2) analyzing expression levels of individual proteins, (3) determining altered expression patterns, and (4) suggesting drugs based on specific protein changes. This segmentation allows each component to be processed independently with specialized methods, improving overall accuracy while managing analysis complexity through modular processing.
Data Source
AI summary
A method includes: receiving information containing a name of a target protein associated with a disease, a name of a variant of the target protein, and a type of mutation associated with a disease; deriving, based on the information, a plurality of lists comprising: a first list containing protein names, a second list containing names of the genetic variant and protein post-translational modification type, a third list containing domain name, and a fourth list containing region name; generating search queries based on combinations of contents of the lists; gathering a plurality of text descriptions which satisfy the search queries, where the text descriptions include descriptions of relations of the target protein with other proteins; and identifying, based on processing the text descriptions, a suggested drug for treating the disease, where the suggested drug is associated with at least one of the other proteins.


