Automated Genomic Data Mining for Clinical Decision Support
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for interpreting genomic data from medical literature are tedious, time-consuming, and lack a widely accepted technique to efficiently extract clinically relevant information, hindering rapid diagnosis and treatment in personalized medicine.
Innovation Solution
A computer-implemented method for electronically mining genomic data from medical literature sources, involving data mining to determine disease-gene and disease-gene-mutation associations, prioritized by the strength of evidence from published articles, allowing for automated knowledge creation and efficient information retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual interpretation of genomic data from medical literature is performed, then accuracy and expertise in disease-gene associations are improved, but time consumption and labor intensity increase significantly
Solution Approach 1:
The patent replaces manual mechanical interpretation of genomic data with an automated computational system that uses natural language processing and data mining algorithms to extract disease-gene associations from medical literature, thereby eliminating time-consuming manual review while maintaining accuracy through systematic computational analysis
Solution Approach 2:
The system enables self-service by automatically performing data mining, text extraction, and association identification without requiring expert intervention, allowing the computational system to independently process genomic data and generate disease-gene associations through automated pipelines
2Quantity of substance
If comprehensive genomic data from multiple medical literature sources is analyzed, then completeness and coverage of disease-gene associations are improved, but system complexity and data processing requirements increase
Solution Approach 1:
The patent implements a universal data processing platform that can handle multiple medical literature sources, different data formats, and various types of genomic data simultaneously, allowing the same system architecture to process diverse inputs through standardized pipelines without requiring separate specialized systems for each data source
Solution Approach 2:
The system introduces intermediary components including natural language processing layers and data normalization modules that act as mediators between raw medical literature data and the core analysis engine, simplifying the processing of complex multi-source data by transforming it into standardized formats before further analysis
3Productivity
If automated data mining techniques are implemented, then productivity and speed of information retrieval are improved, but reliability and accuracy of disease-gene associations may deteriorate
Solution Approach 1:
The patent incorporates feedback mechanisms where the system continuously evaluates the quality and reliability of extracted associations, using confidence scoring and validation against known databases to filter results, thereby maintaining high reliability while operating at automated speeds through iterative refinement of extraction criteria
Solution Approach 2:
The system performs preliminary actions by pre-processing medical literature data through text normalization, entity recognition, and relationship extraction before main analysis, and by pre-validating extracted associations against established genomic databases, ensuring data quality and reliability are established before final interpretation occurs
Data Source
AI summary
A data analysis method and computer system electronically mines published articles from existing medical literature sources to discover associations that may exist between various diseases and various genes and/or gene mutations or other genetic changes. The method and system then organizes, categorizes and prioritizes the discovered associations in accordance with the strength of evidence supporting these associations. The resulting information can then be integrated into the processing of genome sequencing data to more quickly determine what genome sequencing data is of most relevance for clinical decision makings.


