Gene Expression Data Indexing for Alternative Indication Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches to discovering alternative indications for pharmaceuticals often exclude gene expression data and fail to cross-reference vast amounts of data from siloed environments, limiting the identification of additional disease treatments.
Innovation Solution
A method and system that utilize gene expression information and medical literature to identify alternative indications by converting normalized gene expression values to binary expressions, generating indexable documents, and associating disease names with expressed genes based on correlation thresholds, facilitating cross-referencing of genomic data with medical text data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current approaches to discover treatments are used (studying protein and tumor properties), then the process is simple and focused, but the identification of alternative indications is limited and less accurate
Solution Approach 1:
The patent merges gene expression data with medical literature data by creating indexable documents that combine genomic probe data with associated disease information. This integration allows the system to cross-reference vast amounts of data from siloed environments (genomic databases and medical literature), enabling more accurate identification of alternative indications while maintaining a unified processing framework.
Solution Approach 2:
The patent transforms normalized gene expression values into binary expressions (expressed/unexpressed), adding a dimensional transformation that simplifies the data structure for correlation analysis. This binary conversion enables straightforward association with disease names through correlation thresholds, enhancing measurement precision without proportionally increasing system complexity.
2Adaptability or versatility
If gene expression data is excluded from treatment discovery approaches, then the data processing system remains simple, but the identification of alternative indications is limited
Solution Approach 1:
The patent performs preliminary processing of gene expression data by normalizing values and converting them to binary expressions before association with disease names. This preliminary action prepares the data in advance for correlation analysis, ensuring that genomic information is fully utilized to expand the range of identified treatable diseases without creating complex real-time processing requirements.
Solution Approach 2:
The patent creates a universal processing framework that handles both gene expression data and medical literature data through the same indexable document structure. This multi-functional approach allows the system to process diverse data types uniformly, maximizing the utilization of genomic data while maintaining system simplicity through standardized processing procedures.
3Productivity
If data from siloed environments is not cross-referenced, then the system complexity is low, but the discovery of alternative indications is restricted
Solution Approach 1:
The patent segments the data processing into distinct modules: gene expression data processing, medical literature data processing, and correlation analysis. Each module processes specific data types independently before integrating results through correlation thresholds. This segmentation enables efficient cross-referencing of data from siloed environments while maintaining manageable system complexity through modular architecture.
Solution Approach 2:
The patent introduces indexable documents as intermediary structures that bridge gene expression data and medical literature data. These documents serve as mediators that standardize the format and enable efficient cross-referencing between different data sources, accelerating the discovery of alternative indications through standardized intermediate representations that facilitate rapid correlation analysis.
Data Source
AI summary
The present invention relates to a method and system for associating gene expression data with a disease name. A first data set associated with a plurality of genetic probes for a plurality of biological samples may be received. The first data set may be sorted based on a normalized gene expression values for the plurality of genetic probes. A largest value gap of the normalized gene expression values may be identified. A set of expressed genes within the first data set may be identified. An indexable document may be generated for a biological sample of the plurality of biological samples comprising data associated with the set of expressed genes. A second data set associated with an expressed gene of the set of expressed genes may be searched. A disease name may be associated with an expressed gene based on a threshold correlation between the disease name and the expressed gene.


