Node Importance Scoring Framework for Gene Network Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gene expression profiling techniques face challenges in differentiating meaningful genetic relationships from false discoveries, particularly in identifying targetable master regulators within large gene networks, due to high false discovery rates and limitations in node importance scoring frameworks.
Innovation Solution
A computational pipeline is developed to generate a reliable regulatory network from gene expression data using a bootstrap procedure to create consensus adjacency matrix files, determining significance thresholds, and filtering out low-significance edges, along with a node importance scoring framework (nSCORE) to iteratively refine combined node scores and identify key regulators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If gene expression profiling techniques are used to infer gene networks, then information on genetic relationships can be obtained, but false discovery rate increases
Solution Approach 1:
The patent applies preliminary action by performing a bootstrap procedure before final network inference. Multiple bootstrap samples are generated and processed through the network inference algorithm to establish a distribution baseline. This preliminary sampling and analysis allows the method to differentiate true genetic relationships from false discoveries by comparing results against the bootstrap distribution, thereby reducing false discovery rate while preserving information on genetic relationships.
2Productivity
If node importance scoring is applied to identify master regulators, then potential targets can be ranked, but scoring accuracy is limited
Solution Approach 1:
The patent implements feedback by using the bootstrap procedure to generate a distribution of node importance scores. The final scoring incorporates feedback from the bootstrap distribution to adjust and refine the importance scores. This feedback mechanism allows the system to account for variability and uncertainty in the data, improving scoring accuracy while maintaining the ability to rank genes effectively for identifying master regulators.
Data Source
AI summary
Example embodiments of the present invention address problems in computational biology and network theory. As mentioned above, example embodiments enable the determination of where to set a threshold cutoff level to maximize sensitivity of a gene expression profile while minimizing the false discovery rate (FDR). Some example embodiments exploit the filtered gene expression profile to discover potential master regulators that may be targetable by various perturbagen-based treatments. Other example embodiments facilitate the harvesting of meaningful data from other type of network and node statistics inputs, even those not related specifically to computational biology or the exploitation of inferred gene expression profiles.


