Coherent Voting Network for Breast Cancer Survival Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current molecular prognostic tests for breast cancer, such as Mammaprint and Oncotype DX, suffer from high false positive rates, necessitating a more efficient method to predict long-term survival by selecting a reduced panel of biomarkers that can be measured at a lower cost.
Innovation Solution
A method is developed to create a coherent voting network using a workstation with algorithms and statistical tests to select a panel of biomarkers from 24,000 mRNA molecules, reducing the number to a few dozen, which can predict survival after five years post-tumor removal, by utilizing data from the METABRIC consortium and applying graph-theoretic methods to identify key gene expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large panel of molecular biomarkers (about 25,000 mRNA) is used for prediction, then the predictive accuracy is improved, but the measurement cost and complexity increase significantly
Solution Approach 1:
The patent extracts and selects a reduced panel of about 50-100 most informative biomarkers from the initial 25,000 mRNA molecules. This extraction is achieved through a multi-step filtering process that evaluates biomarker importance using statistical tests and graph-theoretic methods, ultimately identifying a subset that maintains predictive accuracy while significantly reducing measurement complexity and cost.
Solution Approach 2:
The patent segments the large set of biomarkers into a manageable panel by dividing the selection process into distinct stages: initial filtering based on statistical significance, intermediate selection using graph theory to identify key nodes in the biomarker network, and final validation. This segmentation allows systematic reduction from 25,000 to approximately 50-100 biomarkers while preserving predictive power.
2Ease of manufacture
If existing molecular prognostic tests (Mammaprint, Oncotype DX) are used, then the prediction process is standardized, but the false positive rate remains high
Solution Approach 1:
The patent implements feedback mechanisms through iterative validation processes where the selected biomarker panel is continuously tested against patient outcomes. The graph-theoretic approach allows dynamic adjustment of biomarker selection based on predictive performance metrics, enabling the system to learn from results and refine its predictions to reduce false positives while maintaining standardization.
Solution Approach 2:
The patent changes the selection parameters and weighting factors used in evaluating biomarkers. By using graph-theoretic methods to identify key nodes and adjusting statistical thresholds, the system optimizes the balance between standardization and accuracy, reducing false positives compared to conventional tests while maintaining a standardized prediction framework.
3Device complexity
If a reduced panel of biomarkers (few dozen mRNA) is selected, then the measurement cost is reduced, but the predictive capability may be compromised
Solution Approach 1:
The patent performs preliminary actions through extensive pre-processing and feature selection using the METABRIC dataset before final biomarker selection. This preliminary work includes statistical filtering, graph construction, and identification of critical biomarker interactions, ensuring that the reduced panel of 50-100 biomarkers contains the most predictive elements, thereby maintaining high predictive capability despite the reduced number of markers.
Solution Approach 2:
The patent creates a composite predictive model that integrates multiple types of information about biomarkers (expression levels, statistical significance, graph network positions, interaction patterns) into a unified prediction system. This composite approach allows the reduced panel to achieve high predictive accuracy by synthesizing multiple data dimensions, effectively compensating for the reduced number of individual biomarkers measured.
Data Source
Figure 1
Figure 2
AI summary
Method for creating a coherent voting network including the steps of: a) organizing (1) predetermined data from a cohort of patients in a master matrix having in its row the list of patients and having in its column the expression value of a panel of genes; b) applying (2) a predetermined statistical test to each gene to evaluate which genes better discriminates survival or not-survival classes for patients, thus obtaining a first candidate panel of genes; c) discretizing (4) the expression value of each of the genes belonging to said first candidate panel; d) converting (6) the quantized master matrix in a first bipartite graph (G); e) applying (8) a predetermined algorithm to the bipartite graph (G) to obtain a collection of bipartite communities; f) applying (10) a predetermined algorithm to the communities, thus obtaining a second candidate panel; g) creating (12) an updated bipartite graph (G') comprising nodes-patients and nodes-gene whose genes belongs to the second candidate panel; h) repeating (14) step e) on the updated bipartite graph (G') ; i) applying (16), at each community, a decision function on each patient belonging to it, to determine whether to assign her at the survival or not-survival class,; j) checking (18), for each patient belonging to the various communities of step i), whether the class assigned is the same as the class of the master matrix; k) checking (20) whether the percentage of coherent patients in the voting network is greater than a predetermined threshold, thus obtaining a coherent coting network.