Group-Sparse Non-Negative CCA for Prostate Cancer Feature Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for selecting discriminative features from multiple feature views in disease prognosis and diagnosis, such as SMVCCA, are sub-optimal due to negatively correlated latent components, increased complexity, and neglect of modality-specific information, leading to reduced accuracy and interpretability.
Innovation Solution
The implementation of Group-Sparse Non-Negative Supervised Canonical Correlation Analysis (GNCCA) with a Variable Importance in Projections (VIP) score, which incorporates non-negativity and group-sparsity constraints to ensure positive correlations and simultaneous view association, enabling more accurate and efficient selection of discriminative features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional SMVCCA methods are used to select discriminative features from multiple feature views, then feature selection is performed, but negatively correlated latent components reduce accuracy and interpretability
Solution Approach 1:
The patent changes the parameter constraints of the canonical correlation analysis by imposing non-negativity constraints on the loading matrices. This transforms the conventional CCA formulation into a non-negative CCA (NN-CCA) framework, which ensures that all latent components are non-negative and eliminates negative correlations, thereby improving both accuracy and interpretability of the selected features.
2Quantity of substance
If pre-feature selection step is applied to reduce redundant features, then feature redundancy is reduced, but system complexity and computation time increase
Solution Approach 1:
The patent merges the feature selection step with the canonical correlation analysis by incorporating sparsity-inducing regularization terms directly into the CCA objective function. This integration eliminates the need for separate pre-feature selection steps, reducing system complexity while maintaining the ability to select discriminative features from multiple views simultaneously.
Solution Approach 2:
The patent applies sparsity constraints as preliminary conditions in the optimization formulation, which automatically performs feature selection during the CCA computation process. This preliminary action is embedded in the mathematical formulation through L1 regularization or similar sparsity-inducing penalties, allowing the system to select features without requiring additional preprocessing steps.
3Loss of information
If conventional CCA methods are used to fuse features from multiple modalities, then correlated metaspace is found, but modality-specific information is neglected due to bias towards modalities with greater number of features
Solution Approach 1:
The patent applies local quality by allowing different regularization parameters and constraints for each modality-specific loading matrix. This enables the model to adaptively weigh the importance of each modality based on its own characteristics rather than being biased by the number of features. Each modality can maintain its unique information while contributing to the shared latent space.
Data Source
AI summary
Methods, apparatus, application specific integrated circuits (ASIC)s and other embodiments associated with analyzing a cancerous prostate using group-sparse non-negative canonical correlation analysis (GNCCA) with a variable importance in the projections (VIP) score are described. One example apparatus includes a set of logics that acquires a set of features from a plurality of feature views of a region of tissue demonstrating cancerous pathology, produces a ranked set of discriminative features using GNCCA with the VIP score, optimizes computation of the GNCCA using a vector-block coordinate descent (BCD) approach, and provides a prostate cancer (CaP) grade or a biochemical recurrence (BcR) score based on the set of discriminative features. Embodiments of example apparatus may generate and display the CaP grade, BcR score, or set of discriminative features.


