Phylogenetic Analysis Platform for Proteomic Data Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for classifying specimens based on proteomic analyses are limited by the lack of universal, broadly acceptable methods, leading to poor comparability across different locations and limited clinical relevance, especially in cancer diagnosis and typing, due to reliance on cancer type-specific sorting algorithms with low specificity and failure to utilize all data variability.
Innovation Solution
The development of a universal phylogenetic analysis platform, phyloproteomics and phyloarray, which uses a new data-mining parsing algorithm (UNIPAL) and a publicly available phylogenetic algorithm (MIX) to analyze mass spectrometry and gene expression data, allowing for biologically meaningful groupings of specimens by identifying derived and ancestral states, enabling high clinical relevance and interplatform comparability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If phenetic approaches such as clustering are used to classify specimens based on overall similarity, then classification can be performed, but comparability of proteomic analyses performed in diverse locations becomes unattainable and diagnostic specificity is reduced
Solution Approach 1:
The patent transforms the analysis from phenetic (overall similarity) to phylogenetic (evolutionary relationships) by changing the fundamental parameter of classification. It uses derived versus ancestral character states to create a classification system that is both comparable across locations and specific for diagnosis, resolving the contradiction between adaptability and reliability
Solution Approach 2:
The patent introduces an intermediary outgroup (normal specimens) that mediates the comparison between test specimens. By comparing all specimens against this common reference framework, it enables comparability across different locations and platforms while maintaining diagnostic specificity through the derived/ancestral classification
2Adaptability or versatility
If cancer type-specific sorting algorithms are used, then classification can be achieved for specific cancer types, but the algorithms do not apply well across other cancer types and specificity remains below 95%
Solution Approach 1:
The patent creates a universal phylogenetic classification system that can be applied across all cancer types using the same derived/ancestral framework. The outgroup comparison method serves multiple functions: it classifies different cancer types, identifies transitional cases, and provides a common reference framework, achieving both broad applicability and high specificity
Solution Approach 2:
The patent changes the classification parameter from cancer-type-specific patterns to evolutionary derived/ancestral states. This parameter transformation enables a single algorithm to accurately classify multiple cancer types while maintaining above 95% specificity through the phylogenetic relationships revealed by the analysis
3Productivity
If statistical methods are used to analyze proteomic data, then analysis can be performed, but not all potentially useful variability within the data is utilized
Solution Approach 1:
The patent performs preliminary outgroup comparison to establish derived versus ancestral states for all data points before final classification. This preliminary action identifies and preserves all useful variability in the data, including transitional cases that statistical methods often discard, maximizing data utilization while maintaining information integrity
Data Source
AI summary
A universal data-mining platform is provided capable of analyzing mass spectrometry (MS) serum proteomic profiles and/or gene array data to produce biologically meaningful classification; i.e., group together biologically related specimens into clades. This platform utilizes the principles of phylogenetics, such as parsimony, to reveal susceptibility to cancer development (or other physiological or pathophysiological conditions), diagnosis and typing of cancer, identifying stages of cancer, as well as post-treatment evaluation. By outgroup comparison, the parsing algorithm identifies under and/or overexpressed gene values or in the case of sera, (i) novel or (ii) vanished MS peaks, and peaks signifying (iii) up or (iv) down regulated proteins, and scores the variations as either derived (do not exist in the outgroup set) or ancestral (exist in the outgroup set); the derived is given a score of “1”, and the ancestral a score of “0”—these are called the polarized values.


