Phylogenetic Tree Creation Using Mutation Group Frequency Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for creating phylogenetic trees from cancer samples with mixed aggregations of cancer cells and normal cells are inaccurate due to the difficulty in obtaining large numbers of mutation sequences and calculating mixture ratios, especially when DNA sequences of whole genomes are unavailable.
Innovation Solution
A phylogenetic tree creation apparatus that groups mutations by frequency and uses a parent-child relation determination section to create a tree structure based on correlation coefficients, allowing for accurate representation of evolutionary relations among clones despite limited sequence data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If next generation sequencing is used to obtain vast amounts of sequence data, then the quantity of sequence data is improved, but the cost and complexity of analyzing mixed aggregations of cancer cells worsens
Solution Approach 1:
The patent segments the complex analysis problem into distinct modules: a graph creation section that builds parent-child graphs from mutation group frequency data, and a parent-child relation determination section that evaluates correlations. This segmentation allows the system to handle vast amounts of sequence data by processing it through specialized, manageable components rather than attempting comprehensive analysis all at once.
Solution Approach 2:
The patent introduces mutation group frequency data as an intermediary representation that bridges the gap between raw sequence data and phylogenetic tree construction. By grouping mutations and calculating their frequencies across samples, the system creates a simplified intermediate dataset that captures essential evolutionary information while reducing the complexity of the original vast sequence data.
2Measurement precision
If mutation group frequency data is used to create parent-child graphs, then the accuracy of phylogenetic trees is improved, but the computational requirements worsen
Solution Approach 1:
The patent performs preliminary actions by pre-calculating mutation group frequencies and constructing parent-child graphs before the actual phylogenetic tree evaluation. This preliminary processing organizes the data into a structured format that facilitates more efficient subsequent analysis, reducing the computational burden during the tree construction phase while maintaining high accuracy.
Solution Approach 2:
The patent changes the parameters of analysis by focusing on mutation group frequencies rather than individual mutation details. This parameter transformation reduces the dimensionality of the problem from analyzing every possible mutation combination to working with aggregated frequency data, thereby improving accuracy while reducing computational requirements.
3Measurement precision
If complete genome sequencing is performed on cancer samples, then the precision of mutation detection is improved, but the cost and time required worsens
Solution Approach 1:
The patent extracts only the essential information needed for phylogenetic analysis by focusing on mutation group frequencies rather than complete genome sequences. This extraction approach isolates the critical evolutionary signals from the vast amount of redundant genomic data, achieving sufficient precision for detecting mutations and their evolutionary relationships without the time and cost burden of complete genome sequencing.
Data Source
AI summary
According to the present invention, a phylogenetic tree can be created on the basis of frequency data regarding a large number of mutations detected from the samples of a cancer. Each sample to be analyzed contains a mixture of plural clones having different genomes. Mutations having about the same frequencies are grouped to make plural groups, and an analysis is executed based on data listing the mutation frequencies of individual groups (called mutation group frequency data). It is assumed that pairs of clones corresponding respectively to mutation groups such that frequencies of one group is equal to or greater than that of another in all the samples have parent-child relations, and a graph structure having the clones as vertices and the parent-child relations as edges is created. In this graph, parent-child relations contradictory to the mutation group frequency data are removed, and a clone to become a parent is selected in consideration of correlation coefficients among the mutation group frequencies in the samples.


