Mutational Signature Detection From Sparse Targeted Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting mutational signatures and exposures in cancer genomes are limited by their requirement for whole-genome sequencing data and struggle with sparse data from targeted sequencing assays, particularly in predicting homologous recombination deficiencies.
Innovation Solution
A method that clusters samples and dynamically updates mutational signatures and exposure vectors using an optimization procedure, such as Expectation-Maximization, to identify mutational signatures and their exposures in sparse sequencing data, allowing for efficient detection without relying on whole-genome sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If whole-genome sequencing is used for detecting mutational signatures, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts only the necessary mutation information from whole-genome sequencing data, focusing on identifying mutational signatures through targeted analysis of mutation patterns rather than processing complete genomic sequences. This allows detection of mutational signatures using reduced sequencing depth or targeted sequencing approaches.
Solution Approach 2:
The patent creates a simplified representation of genomic data by extracting mutation count matrices and signature probability distributions, which can be analyzed without processing the full genomic sequences. This copying approach enables signature detection from condensed data representations.
2Ease of operation
If targeted sequencing assays are used, then ease of operation and cost are improved, but measurement precision deteriorates
Solution Approach 1:
The patent performs preliminary clustering of samples based on mutation patterns before signature detection. This pre-processing step organizes the sparse targeted sequencing data into meaningful groups, enabling accurate signature detection even from limited mutation counts by leveraging patterns within and between clusters.
Solution Approach 2:
The patent employs iterative optimization procedures that use feedback from initial signature estimates to refine and improve detection accuracy. The algorithm continuously adjusts signature probabilities and exposure estimates based on observed mutation data, enhancing precision from sparse targeted sequencing results.
3Ease of manufacture
If sparse data from targeted sequencing is used, then ease of manufacture and cost are improved, but reliability deteriorates
Solution Approach 1:
The patent merges information from multiple sources including mutation count data, sample clustering results, and signature probability estimates into a unified predictive model. This integration allows reliable detection of homologous recombination deficiency by combining signals from sparse targeted sequencing data with pattern recognition from clustered samples.
Solution Approach 2:
The patent uses dynamic optimization procedures that adapt to the sparsity and variability of targeted sequencing data. The algorithm dynamically adjusts its analysis based on the specific characteristics of each sample and cluster, improving reliability by tailoring the detection approach to the actual data quality and distribution.
Data Source
AI summary
A method of detecting mutational signatures of a sample and their exposures in a collection of samples, each being characterized by nucleic acid sequencing information describing at least one mutation, comprises: clustering the samples to provide clusters and respective exposure vectors, where each exposure vector describes prior probabilities for a plurality of signatures to emit a mutation. An optimization procedure is applied to dynamically re-cluster the samples and to dynamically update the signatures and exposure vectors. Optionally, the mutational signatures in the sample are determined using an exposure vector of one or more clusters associated with the sample.


