Tandem-Repeat Genotyping Using EM and PCR Stutter Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sequencing systems face inaccuracies in tandem-repeat genotyping due to PCR stutter errors and inefficiencies in processing, particularly when relying on population-scale data and requiring excessive computational resources, limiting their ability to accurately predict genotypes for individual samples.
Innovation Solution
A tandem-repeat genotype sequencing system that utilizes spanning reads and a stutter model within an Expectation-Maximization (EM) algorithm to iteratively predict and update genotype probabilities, accounting for PCR stutter artifacts and reducing reliance on population data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing sequencing systems use conventional alignment methods and variant callers to determine tandem-repeat genotypes, then they can process genomic data through standard pipelines, but they produce inaccurate genotype predictions due to alignment errors and PCR stutter artifacts
Solution Approach 1:
The patent extracts and isolates the tandem-repeat regions from the rest of the genome for specialized processing. Instead of applying general-purpose alignment and variant calling methods to these problematic regions, the system separates them out and applies dedicated algorithms that account for their unique characteristics, including PCR stutter artifacts and alignment challenges.
Solution Approach 2:
The patent introduces an intermediary computational layer between raw sequencing reads and final genotype calls. This intermediary processing stage includes specialized algorithms that model PCR stutter patterns and use probabilistic approaches to distinguish true genetic variants from artifacts, thereby mediating between the noisy input data and reliable genotype predictions.
2Measurement precision
If existing tandem-repeat genotyping systems rely on population-scale sequencing data and phased SNP haplotypes to predict genotypes, then they can leverage extensive computational resources and large datasets, but they cannot accurately genotype individual samples and require excessive data input
Solution Approach 1:
The patent segments the genotyping problem into two distinct approaches: one for population-scale data and another for individual samples. For individual samples, it uses a simplified model that processes only the necessary spanning reads without requiring extensive population data or phased haplotypes, thereby reducing data input requirements while maintaining accuracy.
Solution Approach 2:
The patent applies partial action by using only the subset of reads that are necessary for tandem-repeat genotyping (spanning reads that cover the repeat regions) rather than processing all sequencing data. This selective approach reduces computational burden and data requirements while achieving accurate individual sample genotyping.
3Measurement precision
If existing systems process all sequencing reads through comprehensive analysis pipelines including SNP calling and phasing, then they can utilize maximum computational resources for analysis, but they incur prohibitive computational costs and processing time
Solution Approach 1:
The patent extracts only the essential information needed for tandem-repeat genotyping from the sequencing data, specifically the spanning reads that cover the repeat regions. By eliminating unnecessary processing steps such as full SNP calling and phasing for this specific application, the system achieves significant computational efficiency gains while maintaining genotyping accuracy.
Solution Approach 2:
The patent performs partial processing by applying a streamlined analysis pipeline that focuses exclusively on tandem-repeat regions rather than processing the entire genome through comprehensive SNP calling and phasing. This partial action approach reduces computational costs and processing time while achieving the specific goal of accurate tandem-repeat genotyping.
Data Source
AI summary
This disclosure describes methods, non-transitory-computer readable media, and systems that can accurately generate genotypes for tandem-repeat regions of a genomic sample by utilizing an expectation-maximization (EM) algorithm and a stutter model. The disclosed system can extract spanning nucleotide reads that comprise whole tandem-repeat regions. The disclosed system may perform an expectation stage of an EM algorithm and utilize a stutter model to predict expected genotype probabilities of tandem-repeat genotypes given a distribution of spanning reads. In some implementations, the disclosed system further performs a maximization stage of the EM algorithm to adjust parameters of the stutter model based on the expected genotype probabilities to maximize a total probability of the expected genotype probabilities. The disclosed system can repeat the expectation and maximization stages until the total probability of the expected genotype probabilities converges. The disclosed system may predict a genotype for the tandem repeat based on the converged genotype probabilities.


