Bayesian Genotyping System for Microsatellite Repeat Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for genotyping microsatellite repeats using next-generation sequencing technology are inadequate due to challenges in accurately identifying repeat mutations, particularly in repetitive sequences, as they fail to account for error rates and are blind to homopolymers, leading to difficulties in distinguishing true alleles from false heterozygosity and identifying genetic variations genome-wide.
Innovation Solution
A system and method that combines a repeat-aware approach with a Bayesian genotyping method using an empirically-derived error model, which selects flanking reads that span repeats and associates sequence attributes with genotyping error rates to calculate genotype probabilities, enabling accurate genotyping of microsatellite repeats in whole genome sequencing data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current indel genotyping methods are used on repeat regions, then indels within repeat regions can be revealed, but the methods fail to accurately identify true repeat alleles and produce false heterozygosity calls
Solution Approach 1:
The patent changes the parameters used for genotyping by incorporating repeat-specific error rates that account for repeat length, repeat unit size, and sequence purity. Instead of using uniform error models, the system adjusts error parameters based on the specific characteristics of each repeat region, allowing accurate distinction between true alleles and sequencing errors.
Solution Approach 2:
The patent performs preliminary actions by pre-calculating error rates for different repeat types and incorporating this information before genotyping. The system pre-processes the data to identify repeat regions and assigns appropriate error models based on repeat properties, enabling more accurate subsequent genotyping calls.
2Adaptability or versatility
If lobSTR method is used for genome-wide microsatellite calls, then microsatellite genotypes can be identified, but the method is blind to homopolymer runs
Solution Approach 1:
The patent creates a universal genotyping system that handles multiple types of repetitive sequences including microsatellites and homopolymers. By using a flexible error model that adapts to different repeat types based on their intrinsic properties, the system achieves genome-wide applicability while detecting previously missed variants like homopolymer runs.
Solution Approach 2:
The patent applies local quality by tailoring the error model to the specific characteristics of each repeat type. Different error rates are assigned based on local properties such as repeat length, unit size, and purity, allowing the system to optimize detection for each region type including homopolymers.
3Ease of manufacture
If repeat-agnostic indel callers are used, then processing is simpler, but error rates of different repeat types are not accounted for
Solution Approach 1:
The patent implements self-service by having the system automatically determine the appropriate error model for each repeat region based on its intrinsic properties. The genotyping method autonomously identifies repeat characteristics and selects corresponding error rates without requiring manual intervention, maintaining simplicity while improving accuracy.
Data Source
AI summary
A system and method for genotyping tandem repeats in sequencing data. The invention uses Bayesian model selection guided by an empirically-derived error model that incorporates properties of sequence reads and reference sequences to which they map.


