Transcript Abundance Estimation Using Mix 2< Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for transcriptome analysis using next-generation sequencing (NGS) struggle to accurately differentiate between transcript variants and estimate their abundances due to unrealistic assumptions about read distributions, leading to inaccurate abundance estimates.
Innovation Solution
A statistical model called the Mix 2< model is employed, which learns the bias of read distribution and transcript abundances simultaneously by using mixtures of functions trained together, allowing for more accurate estimation of transcript abundances through a maximum likelihood framework and expectation maximization algorithm.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the Cufflinks method is used to estimate transcript abundances, then the computational framework is established and basic abundance estimation is achieved, but the accuracy of transcript abundance estimates deteriorates due to unrealistic assumptions about read distributions
Solution Approach 1:
The patent changes the parameters of the statistical model by introducing a more realistic read distribution model that accounts for fragment length biases and positional biases. Instead of assuming uniform read distribution, the model uses empirically derived parameters to describe the actual distribution patterns of reads across transcripts, thereby improving the accuracy of abundance estimates while maintaining computational feasibility.
Solution Approach 2:
The patent replaces the simple uniform distribution assumption with a sophisticated statistical model that incorporates multiple bias factors. This substitution transforms the mechanical counting approach into a probabilistic framework that better reflects the biological and technical realities of RNA-Seq data generation, resolving the contradiction between model simplicity and realism.
2Measurement precision
If multiple transcripts are assembled from sequence reads, then the ability to differentiate transcript variants is improved, but the difficulty of correctly aligning short reads to specific transcript variants increases
Solution Approach 1:
The patent employs an iterative expectation-maximization algorithm that uses feedback from the observed read data to refine the abundance estimates of multiple transcripts. The model continuously adjusts the predicted read distributions based on the current abundance estimates and compares them with the actual observed data, using this feedback to improve alignment accuracy and transcript differentiation in subsequent iterations.
Solution Approach 2:
The patent introduces dynamic parameter estimation where the read distribution parameters are not fixed but are estimated from the data itself. This allows the model to adapt to the specific characteristics of each transcript and dataset, improving the ability to differentiate variants while accounting for the dynamic nature of read alignment challenges.
3Measurement precision
If the number of transcripts is increased to include all expected isoforms, then the completeness of transcript coverage is improved, but the complexity of the statistical model and computational burden increases
Solution Approach 1:
The patent creates a universal statistical framework that can handle any number of transcripts simultaneously through the use of shared bias parameters. The model uses a common set of fragment length and positional bias parameters that apply across all transcripts, allowing the system to scale to multiple isoforms without proportionally increasing model complexity. This multi-functional approach enables comprehensive transcript coverage while maintaining computational efficiency.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
The present invention relates to a method of estimating transcript abundances comprising the steps of a) obtaining transcript fragment sequencing data from a potential mixture of transcripts of a genetic locus of interest, b) assigning said fragment sequencing data to genetic coordinates of said locus of interest thereby obtaining a data set of fragment genetic coordinate coverage, said coverage for each genetic coordinate combined forming a coverage envelope curve, c) setting a number of transcripts of said mixture, d) pre-setting a probability distribution function of modelled genetic coverage for each transcript i, with i denoting the numerical identifier for a transcript, wherein said probability distribution function is composed of the mathematical product of a weight factor αi of said transcript i and the sum of at least 2 probability subfunctions j, with j denoting the numerical identifier for an probability subfunction, each probability subfunction j being independently weighted by a weight factor βi, j, e) adding the probability distribution functions for each transcript to obtain a sum function, f) fitting the sum function to the coverage envelope curve thereby optimizing the values for αi and βi, j to increase the fit, g) repeating steps e) and f) until a pre-set convergence criterion has been fulfilled, thereby obtaining the estimated transcript abundance for each transcript of the mixture given by the weight factor αi as optimized after the convergence criterion has been fulfilled.