Genome Sequencing Base-Calling and Alignment via Probabilistic Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genome sequencing technologies face limitations in producing long, accurate sequence reads with contextual information, leading to challenges in haplotypic analysis and polymorphism detection, particularly in biomedical applications, due to short read lengths and lack of contextual data.
Innovation Solution
The development of methods and systems for base calling and alignment that utilize a score function based on sequencing platform information, incorporating branch-and-bound processes and probabilistic frameworks to analyze polymorphisms and generate accurate base-call interpretations and alignments from raw sequencing data, including reference sequences to correct errors and extend read lengths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short-read sequencing platforms are used to reduce cost and increase throughput, then productivity and cost-effectiveness are improved, but read length and contextual information are reduced
Solution Approach 1:
The patent combines multiple short reads that overlap in genomic space to reconstruct longer contiguous sequences (contigs). By merging overlapping reads and using probabilistic models to resolve ambiguities, the system effectively extends read length while maintaining the high throughput advantage of short-read platforms.
Solution Approach 2:
The patent transitions from analyzing individual short reads in one dimension to analyzing the collective information across multiple overlapping reads in a higher-dimensional space. By considering reads as vectors in a probabilistic framework and using matrix operations to integrate information across multiple dimensions (multiple reads, multiple positions, multiple possible alignments), the system recovers long-range contextual information.
2Productivity
If short reads are used for sequencing, then cost is reduced and throughput is increased, but measurement precision and reliability of polymorphism detection deteriorate
Solution Approach 1:
The patent implements an iterative refinement process where initial alignment estimates are used to improve base-calling, which in turn improves alignment accuracy, creating a positive feedback loop. The probabilistic framework continuously refines estimates of the true sequence by incorporating feedback from multiple overlapping reads and adjusting probabilities based on observed data consistency.
Solution Approach 2:
The patent performs sequencing with excessive coverage (depth) beyond what would be needed for a single long read, using many overlapping short reads to compensate for individual read limitations. This excessive sampling ensures that each genomic position is covered by multiple independent measurements, improving polymorphism detection accuracy through statistical aggregation.
3Measurement precision
If complex algorithms are used to improve base-calling accuracy, then measurement precision is improved, but device complexity and computational requirements increase
Solution Approach 1:
The patent replaces complex iterative mechanical/base-calling algorithms with a probabilistic framework that uses matrix operations and statistical inference. Instead of relying on sequential refinement algorithms, the system formulates base-calling as a probability estimation problem that can be solved more efficiently using linear algebra and information theory principles.
Solution Approach 2:
The patent changes the fundamental parameters of the base-calling problem from deterministic threshold-based decisions to probabilistic estimates with confidence scores. By transforming the problem from finding a single best answer to estimating probability distributions, the system achieves improved accuracy while enabling more efficient computational approaches using statistical methods.
4Reliability
If multiple reference sequences are used for alignment, then reliability of polymorphism detection is improved, but device complexity and computational load increase
Solution Approach 1:
The patent performs preliminary actions by pre-processing the multiple reference sequences to identify and catalog known polymorphic positions before the actual alignment process. By preparing reference data structures in advance that encode expected variations, the system reduces the computational complexity during alignment while maintaining the ability to detect polymorphisms reliably across multiple references.
Data Source
AI summary
Exemplary methods, procedures, computer-accessible medium, and systems for base-calling, aligning and polymorphism detection and analysis using raw output from a sequencing platform can be provided. A set of raw outputs can be used to detect polymorphisms in an individual by obtaining a plurality of sequence read data from one or more technologies (e.g., using sequencing-by-synthesis, sequencing-by-ligation, sequencing-by-hybridization, Sanger sequencing, etc.). For example, provided herein are exemplary methods, procedures, computer-accessible medium and systems, which can include and/or be configured for obtaining raw output from a sequencing platform configured to be used for reading fragment(s) of genomes, obtaining reference sequences for the genomes obtained independently from the raw output, and generating a base-call interpretation and/or alignment using the raw output and the reference sequences. For example, a score function can be determined based on information associated with the sequencing platform that can be used to analyze polymorphisms based on the base-call interpretation and/or alignment.


