Genome Sequencing Base-Calling and Alignment via Probabilistic Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genome sequencing technologies face limitations in producing long, accurate sequence reads with contextual information, leading to challenges in haplotypic analysis and polymorphism detection, particularly in biomedical applications, due to short read lengths and lack of contextual data.

Innovation Solution

The development of methods and systems for base calling and alignment that utilize a score function based on sequencing platform information, incorporating branch-and-bound processes and probabilistic frameworks to analyze polymorphisms and generate accurate base-call interpretations and alignments from raw sequencing data, including reference sequences to correct errors and extend read lengths.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short-read sequencing platforms are used to reduce cost and increase throughput, then productivity and cost-effectiveness are improved, but read length and contextual information are reduced

Engineering Contradiction:
ImprovethroughputVSAvoidread length
Core Design Contradiction:
ProductivityVSLength of moving object

Solution Approach 1:

The patent combines multiple short reads that overlap in genomic space to reconstruct longer contiguous sequences (contigs). By merging overlapping reads and using probabilistic models to resolve ambiguities, the system effectively extends read length while maintaining the high throughput advantage of short-read platforms.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from analyzing individual short reads in one dimension to analyzing the collective information across multiple overlapping reads in a higher-dimensional space. By considering reads as vectors in a probabilistic framework and using matrix operations to integrate information across multiple dimensions (multiple reads, multiple positions, multiple possible alignments), the system recovers long-range contextual information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If short reads are used for sequencing, then cost is reduced and throughput is increased, but measurement precision and reliability of polymorphism detection deteriorate

Engineering Contradiction:
ImprovethroughputVSAvoidpolymorphism detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements an iterative refinement process where initial alignment estimates are used to improve base-calling, which in turn improves alignment accuracy, creating a positive feedback loop. The probabilistic framework continuously refines estimates of the true sequence by incorporating feedback from multiple overlapping reads and adjusting probabilities based on observed data consistency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs sequencing with excessive coverage (depth) beyond what would be needed for a single long read, using many overlapping short reads to compensate for individual read limitations. This excessive sampling ensures that each genomic position is covered by multiple independent measurements, improving polymorphism detection accuracy through statistical aggregation.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If complex algorithms are used to improve base-calling accuracy, then measurement precision is improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improvebase-calling accuracyVSAvoidalgorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex iterative mechanical/base-calling algorithms with a probabilistic framework that uses matrix operations and statistical inference. Instead of relying on sequential refinement algorithms, the system formulates base-calling as a probability estimation problem that can be solved more efficiently using linear algebra and information theory principles.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the base-calling problem from deterministic threshold-based decisions to probabilistic estimates with confidence scores. By transforming the problem from finding a single best answer to estimating probability distributions, the system achieves improved accuracy while enabling more efficient computational approaches using statistical methods.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If multiple reference sequences are used for alignment, then reliability of polymorphism detection is improved, but device complexity and computational load increase

Engineering Contradiction:
Improvepolymorphism detection reliabilityVSAvoidalignment system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing the multiple reference sequences to identify and catalog known polymorphic positions before the actual alignment process. By preparing reference data structures in advance that encode expected variations, the system reduces the computational complexity during alignment while maintaining the ability to detect polymorphisms reliably across multiple references.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10964408B2Method, computer-accessible medium and system for base-calling and alignment
Publication Date: 2021.03.30 NEW YORK UNIV
  • US10964408B2 patent drawing
  • US10964408B2 patent drawing
  • US10964408B2 patent drawing

AI summary

Exemplary methods, procedures, computer-accessible medium, and systems for base-calling, aligning and polymorphism detection and analysis using raw output from a sequencing platform can be provided. A set of raw outputs can be used to detect polymorphisms in an individual by obtaining a plurality of sequence read data from one or more technologies (e.g., using sequencing-by-synthesis, sequencing-by-ligation, sequencing-by-hybridization, Sanger sequencing, etc.). For example, provided herein are exemplary methods, procedures, computer-accessible medium and systems, which can include and/or be configured for obtaining raw output from a sequencing platform configured to be used for reading fragment(s) of genomes, obtaining reference sequences for the genomes obtained independently from the raw output, and generating a base-call interpretation and/or alignment using the raw output and the reference sequences. For example, a score function can be determined based on information associated with the sequencing platform that can be used to analyze polymorphisms based on the base-call interpretation and/or alignment.