Regional Barcoding for DNA Sequencing Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation sequencing (NGS) technologies face challenges in accurately aligning DNA sequence reads due to errors and complexities, leading to incorrect variant detection and increased costs in sequencing entire genomes.

Innovation Solution

A method involving the use of probes with linked payload polynucleotides and barcoding agents to fragment and barcode DNA samples, allowing for the decoupling and sequencing of regions, which facilitates accurate alignment and reconstruction of DNA sequences through regional barcoding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If next-generation sequencing technologies are used to perform parallel DNA fragment reading, then the cost and time of sequencing is significantly reduced, but errors in the read process occur which adversely impact alignment accuracy and lead to incorrect variant detection

Engineering Contradiction:
Improvesequencing throughputVSAvoidalignment accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces long-range linkage information as an intermediary element that mediates between short read fragments and the reference genome. These linkages connect distant genomic regions, providing contextual information that helps correctly position reads in repetitive or complex regions, thereby improving alignment accuracy without reducing sequencing throughput

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary assembly of contigs from read fragments before final alignment to the reference genome. By pre-assembling connected fragments into larger contiguous sequences with established relative positions, the system reduces alignment errors and improves variant detection accuracy before the reads are mapped to the reference

Inventive Principle:
Principle #10Preliminary action

2Productivity

If computational techniques are used to align sequence reads, then parallel processing is enabled, but misalignment of sequence reads occurs leading to incorrect variant detection

Engineering Contradiction:
Improveread alignment speedVSAvoidvariant detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the alignment process into multiple stages: initial read fragmentation, contig assembly with long-range linkages, and final reference alignment. This segmentation allows parallel processing at each stage while using the structural information from linkages to guide and verify alignments, preventing misalignment errors that would otherwise occur in single-stage computational alignment

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where long-range linkage information is used to validate and correct alignment positions. When reads are aligned computationally, the system checks whether the inferred distances and orientations are consistent with the observed long-range linkages, providing feedback that corrects misalignments and improves variant detection accuracy

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240344112A1Iterative oligonucleotide barcode expansion for labeling and localizing many biomolecules
Publication Date: 2024.10.17 GOOGLE LLC
  • US20240344112A1 patent drawing
  • US20240344112A1 patent drawing
  • US20240344112A1 patent drawing

AI summary

Contemporary gene sequencing techniques, including “Next. Generation Sequencing” techniques, can include sequencing a plurality of fragments of a target polynucleotide. However, the limitations of existing sequencing techniques means that it can be difficult and/or expensive to align the generated read fragments. Methods provided herein include inserting dual polynucleotide ‘bar-codes’ into a target poly nucleotide that remain mechanically connected via a Tinker.’ Tire barcodes can then be ‘grown’ via. a. pool-split-pool process such that polynucleotide fragments that are linked by linkers exhibit the same complete barcode sequence that is different, from the complete barcode sequence exhibited by non-linked polynucleotide fragments. The joined fragments can then be separated and sequenced. Each read sequence thus begins with a regionally-specific barcode that can be used to associate fragments from the region together, allowing for increased accuracy and reduced computational cost in aligning the read fragments and/or performing other sequencing processes on the read fragments.