Minimizer-Based Sequence Read Assembly for Repeat Regions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current next-generation sequencing techniques face challenges in accurately sequencing nucleic acid molecules with repeat regions and structural variants, as they often generate short sequence reads that are difficult to assemble and distinguish between repeats and replicates, leading to uncertainty in determining the origin of similar sequence reads.

Innovation Solution

A computer-implemented method that processes mutated sequence reads by applying a common minimizer function to determine positions of minimizers and mutations, allowing for the counting of matching and mismatching mutations to calculate the probability that two sequence reads derive from the same sequence, thereby improving the assembly of sequences from short reads.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If next generation sequencing techniques are used to sequence longer nucleic acid sequences, then sequencing capability is improved, but cost and difficulty increase

Engineering Contradiction:
Improvesequence read lengthVSAvoidsequencing complexity and cost
Core Design Contradiction:
Length of moving objectVSDevice complexity

Solution Approach 1:

The patent segments the sequencing problem by using short sequence reads combined with mutagenesis to achieve long sequence assembly. Instead of attempting to sequence long molecules directly, the method breaks them into short reads and uses mutation patterns to assemble them correctly, resolving the contradiction between read length and sequencing complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces mutations as an intermediary mechanism to solve the sequencing problem. By introducing known mutation patterns into the template nucleic acid molecules before sequencing, the mutations serve as landmarks that guide the assembly process, enabling accurate reconstruction of long sequences from short reads without requiring expensive long-read sequencing technology

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If short sequence reads are generated from nucleic acid molecules with repeat regions, then sequencing is simplified, but ability to distinguish between repeats and replicates deteriorates

Engineering Contradiction:
Improvesequencing simplicityVSAvoidinformation about sequence origin
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies local quality by introducing mutations at specific locations within the nucleic acid molecules. These localized mutations create unique mutation patterns at different positions, allowing the system to distinguish between otherwise identical repeat regions based on their specific mutation profiles, thereby preventing information loss about sequence origin

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses mutation patterns as a form of molecular 'color coding' to distinguish between different repeat regions and replicates. By introducing mutations that create distinctive patterns (analogous to color changes), the system can identify and differentiate between sequences that would otherwise appear identical in short reads, resolving the ambiguity of sequence origin

Inventive Principle:
Principle #32Color changes

3Measurement precision

If mutation patterns are introduced to assist sequence assembly, then assembly accuracy is improved, but computational processing complexity increases

Engineering Contradiction:
Improvesequence assembly accuracyVSAvoidcomputational processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by introducing mutations into the nucleic acid molecules before sequencing. This pre-processing step creates the mutation patterns that will later guide assembly, performing the complex information encoding before the sequencing and computational analysis steps, thereby reducing the computational burden during data processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses multiple copies of the template nucleic acid molecule with different mutation patterns. By sequencing these copies and comparing their mutation profiles, the system can accurately assemble the original sequence through consensus, reducing computational complexity by distributing the information across multiple simpler sequences rather than analyzing a single complex molecule

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230044570A1Method for determining a measure correlated to the probability that two mutated sequence reads derive from the same sequence comprising mutations
Publication Date: 2023.02.09 ILLUMINA SINGAPORE PTE LTD
  • US20230044570A1 patent drawing
  • US20230044570A1 patent drawing
  • US20230044570A1 patent drawing

AI summary

Disclosed is a computer-implemented method for determining a measure correlated to the probability that two mutated sequence reads derive from the same sequence comprising mutations. The method comprises receiving mutated sequence reads each corresponding to a subsequence of a sequence comprising mutations compared to a sequence not comprising mutations, applying a common minimizer function to each mutated sequence read, to determining minimizers for each mutated sequence read, determining positions of the one or more minimizers in each mutated sequence read, determining positions of mutations in each mutated sequence read, and for at least two mutated sequence reads with a common minimizer, counting the number of mutations with matching position and/or mismatching position when the respective minimizers are aligned. Also disclosed is a corresponding method for determining at least a portion of a sequence of at least one target template nucleic acid molecule.