Genome Assembly via Mutagenic Diversity for Repetitive Regions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genome sequencing technologies face challenges in assembling complex genomes due to repetitive and low-information regions, which are exacerbated by high error rates in long-read platforms and the requirement for large DNA input, leading to incomplete and inaccurate genome assemblies.

Innovation Solution

Introducing random mutations into nucleic acid samples to increase sequence diversity, allowing for the assembly of previously difficult-to-assemble regions using a combination of mutagenic reactions, size selection, and amplification techniques, such as multiple strand displacement amplification with nucleotide analogs, to generate sequencing libraries that can be accurately assembled.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If long-read sequencing platforms are used to increase read length for assembling repetitive regions, then the ability to span repetitive regions is improved, but the error rate increases significantly

Engineering Contradiction:
Improveread lengthVSAvoiderror rate
Core Design Contradiction:
Length of moving objectVSReliability

Solution Approach 1:

The patent segments the genome assembly problem into two distinct phases: first, long reads are used to span repetitive regions and establish the overall structure; second, short accurate reads are used to polish and correct errors in the assembled sequence. This segmentation allows each sequencing technology to be used for its strength while mitigating their respective weaknesses.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges the outputs of long-read and short-read sequencing assemblies by using the long-read assembly as a scaffold and incorporating accurate short-read data to correct errors. This combination produces a final assembly that benefits from both the long span capability and high accuracy of the two approaches.

Inventive Principle:
Principle #5Merging (Combining)

2Loss of information

If long-read sequencing is used to assemble complex genomes, then coverage of repetitive regions is improved, but the requirement for large amounts of input DNA increases

Engineering Contradiction:
Improvecoverage of repetitive regionsVSAvoidinput DNA amount
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies partial action by using a small subset of long reads (only enough to span repetitive regions and establish structure) rather than requiring full coverage with long reads. The remaining gaps and errors are filled by short reads, reducing the total long-read input DNA requirement while still achieving complete coverage of repetitive regions.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If short-read sequencing platforms are used for high accuracy, then the error rate is reduced, but the ability to assemble repetitive regions is lost

Engineering Contradiction:
ImproveaccuracyVSAvoidassembly of repetitive regions
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent performs preliminary assembly using long reads to span repetitive regions and establish the overall genomic structure before applying short reads. This preliminary action creates a framework that guides the subsequent accurate polishing with short reads, ensuring that repetitive regions are correctly positioned before accuracy optimization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The long-read assembly serves as an intermediary scaffold that bridges the gap between short reads and the final accurate assembly. It provides the structural framework that short reads alone cannot generate, while the short reads provide the accuracy that the long-read assembly lacks. The intermediary scaffold enables the combination of both approaches.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If traditional assembly methods are used for complex genomes, then the process is simpler, but the assembly completeness and accuracy deteriorate

Engineering Contradiction:
Improveassembly process complexityVSAvoidassembly accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent segments the assembly process into distinct phases with different objectives: structure establishment using long reads, error identification, and accuracy polishing using short reads. This segmentation transforms a single complex accurate assembly problem into manageable stages, improving accuracy without overwhelming complexity.

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables the assembly of larger, more accurate contigs with reduced error rates and increased sequence coverage, overcoming the limitations of existing technologies by generating tunable mutation rates and improving the assembly of complex genomic regions.

Implementation Method 1

the multiple strand displacement amplification reaction uses Phi29 DNA polymerase

Methodology Applied
Scientific EffectMultiple strand displacement amplification:

Implementation Method 2

the introducing mutations step of the method includes performing a multiple displacement amplification reaction using a nucleotide analog

Methodology Applied
Scientific EffectNucleotide analog incorporation:

Data Source

PatentEP3870718B1Methods and uses of introducing mutations into genetic material for genome assembly
Publication Date: 2023.11.29 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • EP3870718B1 patent drawingFigure 1
  • EP3870718B1 patent drawingFigure 2
  • EP3870718B1 patent drawingFigure 3

AI summary

Methods of sequencing and assembling a nucleic acid sequence from a nucleic acid sample containing repetitive or low-information regions, which are typically difficult to sequence and/or assemble are provided. The methods of sequencing and assembling introduce mutations into the sample to increase sequence diversity between various repetitive regions present in the nucleic acid sample. This sequence diversity allows various segments to assemble independently of different, but similar sequences present in the nucleic acid sample.