Reference-Guided Genome Sequencing Memory Partitioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sample handling methods for genome sequencing are inefficient in terms of computational cost, memory resources, and scalability, particularly in de novo and reference-aligned sequencing, due to the need for large memory storage and limited compute threads, which restricts the processing capacity and scalability.

Innovation Solution

A system that includes a reference-guided device and hosts for pre-processing sample reads into probabilistically localized groups, using overlapping reference sequences stored in arrays, allowing for efficient sorting and alignment of sample reads, reducing memory requirements and enhancing scalability by distributing sample reads across multiple smaller memories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If the full reference genome is stored in a single shared memory for reference-aligned sequencing, then memory access is centralized, but the number of compute cores and compute threads that can access the shared memory is limited

Engineering Contradiction:
Improvememory architectureVSAvoidnumber of compute threads
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The reference genome is divided into multiple partitions, with each partition stored in a separate memory unit. This segmentation allows multiple compute cores to access different memory units simultaneously, increasing the number of compute threads from limited shared memory access to many distributed memory accesses, while reducing the memory burden on each individual compute unit.

Inventive Principle:
Principle #1Segmentation

2Productivity

If sample reads are randomly partitioned into groups for processing, then parallel processing is enabled, but each compute thread requires a large dedicated memory to store the entire reference genome

Engineering Contradiction:
Improveparallel processing capacityVSAvoidmemory resources per compute thread
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

Each compute unit is assigned a specific partition of the reference genome rather than requiring access to the entire reference genome. This local quality approach allows compute threads to process sample reads efficiently with reduced memory resources, as each compute unit only needs to store and access its assigned reference partition rather than the full reference genome.

Inventive Principle:
Principle #3Local quality

3Reliability

If a large group of sample reads is stored in shared memory for de novo sequencing, then all reads can be processed together, but the memory cost is expensive and the number of compute threads is limited

Engineering Contradiction:
Improvesequencing accuracyVSAvoidmemory resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Sample reads are divided into multiple groups, with each group associated with a specific reference genome partition. Each compute unit processes its assigned group of reads against its assigned reference partition in parallel. This segmentation enables distributed processing across many compute threads with reduced memory requirements per unit, while maintaining sequencing accuracy through comprehensive coverage of the entire reference genome across all compute units.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11837330B2Reference-guided genome sequencing
Publication Date: 2023.12.05 WESTERN DIGITAL TECHNOLOGIES INC
  • US11837330B2 patent drawing
  • US11837330B2 patent drawing
  • US11837330B2 patent drawing

AI summary

Methods and systems for processing a plurality of sample reads for genome sequencing include, for each sample read of the plurality of sample reads, comparing substring sequences from the sample read to reference sequences representing different portions of a reference genome. One or more reference sequences are identified that match one or more of the compared substring sequences, and a probabilistic location within the reference genome is determined for the sample read based on the one or more identified reference sequences. The reference genome is partitioned for reference-aligned genome sequencing based on the determined probabilistic locations of the respective sample reads.