Reference-Guided Genome Sequencing Memory Partitioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sample handling methods for genome sequencing are inefficient in terms of computational cost, memory resources, and scalability, particularly in de novo and reference-aligned sequencing, due to the need for large memory storage and limited compute threads, which restricts the processing capacity and scalability.
Innovation Solution
A system that includes a reference-guided device and hosts for pre-processing sample reads into probabilistically localized groups, using overlapping reference sequences stored in arrays, allowing for efficient sorting and alignment of sample reads, reducing memory requirements and enhancing scalability by distributing sample reads across multiple smaller memories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the full reference genome is stored in a single shared memory for reference-aligned sequencing, then memory access is centralized, but the number of compute cores and compute threads that can access the shared memory is limited
Solution Approach 1:
The reference genome is divided into multiple partitions, with each partition stored in a separate memory unit. This segmentation allows multiple compute cores to access different memory units simultaneously, increasing the number of compute threads from limited shared memory access to many distributed memory accesses, while reducing the memory burden on each individual compute unit.
2Productivity
If sample reads are randomly partitioned into groups for processing, then parallel processing is enabled, but each compute thread requires a large dedicated memory to store the entire reference genome
Solution Approach 1:
Each compute unit is assigned a specific partition of the reference genome rather than requiring access to the entire reference genome. This local quality approach allows compute threads to process sample reads efficiently with reduced memory resources, as each compute unit only needs to store and access its assigned reference partition rather than the full reference genome.
3Reliability
If a large group of sample reads is stored in shared memory for de novo sequencing, then all reads can be processed together, but the memory cost is expensive and the number of compute threads is limited
Solution Approach 1:
Sample reads are divided into multiple groups, with each group associated with a specific reference genome partition. Each compute unit processes its assigned group of reads against its assigned reference partition in parallel. This segmentation enables distributed processing across many compute threads with reduced memory requirements per unit, while maintaining sequencing accuracy through comprehensive coverage of the entire reference genome across all compute units.
Data Source
AI summary
Methods and systems for processing a plurality of sample reads for genome sequencing include, for each sample read of the plurality of sample reads, comparing substring sequences from the sample read to reference sequences representing different portions of a reference genome. One or more reference sequences are identified that match one or more of the compared substring sequences, and a probabilistic location within the reference genome is determined for the sample read based on the one or more identified reference sequences. The reference genome is partitioned for reference-aligned genome sequencing based on the determined probabilistic locations of the respective sample reads.


