Hierarchical DNA Alignment Index for Fast Short-Read Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The computational intensity of aligning small DNA samples (short reads) to large reference datasets in DNA sequencing is exponentially increased by the growing size of reference data sets, hindering efficient and accurate genetic information processing for healthcare, agriculture, and crime solving.

Innovation Solution

A hierarchical inverted index table is constructed to efficiently map data patterns onto larger data sets, utilizing a 2-phase approach of candidate selection and evaluation, where candidate locations are identified using a hierarchical index table based on reference data, and further refined through insertion or deletion analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional alignment methods are used to match short read sequences to reference genomes, then alignment accuracy can be maintained, but computational time and resources increase exponentially as reference data size grows

Engineering Contradiction:
Improvealignment accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The reference genome is divided into smaller k-mer segments (subsequences of length k), and an index table is constructed that maps each k-mer to its occurrence positions in the reference genome. This segmentation allows the alignment algorithm to quickly locate candidate regions by breaking down the large reference genome into manageable indexed units, thereby reducing computational time while maintaining alignment accuracy through precise position mapping.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The index table is pre-computed and stored before the actual alignment process begins. This preliminary action involves creating a comprehensive mapping of all k-mers to their positions in the reference genome, which can then be rapidly queried during alignment without re-processing the entire reference data. This pre-computation significantly reduces the time required for each alignment operation while preserving accurate position information.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the reference genome size increases to improve completeness of genetic information, then more comprehensive data is available, but the computational complexity of alignment increases exponentially

Engineering Contradiction:
Improvecompleteness of genetic informationVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The reference genome is segmented into k-mers and indexed, transforming the complexity from linear scanning of the entire genome to efficient hash table lookups. This segmentation approach allows the system to handle larger, more complete reference genomes by converting the alignment problem from O(n) complexity to O(1) average case complexity per query, where n is the reference genome size.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An index table serves as an intermediary data structure between the reference genome and the alignment algorithm. This intermediary pre-organizes the reference data into a query-optimized format, allowing rapid retrieval of candidate positions without directly processing the entire reference genome during alignment. The index table mediates between the large reference dataset and the alignment requirements, reducing computational complexity while maintaining access to complete genetic information.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If exhaustive indexing of all subsequences is performed to ensure complete matching, then matching completeness is improved, but memory requirements and construction time increase significantly

Engineering Contradiction:
Improvematching completenessVSAvoidmemory requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of uniformly indexing all possible subsequences, the method selectively indexes only those k-mers that actually appear in the reference genome. This local quality approach means that memory is allocated proportionally to the actual content distribution rather than the theoretical maximum, reducing memory requirements while maintaining matching completeness for all present sequences. The index table size becomes proportional to the number of unique k-mers in the reference rather than 4^k.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The k-mer length parameter k is optimized to balance between matching specificity and index size. By adjusting k, the system can control both the memory requirements (which increase with larger k) and the matching completeness (which improves with larger k). This parameter optimization allows finding the sweet spot where sufficient matching completeness is achieved without excessive memory consumption, adapting to different reference genome sizes and available resources.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12626781B2DNA alignment using a hierarchical inverted index table
Publication Date: 2026.05.12 HYPERX HOLDINGS LLC
  • US12626781B2 patent drawing
  • US12626781B2 patent drawing
  • US12626781B2 patent drawing

AI summary

System and method for constructing a hierarchical index table usable for matching a search sequence to reference data. The index table may be constructed to contain entries associated with an exhaustive list of all subsequences of a given length, wherein each entry contains the number and locations of matches of each subsequence in the reference data. The hierarchical index table may be constructed in an iterative manner, wherein entries for each lengthened subsequence are selectively and iteratively constructed based on the number of matches being greater than each of a set of respective thresholds. The hierarchical index table may be used to search for matches between a search sequence and reference data, and to perform misfit identification and characterization upon each respective candidate match.