Hamming Code Indexing for Rapid Oligonucleotide Sequence Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current nucleotide sequencing methods are inefficient for rapid identification of unique oligonucleotide sequences in genomic datasets, requiring substantial time and computational resources, especially when aligning large numbers of sequences against custom or indexed genomic markers.

Innovation Solution

A computer-implemented method using perfect Hamming code indexing and a four-level trie structure to efficiently align 20mer oligonucleotide sequences with nucleotide datasets, allowing for rapid comparison and identification of unique sequences by mapping 20mer fragments to equivalence classes, thereby reducing the search space and computational time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current alignment tools (bowtie, bwa) are used to align oligonucleotide sequences, then alignment accuracy is maintained, but computational time and resource requirements increase substantially

Engineering Contradiction:
Improvealignment accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the genomic sequence into non-overlapping 5mer fragments and organizes them into a hierarchical trie structure with four levels. Each level represents a specific position in the 20mer oligonucleotide, containing only those 5mers that are present in the genomic sequence. This segmentation reduces the search space from examining all possible 5mers to only those actually present in the genome, enabling rapid alignment while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

2Productivity

If pre-built indices are created for standard genomes, then alignment speed is improved, but adaptability to custom genomes is lost

Engineering Contradiction:
Improvealignment speedVSAvoidcustom genome compatibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal alignment method that functions equally well for standard genomes and custom genomes. The trie structure can be built for any genomic sequence, making the alignment tool universally applicable. The method does not require pre-built indices for specific genomes but can generate the necessary structure dynamically for any target genome, thus maintaining both speed and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If exhaustive search methods are used to align all oligonucleotides, then comprehensive coverage is achieved, but computational complexity increases

Engineering Contradiction:
Improvesearch completenessVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by pre-processing the genomic sequence into a hierarchical trie structure before alignment. During this preliminary step, all 5mer fragments are extracted, organized into equivalence classes, and arranged in the trie structure with levels corresponding to positions in the 20mer. This pre-computation enables the actual alignment to proceed rapidly by simply traversing the pre-built structure, achieving both completeness and efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2923293B1Efficient comparison of polynucleotide sequences
Publication Date: 2021.05.26 ILLUMINA INC
  • EP2923293B1 patent drawingFigure 1
  • EP2923293B1 patent drawingFigure 2
  • EP2923293B1 patent drawingFigure 3

AI summary

The disclosure relates to rapid detection of oligonucleotide sequence in a nucleic acid sequence database through the configuration of the database into rapidly searchable index classes built around perfect Hamming code oligonucleotides.