Nucleic Acid Tokenization Using SPLASH k-Mer Anchor-Target Pairs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI models, particularly transformer models, face inefficiencies in training and performance due to conventional tokenization methods that do not effectively capture the complexity and context of genetic sequence data, leading to suboptimal analysis of nucleic acid data.
Innovation Solution
The use of Statistically Primary alignment Agnostic Sequence Homing (SPLASH) technique to generate k-mer anchors and targets, which are then converted into tokens by appending the most abundant k-mer targets, providing a more compact and information-rich representation of genetic data for improved signal compression and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional tokenization methods are used for genetic sequence data, then the data can be processed by transformer models, but the training efficiency and performance are suboptimal due to inability to capture sequence complexity and context
Solution Approach 1:
The patent applies segmentation by dividing genetic sequences into k-mers (subsequences of length k) and further segmenting them into anchor-target pairs. This segmentation allows the model to process sequences in manageable units while preserving local sequence context and relationships, thereby improving both training efficiency and analysis accuracy simultaneously
Solution Approach 2:
The patent transforms the one-dimensional sequence data into a two-dimensional anchor-target structure by identifying anchor k-mers and their corresponding target k-mers at fixed offsets. This dimensional transformation enables the model to capture sequence context and complexity more effectively, resolving the contradiction between processing efficiency and analysis precision
2Measurement precision
If conventional tokenization is used, then processing is simpler, but signal compression and resolution are insufficient for complex genetic data
Solution Approach 1:
The patent applies preliminary action by pre-processing genetic sequences to identify and extract anchor-target pairs before feeding them to the transformer model. This pre-processing step prepares the data in an optimized format that enhances signal resolution while managing complexity through systematic sequence analysis rather than ad-hoc processing
Solution Approach 2:
The patent introduces anchor-target pairs as an intermediary representation between raw genetic sequences and model input tokens. This intermediary structure captures sequence context and relationships, improving signal resolution while the systematic methodology for generating these pairs manages the overall complexity of the tokenization process
3Reliability
If more sophisticated model architectures are used, then model capability improves, but training data requirements and processing time increase
Solution Approach 1:
The patent segments genetic sequences into anchor-target pairs that capture local sequence context and relationships. This segmentation enables sophisticated transformer models to process genetic data more efficiently by focusing on relevant local patterns rather than entire sequences, thereby improving model performance while reducing training time
Solution Approach 2:
The patent extracts the most relevant features from genetic sequences by identifying anchor k-mers and their corresponding target k-mers. This extraction process distills complex sequence information into essential anchor-target relationships, allowing sophisticated models to achieve high performance with reduced processing time and data requirements
Data Source
AI summary
Systems and methods for nucleic acid data tokenization in accordance with embodiments of the invention are illustrated. One embodiment includes a method for tokenizing genetic sequence data, comprising obtaining genetic sequence data, extracting k-mers from the genetic sequence data as a plurality of k-mer anchors, appending at least one most abundant k-mer target to each k-mer anchor, and generating tokens based on the appended k-mer anchors and targets. In a further embodiment, extracting k-mers from the genetic sequence data includes using a Statistically Primary alignment Agnostic Sequence Homing (SPLASH) technique. In still another embodiment, the method further includes steps for appending a count for each appended k-mer target to the k-mer anchor.


