Nucleic Acid Tokenization Using SPLASH k-Mer Anchor-Target Pairs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI models, particularly transformer models, face inefficiencies in training and performance due to conventional tokenization methods that do not effectively capture the complexity and context of genetic sequence data, leading to suboptimal analysis of nucleic acid data.

Innovation Solution

The use of Statistically Primary alignment Agnostic Sequence Homing (SPLASH) technique to generate k-mer anchors and targets, which are then converted into tokens by appending the most abundant k-mer targets, providing a more compact and information-rich representation of genetic data for improved signal compression and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional tokenization methods are used for genetic sequence data, then the data can be processed by transformer models, but the training efficiency and performance are suboptimal due to inability to capture sequence complexity and context

Engineering Contradiction:
Improvetraining efficiencyVSAvoidsequence analysis accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing genetic sequences into k-mers (subsequences of length k) and further segmenting them into anchor-target pairs. This segmentation allows the model to process sequences in manageable units while preserving local sequence context and relationships, thereby improving both training efficiency and analysis accuracy simultaneously

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the one-dimensional sequence data into a two-dimensional anchor-target structure by identifying anchor k-mers and their corresponding target k-mers at fixed offsets. This dimensional transformation enables the model to capture sequence context and complexity more effectively, resolving the contradiction between processing efficiency and analysis precision

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If conventional tokenization is used, then processing is simpler, but signal compression and resolution are insufficient for complex genetic data

Engineering Contradiction:
Improvesignal resolutionVSAvoidtokenization complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-processing genetic sequences to identify and extract anchor-target pairs before feeding them to the transformer model. This pre-processing step prepares the data in an optimized format that enhances signal resolution while managing complexity through systematic sequence analysis rather than ad-hoc processing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces anchor-target pairs as an intermediary representation between raw genetic sequences and model input tokens. This intermediary structure captures sequence context and relationships, improving signal resolution while the systematic methodology for generating these pairs manages the overall complexity of the tokenization process

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If more sophisticated model architectures are used, then model capability improves, but training data requirements and processing time increase

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments genetic sequences into anchor-target pairs that capture local sequence context and relationships. This segmentation enables sophisticated transformer models to process genetic data more efficiently by focusing on relevant local patterns rather than entire sequences, thereby improving model performance while reducing training time

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the most relevant features from genetic sequences by identifying anchor k-mers and their corresponding target k-mers. This extraction process distills complex sequence information into essential anchor-target relationships, allowing sophisticated models to achieve high performance with reduced processing time and data requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20260018248A1Systems and Methods for Nucleic Acid Data Tokenization
Publication Date: 2026.01.15 THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
  • US20260018248A1 patent drawing
  • US20260018248A1 patent drawing
  • US20260018248A1 patent drawing

AI summary

Systems and methods for nucleic acid data tokenization in accordance with embodiments of the invention are illustrated. One embodiment includes a method for tokenizing genetic sequence data, comprising obtaining genetic sequence data, extracting k-mers from the genetic sequence data as a plurality of k-mer anchors, appending at least one most abundant k-mer target to each k-mer anchor, and generating tokens based on the appended k-mer anchors and targets. In a further embodiment, extracting k-mers from the genetic sequence data includes using a Statistically Primary alignment Agnostic Sequence Homing (SPLASH) technique. In still another embodiment, the method further includes steps for appending a count for each appended k-mer target to the k-mer anchor.