Sequence Identification Using Minimum Description Length Heuristics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques are inadequate for efficiently and accurately identifying sequences of interest within large data series, such as genomes, due to computational inefficiencies and failure to recognize biologically significant sequences.
Innovation Solution
A method involving statistical heuristics and minimum description length principles is employed to analyze data series using a grammar-based approach, where sub-sequences are identified as sequences of interest based on comparisons with reference conditions, and the grammar and data series are updated with symbols representing these sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional techniques are used to analyze large data series such as genomes, then the analysis can be performed with simple methods, but the computational efficiency is insufficient and meaningful sequences cannot be accurately identified
Solution Approach 1:
The patent transforms the sequence identification problem into a parameter optimization problem by defining a score function that evaluates sequences based on multiple parameters including repetition frequency, sequence length, and information content. This allows systematic identification of meaningful sequences through parameter thresholding rather than traditional pattern matching
Solution Approach 2:
The patent segments the large data series into overlapping windows or substrings of fixed length, then independently evaluates each segment using the score function. This division allows parallel processing and reduces the computational complexity of analyzing entire genomes at once while maintaining identification accuracy
2Measurement precision
If the data series is analyzed in detail to identify all meaningful sequences, then the identification accuracy improves, but the computational resources and time required increase significantly
Solution Approach 1:
The patent employs a scoring system that ranks sequences by their likelihood of being meaningful, then applies a threshold to select only those sequences exceeding the threshold. This partial action approach identifies the most significant sequences without exhaustively analyzing every possible substring, thereby reducing analysis time while maintaining high identification accuracy for biologically relevant sequences
Data Source
AI summary
The present technique provides for the analysis of a data series to identify sequences of interest within the series. The analysis may be used to iteratively update a grammar used to analyze the data series or updated versions of the data series. Furthermore, the technique provides for the calculation of a minimum description length heuristic, such as a symbol compression ratio, for each sub-sequence of the analyzed data sequence. The technique may then compare a selected heuristic value against one or more reference conditions to determine if additional iteration is to be performed. The grammar and the data sequence may be updated between iterations to include a symbol representing a string corresponding to the selected heuristic value based upon a non-termination result of the comparison. Alternatively, the string corresponding to the selected heuristic value may be identified as a sequence of interest based upon a termination result of the comparison.


