RNA Sequence and Structure Encoding Method
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current RNA data storage systems lack efficient representation of both nucleotide sequence and secondary structure information, leading to increased storage needs and computational burdens due to the absence of structural data in public repositories, requiring separate prediction steps that consume resources and time.
Innovation Solution
The encoding method combines nucleotide sequence and secondary structure information into a single encoded sequence string by affixing structural characters to nucleotide stretches and removing redundant information, allowing for efficient storage and decoding of RNA data without separate structure prediction steps.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If structural information is stored separately in public databases, then storage capacity for sequence data is sufficient, but additional storage space and computational resources are required for structure data
Solution Approach 1:
The patent merges nucleotide sequence information and secondary structure information into a single encoded sequence string. Structural characters are affixed to nucleotide stretches, creating a unified data representation that eliminates the need for separate storage of sequence and structure data, thereby reducing overall storage requirements and data management complexity
Solution Approach 2:
The encoded sequence string serves multiple functions simultaneously: it stores both the nucleotide sequence and the secondary structure information, and enables both sequence analysis and structure prediction capabilities in a single data structure, making the system more versatile and efficient
2Loss of information
If separate structure prediction steps are performed, then structural information can be obtained, but computational resources and time are consumed
Solution Approach 1:
The patent performs the structure prediction action in advance by pre-computing and storing structural characters alongside nucleotide sequences in the encoded format. This preliminary action eliminates the need for separate structure prediction steps when analyzing the data, as the structural information is already available in the encoded sequence string
Solution Approach 2:
By combining sequence and structure information into a single encoded format, the patent eliminates redundant computational steps. Researchers can analyze both sequence and structure simultaneously from the encoded string, improving computational efficiency while maintaining complete structural information availability
3Loss of information
If structural characters are affixed to all nucleotide stretches, then complete structure information is maintained, but data size increases
Solution Approach 1:
The patent applies structural characters selectively based on local characteristics of nucleotide stretches. Structural characters are affixed to stretches that require structural annotation, while redundant characters are removed from stretches where structure information is already implied or redundant, achieving optimal data compression while maintaining completeness
Solution Approach 2:
The encoding process discards redundant structural information that can be inferred from the sequence context or is unnecessarily repetitive, while recovering and preserving essential structural characters that are needed for accurate structure representation. This selective approach reduces data size while maintaining information completeness
Data Source
AI summary
Systems and methods to enable representation of sequence as well as structural information of an RNA molecule in the form of a single encoded string are described. The encoding steps are based on identifying one or more contiguous stretches of ribonucleotide bases having similar structural attributes and base-pairing patterns. In the encoded string, each of the identified contiguous structure stretches is represented by a single character that indicates the corresponding structural attribute. Appending these structural characters to the corresponding contiguous ribonucleotide character stretches, and subsequently eliminating redundant ribonucleotide characters based on standard base-pairing rules results in generating the final encoded string. Such concomitant representation of sequence and structural information of a given RNA molecule in a single encoded string enables efficient storage and easy dissemination of RNA data.


