RNA Sequence and Structure Encoding Method

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current RNA data storage systems lack efficient representation of both nucleotide sequence and secondary structure information, leading to increased storage needs and computational burdens due to the absence of structural data in public repositories, requiring separate prediction steps that consume resources and time.

Innovation Solution

The encoding method combines nucleotide sequence and secondary structure information into a single encoded sequence string by affixing structural characters to nucleotide stretches and removing redundant information, allowing for efficient storage and decoding of RNA data without separate structure prediction steps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If structural information is stored separately in public databases, then storage capacity for sequence data is sufficient, but additional storage space and computational resources are required for structure data

Engineering Contradiction:
Improvestorage capacityVSAvoiddata management complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent merges nucleotide sequence information and secondary structure information into a single encoded sequence string. Structural characters are affixed to nucleotide stretches, creating a unified data representation that eliminates the need for separate storage of sequence and structure data, thereby reducing overall storage requirements and data management complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The encoded sequence string serves multiple functions simultaneously: it stores both the nucleotide sequence and the secondary structure information, and enables both sequence analysis and structure prediction capabilities in a single data structure, making the system more versatile and efficient

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If separate structure prediction steps are performed, then structural information can be obtained, but computational resources and time are consumed

Engineering Contradiction:
Improvestructural information availabilityVSAvoidcomputational efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent performs the structure prediction action in advance by pre-computing and storing structural characters alongside nucleotide sequences in the encoded format. This preliminary action eliminates the need for separate structure prediction steps when analyzing the data, as the structural information is already available in the encoded sequence string

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By combining sequence and structure information into a single encoded format, the patent eliminates redundant computational steps. Researchers can analyze both sequence and structure simultaneously from the encoded string, improving computational efficiency while maintaining complete structural information availability

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If structural characters are affixed to all nucleotide stretches, then complete structure information is maintained, but data size increases

Engineering Contradiction:
Improvestructure information completenessVSAvoiddata size
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent applies structural characters selectively based on local characteristics of nucleotide stretches. Structural characters are affixed to stretches that require structural annotation, while redundant characters are removed from stretches where structure information is already implied or redundant, achieving optimal data compression while maintaining completeness

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The encoding process discards redundant structural information that can be inferred from the sequence context or is unnecessarily repetitive, while recovering and preserving essential structural characters that are needed for accurate structure representation. This selective approach reduces data size while maintaining information completeness

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS11101018B2Encoding and decoding of RNA data
Publication Date: 2021.08.24 TATA CONSULTANCY SERVICES LTD
  • US11101018B2 patent drawing
  • US11101018B2 patent drawing
  • US11101018B2 patent drawing

AI summary

Systems and methods to enable representation of sequence as well as structural information of an RNA molecule in the form of a single encoded string are described. The encoding steps are based on identifying one or more contiguous stretches of ribonucleotide bases having similar structural attributes and base-pairing patterns. In the encoded string, each of the identified contiguous structure stretches is represented by a single character that indicates the corresponding structural attribute. Appending these structural characters to the corresponding contiguous ribonucleotide character stretches, and subsequently eliminating redundant ribonucleotide characters based on standard base-pairing rules results in generating the final encoded string. Such concomitant representation of sequence and structural information of a given RNA molecule in a single encoded string enables efficient storage and easy dissemination of RNA data.