RNA Foundation Model for Structure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting RNA structure and function are limited by the scarcity of annotated data, particularly for non-coding RNAs, and lack effective end-to-end deep learning-based approaches for 3D structure modeling and functional group prediction.

Innovation Solution

A language model, referred to as the RNA foundation model, is trained on a large-scale dataset of unannotated RNA sequences to generate embeddings that can be used by downstream neural networks to predict structural and functional characteristics, such as secondary structure and RNA-protein interactions, without relying on hand-annotated information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing DL-based approaches are used for RNA structure prediction, then prediction accuracy can be improved, but the model cannot generalize well to unknown RNA types due to task-specific architecture design

Engineering Contradiction:
Improveprediction accuracyVSAvoidgeneralization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a task-agnostic RNA foundation model that can serve multiple downstream tasks. The model architecture is not specialized for any particular prediction task but is instead designed to learn general RNA sequence representations that can be applied to various structure and function prediction tasks, enabling both accurate predictions and good generalization to unknown RNA types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies preliminary action through pre-training the model on a large-scale dataset of RNA sequences before applying it to specific downstream tasks. This pre-training phase allows the model to learn fundamental RNA sequence patterns and representations in advance, which then improves its performance and generalization capability when applied to various prediction tasks.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If hand-annotated information is used for training, then model training can be performed, but the ability to generalize is limited due to scarcity of annotated data

Engineering Contradiction:
Improvemodel training feasibilityVSAvoidgeneralization ability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent applies self-service by implementing a self-supervised learning approach where the model learns from unannotated RNA sequences through masking and reconstruction tasks. Instead of relying on external hand-annotated data, the model generates its own training signals by predicting masked nucleotides, thereby eliminating the bottleneck of scarce annotated data while maintaining training feasibility.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies parameter changes by shifting the training paradigm from supervised learning (requiring annotated data) to self-supervised learning (using unannotated data). This fundamental change in the learning objective and data requirements allows the model to leverage the abundant unannotated RNA sequence data while still achieving effective training and generalization.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If thermodynamic methods are used for RNA secondary structure prediction, then computational simplicity is maintained, but important tertiary interaction information is missed

Engineering Contradiction:
Improvecomputational simplicityVSAvoidtertiary interaction information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies mechanics substitution by replacing traditional thermodynamic methods with a deep learning-based approach. Instead of using physics-based thermodynamic calculations that are computationally simple but miss tertiary interactions, the model uses neural networks to learn complex patterns from sequence data, capturing both secondary and tertiary interaction information while maintaining computational efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240331798A1Interpretable RNA foundation model for RNA structure and function predictions
Publication Date: 2024.10.03 THE CHINESE UNIVERSITY OF HONG KONG
  • US20240331798A1 patent drawing
  • US20240331798A1 patent drawing
  • US20240331798A1 patent drawing

AI summary

A foundation model for analysis of RNA sequences, including ncRNA sequences, can be trained to provide output embeddings (in a high-dimensional space) corresponding to input RNA sequences. Training of the RNA foundation model can use a large-scale dataset of RNA sequences without any annotation as to structure or function. The trained RNA foundation model can thereafter be used to produce embeddings that can be used as input features in downstream task-specific machine-learning models (or other computer models) that can learn to predict particular aspects of structure and/or function for a given RNA sequence.