Latent diffusion models for controllable RNA sequence generation

The RNA-Diffusion model addresses the challenges of RNA sequence generation by using a sequence autoencoder and guided diffusion with reward networks to produce RNA sequences with enhanced functional properties, outperforming existing methods in controllability and scalability.

WO2026060030A1PCT designated stage Publication Date: 2026-03-19THE TRUSTEES OF PRINCETON UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-10
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Conventional RNA sequence generation methods, including heuristic rules and basic statistical modeling, struggle with capturing the intricate and nonlinear dependencies of RNA sequences, particularly in terms of controllability, scalability, and precise functional targeting, while machine learning-based models like VAEs and GANs face challenges in encoding variable-length, symbol-based data into a continuous latent space.

Method used

A latent diffusion model, RNA-Diffusion, is developed, utilizing a sequence autoencoder with a pretrained RNA language model and Querying Transformer to map RNA sequences into a fixed-length latent space, combined with a guided diffusion model and reward network to predict functional properties like Mean Ribosome Loading and Translation Efficiency, enabling controlled generation of RNA sequences.

Benefits of technology

The model generates RNA sequences that closely resemble natural non-coding RNAs, achieving significant improvements in functional metrics such as 166.7% higher Translation Efficiency and 52.6% higher Mean Ribosome Loading compared to baselines, and supports scalable generation of varying-length sequences without padding or truncation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045798_19032026_PF_FP_ABST
    Figure US2025045798_19032026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments herein provide for generating synthetic RNA sequences. In one embodiment, a system comprises an interface operable to communicatively couple to a database comprising a dataset of training RNA sequence, and to receive the dataset of training RNA sequences. Each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations. The system also comprises a processor that is operable to train a generative machine learning model on the dataset of training RNA sequences, to encode the training RNA sequences into a latent representation using an encoder network of the generative machine learning model, and to decode the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences. The one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.9903.019WO1 LATENT DIFFUSION MODELS FOR CONTROLLABLE RNA SEQUENCE GENERATION

[0001] The present application claims priority to, and thus the benefit of an earlier filing date from, U.S. Provisional Application No.63 / 693,285 (filed on September 11, 2024), the contents of which are incorporated herein by reference in their entirety. Field

[0002] The disclosure relates to genetics, and more specifically to the field of RNA sequence generation. Background

[0003] The field of synthetic biology and bioinformatics has witnessed remarkable advancements in recent years, particularly in the area of RNA sequence generation and analysis. Ribonucleic acid (RNA) molecules, especially non-coding RNAs and untranslated regions (UTRs) of mRNA, play critical roles in various cellular functions including gene regulation, translation efficiency, and protein synthesis. However, the inherent complexity of RNA sequences, such as variable lengths, lack of well-resolved tertiary structures, and non-linear sequence-function relationships, poses significant challenges for traditional generative models.

[0004] Conventional approaches to RNA sequence design often rely on heuristic rules, exhaustive search, or basic statistical modeling, which are inadequate for capturing the intricate and nonlinear dependencies inherent in biological sequences. Recently, machine learning-based generative models, such as variational autoencoders (VAEs) and generative adversarial networks (GANs), have been explored in biological domains, but they still struggle with controllability, scalability to diverse sequence lengths, and precise functional targeting.

[0005] Diffusion probabilistic models, originally developed for continuous data generation tasks in fields such as computer vision and time-series modeling, have emerged as powerful generative techniques with enhanced controllability and sampling stability. However, their application to discrete biological sequences, particularly RNA, remains underexplored due to the challenges of encoding variable-length, symbol-based data into a continuous latent space suitable for diffusion.Attorney Docket No.9903.019WO1 Summary

[0006] Embodiments described herein provide RNA-Diffusion, a novel latent diffusion framework designed for the generation and optimization of discrete RNA sequences with controllable biochemical properties. This system leverages a latent space approach, inspired by advances in computer vision diffusion models, and adapts it to accommodate the unique characteristics of RNA.

[0007] In one embodiment, the disclosed system comprises an RNA sequence autoencoder that maps variable-length RNA sequences to a fixed-length latent space using a pretrained RNA language model (RNA-FM) and a Querying Transformer, or “Q-Former”. This allows for efficient encoding of biologically meaningful features. The system also comprises a guided diffusion model trained in the latent space using a score-based denoising network, capable of sampling new RNA latent embeddings under external guidance. The system also comprises a latent reward model trained to predict functional properties such as Mean Ribosome Loading (MRL) and Translation Efficiency (TE), which are critical indicators of mRNA translation performance. This model provides gradient-based guidance to steer the diffusion process toward sequences with desired biological functionality.

[0008] The system supports plug-and-play optimization through external guidance signals, enabling the generation of RNA sequences, such as 5’ UTRs that are fine-tuned for high translation efficiency and ribosome loading. Empirical results demonstrate that sequences generated by RNA-Diffusion closely resemble natural non-coding RNAs across multiple biological indicators and significantly outperform baselines in functional metrics.

[0009] Additionally, the disclosed model accommodates the generation of varying- length sequences without the need for padding or truncation, making it highly scalable. The system, including model architecture, training and inference code, and pretrained weights, may be commercialized through APIs or licensing frameworks for applications in mRNA therapeutic design, synthetic biology, and RNA functional analysis.

[0010] In one embodiment, a method for generating synthetic RNA sequences comprises training a generative machine learning model on a dataset of training RNA sequences. Each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations. The method also comprises encoding the training RNA sequences into a latent representation using an encoder network of the generative machine learning model, andAttorney Docket No.9903.019WO1 decoding the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences. The one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure. The one or more biological annotations may include, among other things, a minimal Levenshtein distance, a minimal 4-mer distance, a guanine-cytosine content, a minimum free energy, a structure property, and / or a functional property of the training RNA sequences. And the generative machine learning model may include, among other things, at least one of a variational autoencoder, a transformer-based autoregressive model, or a generative adversarial network.

[0011] In some embodiments, the method also includes training the generative machine learning model using a loss function comprising a reconstruction loss and / or estimating a functional property of RNA with a trained reward network of the generative machine learning model. In some embodiments, the training RNA sequences comprise variable lengths, and the method further comprises training the reward network directly in a fixed-length latent space that is mapped to the variable length training RNA sequences. The method may also comprise using a guided diffusion model to generate latent RNA embeddings based at least on guidance from the reward network.

[0012] In another embodiment, a non-transitory computer readable medium comprises instructions for generating synthetic RNA sequences that, when executed by a processor, direct the processor to train a generative machine learning model on a dataset of training RNA sequences. Each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations. The instructions may also direct the processor to encode the training RNA sequences into a latent representation using an encoder network of the generative machine learning model, and to decode the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences. The one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.

[0013] In another embodiment, a processing system is operable to generate synthetic RNA sequences. The processing system comprises an interface operable to communicatively couple to a database comprising a dataset of training RNA sequence, and to receive the dataset of training RNA sequences. Each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations. The processing system also comprises a processor that is operable to train a generative machine learning model on the dataset of training RNAAttorney Docket No.9903.019WO1 sequences, to encode the training RNA sequences into a latent representation using an encoder network of the generative machine learning model, and to decode the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences. The one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.

[0014] Other illustrative embodiments (e.g., methods and computer-readable media relating to the foregoing embodiments) may be described below. The features, functions, and advantages that have been discussed can be achieved independently in various embodiments or may be combined in yet other embodiments, further details of which can be seen with reference to the following description and drawings. Description of the Drawings

[0015] Some embodiments of the present disclosure are now described, by way of example only, and with reference to the accompanying drawings. The same reference number represents the same element or the same type of element on all drawings.

[0016] FIG.1 is a block diagram of an RNA sequence autoencoder 100, in one exemplary embodiment.

[0017] FIG.2 illustrates a block diagram of a latent diffusion module 200, in one exemplary embodiment.

[0018] FIG.3 is a graph illustrating a distribution of generated RNA sequences having varying lengths that closely align with natural ncRNAs, in one exemplary embodiment

[0019] FIG.4 illustrates minimum 4-mer distances of generated RNA sequences, in one exemplary embodiment.

[0020] FIG.5 illustrates minimum sequence Levenshtein distances of generated RNA sequences, in one exemplary embodiment

[0021] FIG.6 illustrates GC content ratios of generated RNA sequences, in one exemplary embodiment.

[0022] FIG.7 illustrates minimum free energy of generated RNA sequences, in one exemplary embodiment.

[0023] FIG.8 illustrates minimum Levenshtein distances of RNA secondary structure, in one exemplary embodiment.Attorney Docket No.9903.019WO1

[0024] FIG.9 illustrates t-SNE visualizations of RNAs in latent embedding space, in one exemplary embodiment.

[0025] FIG.10 illustrates sequence length histograms that compare generated sequences to natural 5’-UTRs, in one exemplary embodiment.

[0026] FIGS.11 and 12 are graphs illustrating performances of guided diffusion, in one exemplary embodiment.

[0027] FIG.13 is a flowchart depicting a method for generating synthetic RNA sequences, in one exemplary embodiment.

[0028] FIG.14 depicts an illustrative computing system operable to execute programmed instructions embodied on a computer readable medium. Detailed Description

[0029] The figures and the following description depict specific illustrative embodiments of the disclosure. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody the principles of the disclosure and are included within the scope of the disclosure. Furthermore, any examples described herein are intended to aid in understanding the principles of the disclosure, and are to be construed as being without limitation to such specifically recited examples and conditions. As a result, the disclosure is not limited to the specific embodiments or examples described below, but by the claims and their equivalents.

[0030] The embodiments herein provide for RNAdiffusion, a latent diffusion model for generating and optimizing discrete RNA sequences. RNA is a particularly dynamic and versatile molecule in biological processes. RNA sequences exhibit high variability and diversity amongst biological sequences, due to their variable lengths, often unresolved three-dim structures, diverse functions and alternative splicing. The embodiments herein utilize pretrained BERT-type models (Bidirectional Encoder Representations from Transformers) to encode raw RNAs into token- level semantic representations. A Q-Former is employed to compress these representations into a fixed-length set of latent vectors, with an autoregressive decoder trained to reconstruct RNA sequences from these latent variables. Then, a continuous diffusion model within this latent space is developed. To enable optimization, reward networks are trained to estimate functional properties of RNA from the latent variables. A gradient-based guidance during the backwardAttorney Docket No.9903.019WO1 diffusion process is used, aiming to generate RNA sequences that are optimized for higher rewards. Empirical experiments confirm that RNAdiffusion generates non-coding RNAs that align with natural distributions across various biological indicators. The diffusion model on untranslated regions (UTRs) of mRNA is fine-tuned and sample sequences for protein translation efficiencies are optimized. The guided diffusion model effectively generates diverse UTR sequences with high Mean Ribosome Loading (MRL) and Translation Efficiency (TE). These results hold promise for studies on RNA sequence-function relationships, protein synthesis, and enhancing therapeutic RNA design.

[0031] Diffusion models demonstrate exceptional performances in modelling continuous data, with applications in images synthesis, point clouds generation, video synthesis, reinforcement learning, time series, and molecule structure generation. An important advantage of diffusion models is that their generation process can be controlled to achieve specific objectives via incorporating additional guidance signal. The guidance can steer the backward process toward generating samples with desired properties without additional training.

[0032] Beyond continuous domains, people have explored using diffusion models to generate discrete sequences to inherit the benefits of controllability through guidance. For example, in the domain of text generation, state-of-the-art performances are achieved by scaling up auto-regressive models. However, for generating biological sequences such as proteins, DNAs, and RNAs, language models have to be adapted to handle their specific sequence characteristics.

[0033] Generative models for RNAs introduce distinct challenges in machine learning due to their inherent complexity. Unlike DNA, which is often characterized by stable and predictable regions such as coding sequences, RNA sequences demonstrate higher variability and diversity. They can have missing tokens and vary dramatically in length due to processes like alternative splicing, which allows splicing a gene sequence in multiple ways to produce different isoforms. Additionally, unlike proteins, the three-dimensional structures of many RNAs remain unresolved, further complicating the modeling of their functions.

[0034] Generative models for RNA can be particularly useful for designing untranslated regions (UTR) of mRNA vaccines, which play a central role in regulating gene expression and mRNA stability. The ability to generate optimized UTR sequences offers significant potential forAttorney Docket No.9903.019WO1 controlling gene expression and advancing targeted therapies. The embodiments herein provide a robust, powerful language model for generating non-coding RNAs and optimizing UTRs.

[0035] The functional properties of RNAs, especially non-coding RNAs, involve complex nonlinear inter-actions within the sequence, which do not conform to a simple left-to- right progression. Thus, autoregressive models are inadequate for capturing their patterns and dynamics. Diffusion models, on the other hand, allow for the holistic modeling and generation of RNA sequences. Raw tokens (i.e., A,U,G,C) are generally too low-level and they need to be grouped (e.g., into codons or structured motifs) to exhibit biological functions. Thus, a denoising network needs to account for both the high-level picture of the entire sequence and the low-level syntax. Additionally, RNAs have dramatically different sequence lengths, adding difficulty to efficient training and generalization. Although transformer denoisers can handle variable-length inputs, current systems either pad / truncate inputs to a fixed length or build an additional length prediction module to ensure robust practical performances.

[0036] To overcome these challenges, diffusion models herein operate on a latent space for RNA sequences. In one embodiment, a sequence autoencoder maps raw input sequences into a latent embedding space. And a score-based diffusion model is trained to learn this latent distribution. One example of the auto encoder is illustrated in FIG.1. This latent-space approach is particularly effective for RNA sequences for several reasons. First, it delegates the generation of the intricate details of raw RNA sequences to the decoder, allowing the diffusion model to concentrate on capturing the coarse-grained distribution in the meaningful latent space. Second. the latent space is designed to be of fixed size, simplifying both the design and training of the diffusion model by eliminating the need for additional modules or adaptations to accommodate varying sequence lengths.

[0037] For the sequence autoencoder, a pretrained RNA language model (RNA-FM) is adopted to map raw sequences to biologically-meaningful token-level representations. Then a trained Q-Former is used to summarize the token representations into a fixed-length list of latent vectors. Next, a standard score-based denoising diffusion model is trained to model the distribution of RNAs in the latent space. For validation, the set of generated sequences is analyzed and a range of biological indicators is examined, such as minimal Levenshtein distance, minimal 4-mer distance, GC content (guanine-cytosine content), minimum free energy, secondary structure, etc. Empirical experiments confirm that the generated sequences align withAttorney Docket No.9903.019WO1 natural distributions of non-coding RNAs across these biological indicators. In other words, the RNA latent diffusion models herein can generate novel RNA sequences that are highly similar to natural sequences.

[0038] As latent variables from Q-Formers provide biologically-meaningful summarization, reward networks are trained directly in the latent space to predict RNA’s functional properties. The rewards of interests are the Mean Ribosome Loading (MRL) and the Translation Efficiency (TE), which measure the mRNA-to-protein production levels and efficiencies. These functional properties of mRNA are largely determined by UTRs. By using trained reward models, their gradients are computed and guidance for the diffusion model is constructed. Thus, optimized UTR sequences can be generated toward higher values of MRL and TE by adding this guidance to the backward generation process of the pretrained diffusion model. Adding guidance may be done in a plug-and-play manner, without additional training. Empirically, the generated UTRs from guided diffusion achieve a 166.7% improvement of TE and a 52.6% improvement of MRL compared to mRNA baselines.

[0039] To summarize, the first latent diffusion model for RNAs is presented herein. The sequence autoencoder involves a novel adaptation of Q-Formers that summarizes the varying- length output from pretrained encoders into fixed-length embeddings. Empirical study confirmed that the model can generate non-coding RNAs that align with natural distributions across various biological indicators. Q-Formers have proven effective for retaining biological information in the latent space. In particular, reward models are trained to predict the Mean Ribosome Loading (MRL) and Translation Efficiency (TE) of 5’-UTR sequences with Spearman R of 0.69 and 0.56, respectively. The latent diffusion model allows controlled generation via guidance. 1. Methodology

[0040] RNAdiffusion generally has three parts: (1) a sequence autoencoder that maps between a space of variable-length discrete RNA sequences and a continuous latent space; (2) a continuous diffusion model that is trained to model the distributions in the latent space, with optionally added guidance; and (3) a reward model trained on top of the latent space for computing guidance. 1.1 RNA Sequence Autoencoder

[0041] FIG.1 illustrates a block diagram of an RNA sequence autoencoder 100, in one exemplary embodiment. The RNA sequence auto encoder 100 is operable to remove informationAttorney Docket No.9903.019WO1 redundancy in the raw sequence of tokens and compress them into a fixed-length set of continuous embedding vectors. The RNA sequence autoencoder 100 comprises an RNA-FM model (encoder) 104, a Q-Former 102, and a decoder 110. The Q-Former 102 is operable to translate between the sequence space and the latent space 106.

[0042] Encoder 104: A pretrained RNA-FM (e.g., a BERT-type encoder-only model) is the first component of the sequence autoencoder 100. BERT-type encoders are pretrained via masked language modeling task to predict randomly masked tokens. The encoder 104 maps a sequence to a sequence of contextualized token embeddings of the same length L. fixed during the training of the sequence autoencoder 100.

[0043] Q-Former 102: Next, the Q-Former 102, denoted by , is used to summarize the embeddings of the pre-trained encoder into K latent, where K is a fixed number independent of L. This achieves three goals:in latent vectors, making it easy to train the decoder 110 for raw sequence reconstruction; (2) it uses latent vectors later for predicting RNA functional properties, via training additional prediction models on top of these embedding; and (3) it reduces the dimensionality of raw sequences, so that diffusion models can focus on learning intrinsic structure of latent RNA vectors.

[0044] The Q-Former 102 takes K trainable query token embeddings as input. The query tokens go through the transformer with cross-attentions tosequence given by the encoder and progressively summarize the original token embeddings of varying lengths into a fixed-size sequence of embedding vectors. This is a flexible approach compared to using only the <cis> token embedding. When K is increased, more information in the latent vectors is retained and allows accurate reconstruction, which helps to mitigate the KL- vanishing problem.

[0045] Decoder 110: A standard causal transformer is used as the decoder 110, denoted by . A simple linear projection layer is built to convert the K latent vectors into soft prompt embeddings 108 of equal size for the decoder 110 to condition on. Then the decoder 110 autoregressively reconstructs the original input sequence given the latent vectors.

[0046] Training Objective: With the pre-trained encoder fixed, the Q-Former 102 and the decoder 110 are trained from scratch in an end-to-end manner. The training objective is toAttorney Docket No.9903.019WO1 reconstruct the original sequence from the embeddings given by the fixed pretrained encoder model. The loss function is given by

[0047] where query decoder 110, respectively. A small (KL-regularization term e.g., 1e-6) in latent diffusion models is typically used, and any KL-regularization is excluded, simplifying the design and helping to avoid KL-vanishing. Empirically, the training works well without the KL term. 1.2 Latent Diffusion Module 200

[0048] FIG.2 illustrates a block diagram of a latent diffusion module 200, in one exemplary embodiment. The latent diffusion module 200 includes a classical continuous diffusion model 204 built to model the distribution of the latent variables z, which generates samples from the target distribution by a series of noise removal processes. Specifically, a denoising network is adopted to predict the added noise from the noisy input , where t is a time step of the forward process, is thestrength level, and is the added Gaussian noise.

[0049] The guided diffusion model 204, with a pre-trained score network, generates latent RNA embeddings under external guidance (206). A latent reward model 202 is trained on the latent space 206 to predict functional properties of RNA, for computing guidance of diffusion.

[0050] For the denoising score network denoiser Θ, transformer denoising networks are used. The prediction scheme is used in Denoising Diffusion Probabilistic Models (DDPM) and adopts the simplified training objective given by: .easy to train diffusion models and treat data as if it is continuous like images. Unlike other discrete diffusion models, this framework does not require truncating / padding sequences to a fixed length, and it does not require training an additional length predictor. An auto-regressive decoder has been trained to generate sequences with variable lengths.Attorney Docket No.9903.019WO1 1.3 Guided Generation in the Latent Space

[0053] To generate RNA sequences of desired properties, the reward model 202 for predicting the property of new sequences is used. For example, there is a labeled dataset , where r is the measured reward of interest (e.g., the translation efficiency). A reward is trained on top of the latent space 206 to predict r given the input sequence x. Theis

[0054] where is the embedding obtained from the Q-Former 102, and are200 and to.

[0055] Finally, the reward model 200 and to is used to compute a gradient-based guidance (e.g., via universal guidance) and add it to the backward generation process. This guidance steers the random paths of the backward generation process of the diffusion model 204 toward a higher reward region in a plug-and-play manner. Specifically, at the step t of the backward diffusion process, given the current noisy sample zt, a clean sample is first estimated as: .noisy sample zt. The gradientinto the current predicted noise as follows, with r* being theguidance strength. For example, . Controlling the generation for decoder-only modelsgenerally hard without fine-tuning the model. 2. RNA Sequence Autoencoder 2.1 Experiment Setup

[0057] Models: The pretrained RNA-FM 104 is adopted as the encoder. RNA-FM 104 is an encoder-only model trained with masked language modelling on non-coding RNAs (ncRNAs), which can extract biologically meaningful embeddings from ncRNAs and can be adapted for various downstream tasks, such as UTR function prediction, RNA-protein interactionAttorney Docket No.9903.019WO1 prediction, secondary structure prediction, and etc. For the Q-Former 102, the ESM-2 model architecture is modified by adding K trainable query token embeddings and adding cross- attention layers in each transformer block that attends to the embeddings from the RNA-FM 104. For the decoder 110 model, the same model architecture of ProGen2-small is adopted our customized tokenizers for RNAs.

[0058] Dataset: The RNA sequence autoencoder 100 is pre-trained with the ncRNA subset of Ensembl database downloaded from RNAcentral, which contains collected ncRNA sequences from a variety of vertebrate genomes and model organisms. The sequences are selected to be shorter than 768 bp to avoid exceeding RNA-FM’s maximum length limit and exclude those with rare tokens more than {A, C, G, U}, curating a dataset of 1.1M samples for training and testing. 2.2 Results

[0059] A sweep over hyperparameters is performed and obtains a set of RNA sequence autoencoders with high reconstruction capabilities. Specifically, there are two hyperparameters that influence the dimension of the latent space – the number of query tokens K and the hidden dimension D of each latent vector. K is tuned among {16, 32} and D among {40, 80, 160, 320}. For each input sequence x of length L, the length-normalized reconstruction loss (NLL) iscalculated as: . The length-normalized editdistanceand the a held-outhidden dimension D reduces the reconstruction errors and the error rates. Increasing the dimension of the latent space retains more information in the latent embeddings. Hence, it is easier to reconstruct the input sequence.Attorney Docket No.9903.019WO13. Latent Diffusion Models for RNAs 3.1 Experimental Set up

[0060] Models and Datasets: A 24-layer Transformer model with hidden dimension 2048 is built as the denoising network for the latent diffusion model. As the latent vectors are continuous, there is need any additional embedding layer or rounding operation. The same Ensembl ncRNA dataset is used and applies the same preprocessing as the RNA sequence autoencoder 100. The diffusion model 204 for 10 epochs with the sequence autoencoder frozen.

[0061] Generation Details: The DDIM sampling method with η= 0 (deterministic) and 50 denoising steps to get samples from our latent diffusion models. Then, the latent vectors are transformed into soft prompts (e.g., element 108 of FIG.1) for the decoder 110 and auto- regressively generate a sample sequence nucleus-by-nucleus with top-p = 0.95 and temperature 1. Any invalid sample sequence containing special tokens in the middle (e.g., <bos>) is removed.

[0062] Evaluation Metrics: To evaluate the performance of the latent diffusion model, the weighted loss is calculated on the test dataset according to:, which is an upper. visualized by using a t- SNE map, and comparing it with reference natural RNAs and randomly generated sequences, which match the length distribution of natural data. The generated latent vectors are fed through the decoder 110 to get the RNA sequences.

[0063] Minimum Levenshtein Distance: The Levenshtein distance is defined as the minimum number of edits (e.g., insertions, deletions, and / or substitutions) required to transform one sequence into another. For each generated RNA sequence, the smallest Levenshtein distanceAttorney Docket No.9903.019WO1 to all sequences in a reference dataset of natural RNAs is calculated, and the distance is normalized by the length of the generated sequence. By looping over all the generated RNA sequences, a distribution of minimum normalized Levenshtein distances is obtained.

[0064] Minimum 4-mer Distance: The frequencies of 4-mers of each sequence are calculated. The 4-mer distance of two sequences is defined as the Euclidean distance between the two frequency vectors. Similarly, for each generated RNA sequence, the minimum 4-mer distance against a reference natural RNA dataset is calculated. The distribution of minimum 4- mer distances for the generated RNA sequences is then obtained.

[0065] GC content: GC content, defined as the percentage of guanine (G) and cytosine (C), is an important indicator of molecular stability. A higher GC content generally implies higher stability due to the stronger hydrogen bonds between G and C compared to A and T.

[0066] Minimum Free Energy (MFE): MFE represents the minimum energy required for the molecule to maintain its conformation. The MFE is crucial measure of an RNA segment’s structural stability, reflecting how tightly the RNA molecule can fold based on its nucleotide composition and sequence arrangement. A lower MFE suggests a more stable and tightly folded RNA structure. The MFE of generated RNA sequences is calculated using a ViennaRNA Package.

[0067] Minimum Levenshtein Distance of RNA Secondary Structure: RNA secondary structure, which involves local interactions like hairpins and stem-loops, plays an important role in the RNA’s biological functions. Dot-bracket notation represents these structures by using dots for unpaired nucleotides and matching parentheses for base pairs, simplifying the visualization and analysis of RNA sequences. The secondary structure of RNA sequences is calculated using the ViennaRNA Package, and then the distribution of the minimum normalized Levenshtein distances of the secondary structures of the generated RNAs is calculated to those of natural referential RNAs. 3.2 Results - Samples from RNAdiffusion Resemble Natural ncRNAs

[0068] The test NLL of the latent generated distribution is seen in Table I. Increasing the query token length K and the embedding dimension D generally decreases the test NLL per dimension. This is expected from the information-theoretic view, as increasing the dimension of the latent space means that fewer bits per dimension are needed to represent the uncertainty sourced from the distribution of the discrete sequences. Nevertheless, doubling the totalAttorney Docket No.9903.019WO1 dimension of the latent space does not necessarily decrease the NLL by half, as this also allows the autoencoder to retain more information in the latent space, which increases the entropy of the latent distribution.

[0069] For subsequent experiments, the trained RNA autoencoders have been used that have average normalized edit distance < 1%. They correspond to hyperparameters (K, D) = (16, 320), (32, 160), and (32, 320). The parameter (32, 160) is selected for generation, which achieves the smallest total NLL, i.e., NLL per dimension multiplied by the dimension of latent space (K x D), excluding the identical sequence when calculating the minimum Levenshtein distance and minimum 4-mer distance.

[0070] Generated sequences vary in length, with a distribution closely aligned with that of natural ncRNAs, as illustrated in FIG.3. FIG.4 illustrates minimum 4-mer distances. FIG.5 illustrates minimum sequence Levenshtein distances, FIG.6 illustrates GC content ratios. FIG.7 illustrates minimum free energy, FIG.8 illustrates minimum Levenshtein distances of RNA secondary structure. And FIG.9 illustrates t-SNE visualizations of RNAs in latent embedding space.

[0071] The distributions of the aforementioned metrics of the generated sequences are plotted and compared to a held-out natural ncRNA reference set. Additionally, a curated random sequence set is created by replacing each token in the reference sequences with nucleotide randomly sampled at a probability of 0.5 for GC, approximating the average GC content ratio of natural ncRNAs, while maintaining the original sequence length.

[0072] The sequences generated by the latent diffusion model exhibit greater similarity to the natural sequences compared to the random sequences shown in FIGS.4-9. These figures show the generated samples from RNAdiffusion compared to natural ncRNAs and random sequences, each including 9000 sequences.

[0073] The Levenshtein and 4-mer distances underscore the proximity to the closest natural sequences. The results reveal that both generated and natural sequences exhibit similar modes in terms of distances to the nearest natural sequence whereas randomly generated samples display a distinct distribution. This implies that the generative model not only captures the contextual characteristics of natural RNA sequences but also produces novel sequences that maintain a minimum distance from natural samples. Additionally, generated sequences demonstrate a mean GC content similar to that of natural samples while displaying a broaderAttorney Docket No.9903.019WO1 range of GC content compared to random samples. The generated RNA sequences further exhibit distributions of Minimum Free Energies (MFEs) and secondary structures that considerably resemble natural patterns. This suggests that the latent diffusion model can generate stable RNA molecules with plausible secondary structures. Moreover, in the latent space shown in t-SNE maps, the generated sequences predominantly overlap with natural sequences and more accurately capture several separated modes presented in natural sequences. In contrast, random sequences fail to capture these distinct modes and display a larger shift from the natural distribution.

[0074] The consensus secondary structures for various subsets of the generated RNA sequences are generated and compared against the natural sequences. The generated RNA sequences adopt secondary structures similar to those conserved in some common non-coding RNA families like tRNA, rRNA, and snRNA. Notably, the model does not explicitly condition on any specific families during generation. 4 Optimizing 5’-UTR Sequences with Guided Diffusion

[0075] The 5’ untranslated region (5’-UTR), located at the beginning of mRNA before the coding sequence, plays an important role in regulating protein synthesis by impacting mRNA stability, localization, and translation. Here, the focus is on generating 5’-UTR sequences with high protein production levels through guided diffusion.

[0076] Consider two functional properties of 5’-UTRs that correlate with protein production levels - Mean Ribosome Loading (MRL) and Translation Efficiency (TE). MRL measures the number of ribosomes actively translating an mRNA, reflecting its translation efficiency. In molecular biology, techniques like polysome profiling (Ribo-seq) are employed to determine MRL. TE represents the rate at which mRNA is translated into protein. It is determined by dividing the RNA-seq RPKM, which shows ribosomal footprints, by the RNA-seq RPKM that measures the mRNA’s abundance in the cell. This calculation helps gauge how efficiently an mRNA is being translated, independent of its expression level. 4.1 Fine-tuning RNAdiffusion for 5’-UTRs

[0077] Although both UTRs and ncRNAs share similarities in their non-coding properties, UTRs contain important functional elements that are irrelevant to ncRNAs. Thus, the pretrained RNA diffusion could not be directly applied to generate UTR sequences due to the distributional shift between UTRs and ncRNAs. To adapt RNA diffusion to generate 5’-UTRAttorney Docket No.9903.019WO1 sequences, a score function of RNA diffusion is fine-tuned using the five-species 5’-UTR dataset from UTR-LM, where sequences with lengths less than 768 and get 205K samples were elected. And the diffusion model for 1 epoch on top of the frozen sequence autoencoder with (K, D) = (32, 160) is fine-tuned.

[0078] Upon testing the autoencoder on UTR test set, it shows considerable low reconstruction errors (e.g., NLL= 0.0001, NED=0.01% ± 0.31%), confirming that the autoencoder pretrained with ncRNAs could be transferred onto UTRs. To compare the generated sequences of the fine-tuned diffusion model with natural 5’-UTRs, the sequence length histograms are illustrated in FIG.10 and the metrics are computed. The results demonstrate that they align closely with the natural distribution. 4.2 Latent Reward Modeling for MRL and TE

[0079] Datasets: For the TE prediction task, three endogenous human 5’-UTR datasets were analyzed. These contain 28.246 sequences and their measured TE.10% of the dataset was reserved for testing, 10% was reserved for validation, and the remaining 80% was used for training. For the MRL prediction task, following the length-based held-out testing approach, the model was trained on a collection of 83,919 random 5’-UTR sequences ranging from 25 to 100 bp with MRL measurement.10% of the data was reserved for validation. Subsequently, the model was tested on 7,600 human 5’-UTR sequences.

[0080] Model Architecture: The latent reward networks take as input a list of K latent vectors, obtained by the sequence autoencoder. A 6-layer Convolutional Residual Network (ResNet) was constructed specifically designed for 1D convolution, applicable to both the MRL and TE prediction tasks.

[0081] Results: For each task, two reward models were trained on top of the latent spaces from the pretrained RNA sequence autoencoder ((K, D) = (32, 160)), each initialized under distinct random seeds.

[0082] One model is utilized for guided generation as described in Section 4.3, while the other is employed for validation assessment. The Spearman and Pearson correlation coefficients were calculated between the predicted rewards and the actual rewards as shown in Table 2. Notably, the Spearman R values on the held-out test sets are 0.56 for the TE task and 0.69 for the MRL task, achieving slightly inferior performance compared with baselines and the state-of-the- art predictor built upon UTR-LM. The results support that the embeddings after Q-FormerAttorney Docket No.9903.019WO1 remain useful for downstream tasks, allowing for effective reward models to be built based on them.the on same 4.3 Reward-Guided Diffusion

[0083] Gradient-based forward universal guidance in the latent space was computed with the reward models trained in Section 4.2 to shift the latent backward sample paths towards generating high MRLs1TEs sequences. The loss function l was chosen to be L2loss, and the target reward r* and the guidance strength λ were varied. The latent reward models were trained with one random seed for TE and MRL optimization. For cross-validation, the reward models trained with the other random seed were employed as the oracle to evaluate the rewards of generated sequence.

[0084] Results: The guidance was tuned by varying the value of guidance strengths λ and target reward values r*. Performances of output from the guided diffusion are shown in FIGS.11 and 12. At a mild guidance strength, increasing the target TE or MRL values generally brings the reward improvements of generated sequences, demonstrating the efficacy of the reward-guided generation approach. Among all the guidance settings, the best validation TE (1.84 ± 0.82) is achieved with the moderate guidance strength as 800 and the largest log-target value as 5, showing 166.7% improvement compared to the validation TE of unguided generation (0.69 ± 0.22). Under the MRL-guidance setting of λ = 800 and log r* = 4 , mean validation MRL of the generated sequences increases from 3.92 ± 1.31 to 5.98 ± 0.90 with maximum relative increase of 52.6%.

[0085] Guidance Tradeoffs: For higher guidance strength, the improvements can be artificially amplified when evaluated with the same reward model used for guidance. Moreover, setting λ and r* too large may eventually hurt the cross-validation performance since theAttorney Docket No.9903.019WO1 generation process is over-adapted towards the imperfect guidance reward model, which is a general phenomenon known as reward hacking.

[0086] To get insights into the phenomenon, the generated distributions were compared under different guidance strengths in terms of biological metrics. With the λ and r* increasing, the disparities in GC content (e.g., lower mean) and Minimum Free Energy (e.g., higher mean) between the generated sequences and natural sequences become larger, which might be attributed to the distribution shift of the generated sequences. Adopting moderate λ’s and r*’s may be important for reward-guided generation in terms of the trade-off between achieving higher reward values and maintaining generation qualities.

[0087] More concretely, both reward evaluation and Minimum Free Energy (MFE) calculation were combined into Pareto front curves and compared with two competitive baselines. A gradient-guided latent diffusion model achieves higher rewards with the same MFE level compared to Random and UTRGAN. Although the sequences optimized by UTRGAN could achieve high rewards, they consistently exhibit high free energy. This demonstrates that the method has a better reward-structural stability trade-off compared to other baselines. 5. Conclusion

[0088] The embodiments herein provide for latent diffusion models on RNA generation. The pipeline involves a novel adaptation of the Q-Former to handle the inherently discrete, variable-length characteristics of RNA sequences. Through the incorporation of the gradient guidance of the latent reward models, the latent diffusion model can generate novel 5’-UTRs with 166.7% higher mean TEs and 52.6% higher mean MRLs. These results provide a biologically meaningful demonstration for RNA generation, a less explored yet exciting area of research.

[0089] This work aims to generate 5’-UTRs with optimized protein production levels, which could significantly enhance the understanding and engineering of gene expression. This can lead to breakthroughs in developing new therapies for diseases such as cancer and genetic disorders. It can also improve synthetic biology systems that produce valuable compounds, such as pharmaceuticals and bio-fuels, more efficiently and sustainably.

[0090] FIG.13 is a flowchart of a method 300 for generating synthetic RNA sequences, in one exemplary embodiment. In this embodiment, a generative machine learning model 100 is trained on a dataset of training RNA sequences, in the process element 302. Each training RNAAttorney Docket No.9903.019WO1 sequence is represented by a nucleotide sequence and by one or more biological annotations. Then, an encoder of the generative machine learning model, such as the RNA-FM encoder 104 of FIG.1, encodes the training RNA sequences into a latent representation, in the process element 304. Then a decoder of the generative machine learning model, such as the decoder 110 of FIG.1, decodes the latent representation to generate one or more novel RNA sequences, in the process element 306. Each of the novel RNA sequences may exhibit at least one of a desired biological characteristic or a desired physical structure.

[0091] Any of the various computing and / or control elements shown in the figures or described herein may be implemented as hardware, as a processor implementing software or firmware, or some combination of these. For example, an element may be implemented as dedicated hardware. Dedicated hardware elements may be referred to as “processors,” “controllers,” or some similar terminology. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, a network processor, application specific integrated circuit (ASIC) or other circuitry, field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), non-volatile storage, logic, or some other physical hardware component or module.

[0092] In one embodiment, instructions stored on a computer readable medium direct a computing system of any of the devices and / or servers discussed herein, such as health server 220, to perform the various operations disclosed herein. In some embodiments, all or portions of these operations may be implemented in a networked computing environment, such as a cloud computing system. Cloud computing often includes on-demand availability of computer system resources, such as data storage (cloud storage) and computing power, without direct active management by an entity. Cloud computing relies on the sharing of resources, and generally includes on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service.

[0093] FIG.14 depicts one illustrative cloud computing system 500 operable to perform the above operations by executing programmed instructions tangibly embodied on one or moreAttorney Docket No.9903.019WO1 computer readable storage mediums. The cloud computing system 500 generally includes the use of a network of remote servers hosted on the internet to store, manage, and process data, rather than a local server or a personal computer (e.g., in the computing systems 502-1 - 502-N). Cloud computing enables users to use infrastructure and applications via the internet, without installing and maintaining them on-premises. In this regard, the cloud computing network 520 may include virtualized information technology (IT) infrastructure (e.g., servers 524-1 - 524-N, the data storage module 522, operating system software, networking, and other infrastructure) that is abstracted so that the infrastructure can be pooled and / or divided irrespective of physical hardware boundaries. In some embodiments, the cloud computing network 520 can provide users with services in the form of building blocks that can be used to create and deploy various types of applications in the cloud on a metered basis.

[0094] Various components of the cloud computing system 500 may be operable to implement the above operations in their entirety or contribute to the operations in part. For example, a computing system 502-1 may be used to perform analysis of gene sequencing data, and then store that analysis along with the gene sequencing data in a data storage module 522 (e.g., a database) of a cloud computing network 520. Various computer servers 524-1 - 524-N of the cloud computing network 520 may be used to operate on the gene sequencing data and / or transfer the gene sequencing analysis and / or the gene sequencing data to another computing system 502-N.

[0095] Some embodiments disclosed herein may utilize instructions (e.g., code / software) accessible via a computer-readable storage medium for use by various components in the cloud computing system 500 to implement all or parts of the various operations disclosed hereinabove. Examples of such components include the computing systems 502-1 - 502-N.

[0096] Exemplary components of the computing systems 502-1 - 502-N may include at least one processor 504, a computer readable storage medium 514, program and data memory 506, input / output (I / O) devices 508, a display device interface 512, and a network interface 510. For the purposes of this description, the computer readable storage medium 514 comprises any physical media that is capable of storing a program for use by the computing system 502. For example, the computer-readable storage medium 514 may be an electronic, magnetic, optical, electromagnetic, infrared, semiconductor device, or other non-transitory medium. Examples ofAttorney Docket No.9903.019WO1 the computer-readable storage medium 514 include a solid-state memory, a magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk, and an optical disk. Some examples of optical disks include Compact Disk - Read Only Memory (CD-ROM), Compact Disk - Read / Write (CD-R / W), Digital Versatile Disc (DVD), and Blu-Ray Disc.

[0097] The processor 504 is coupled to the program and data memory 506 through a system bus 516. The program and data memory 506 include local memory employed during actual execution of the program code, bulk storage, and / or cache memories that provide temporary storage of at least some program code and / or data in order to reduce the number of times the code and / or data are retrieved from bulk storage (e.g., a hard disk drive, a solid state drive, or the like) during execution.

[0098] Input / output or I / O devices 508 (including but not limited to keyboards, displays, touchscreens, microphones, pointing devices, etc.) may be coupled either directly or through intervening I / O controllers. Network adapter interfaces 510 may also be integrated with the system to enable the computing system 502 to become coupled to other computing systems or storage devices through intervening private or public networks. The network adapter interfaces 510 may be implemented as modems, cable modems, Small Computer System Interface (SCSI) devices, Fibre Channel devices, Ethernet cards, wireless adapters, etc. Display device interface 512 may be integrated with the system to interface to one or more display devices, such as screens for presentation of data generated by the processor 504.

Claims

Attorney Docket No.9903.019WO1 CLAIMS What is claimed is:

1. A method for generating synthetic RNA sequences, comprising: training a generative machine learning model on a dataset of training RNA sequences, wherein each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations; encoding the training RNA sequences into a latent representation using an encoder network of the generative machine learning model; and decoding the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences, wherein the one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.

2. The method of claim 1, wherein: the one or more biological annotations include at least one of a minimal Levenshtein distance, a minimal 4-mer distance, a guanine-cytosine content, a minimum free energy, a structure property, or a functional property of the training RNA sequences.

3. The method of claim 1, wherein: the generative machine learning model comprises at least one of a variational autoencoder, a transformer-based autoregressive model, or a generative adversarial network.

4. The method of claim 1, further comprising: training the generative machine learning model using a loss function comprising a reconstruction loss.

5. The method of claim 1, further comprising: estimating a functional property of RNA with a trained reward network of the generative machine learning model.Attorney Docket No.9903.019WO1 6. The method of claim 5, wherein: the training RNA sequences comprise variable lengths; and the method further comprises training the reward network directly in a fixed-length latent space that is mapped to the variable length training RNA sequences.

7. The method of claim 5, further comprising: using a guided diffusion model to generate latent RNA embeddings based at least on guidance from the reward network.

8. A non-transitory computer readable medium comprising instructions for generating synthetic RNA sequences that, when executed by a processor, direct the processor to: train a generative machine learning model on a dataset of training RNA sequences, wherein each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations; encode the training RNA sequences into a latent representation using an encoder network of the generative machine learning model; and decode the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences, wherein the one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.

9. The computer readable medium of claim 8, wherein: the one or more biological annotations include at least one of a minimal Levenshtein distance, a minimal 4-mer distance, a guanine-cytosine content, a minimum free energy, a structure property, or a functional property of the training RNA sequences.

10. The computer readable medium of claim 8, wherein: the generative machine learning model comprises at least one of a variational autoencoder, a transformer-based autoregressive model, or a generative adversarial network.Attorney Docket No.9903.019WO1 11. The computer readable medium of claim 8, further comprising instructions that direct the processor to: train the generative machine learning model using a loss function comprising a reconstruction loss.

12. The computer readable medium of claim 8, further comprising instructions that direct the processor to: estimate a functional property of RNA with a trained reward network of the generative machine learning model.

13. The computer readable medium of claim 12, wherein: the training RNA sequences comprise variable lengths; and the computer readable medium further comprises instructions that direct the processor to train the reward network directly in a fixed-length latent space that is mapped to the variable length training RNA sequences.

14. The computer readable medium of claim 12, further comprising instructions that direct the processor to: use a guided diffusion model to generate latent RNA embeddings based at least on guidance from the reward network.Attorney Docket No.9903.019WO1 15. A processing system for generating synthetic RNA sequences, the processing system comprising: an interface operable to communicatively couple to a database comprising a dataset of training RNA sequence, and to receive the dataset of training RNA sequences, wherein each training RNA sequence is represented by a nucleotide sequence and by one or more biological annotations; and a processor operable to train a generative machine learning model on the dataset of training RNA sequences, to encode the training RNA sequences into a latent representation using an encoder network of the generative machine learning model, and to decode the latent representation using a decoder network of the generative machine learning model to generate one or more novel RNA sequences, wherein the one or more novel RNA sequences exhibit at least one of a desired biological characteristic or a desired structure.

16. The processing system of claim 15, wherein: the one or more biological annotations include at least one of a minimal Levenshtein distance, a minimal 4-mer distance, a guanine-cytosine content, a minimum free energy, a structure property, or a functional property of the training RNA sequences.

17. The processing system of claim 15, wherein: the generative machine learning model comprises at least one of a variational autoencoder, a transformer-based autoregressive model, or a generative adversarial network.

18. The processing system of claim 15, wherein: the processor is further operable to train the generative machine learning model using a loss function comprising a reconstruction loss.

19. The processing system of claim 15, wherein: the processor is further operable to estimate a functional property of RNA with a trained reward network of the generative machine learning model.Attorney Docket No.9903.019WO1 20. The processing system of claim 19, wherein: the training RNA sequences comprise variable lengths; and the processor is further operable to train the reward network directly in a fixed-length latent space that is mapped to the variable length training RNA sequences.

21. The processing system of claim 19, wherein: the processor is further operable to use a guided diffusion model to generate latent RNA embeddings based at least on guidance from the reward network.

Citation Information

Patent Citations

  • Methods and systems for determining gene expression profiles and cell identities from multi-omic imaging data

    US20220180975A1

  • Method and apparatus using machine learning for evolutionary data-driven design of proteins and other sequence defined biomolecules

    US20220348903A1

  • Live-cell label-free prediction of single-cell omics profiles by microscopy

    WO2023091970A1

  • Structural and transformer based machine-learning models for design of engineered guide systems for adenosine deaminase acting on RNA editing

    WO2023240209A1

  • Native expansion of a sparse training dataset into a dense training dataset for supervised training of a synonymous variant sequence generator

    WO2024086143A1