Method for CDS- / UTR-based RNA design and RNA obtained therefrom

GEMORNA, a deep learning-based model, addresses the challenges of mRNA sequence design by optimizing CDSs and UTRs, achieving higher expression and stability, as evidenced by improved in vitro and in vivo performance.

WO2025218758A1PCT designated stage Publication Date: 2025-10-23BEIJING HENGYU BIOTECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/089654
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-17
Filing Date
2025-04-17
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

The design of mRNA sequences for therapeutic applications faces challenges due to the vast potential sequence space and complex cellular environment, which affects translation efficiency and stability, necessitating improved mRNA sequence design for enhanced expression and durability.

Method used

The development of a generative model named GEMORNA, which utilizes deep learning to design coding sequences (CDSs) and untranslated regions (UTRs) for mRNA sequences, optimizing translation capacity and stability through the extraction of implicit rules from natural sequences.

Benefits of technology

GEMORNA-generated sequences exhibit higher expression levels and durability, demonstrated by improved in vitro and in vivo results, including enhanced immune responses in COVID-19 vaccine experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025089654_23102025_PF_FP_ABST
    Figure CN2025089654_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method for CDS-based RNA design, a method for UTR-based RNA design, and the RNA obtained therefrom.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR CDS- / UTR-BASED RNA DESIGN AND RNA OBTAINED THEREFROM

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to PCT Application No. PCT / CN2024 / 088336, filed on April 17, 2024. The disclosure of the prior application is considered part of the disclosure of this application and is incorporated in its entirety into this application.

[0003] REFERENCE TO THE SEQUENCE LISTING

[0004] The Sequence Listing titled DCF240247WO-sequence listing, which was created on April 12, 2024 and is 55, 329 bytes in size, is hereby incorporated by reference in its entirety.TECHNICAL FIELD

[0005] Disclosed herein are novel RNA sequences that can be used for enhanced gene expression in various contexts. Also disclosed herein are methods of designing CDS sequences, UTR sequences, and full-length mRNA sequences.BACKGROUND

[0006] Messenger RNA (mRNA) vaccines have established their potentials as an effective approach to prevent severe COVID-19 diseases. There are numerous efforts to extend mRNA therapeutics to other indications, including cancers and rare diseases. Stronger and longer-lasting protein expressions are necessary for the successful applications of mRNA therapeutics, yet that are challenging because mRNA designs exhibit wide differences in expression-related properties. mRNA sequence components, such as coding sequences (CDSs) and untranslated regions (UTRs) , are crucial determinants of therapeutic activity. Therefore, designing optimal mRNA sequences is crucial for unlocking the full potential of mRNA therapeutics. However, it remains an extremely challenging task because of the vast potential mRNA sequence space (Fig. 1) .SUMMARY

[0007] Despite the tremendous success of mRNA COVID-19 vaccines, the extension of this modality to a broader spectrum of diseases necessitates substantial enhancements, particularly in the design of mRNAs with elevated expression levels and extended durability.

[0008] Disclosed herein is a comprehensive strategy for the design of mRNA sequences with enhanced translational efficiency and stability. This strategy utilizes deep generative models developed by the inventors, termed GEMORNA, which enable the generation of novel coding sequences (CDSs) , untranslated regions (UTRs) , and full-length mRNAs with superior translation capacity.

[0009] The inventors experimentally tested GEMORNA-generated CDSs and UTRs for linear mRNAs and confirmed that these CDSs and UTRs achieved higher expression levels of the related mRNAs. It was demonstrated that by combining such GEMORNA-generated RNA elements, the inventors were able to design full-length therapeutic mRNA sequences, including COVID-19 mRNA sequences, with improved translation capability and optimized immunogenicity.

[0010] The analyses and experimental results suggest that the generative AI platform and the strategy described in the present application offer substantial benefits for therapeutic mRNA design.

[0011] In a first aspect of the present application, there provides an RNA, comprising at least one nucleotide sequence with or without chemical modifications, wherein the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-41 and 53-57, any of truncated SEQ ID NOs: 1-41 and 53-57, or any sequence having at least 70%sequence identity of SEQ ID NOs: 1-41 and 53-57.

[0012] In a second aspect of the present application, there provides a nucleic acid, capable of generating the RNA according to the first or eleventh aspect of the present application.

[0013] In a third aspect of the present application, there provides a synthetic construct, comprising the RNA according to the first or eleventh aspect of the present application, or the nucleic acid according to the second aspect of the present application.

[0014] In a fourth aspect of the present application, there provides a cell, comprising the RNA according to the first or eleventh aspect of the present application, the nucleic acid according to the second aspect of the present application, or the synthetic construct according to the third aspect of the present application.

[0015] In a fifth aspect of the present application, there provides a conjugate, comprising the RNA according to the first or eleventh aspect of the present application, the nucleic acid according to the second aspect of the present application, or the synthetic construct according to the third aspect of the present application.

[0016] In a sixth aspect of the present application, there provides a composition, comprising the RNA according to the first or eleventh aspect of the present application, the nucleic acid according to the second aspect of the present application, the synthetic construct according to the third aspect of the present application, the cell according to the forth aspect of the present application, or the conjugate according to the fifth aspect of the present application.

[0017] In a seventh aspect of the present application, there provides a method for preventing or treating a disease or condition, comprising administrating the RNA according to the first or eleventh aspect of the present application, the nucleic acid according to the second aspect of the present application, the synthetic construct according to the third aspect of the present application, the cell according to the forth aspect of the present application, the conjugate according to the fifth aspect of the present application, or the composition according to the sixth aspect of the present application to the subject in need thereof.

[0018] In an eighth aspect of the present application, there provides a use of the RNA according to the first or eleventh aspect of the present application, the nucleic acid according to the second aspect of the present application, the synthetic construct according to the third aspect of the present application, the cell according to the forth aspect of the present application, the conjugate according to the fifth aspect of the present application or the composition according to the sixth aspect of the present application in the manufacture of a drug for preventing or treating a disease or condition.

[0019] In a ninth aspect of the present application, there provides for generating a CDS-based RNA design, comprising the following steps in order:

[0020] (a) obtaining a source protein sequence to be transformed;

[0021] (b) generating, based on a first predetermined embedding dimension, first vectors corresponding to amino acids in the source protein sequence;

[0022] (c) inputting the first vectors into a trained neural network to generate a first representation of the source protein sequence;

[0023] (d) initializing a target RNA sequence with a start token;

[0024] (e) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence generated so far;

[0025] (f) inputting the vectors obtained from (e) into a trained neural network to obtain a representation of the prefix of the target RNA sequence;

[0026] (g) generating, based on the representations of the prefix of the target RNA sequence and the source protein sequence, an updated prefix of the target RNA sequence;

[0027] (h) repeating the steps of (e) - (g) until a stop condition is met; and

[0028] (i) outputting a final target RNA sequence, that is the CDS-based RNA.

[0029] In a tenth aspect of the present application, there provides a method for generating a UTR-based RNA design, comprising the following steps in order:

[0030] (a) initializing a target RNA sequence with a start token;

[0031] (b) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence;

[0032] (c) inputting the vectors into a trained neural network to generate an updated prefix of the target RNA sequence;

[0033] (d) repeating the steps of (b) - (c) until a stop condition is met; and

[0034] (e) outputting a final target RNA sequence, that is the UTR-based RNA.

[0035] In an eleventh aspect of the present application, there provides an RNA, comprising at least one RNA sequence obtained by the method according to the ninth aspect of the present application, and / or at least one RNA sequence obtained by the method according to the tenth aspect of the present application.

[0036] In a twelfth aspect of the present application, there provides a deep learning-based model for CDS-based RNA design, named GEMORNA-CDS, utilizing the method according to the ninth aspect of the present application to generate CDS-based RNAs.

[0037] In a thirteenth aspect of the present application, there provides a deep learning-based model for UTR-based RNA design, named GEMORNA-UTR, utilizing the method according to the tenth aspect of the present application to generate UTR-based RNAs.BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The following is a brief description of the drawings, which are presented for the purposes of illustrating the exemplary embodiments disclosed herein and not for the purposes of limiting the scope of the claims.

[0039] Figure 1 shows a visualization of the mRNA components, and the exponentially increased design space.

[0040] Figure 2 shows the platform according to the present application for mRNA sequence generation, screening and validation.

[0041] Figure 3 shows the frequency of unwanted codon pairs in CAI (Codon Adaptation Index) -optimized, natural, and GEMORNA sequences.

[0042] Figure 4 shows the frequency of slippery motifs in randomly generated, CAI-optimized, natural, and GEMORNA sequences.

[0043] Figure 5 shows Codon stability coefficients measured by ORFome-assay (first row) , and codon relative adaptiveness calculated from the codon frequencies of GEMORNA CDSs (second row) or CAI-optimized CDSs (third row) .

[0044] Figure 6 shows the distributions of max identity between natural and GEMORNA 5’ UTRs; a small max identity indicates a low sequence similarity between GEMORNA-generated UTRs and natural UTRs.

[0045] Figure 7 shows Mean Ribosome Loading (MRL) and Minimum Free Energy (MFE) distributions of natural and GEMORNA-generated 5’ UTRs.

[0046] Figure 8 shows expression levels of different Fluc2P CDSs at 48 h in HEK293T cells.

[0047] Figure 9 shows expression levels of different Fluc2P CDSs at 48 h in HepG2 cells.

[0048] Figure 10 shows in vitro firefly luminescence activities of the benchmark 5’ UTRs and GEMORNA 5’ UTRs.

[0049] Figure 11 shows normalized firefly luminescence of GEMORNA 3’ UTRs compared to BNT’s 3’ UTR at 48 hours in HEK293T cells.

[0050] Figure 12 shows normalized firefly luminescence activities of GEMONRA full-length mRNAs encoding the Fluc2P reporter against benchmarks. The left panel presents firefly luminescence activities over time in HEK293T cells, while the right panel compares results across two cell lines.

[0051] Figure 13 shows a comparison of normalized hEPO activities of GEMORNA full-length mRNAs against the benchmark in HEK293T cells.

[0052] Figure 14 shows a comparison of hEPO expression in mice over time between GEMORNA mRNAs and the benchmark.

[0053] Figure 15 shows in vivo anti-spike IgG comparisons of COVID-19 mRNA vaccine sequences between GEMORNA sequences and the benchmarks.DETAILED DESCRIPTION

[0054] The present application is explained in greater detail below. This description is not intended to be a detailed catalog of all the different ways in which the present application may be implemented, or all the features that may be added to the present application. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein will be apparent to those skilled in the art in light of the instant disclosure which do not depart from the present application. Hence, the following description is intended to illustrate some particular embodiments of the present application, and not to exhaustively specify all permutations, combinations and variations thereof.

[0055] Unless the context indicates otherwise, it is specifically intended that the various features of the present application described herein can be used in any combination. Moreover, the present application also contemplates that in some embodiments of the present application, any feature or combination of features set forth herein can be excluded or omitted. To illustrate, if the description states that a complex comprises components A, B and C, it is specifically intended that any of A, B or C, or a combination thereof, can be omitted and disclaimed singularly or in any combination.

[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure pertains. Although any methods and materials similar or equivalent to those described herein may be used in the practice for testing of the present application, the preferred materials and methods are described herein. In describing and claiming the present application, the following terminology will be used.

[0057] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.

[0058] The singular forms “a, ” “an” , and “the” include plural referents unless the context clearly dictates otherwise.

[0059] As used in the description and in the claims, the open-ended transitional phrases “comprise (s) ” , “comprising” , “include (s) ” , “including” , “have” , “has” , “having” , “contain (s) ” , “containing” and variants thereof require the presence of the named ingredients / steps and permit the presence of other ingredients / steps. These phrases should also be construed as disclosing the closed-ended phrases “consist of” or “consist essentially of” that permit only the named ingredients / steps and unavoidable impurities and exclude other ingredients / steps.

[0060] Numerical values in the description and claims of this application should be understood to include numerical values which are the same when reduced to the same number of significant figures and numerical values which differ from the stated value by less than the experimental error of conventional measurement technique of the type described in the present application to determine the value.

[0061] As mentioned above, mRNA sequence design problem is still challenging for two reasons. On one hand, the design space is extremely vast, which prohibits exhaustive testing on a one-by-one basis. On the other hand, many factors in the cellular environment, such as ribosomes and various RNA binding proteins (RBP) , and different innate immune response pathways, may have a synergistic influence on mRNA molecules, which prohibits direct optimization with an explicit and comprehensive objective.

[0062] Here, the inventors resort to creating generative models to design mRNA sequences with improved translation capacity. Generative models have been used successfully for text generation. Drawing parallels between human language and genetic language, the inventors adapted the concept and architecture of these generative models to mRNA sequences. Specifically, designing a CDS for a given protein is analogous to generating a translated sentence from its source language, and designing a UTR is akin to freely composing a poem due to the absence of protein sequence constraints. The inventors developed a generative model named GEMORNA, which stands for generative models for RNA, as the core of the inventors’ mRNA design process.

[0063] Computational analysis showed that the models according to the present application extract implicit rules from massive natural sequences, and in vitro results confirmed that the GEMORNA-generated sequences according to the present application exhibit both more durable expression and higher expression levels. More importantly, the head-to-head COVID vaccine experiment showed improved immune response in vivo over commercial mRNA vaccine. These results showed that the generative AI-based platform according to the present application can potentially benefit the development of mRNA vaccine and other mRNA drugs. Figure 2 shows the platform according to the present application for mRNA sequence generation, screening and validation.

[0064] As used herein, the term “Codon Adaptation Index” or “CAI” refers to a measure of how closely the codon usage of a gene matches the preferred codon usage of highly expressed genes in a given organism, reflecting its potential for efficient translation.

[0065] As used herein, the term “slippery motif” refers to a specific sequence pattern in the mRNA sequence that enables the ribosome to shift the reading frame during the translation process, thus leading to the production of different protein products. This sequence pattern usually has some special base compositions and structural features, making it easier for the ribosome to experience the slipping phenomenon during translation.

[0066] Method of generating RNA designs

[0067] In one aspect of the present application, there provides a method for generating a CDS-based RNA design, comprising the following steps in order:

[0068] (a) obtaining a source protein sequence to be transformed;

[0069] (b) generating, based on a first predetermined embedding dimension, first vectors corresponding to amino acids in the source protein sequence;

[0070] (c) inputting the first vectors into a trained neural network to generate a first representation of the source protein sequence;

[0071] (d) initializing a target RNA sequence with a start token;

[0072] (e) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence generated so far;

[0073] (f) inputting the vectors obtained from (e) into a trained neural network to obtain a representation of the prefix of the target RNA sequence;

[0074] (g) generating, based on the representations of the prefix of the target RNA sequence and the source protein sequence, an updated prefix of the target RNA sequence;

[0075] (h) repeating the steps of (e) - (g) until a stop condition is met; and

[0076] (i) outputting a final target RNA sequence, that is the CDS-based RNA.

[0077] As used herein, the term “prefix of an RNA” refers to the contiguous subsequence consisting of all nucleotides generated from the start token up to the current iteration during sequence generation. It represents the partial RNA sequence constructed by the method for generating a UTR-based RNA design or a CDS-based RNA design disclosed herein.

[0078] In some embodiments, the source protein can be selected from antigen proteins, such as SARS-CoV-2 spike protein.

[0079] In some embodiments, the start token can be “<sos>” .

[0080] In some embodiments, the predetermined embedding dimension is larger than 100.

[0081] In some embodiments, the trained neural network can be an attention-based neural network.

[0082] In some embodiments, the stop condition is met when the length of the generated RNA sequence reaches three times the length of the source protein sequence.

[0083] In another aspect of the present application, there provides a method for generating a UTR-based RNA design, comprising the following steps in order:

[0084] (a) initializing a target RNA sequence with a start token;

[0085] (b) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence;

[0086] (c) inputting the vectors into a trained neural network to generate an updated prefix of the target RNA sequence;

[0087] (d) repeating the steps of (b) - (c) until a stop condition is met; and

[0088] (e) outputting a final target RNA sequence, that is the UTR-based RNA.

[0089] In some embodiments, the start token can be “<sos>” .

[0090] In some embodiments, the predetermined embedding dimension is larger than 100.

[0091] In some embodiments, the trained neural network can be an attention-based neural network.

[0092] In some embodiments, the stop condition is met when the stop token (for example, “<eos>” is generated) or a predetermined length is reached. In some embodiments, the predetermined length is, for example, between 30 and 200.

[0093] Deep learning-based models

[0094] In one aspect of the present application, there provides a deep learning-based model for CDS-based RNA design, named GEMORNA-CDS, wherein the deep learning-based model utilizes a method for generating a CDS-based RNA design, comprising:

[0095] (a) obtaining a source protein sequence to be transformed;

[0096] (b) generating, based on a first predetermined embedding dimension, first vectors corresponding to amino acids in the source protein sequence;

[0097] (c) inputting the first vectors into a trained neural network to generate a first representation of the source protein sequence;

[0098] (d) initializing a target RNA sequence with a start token;

[0099] (e) iteratively generating, based on a predetermined embedding dimension, vectors corresponding to the target RNA sequence;

[0100] (f) inputting the vectors obtained from (e) into a trained neural network to obtain a representation of the target RNA sequence;

[0101] (g) generating, based on the representations of the target RNA sequence and the source protein sequence, an updated RNA sequence;

[0102] (h) repeating the steps of (e) - (g) until a stop condition is met; and

[0103] (i) outputting a final RNA sequence.

[0104] An amino acid is normally encoded by multiple synonymous codons, and it is essential to choose specific codons not in isolation, but in consideration of their contextual relationships with other codons in CDS design. The inventors developed a deep learning-based generative model, named GEMORNA-CDS, for the generation of CDS with the consideration of complex codon interactions. Specifically, the model takes a protein sequence as the input, and outputs artificial CDS sequences with better codon usage.

[0105] GEMORNA-CDS model demonstrates better codon usage regarding codon pairs and unwanted motifs. It has been shown that codon pairs significantly influence protein expression levels. Although two individual codons may each have moderate frequencies, their joint occurrence could be infrequent due to inhibitory interactions. Notably, natural sequences tend to have low frequencies of unwanted codon pairs, while sequences designed with widely used CAI-optimizing algorithms contain more unwanted codon pairs (Fig. 3) . In contrast, the GEMORNA-generated CDSs exhibited a marked reduction in the incidence of unwanted codon pairs. Furthermore, the GEMORNA-designed CDSs contain fewer slippery motifs, reducing the +1 ribosomal frameshifting frequency (Fig. 4) . Such frameshifting may significantly impact translation levels. Typically, frameshifting leads to the synthesis of proteins with altered amino acid sequences (Thomas E Mulroney, Tuija  Juan Carlos Yam Puc, Maria Rust, Robert F Harvey, Lajos Kalmar, Emily Horner, Lucy Booth, Alexander P Ferreira, Mark Stoneley, et al. N1-methylpseudouridylation of mRNA causes +1 ribosomal frameshifting. Nature, pages 1–6, 2023. ) , which are not anticipated. These findings suggest that the GEMORNA model has the capability to learn and apply codon usage patterns within a contextual framework.

[0106] Besides these aspects, the inventors were also concerned with the in-cell stability of mRNA sequences. Wu et al. (Qiushuang Wu, Santiago Gerardo Medina, Gopal Kushawah, Michelle Lynn DeVore, Luciana A Castellano, Jacqelyn M Hand, Matthew Wright, and Ariel Alejandro Bazzini. Translation affects mRNA stability in a codon-dependent manner in human cells. elife, 8: e45396, 2019) investigated codon-dependent mRNA decay and provided a relative stability ranking of codons based on the ORFome assay. The inventors compared the codon usage of CAI-optimized sequences with the GEMORNA sequences against the ORFome results and observed that the model according to the present application significantly favored more stable codons (Fig. 5) .

[0107] In another aspect of the present application, there provides a deep learning-based model for UTR-based RNA design, wherein the deep learning-based model utilizes a method for generating a UTR-based RNA design, comprising:

[0108] (a) initializing a target RNA sequence with a start token;

[0109] (b) generating, based on a predetermined embedding dimension, vectors corresponding to the target RNA sequence;

[0110] (c) inputting the vectors into a trained neural network to generate an updated RNA sequence;

[0111] (d) repeating the steps of (b) - (c) until a stop condition is met; and

[0112] (e) outputting a final RNA sequence.

[0113] The UTR sequences that surround the CDSs play a major role in the performance of mRNA therapeutics and vaccines by controlling expression levels and stabilities. The inventors developed a deep learning-based generative model, named GEMORNA-UTR, for the generation of UTR sequences with potentially enhanced translational capacity. The model does not require an input, and outputs artificial UTR sequences.

[0114] RNA generated by deep learning-based models

[0115] The inventors assessed the similarities between GEMORNA-generated UTRs and natural UTRs using a BLAST search. The Maximum Identity Score (MIS) , defined as the longest aligned subsequence divided by the length of the query sequence, was employed as the measure of similarity. The MIS distributions for the 5’ UTRs show that most GEMORNA-generated UTRs exhibit low MIS values (Fig. 6) , suggesting that GEMORNA generates a significant number of novel UTRs that are dissimilar to natural UTRs.

[0116] The inventors investigated two in silico measurements, predicted MRL and MFE of the generated 5’ UTRs, since 5’ UTRs with higher MRL and MFE are more efficient in translation initiation. The MRL and MFE distributions of the inventors’ AI-generated 5’ UTRs shifted rightwards to higher values compared to the nature 5’ UTRs (Fig. 7) , which potentially leads to high-efficiency translation initiation.

[0117] In one aspect of the present application, there provides an RNA, comprising at least one nucleotide sequence with or without chemical modifications, wherein the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-41 and 53-57, any of truncated SEQ ID NOs: 1-41 and 53-57, or any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 1-41 and 53-57.

[0118] In some embodiments, the chemical modifications are selected from methylation, pseudouridylation, N1-methylpseudouridine, uridylation, adenylylation, phosphorylation, acetylation, bisulfite conversion.

[0119] Methylation involves the addition of a methyl group to specific nucleotides. For example, N6-methyladenosine (m6A) is a prevalent methylation modification in mRNA. It can affect mRNA stability, translation efficiency, and splicing.

[0120] For pseudouridylation, uridine is converted to pseudouridine (Ψ) or N1-methylpseudouridine (m1Ψ) . Pseudouridylation and N1-methylpseudouridine can influence the folding, stability, expression and immunogenicity of RNA molecules.

[0121] Uridylation involves the addition of one or more uridine residues to the end of an RNA molecule. It can affect the stability and function of the RNA, and is involved in processes such as mRNA decay and microRNA maturation.

[0122] Adenylylation is the addition of adenosine residues to RNA. It can play a role in regulating RNA stability and function.

[0123] Phosphorylation refers to the addition of a phosphate group to the hydroxyl group of the ribose sugar in RNA nucleotides. Phosphorylation can affect the charge and conformation of RNA, and is involved in various cellular processes such as RNA splicing and export.

[0124] Acetylation refers to the addition of an acetyl group to the nitrogenous base or the ribose moiety of RNA. Acetylation can impact RNA-protein interactions and the stability of RNA-RNA duplexes.

[0125] In the process of bisulfite conversion, cytosine residues in RNA can be converted to uracil under the action of bisulfite. This modification is often used in RNA epigenetic studies to detect cytosine methylation.

[0126] In some embodiments, the any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 1-41 and 53-57, or any of truncated SEQ ID NOs: 1-41 and 53-57 is a functional variant of SEQ ID NOs: 1-41 and 53-57. In some embodiments, the functional variant of a parent sequence contains one or more conservative substitutions. For example, the functional variant of SEQ ID NO: 1 contains one or more conservative substitutions.

[0127] As used herein, functional variant refers to a sequence variant of an original sequence that has similar or the same function, activity or other biological properties of the original sequence. In some embodiments, the functional variant has at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%of the function, activity or other biological properties of the original sequence.

[0128] In some embodiments, the “sequence identity” is determined by comparing two optimally aligned sequences, wherein the portion of the sequence may comprise additions or deletions (i.e., gaps) of 20 percent or less as compared to the reference sequences (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid bases or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the reference sequence and multiplying the results by 100 to yield the percentage of sequence identity.

[0129] As used herein, the term “conservative substitution” in the setting of a nucleotide sequence can generally be described as nucleotide substitution in which a nucleotide is replaced with another nucleotide, and the substitution has little or essentially no influence on the function, activity or other biological properties of the original nucleotide sequence, or the substitution has little or essentially no influence on the function, activity or other biological properties of the polypeptide that the original nucleotide sequence translated into.

[0130] In some embodiments, the conservative substitution is determined based on Degeneracy of Bases. In the genetic code, degeneracy refers to the phenomenon that multiple codons can code for the same amino acid. This is related to the structure and function of bases in nucleic acids. For example, the amino acid leucine can be encoded by six different codons: UUA, UUG, CUU, CUC, CUA, and CUG.

[0131] In some embodiments, the RNA is linear.

[0132] In some embodiments, the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-4 and 53-57, any of truncated SEQ ID NOs: 1-4 and 53-57, or any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 1-4 and 53-57, preferably, the RNA is linear.

[0133] In some embodiments, the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 5-16, any of truncated SEQ ID NOs: 5-16, or any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 5-16.

[0134] In some embodiments, the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 17-26, any of truncated SEQ ID NOs: 17-26, or any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 17-26.

[0135] In some embodiments, the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 27-41, any of truncated SEQ ID NOs: 27-41, or any sequence having at least 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%sequence identity of SEQ ID NOs: 27-41.

[0136] In some embodiments, the RNA comprises:

[0137] (1) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or

[0138] (2) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or

[0139] (3) a nucleotide sequence as set forth in SEQ ID NO: 14, a nucleotide sequence of truncated SEQ ID NO: 14, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 14, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or

[0140] (4) a nucleotide sequence as set forth in SEQ ID NO: 14, a nucleotide sequence of truncated SEQ ID NO: 14, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 14, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or

[0141] (5) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or

[0142] (6) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 53, a nucleotide sequence of truncated SEQ ID NO: 53, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 53, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or

[0143] (7) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 55, a nucleotide sequence of truncated SEQ ID NO: 55, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 55, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or

[0144] (8) a nucleotide sequence as set forth in SEQ ID NO: 13, a nucleotide sequence of truncated SEQ ID NO: 13, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 13, a nucleotide sequence as set forth in SEQ ID NO: 54, a nucleotide sequence of truncated SEQ ID NO: 54, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 54, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or

[0145] (9) a nucleotide sequence as set forth in SEQ ID NO: 13, a nucleotide sequence of truncated SEQ ID NO: 13, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 13, a nucleotide sequence as set forth in SEQ ID NO: 56, a nucleotide sequence of truncated SEQ ID NO: 56, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 56, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or

[0146] (10) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 53, a nucleotide sequence of truncated SEQ ID NO: 53, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 53, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or

[0147] (11) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 56, a nucleotide sequence of truncated SEQ ID NO: 56, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 56, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or

[0148] (12) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 57, a nucleotide sequence of truncated SEQ ID NO: 57, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 57, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25.

[0149] In another aspect of the present application, there provides a nucleic acid, capable of generating the RNA disclosed herein.

[0150] In yet another aspect of the present application, there provides a synthetic construct, comprising the RNA disclosed herein, or the nucleic acid disclosed herein.

[0151] In one aspect of the present application, there provides a cell, comprising the RNA disclosed herein, the nucleic acid disclosed herein, or the synthetic construct disclosed herein.

[0152] In another aspect of the present application, there provides a conjugate, comprising the RNA disclosed herein, the nucleic acid disclosed herein, or the synthetic construct disclosed herein.

[0153] In one aspect of the present application, there provides a composition, comprising the RNA disclosed herein, the nucleic acid disclosed herein, the synthetic construct disclosed herein, the cell disclosed herein, or the conjugate disclosed herein.

[0154] In still yet another aspect of the present application, there provides a method for preventing or treating a disease or condition, comprising administrating the RNA disclosed herein, the nucleic acid disclosed herein, the synthetic construct disclosed herein, the cell disclosed herein, the conjugate disclosed herein, or the composition disclosed herein to the subject in need thereof.

[0155] In some embodiments, the subject can be selected from vertebrates, including mammals, such as humans, rats, mice, monkeys, dogs, cats among others.

[0156] In one aspect of the present application, there provides a use of the RNA disclosed herein, the nucleic acid disclosed herein, the synthetic construct disclosed herein, the cell disclosed herein, the conjugate disclosed herein, or the composition disclosed herein in the manufacture of a drug for preventing or treating a disease or condition.

[0157] In some embodiments, the disease is selected from infectious diseases, tumors, cancers, inflammatory diseases, immunological diseases. In some preferred embodiments, the disease is a virus infection, such as COVID-19.

[0158] In one aspect, the present disclosure provides an RNA, comprising at least one RNA sequence obtained by the method for generating a CDS-based RNA design disclosed herein, and / or at least one RNA sequence obtained by the method for generating a UTR-based RNA design disclosed herein. In some embodiments, the RNA comprises at least one of the RNA disclosed herein, such as, an RNA, comprising at least one nucleotide sequence with or without chemical modifications, wherein the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-41 and 53-57, any of truncated SEQ ID NOs: 1-41 and 53-57, or any sequence having at least 70%sequence identity of SEQ ID NOs: 1-41 and 53-57.

[0159] The technical solution and technical effects of the present application will be more clearly and explicitly described by way of illustration in combination with examples. It should be understood that these examples are only for illustrative purposes and not intended to limit the protection scope of the present application. The protection scope of the present application is only defined by the appended claims.

[0160] EXAMPLES

[0161] Methods

[0162] Cell culture and transfections. HEK293T and HepG2 cells were cultured in DMEM (Thermo Fisher) supplemented with 10%FBS and 100 units / mL penicillin and streptomycin. The transfection was performed with Lipofectamine MessengerMax according to the manufacturer’s manual.

[0163] mRNA synthesis. The IVT templates were amplified via PCR and column purified. The mRNAs were then synthesized using T7 Co-transcription RNA Synthesis Kit with UTP fully replaced with N1-Me-Pseudo UTP.

[0164] Firefly luciferase assays. To detect luminescence from Firefly luciferase, 100 μL of luciferase reagent was added to each well, mixed, and incubated for 5 min. 100 μL of the culture medium and luciferase reagent mix was then transferred to a flat-bottomed white opaque plate and luminescence was detected microplate reader.

[0165] hEPO ELISA assay. Erythropoietin was detected by ELISA essentially according to the manufacturer’s instructions except cell culture supernatant at indicated time points post-transfection were used.

[0166] LNP encapsulation of the RNAs. mRNA was synthesized in vitro by T7 RNA polymerase–mediated transcription, using 100%modified N1-methylpseudouridine (m1Ψ) in place of uridine, from a linearized DNA template. The resulting mRNA was collum purified and then remaining mRNA was further magnetic bead purified before use.

[0167] The mRNAs were encapsulated with lipid nanoparticles (LNPs) . First, the RNAs were diluted with citrate buffer (pH 4.0) to a final concentration of 90 μg / ml. Then, the lipids (The molar ratio of lipids in this formulation is ALC-0315: DSPC : Cholesterol : ALC-0159 = 46.3 : 9.4 : 42.7 : 1.6) were mixed with the mRNA solution at the volume ratio of 1: 3 using INano L+ (Micro&Nano) . Then the LNP-mRNA formulations were diluted by 40-fold with 1×PBS buffer (pH 7.2-7.4) and concentrated via ultrafiltration with Amicon Ultra Centrifugal Filter Unit (Merck) . The concentration and encapsulation efficiency of mRNAs were measured by the Quant-it RiboGreen RNA Assay Kit (Invitrogen, R11490) . The size of LNP-mRNA particles was measured using dynamic light scattering on a Zetasizer Lab (Malvern) .

[0168] In vivo delivery of mRNA. For mouse vaccination, groups of 6-to 8-week-old female BALB / c mice were intramuscularly immunized with LNP-mRNAs as indicated or a placebo (LNP only) using a 1-ml sterile syringe, and 2 or 3 weeks later, a second dose was administered to boost the immune responses. The sera of immunized mice were collected for the detection the SARS-CoV-2-specific IgG antibody by ELISA assay as described below. All experiments using mice were conducted under the ethical regulations and were approved by local ethical committees.

[0169] For in vivo EPO expression studies, groups of 6-to 8-week-old female BALB / c mice were injected with LNP-mRNAs intravenously via the tail vein as indicated or a placebo (LNP only) using a 1-ml sterile syringe. Sera of LNP-mRNA injected mice were collected at different time points and the EPO concentrations were measured by ELISA assay.

[0170] IgG antibody ELISA. All immunized mouse serum samples were heat-inactivated, added to 96-well plates coated with recombinant SARS-CoV-2 S antigen and incubated for 2 h at room temperature. After three washes with the wash buffer, horseradish peroxidase HRP-conjugated goat anti-mouse IgG antibody was added to the plates and incubated for 0.5 h at room temperature. Then, the plates were washed four times with the wash buffer. After washing, the HRP substrates and stop solution were added, and the absorbance at 450 nm was measured with 630 nm as a reference with a TECAN Spark microplate reader. Endpoint titers were calculated as the reciprocal serum dilution.

[0171] Example 1: GEMORNA-CDS sequences achieve stronger and long-lasting protein expression on reporter protein

[0172] The inventors validated GEMORNA-generated CDSs with cell-based assays, using a destabilized firefly luciferase (denoted as Fluc2P) as the reporter. GMR-FL1 to GMR-FL4 were four GEMORNA-generated CDSs. Three controls were included for comparison: codon-optimized CDSs from commercial sources (GenScript and IDT) , and a CDS from a commercially available vector (pGL4.11) . The CDS sequences utilized here are listed in Table 1.

[0173] Table 1. CDS sequence List

[0174] The four GMR-FL1 to GMR-FL4 sequences according to the present application exhibited up to a 4.8-fold over pGL4.11 48 hours after transfection (Fig. 8) . Given the rapid degradation of destabilized luciferase post-production, the Fluc activity ratio at 48 hours relative to 24 hours was used as a metric to evaluate the in-cell stability of mRNA constructs. These experiments performed in HepG2 cells showed consistent results with those generated in HEK293T cells (Fig. 9) .

[0175] Example 2: GEMORNA-generated 5’ and 3’ UTRs enhance protein expression in vitro

[0176] The inventors compared GEMORNA-generated 5’ UTRs against widely-used natural UTRs and natural derivatives (C. Zeng, X. Hou, J. Yan, C. Zhang, W. Li, W. Zhao, S. Du, Y. Dong, Leveraging mRNA sequences and nanoparticles to deliver SARS-CoV-2 antigens in vivo. Advanced Materials 32, 2004452 (2020) ) . To isolate the effects of the 5’ UTR, the inventors used Fluc2P as the target protein and kept the CDS and 3’ UTR constant across all constructs. The inventors assessed 12 GEMORNA-designed UTRs, that are GMR-5U1, GMR-5U2, GMR-5U3, GMR-5U4, GMR-5U5, GMR-5U6, GMR-5U7, GMR-5U8, GMR-5U9, GMR-5U10, GMR-5U11 and GMR-5U12, for their ability to enhance protein expression and confirmed that GEMORNA 5’ UTRs exhibited higher Fluc activities (Fig. 10) . The controls used are AG, 5AG+G, hHBB, and S27A-44’ .

[0177] The inventors subsequently evaluated ten GEMORNA-generated 3’ UTRs (GMR-3U1, GMR-3U2, GMR-3U3, GMR-3U4, GMR-3U5, GMR-3U6, GMR-3U7, GMR-3U8, GMR-3U9 and GMR-3U10) and benchmarked them against the commercial mRNA drug 3’ UTRs, BNT (Fig. 11) . Eight of the ten GEMORNA 3’ UTRs, that are GMR-3U10, GMR-3U8, GMR-3U4, GMR-3U2, GMR-3U1, GMR-3U6, GMR-3U7, and GMR-3U9, matched or outperformed BNT’s 3’ UTR. The UTR sequences utilized here are listed in Table 2.

[0178] Table 2. UTR sequence List

[0179] Example 3: GEMORNA full-length mRNA reporters achieved enhanced protein expression.

[0180] The inventors combined GEMORNA-generated UTRs and CDSs, and conducted in vitro tests using Fluc2P reporter. For comparative analysis, the control sequence which consists of IDT CDS alongside natural AG UTRs (Benchmark-FL) was used. The inventors tested five combinations of the GEMORNA-generated elements, that are GMR-FL-F1 to GMR-FL-F5, against the control, with results presented in Fig. 12.

[0181] The inventors observed that all GEMORNA-designed sequences achieved significantly higher Fluc activity than Benchmark-FL. Specifically, GMR-FL-F5 demonstrated a 41-fold increase in expression compared to Benchmark-FL. Protein expression improvements were also observed in HepG2 cells (Fig. 12) .

[0182] The mRNA sequences utilized here are listed in Table 3.

[0183] Table 3. mRNA sequence List

[0184] Example 4: GEMORNA can directly design full-length therapeutics mRNAs with strong expression in vitro and in vivo.

[0185] For a therapeutics hEPO protein, seven GEMORNA-created mRNAs, that are GMR-EPO-F1 to GMR-EPO-F7, were compared to a benchmark sequence that consisted of codon-optimized CDS (R. Chen, S.K. Wang, J.A. Belk, L. Amaya, Z. Li, A. Cardenas, B.T. Abe, C. -K. Chen, P. A. Wender, H.Y. Chang, Engineering circular RNA for enhanced protein production. Nat Biotechnol 41, 262–272 (2023) ) and BNT162b2 UTRs; six of the GEMORNA designs, that are GMR-EPO-F2 to GMR-EPO-F7, achieved enhanced hEPO activities in both HEK293T and HepG2 cells (Fig. 13) . The inventors subsequently selected three hEPO designs, that are GMR-EPO-F5 to GMR-EPO-F7, with the highest in vitro expression for in vivo validation. GEMORNA sequences demonstrated stronger expression levels than the benchmark in mice. Notably, GMR-EPO-F7 achieved a 15-fold increase in expression compared to the benchmark at 24 hours (Fig. 14) .

[0186] The mRNA sequences utilized here are listed in Table 4.

[0187] Table 4. mRNA sequence List

[0188] Example 5: GEMORNA-designed mRNA vaccine demonstrated higher antigen expression and immunogenicity.

[0189] The inventors further validated the GEMORNA model on synthetic CDSs for COVID-19 mRNA vaccines. The inventors tested the immunogenicity of GEMORNA full-length designs, that are GMR-CV-F1 to GMR-CV-F3, against the benchmark, that is BNT162b2. GEMORNA-derived full-length mRNA induced higher antibody titers than BNT162b2 mRNAs at multiple time points post-immunization (Fig. 15) . The results confirm the generalizability of GEMORNA model to different proteins, including therapeutically relevant ones.

[0190] The mRNA sequences utilized here are listed in Table 5.

[0191] Table 5. mRNA sequence List

[0192] Although the specific embodiments have been described, for the applicant or a person skilled in the art, the substitutions, modifications, changes, improvements, and substantial equivalents of the above embodiments may exist or cannot be foreseen currently. Therefore, the submitted appended claims and claims that may be modified are intended to cover all such substitutions, modifications, changes, improvements, and substantial equivalents.

Claims

1.An RNA, comprising at least one nucleotide sequence with or without chemical modifications, wherein the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-41 and 53-57, any of truncated SEQ ID NOs: 1-41 and 53-57, or any sequence having at least 70%sequence identity of SEQ ID NOs: 1-41 and 53-57.2.The RNA according to claim 1, wherein the chemical modifications are selected from methylation, pseudouridylation, N1-methylpseudouridine, uridylation, adenylylation, phosphorylation, acetylation, bisulfite conversion.3.The RNA according to claim 1, wherein the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 1-4 and 53-57, any of truncated SEQ ID NOs: 1-4 and 53-57, or any sequence having at least 70%sequence identity of SEQ ID NOs: 1-4 and 53-57, preferably, the RNA is linear.4.The RNA according to claim 1, wherein the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 5-16, any of truncated SEQ ID NOs: 5-16, or any sequence having at least 70%sequence identity of SEQ ID NOs: 5-16.5.The RNA according to claim 1, wherein the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 17-26, any of truncated SEQ ID NOs: 17-26, or any sequence having at least 70%sequence identity of SEQ ID NOs: 17-26.6.The RNA according to claim 1, wherein the RNA comprises at least one nucleotide sequence with or without chemical modification, and the at least one nucleotide sequence is selected from any of SEQ ID NOs: 27-41, any of truncated SEQ ID NOs: 27-41, or any sequence having at least 70%sequence identity of SEQ ID NOs: 27-41.7.The RNA according to claim 1, wherein the RNA comprises:(1) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or(2) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or(3) a nucleotide sequence as set forth in SEQ ID NO: 14, a nucleotide sequence of truncated SEQ ID NO: 14, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 14, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or(4) a nucleotide sequence as set forth in SEQ ID NO: 14, a nucleotide sequence of truncated SEQ ID NO: 14, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 14, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or(5) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 2, a nucleotide sequence of truncated SEQ ID NO: 2, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 2, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or(6) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 53, a nucleotide sequence of truncated SEQ ID NO: 53, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 53, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or(7) a nucleotide sequence as set forth in SEQ ID NO: 11, a nucleotide sequence of truncated SEQ ID NO: 11, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 11, a nucleotide sequence as set forth in SEQ ID NO: 55, a nucleotide sequence of truncated SEQ ID NO: 55, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 55, and a nucleotide sequence as set forth in SEQ ID NO: 17, a nucleotide sequence of truncated SEQ ID NO: 17, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 17; or(8) a nucleotide sequence as set forth in SEQ ID NO: 13, a nucleotide sequence of truncated SEQ ID NO: 13, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 13, a nucleotide sequence as set forth in SEQ ID NO: 54, a nucleotide sequence of truncated SEQ ID NO: 54, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 54, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or(9) a nucleotide sequence as set forth in SEQ ID NO: 13, a nucleotide sequence of truncated SEQ ID NO: 13, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 13, a nucleotide sequence as set forth in SEQ ID NO: 56, a nucleotide sequence of truncated SEQ ID NO: 56, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 56, and a nucleotide sequence as set forth in SEQ ID NO: 18, a nucleotide sequence of truncated SEQ ID NO: 18, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 18; or(10) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 53, a nucleotide sequence of truncated SEQ ID NO: 53, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 53, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or(11) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 56, a nucleotide sequence of truncated SEQ ID NO: 56, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 56, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25; or(12) a nucleotide sequence as set forth in SEQ ID NO: 10, a nucleotide sequence of truncated SEQ ID NO: 10, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 10, a nucleotide sequence as set forth in SEQ ID NO: 57, a nucleotide sequence of truncated SEQ ID NO: 57, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 57, and a nucleotide sequence as set forth in SEQ ID NO: 25, a nucleotide sequence of truncated SEQ ID NO: 25, or a nucleotide sequence having at least 70%sequence identity of SEQ ID NO: 25.8.A nucleic acid, capable of generating the RNA according to any of claims 1-7.9.A synthetic construct, comprising the RNA according to any of claims 1-7, or the nucleic acid according to claim 8.10.A cell, comprising the RNA according to any of claims 1-7, the nucleic acid according to claim 8, or the synthetic construct according to claim 9.11.A conjugate, comprising the RNA according to any of claims 1-7, the nucleic acid according to claim 8, or the synthetic construct according to claim 9.12.A composition, comprising the RNA according to any of claims 1-7, the nucleic acid according to claim 8, the synthetic construct according to claim 9, the cell according to claim 10, or the conjugate according to claim 11.13.A method for preventing or treating a disease or condition, comprising administrating the RNA according to any of claims 1-7, the nucleic acid according to claim 8, the synthetic construct according to claim 9, the cell according to claim 10, the conjugate according to claim 11, or the composition according to claim 12 to the subject in need thereof.14.A use of the RNA according to any of claims 1-7, the nucleic acid according to claim 8, the synthetic construct according to claim 9, the cell according to claim 10, the conjugate according to claim 11, or the composition according to claim 12 in the manufacture of a drug for preventing or treating a disease or condition.15.A method for generating a CDS-based RNA design, comprising the following steps in order:(a) obtaining a source protein sequence to be transformed;(b) generating, based on a first predetermined embedding dimension, first vectors corresponding to amino acids in the source protein sequence;(c) inputting the first vectors into a trained neural network to generate a first representation of the source protein sequence;(d) initializing a target RNA sequence with a start token;(e) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence generated so far;(f) inputting the vectors obtained from (e) into the trained neural network to obtain a representation of the prefix of the target RNA sequence;(g) generating, based on the representations of the prefix of the target RNA sequence and the source protein sequence, an updated prefix of the target RNA sequence;(h) repeating the steps of (e) - (g) until a stop condition is met; and(i) outputting a final target RNA sequence, that is the CDS-based RNA.16.A method for generating a UTR-based RNA design, comprising the following steps in order:(a) initializing a target RNA sequence with a start token;(b) generating, based on a predetermined embedding dimension, vectors corresponding to the prefix of the target RNA sequence;(c) inputting the vectors into a trained neural network to generate an updated prefix of the target RNA sequence;(d) repeating the steps of (b) - (c) until a stop condition is met; and(e) outputting a final target RNA sequence, that is the UTR-based RNA.17.An RNA, comprising at least one RNA sequence obtained by the method according to claim 15, and / or at least one RNA sequence obtained by the method according to claim 16.18.The RNA of claim 17, wherein the RNA comprises at least one of the RNA according to any of claims 1-7.19.A deep learning-based model for CDS-based RNA design, named GEMORNA-CDS, utilizes the method of claim 15 to generate CDS-based RNAs.20.A deep learning-based model for UTR-based RNA design, named GEMORNA-UTR, utilizes the method of claim 16 to generate UTR-based RNAs.

Citation Information

Patent Citations

  • Synthetic nucleic acid molecule and methods of preparation

    WO2006034061A2

  • Modified nucleoside and synthesis method therefor

    WO2021027614A1

  • Systems and methods for producing RNA constructs with increased translation and stability

    WO2022047427A2

  • Modified functional nucleic acid molecules

    WO2022064221A1