Methods for detecting capped rna molecules

By using capped primers to transcribe RNA molecules in vitro and combining reverse transcription and PCR techniques, the problem of identifying and quantifying capped RNA molecules in existing technologies has been solved, thus improving the efficiency of quality control and safety analysis of mRNA vaccines and therapies.

CN122497762APending Publication Date: 2026-07-31THE UNIVERSITY OF QUEENSLAND
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE UNIVERSITY OF QUEENSLAND
Filing Date
2024-11-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and quantify capped and uncapped RNA molecules in samples, particularly in mRNA vaccines and therapies, impacting their activity and safety analysis.

Method used

RNA molecules were transcribed in vitro using a capped primer with the universal form m7GpppN1[N2]m[N3]n. Capped and uncapped RNA molecules were identified and quantified by identifying the presence or absence of 5'-terminal nucleotides. Quantitative analysis was performed using reverse transcription and PCR techniques.

Benefits of technology

This technology enables accurate identification and quantification of capped and uncapped RNA molecules, improving the efficiency of quality control and safety analysis for mRNA vaccines and therapies while reducing costs and time consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This disclosure relates to a method for identifying and / or quantifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in a sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been capped with a capping primer having the universal form m7GpppN1[N2]m[N3]n. In one embodiment, the method may involve identifying and / or quantifying at least the 5' terminal nucleotide of RNA molecules from the sample.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to Australian Provisional Application No. 2023903657, filed on 14 November 2023, the contents of which are incorporated herein by reference in their entirety.

[0003] Cross-referencing of sequence lists

[0004] This application contains a sequence list, which has been submitted electronically as an XML document in ST.26 format and is incorporated herein by reference in its entirety. Technical Field

[0005] This disclosure relates to methods for identifying capped and uncapped RNA molecules in a sample. In one form, it relates to a method for quantifying capped and / or uncapped mRNA molecules in a sample. Background Technology

[0006] Throughout this specification, any discussion of prior art should never be construed as an admission that such prior art is well-known or constitutes part of the general common sense of the field.

[0007] Synthetic RNA compositions (such as mRNA vaccines and therapies) utilize at least some features of endogenous mRNA to enable its translation via intracellular ribosomes. Synthetic mRNA molecules may contain a 5' cap, typically a 7-methyl-guanosine cap (5' m7G cap) or an analogue, which facilitates ribosome recognition and subsequent translation. The 5' cap may also have additional functions, such as regulating nuclear export of the mRNA molecule and preventing its degradation by exonucleases.

[0008] mRNA vaccines typically undergo rigorous analysis to ensure their quality, efficacy, and safety. This includes confirming the presence of the 5' cap. Uncapped mRNA molecules, even if otherwise intact, are generally not effectively bound by ribosomes and translated into functional protein products. The presence of uncapped mRNA molecules in synthetic mRNA pharmaceutical compositions may reduce the activity and function of the composition.

[0009] Furthermore, uncapped mRNA activates RIG-I (a cytoplasmic antiviral innate immune sensor), leading to an interferon response (Hornung et al., 2006). Reversed-phase ion-pair chromatography (IP-RP-HPLC) can be used to verify the presence of the 5' cap. Alternatively, liquid chromatography-mass spectrometry (LC-MS) can also be used for this purpose (Galloway et al., 2020; Beverly et al., 2016). However, both HPLC and MS-based methods are time-consuming and expensive. Moreover, IP-RP-HPLC and LC-MS provide insufficient quantitative assessment of the fraction of 5'-capped mRNA molecules in mixed mRNA samples.

[0010] In this context, a method is needed to identify capped and uncapped RNA molecules in a sample. Furthermore, a method is needed to quantify capped and / or uncapped RNA molecules in a sample. Summary of the Invention

[0011] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in a sample comprising a mixture of capped and uncapped RNA molecules;

[0012] This capped RNA molecule has been used in the universal form m7GpppN. 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0013] m7G is an N7-methylated guanosine or a guanosine analogue;

[0014] PPP stands for triphosphate;

[0015] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0016] m is 0 or 1; and

[0017] n is any integer from 0 to 8;

[0018] The method includes:

[0019] a) Identify at least the 5' terminal nucleotide of the RNA molecule from the sample;

[0020] in:

[0021] Identify N 1 The presence of a 5' terminal nucleotide is used to identify capped RNA molecules; and / or

[0022] Identify N 1 The absence of a 5' terminal nucleotide indicates the identification of uncapped RNA molecules.

[0023] In one aspect, the present invention provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in a sample comprising a mixture of capped RNA molecules and uncapped RNA molecules;

[0024] This capped RNA molecule has been used in the universal form m7GpppN. 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0025] m7G is an N7-methylated guanosine or a guanosine analogue;

[0026] PPP stands for triphosphate;

[0027] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0028] m is 0 or 1; and

[0029] n is any integer from 0 to 8;

[0030] The method includes:

[0031] a) Identify at least the 5' terminal nucleotide of multiple RNA molecules from the sample by the following methods:

[0032] Identify N 1 The presence of a 5'-terminal nucleotide, and then the identification of the 5'-terminal nucleotide as the nucleotide corresponding to the +1 position of the DNA template, thereby identifying capped RNA molecules; and / or

[0033] Identify N 1 The absence of the 5' terminal nucleoside, and then the identification of the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template, thus identifying the uncapped RNA molecule.

[0034] in:

[0035] The nucleoside at the +1 position of the DNA template contains adenosine, and N 1 Contains adenosine;

[0036] The nucleoside at the +1 position of the DNA template contains cytidine, and N 1 Contains cytidine; or

[0037] The nucleoside at the +1 position of the DNA template contains thymidine, and N 1 It is uridine; and

[0038] b) Quantitative N 1 The presence of [something] allows for the quantification of capped RNA molecules in the sample.

[0039] In one embodiment, this disclosure provides a method according to the aspect, wherein the nucleoside at position +1 of the DNA template is adenosine, and N1 is selected from the group consisting of adenosine, N2-methyladenosine, or N6-methyladenosine, and the nucleoside at position +2 of the DNA template comprises guanosine, and N2 comprises guanosine.

[0040] In one embodiment, N 1 It is not a preferred starting nucleoside for RNA polymerase.

[0041] In one embodiment, step a) further includes identifying the length of the RNA molecule, wherein the uncapped RNA molecule is x nucleotides long, and x is any suitable integer; and wherein identifying an RNA molecule that is at least x+1 nucleotides long is identifying a capped RNA molecule; and identifying an RNA molecule that is x nucleotides long is identifying an uncapped RNA molecule.

[0042] In one embodiment, the quantification in step b) includes:

[0043] Determine the absolute amount of capped RNA molecules in the sample;

[0044] Determine the abundance of capped RNA molecules relative to uncapped RNA molecules in a sample; and / or

[0045] Determine the abundance of capped RNA molecules relative to total RNA molecules in the sample.

[0046] In one embodiment, N 1 N 2 and N 3 At least one of each instance is independently any modified or non-natural nucleoside, and step a) further includes identifying the modified or non-natural nucleoside, wherein:

[0047] The identification of capped RNA molecules is achieved by detecting the presence of modified or non-natural nucleosides; and

[0048] The absence of modified or non-natural nucleosides is used to identify uncapped RNA molecules.

[0049] Optionally, the modified or non-natural nucleosides are selected from N2-methyladenosine or N6-methyladenosine.

[0050] In one embodiment, N1 The nucleosides are modified or non-natural, and step a) further includes identifying the modified or non-natural nucleosides, wherein:

[0051] The identification of capped RNA molecules is achieved by detecting the presence of modified or non-natural nucleosides as 5'-terminal nucleosides; and

[0052] The absence of modified or non-natural nucleosides as 5'-terminal nucleosides is used to identify uncapped RNA molecules.

[0053] In one embodiment, m = 1 and n = 0.

[0054] In one embodiment, the capping primer is selected from the group consisting of: 5'm7GpppAN, m7Gpppm2AN, m7Gpppm2Am2N, m7Gpppm6AN, and m7Gpppm6Am2N, m7G(5')ppp(5')G, m7(3')-O-Me-G(5')ppp(5')G), m7G(5')ppp(5')(2'OMeA)pG and m7G(5')ppp(5')(2'OMeA)pU), m7(3'OMeG)(5')ppp('5)m6(2'OMeA)pG), m7Gpppm6A mN, m7GpppN mN and 7mG3OMe5'ppp5'N.

[0055] In one embodiment, the RNA molecule has been transcribed by T7 RNA polymerase.

[0056] In one embodiment, the DNA template comprises:

[0057] The sequences selected are from the group consisting of TAATACGACTCACTATA (SEQ ID NO: 23), TAATACGACTCACTATAAG (SEQ ID NO: 15) and TAATACGACTCACTATAAT (SEQ ID NO: 24), wherein the RNA molecule has been transcribed by T7 RNA polymerase;

[0058] The sequence AATTAACCCTCACTATA (SEQ ID NO: 27) indicates that the RNA molecule has been transcribed by T3 RNA polymerase; or

[0059] The sequence is ATTTAGGTGACACTATA (SEQ ID NO: 28), in which the RNA molecule has been transcribed by Sp6 RNA polymerase.

[0060] In one embodiment, the identification in step a) further includes

[0061] Reverse transcription of RNA molecules from a sample to form complementary DNA (cDNA) molecules, and

[0062] Identifying at least the terminal nucleoside of a cDNA molecule, wherein identifying the terminal nucleoside of a cDNA molecule is equivalent to identifying at least the 5' terminal nucleoside of an RNA molecule.

[0063] In one embodiment, the identification in step a) includes identifying at least one to three 5' nucleotides of the RNA molecule, identifying at least one to ten 5' nucleotides of the RNA molecule, or identifying substantially all nucleotides of the RNA molecule.

[0064] In one embodiment, identification includes methods selected from the group consisting of: reverse transcription polymerase chain reaction (PCR) (RT-PCR), real-time quantitative RT-PCR (qRT-PCR), TaqMan qPCR, quantitative PCR (qPCR), digital PCR (dPCR), digital RT-PCR (dRT-PCR), RT-droplet digital PCR (RT-ddPCR), nested PCR, multiplex PCR, touchdown PCR, hot-start PCR and high-fidelity PCR, RT-PCR combined with high-resolution melting (HRM) analysis, Illumina sequencing, Ion Torrent sequencing, PacBio sequencing, Nanopore sequencing, Sanger sequencing and RNA-seq sequencing; direct RNA sequencing; sequencing-while-synthesizing; ligation or probe hybridization methods; RNA sequencing, RNA ligation, cDNA sequencing, isothermal amplification methods (such as LAMP), RT-Loop-mediated isothermal amplification (RT-LAMP), and cDNA synthesis or hybridization with specific oligonucleotides using primers with a specific template.

[0065] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample, the RNA molecule sample comprising a mixture of capped RNA molecules and uncapped RNA molecules;

[0066] This capped RNA molecule has been used in the universal form m7GpppN. 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0067] m7G is an N7-methylated guanosine or a guanosine analogue;

[0068] PPP stands for triphosphate;

[0069] N 1 It is a modified or non-natural nucleoside, and N2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; or N 2 It is a modified or non-natural nucleoside, and N 1 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0070] m is 0 or 1; and

[0071] n is any integer from 0 to 8;

[0072] The method includes:

[0073] a) Identify whether at least the 5' terminal nucleoside of the RNA molecule from the sample is natural, modified, or non-natural;

[0074] in:

[0075] The identification of capped RNA molecules is achieved by detecting the presence of modified or non-natural nucleosides as 5'-terminal nucleosides; and

[0076] The absence of modified or non-natural nucleosides as 5'-terminal nucleosides is used to identify uncapped RNA molecules.

[0077] In one aspect, this disclosure provides a method for quantifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample, the RNA molecule sample comprising a mixture of capped RNA molecules and uncapped RNA molecules;

[0078] This capped RNA molecule has been used in the universal form m7GpppN. 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0079] m7G is an N7-methylated guanosine or a guanosine analogue;

[0080] PPP stands for triphosphate;

[0081] N 1 It is a modified or non-natural nucleoside, and N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside, or

[0082] N 2 It is a modified or non-natural nucleoside, and N 1 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0083] m is 0 or 1; and

[0084] n is any integer from 0 to 8;

[0085] The method includes:

[0086] To identify whether at least the 5' terminal nucleosides of multiple RNA molecules from a sample are natural, modified, or non-natural nucleosides;

[0087] in:

[0088] The identification of capped RNA molecules is defined as the presence of modified or non-natural nucleosides as 5'-terminal nucleosides; and

[0089] The absence of modified or non-natural nucleosides as 5'-terminal nucleosides is used to identify uncapped RNA molecules.

[0090] The presence of modified or non-natural nucleosides as 5'-terminal nucleosides is quantified, thereby quantifying the capped RNA molecule in the sample.

[0091] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample, the RNA molecule sample comprising a mixture of capped RNA molecules and uncapped RNA molecules;

[0092] This capped RNA molecule has been used in the universal form m7GpppN. 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0093] m7G is an N7-methylated guanosine or a guanosine analogue;

[0094] PPP stands for triphosphate;

[0095] N 1 Any natural, modified, or non-natural nucleoside, in which N 1 It is not the preferred starting nucleotide for this RNA polymerase;

[0096] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; and

[0097] m is 0 or 1; and

[0098] n is any integer from 0 to 8;

[0099] The method includes:

[0100] To determine the length of an RNA molecule, where:

[0101] An uncapped RNA molecule is x nucleotides long, and

[0102] x is any suitable integer; and

[0103] in:

[0104] Identifying an RNA molecule that is at least x+1 nucleotides long constitutes identification of a capped RNA molecule; and

[0105] Identifying an RNA molecule that is x nucleotides long means identifying an uncapped RNA molecule.

[0106] In one aspect, this disclosure provides a pharmaceutical composition comprising a capped RNA molecule identified by the method according to any one of the preceding aspects.

[0107] In one aspect, this disclosure provides a multivalent pharmaceutical composition comprising a plurality of populations of capped RNA molecules identified by the methods of this disclosure. Attached Figure Description

[0108] Figure 1 A schematic diagram is provided, showing an example of (A) a capped primer of the form m7GpppAmG, where Am is N. 1 And G is N 2(B) A DNA template containing a natural T7 promoter sequence (italicized), an unmodified first transcription start site (TSS; underlined), showing the DNA sequence corresponding to the transcription nucleotide (bold), with arrows indicating the expected actual transcription start site; (C) A DNA template containing a T7 promoter sequence (italicized) with a modified first TSS (underlined) to enhance the initiation of in vitro transcription using the capped primer of (A), showing the DNA sequence corresponding to the transcription nucleotide (bold) when transcription is initiated with a capped primer of the form m7GpppAmG, with arrows indicating the expected actual transcription start site; (D) A DNA template containing a T7 promoter sequence (italicized) with a modified TSS (underlined) to enhance the initiation of in vitro transcription using the capped primer of (A), showing the DNA sequence corresponding to the transcription nucleotide (bold) when transcription is initiated with guanosine triphosphate (instead of a capped primer), with arrows indicating the expected actual transcription start site. The shaded boxes in (B), (C), and (D) highlight the positions of two nucleotides that can vary to correspond to the nucleosides in the capped primer. These nucleotides can act as alternative transcription start sites (TSSs), depending on whether transcription is initiated by the capped primer (i.e., the first transcription start site) or by guanosine triphosphate (the second transcription start site). The first terminal transcription start site corresponds to the 5' end +1 nucleoside of the capped RNA molecule (E); and the second terminal transcription start site corresponds to the 5' end +2 nucleoside of the capped RNA molecule (E). (E) shows a capped RNA molecule transcribed from the DNA template at (C), where transcription is initiated at the first TSS. The 5' nucleoside of the capped RNA molecule is a methylated adenosine provided by the capped primer, shown at the +1 position, and the next downstream nucleotide is a guanosine from the capped primer, shown at the +2 position (boxed out). (F) shows an uncapped RNA molecule transcribed from the DNA template at (D), where transcription is initiated at the second TSS. 5' nucleotides are guanosines incorporated from GTP during transcription (boxed out). This corresponds to the +2 position (shown in parentheses) and the second transcription start site on the capped RNA molecule.

[0109] Figure 2 A schematic diagram illustrating the key features of the synthesized mRNA polynucleotide molecule is provided.

[0110] Figure 3A schematic diagram illustrating in vivo transcription and enzyme capping using T7 polymerase is provided, wherein (A) shows a DNA template with a natural T7 promoter sequence (underlined), indicating that the transcription start site (TSS) is guanosine (G; bold) located at the first transcription start site (TSS) (i.e., the +1 position) of the DNA template; and (B) shows an mRNA molecule transcribed from a DNA template with a 5' m7G cap1, with methylated guanosine shown after transcription and enzyme capping (boxed).

[0111] Figure 4 A schematic diagram is provided illustrating the in vitro co-transcriptional capping process using T7 polymerase and capping primers (e.g., CleanCap1 AG reagent (m7Gpppm2AmG)), where (A) is a DNA template with a modified T7 promoter sequence (underlined), indicating adenosine at the first transcription start site (TSS) (i.e., the +1 position) (A; bold); and the mRNA molecule produced from the DNA template after transcription and enzymatic capping, the DNA template having: (B) CleanCap1 AG reagent (m7Gpppm2AmG), showing that the 5' terminal nucleotide (i.e., +1 relative to the DNA template) is 2'OMeA (methylated adenosine at the 2'O position; bold), (C) CleanCap2 AG reagent (m7Gpppm2Am2G) shows that the 5' terminal nucleotide (+1 relative to the DNA template) is 2'OMeA (adenosine methylated at the 2'O position; bold), and the next downstream nucleotide (+2 relative to the DNA template) is 2'OMeG (guanosine methylated at the 2'O position), or (D) CleanCap M6 reagent (m7Gpppm6Am2G) shows that the 5' terminal nucleotide (+1 relative to the DNA template) is 2'OMeA and N 6 -Methylated adenosine (in bold), and its next downstream nucleotide (i.e., +2 relative to the DNA template) is guanosine. Adenosine as a 5' transcription nucleotide was detected and / or methylation was detected indicating the incorporation of a 5' cap analogue for synthesis.

[0112] Figure 5 A schematic diagram illustrating in vitro transcription using T7 polymerase in the absence of a 5' cap is provided, where (A) is a DNA template with an artificial T7 promoter sequence (underlined), indicating that transcription begins at a second transcription start site at the +2 guanosine (G) position (bold); and (B) is an mRNA molecule transcribed from a DNA template without a 5' cap analog, showing that the 5' terminal nucleotide (relative to the +2 of the DNA template) is an unmethylated guanosine (bold). The detection of guanosine as a 5' transcription nucleotide and / or the absence of methylation indicate the absence of a synthetic 5' cap analog.

[0113] Figure 6 provides (A) a schematic diagram of mRNA vaccine samples with capped and uncapped mRNA molecules synthesized in the presence of CleanCap reagent AG of form m7GpppAmG, where transcription (top) begins at the first TSS nucleotide (i.e., A (+1) nucleotide; bold) of the capped RNA molecule, or (bottom) transcription begins at the second TSS nucleotide (i.e., G (+2) nucleotide; bold) of the uncapped RNA molecule; (B) the same molecule incorporated into the library adaptor after library preparation; and (C) a histogram showing the next-generation sequencing results, measuring the 5' end nucleotides, showing the alignment of capped mRNA molecules encoding green fluorescent protein starting with A (+1) in (top) and the alignment of uncapped molecules starting with G (+2) in (bottom).

[0114] Figure 7 Histograms of sequenced mRNA read alignments on the artificial T7 promoter are provided, generated from the following: (top) capped mRNA molecules co-transcribed with a synthetic 5' cap analog (CleanCapAG), with alignment starting at A(+1); (middle) uncapped mRNA molecules (in the absence of a synthetic 5' cap analog and no co-transcriptional capping), with alignment starting at G(+2); and (bottom) capped mRNA molecules co-transcribed after post-transcriptional enzymatic decapping, with alignment starting at A(+1).

[0115] Figure 8 Histograms of sequenced mRNA read alignments at the natural T7 promoter are provided, generated from the following: (top) uncapped mRNA molecules, with alignment starting at G(+1); (middle) capped mRNA molecules transcribed using the Forstovirus capping enzyme, with alignment starting at G(+1); and (bottom) post-transcriptionally enzymatically uncapped and post-transcriptionally capped mRNA molecules, with alignment starting at G(+1).

[0116] Figure 9A schematic diagram is provided illustrating the promoter region of the DNA template and the transcription start site during in vitro transcription using T7 RNA polymerase. (A) A natural T7 promoter sequence with a GG sequence at the transcription start site indicates that transcription begins at the first nucleotide (i.e., guanosine) of the GG sequence; (B) A modified T7 promoter sequence with an AG sequence at the transcription start site indicates that transcription begins at the first nucleotide (adenosine) of the AG sequence of the capped RNA molecule when transcription is initiated with a capped primer of the universal form m7GpppAG; or (C) A modified T7 promoter sequence with an AG sequence at the transcription start site indicates that transcription begins at the second nucleotide (guanosine) of the AG sequence of the uncapped RNA molecule when transcription is not initiated with a capped primer of the universal form m7GpppAG.

[0117] Figure 10 A histogram is provided showing the fraction of mRNA molecules in synthetic mRNA encoding green fluorescent protein (GFP) as measured using next-generation sequencing, with uncapped mRNA being 1186 nucleotides long (grey) and capped mRNA being 1187 nucleotides long, where the sequenced reads are aligned with the reference plasmid sequence.

[0118] Figure 11 A scatter plot is provided showing the direct correlation between the proportion of RNA molecules with adenosine as a 5' terminal nucleotide in the sample and the proportion of capped molecules, as measured using next-generation sequencing.

[0119] Figure 12 illustrates the design of a TaqMan probe for detecting the 5' nucleotide of an RNA molecule. The TaqMan probe was designed to be complementary to the TSO, TSS, and 5' UTR regions of the side-attached RNA molecule sequence. Two probes were designed: (A) with capped mRNA, incorporating the first TSS (i.e., A (+1)) and the second TSS (i.e., G (+2)), or (B) with uncapped mRNA, incorporating the second TSS (i.e., G (+2)) but not the first TSS (i.e., A (+1)). The capped mRNA detection probe incorporated a HEX fluorophore, and the uncapped mRNA detection probe incorporated a FAM fluorophore.

[0120] Figure 13 shows the primer design for the SYBR Green Assay to detect the 5' nucleotide of an RNA molecule. The primers were designed to be complementary to the TSO and TSS regions of the side-attached RNA molecule sequence. The 3' terminal nucleotide of the forward primer is complementary to the TSS of the RNA and terminates at (A) the first TSS in the case of synthetic Cap analog incorporation (i.e., A (+1)), or (B) for uncapped mRNA, the second TSS (i.e., G (+2)), instead of the first TSS (i.e., A (+1)).

[0121] Figure 14 shows the results of the qRT-PCR capping assay. (A) shows the δCt values ​​derived from the SYBR capping analysis, normalized to the δCt values ​​from the cap-independent control primers. The results are presented as a ratio between each sample and the results from 100% capped samples. (B) shows a scatter plot of the δCt analysis from the Taqman capping assay. The δCt values ​​from the Taqman capping analysis were normalized to the δCt values ​​from the cap-independent control primers. The results are presented as a ratio between each sample and the results from 100% capped samples. Detailed Implementation

[0122] The presence of a 5' cap on a synthetic RNA molecule can facilitate the translation of the encoded target polypeptide. The 5' cap can be incorporated into the synthetic RNA molecule during transcription initiation (i.e., co-transcriptional capping) or added after transcription (i.e., post-transcriptional capping). However, the capping process of RNA molecules is not 100% efficient. Therefore, samples of synthetic mRNA vaccines or therapeutics may include a portion of capped and uncapped mRNA molecules. The relative or absolute amount of capped mRNA molecules within a sample of synthetic mRNA molecules (such as mRNA vaccines or therapeutics) is typically assessed.

[0123] This disclosure relates to a method for measuring the presence or absence of a 5' cap on synthetic RNA molecules within a sample, for example, by analyzing the 5' end profile and / or length of the synthetic RNA molecules, thereby inferring the presence or absence of a 5' cap. In one embodiment, this provides a way to determine the relative abundance or absolute amount of capped and uncapped RNA molecules in a sample.

[0124] In one embodiment, this disclosure relates to distinguishing capped RNA molecules from uncapped RNA molecules in a sample comprising a mixture of capped and uncapped RNA molecules transcribed in vitro from a DNA template by RNA polymerase, wherein the capped RNA molecules have been converted to RNA with the universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n Capped primers; where N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; by identifying at least one structural difference between capped and uncapped RNA molecules from a sample. In one embodiment, N 1 It is not a preferred starting nucleoside for RNA polymerase.

[0125] RNA molecules can be transcribed from a DNA template using RNA polymerase in the presence of a capped primer. The DNA template may have a nucleotide sequence containing a promoter sequence suitable for use with RNA polymerase, including a transcription start site that facilitates the binding of perfectly or substantially complementary base pairs between the capped primer and the 3' strand (i.e., the template strand) of the DNA template during transcription.

[0126] When a polymerase selects a capping primer to initiate the transcription of an RNA molecule, the resulting RNA molecule is capped during the transcription reaction. This is called co-transcriptional capping. In this example, the N of the capping primer... 1 Nucleosides are the 5' terminal nucleosides of capped RNA molecules (see...) Figure 1 E). Typically, transcription occurs in the region corresponding to N. 1 The transcription begins at the first transcription start site of the modified promoter at the nucleotide position (compare). Figure 1 A and Figure 1 C).

[0127] When the polymerase does not select a capped primer to initiate transcription of RNA molecules, transcription is initiated by an alternative transcription initiator (e.g., a nucleotide or analogue). Typically, the preferred transcription initiator for this promoter is a nucleoside triphosphate, which produces an uncapped RNA molecule. In this example, the N capped primer... 1 Nucleosides are not the 5' terminal nucleosides of uncapped RNA molecules (compare). Figure 1 (A and 1F). Typically, transcription begins at the first nucleotide of the preferred transcription initiation nucleotide corresponding to RNA polymerase.

[0128] In one embodiment, the N capped primer 1 Nucleosides are not preferred initiating nucleosides for RNA polymerase. In this embodiment, when transcription is initiated by selecting a preferred transcription initiation nucleoside triphosphate (instead of a capped primer), transcription does not correspond to N... 1 The transcription process begins at the nucleotide position (i.e., transcription does not begin at the first transcription start site; the reference DNA template is called the +1 position, such as...). Figure 1 (As shown in D). Transcription typically begins at the second transcription start site (the reference DNA template is called the +2 position, e.g., ...). Figure 1 The transcription start site begins at (as shown in D). Therefore, in this embodiment, the transcription start site of the capped RNA molecule differs from that of the uncapped RNA molecule (compare). Figure 1 C and 1D).

[0129] This disclosure relates to the discovery of identifiable differences between capped and uncapped RNA molecules resulting from differential transcription initiation (e.g., incorporation of a capping primer or nucleotide to initiate transcription).

[0130] In one embodiment, the method includes identifying features selected from a group consisting of:

[0131] Identify N 1 The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N... 1 The absence of the 5' terminal nucleotide identified an uncapped RNA molecule.

[0132] Identifying the 5' terminal nucleotide as the first transcription start site corresponding to the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleotide as the second transcription start site corresponding to the DNA template identifies uncapped RNA molecules.

[0133] Identifying the 5' terminal nucleoside as the nucleoside corresponding to the +1 position of the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template identifies uncapped RNA molecules.

[0134] Identifying modified or non-natural nucleosides identifies capped RNA molecules; and / or identifying the absence of modified or non-natural nucleosides identifies uncapped RNA molecules, in which N 1 N 2 and N 3 At least one of each instance is independently any modified or non-natural nucleoside; and

[0135] The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; RNA molecules with a length of at least x+1 nucleotides are identified, which identifies capped RNA molecules; and RNA molecules with a length of x nucleotides are identified, which identifies uncapped RNA molecules.

[0136] In one form, this disclosure relates to a method for distinguishing and / or identifying capped and uncapped RNA molecules in a sample. In another form, it quantifies the presence of capped and / or uncapped RNA molecules in a sample. In one embodiment, it quantifies the absolute amount of capped and / or uncapped RNA molecules in a sample. In one embodiment, it quantifies the abundance of capped RNA molecules relative to uncapped RNA molecules in a sample. In one embodiment, it quantifies the abundance of capped RNA molecules relative to total RNA molecules in a sample.

[0137] The relative or absolute amounts of capped and / or uncapped RNA molecules can be determined using a number of different techniques, including capillary electrophoresis, PCR-based techniques, or next-generation sequencing-based techniques.

[0138] The method disclosed herein is compatible with other methods for distinguishing capped RNA molecules from uncapped RNA molecules in a sample.

[0139] definition

[0140] In the context of this specification, the term "a / an" is used herein to refer to one or more (i.e., at least one) grammatical object of an article. For example, "an element" means one element or more elements.

[0141] As used herein, the term "about," when applied to a value of interest, refers to a value similar to the stated value. In some embodiments, it should be understood to refer to a range of + / - 10%, preferably + / - 9%, + / - 8%, + / - 7%, + / - 6%, + / - 5%, + / - 4%, + / - 3%, + / - 2%, or + / - 1% of the stated value; or + / - 0.05% or + / - 0.1% of the stated value, unless otherwise stated or otherwise apparent from the context.

[0142] As used herein, the term "capped primer" is used to describe short oligonucleotide molecules comprising, for example, a 5' cap or 5' cap analogue linked to, for example, one to ten nucleotides via a phosphate group (e.g., a triphosphate bridge). Capped primers can be incorporated into the 5' end of an RNA molecule in a transcriptionally initiating and co-transcribed manner to provide a 5' cap analogue of the RNA molecule. In one embodiment, the capped primer has a first nucleotide (N...) linked to via a 5'-5' triphosphate bridge. 1 5',7-methylguanosine (m) 7 G), the first nucleotide can then be linked to a further nucleotide (N). 2 The 5' end of the primer, etc. Capped primers typically have a 3' OH group to facilitate ligation with a 3' nucleoside during transcription.

[0143] As used herein, the term "capped RNA molecule" refers to an RNA molecule having a 5' cap or a 5' cap analogue. In one embodiment, the 5' cap can be bound to the 5' carbon of the 5' terminal nucleotide of the mRNA polynucleotide via a 5' triphosphate bridge. The 5' cap can be a natural 5' cap or a 5' cap analogue.

[0144] The terms “comprise,” “comprised,” or “have” used in this specification and claims are used in an inclusive sense, that is, specifying the presence of the stated feature but not excluding the presence of additional or further features.

[0145] As used herein, the term "complementary" describes the relationship between the first and second nucleotide sequences according to the Watson-Crick base pairing rule, where an adenine (A) base pairs with a uracil (U) base in an RNA molecule or a thymine (T) base in a DNA molecule; and a cytosine (C) base pairs with a guanine (G) base in both RNA and DNA molecules. For example, for a DNA polynucleotide molecule, the sequence "5'-AGTC-3'" is perfectly complementary to the sequence "3'-TCAG-5'"; note that in RNA sequences, uracil (U) is typically used in place of thymine (T).

[0146] The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization. This is particularly important in detection methods that depend on the binding between nucleic acids. Nucleic acid sequences do not need to be "perfectly" (100%) complementary to their target sequences for hybridization to occur. Complementarity can be "partial," where only a few bases in the nucleic acid are matched according to base pairing rules. It should be understood that when two molecules can hybridize under suitable conditions—that is, when two molecules can hybridize and form a double-stranded structure under conditions suitable for a reaction (e.g., ligation, PCR, sequencing, etc.)—the first polynucleotide molecule containing the first nucleotide sequence is specifically complementary to the second polynucleotide molecule containing the second nucleotide sequence; the two sequences are "specifically complementary." The term "specifically complementary" can be used interchangeably with "substantially complementary." It should also be understood that two nucleotide molecules do not need to be complementary along their entire length. For example, a portion of the first polynucleotide molecule may be specifically complementary to and hybridize with a portion of the second polynucleotide molecule. In this example, the two molecules may not hybridize at non-specifically complementary portions. These terms can also be used to refer to individual nucleotides, especially in the context of oligonucleotides. For example, in contrast to or in contrast to the complementarity between the rest of the oligonucleotide and the nucleic acid chain, a particular nucleotide in an oligonucleotide may be noted to have complementarity or lack of complementarity with nucleotides in another nucleic acid chain.

[0147] Hybridization conditions can be stringent, such as 400 mM NaCl, 40 mM PIPES at pH 6.4, 1 mM EDTA, at 50 or 70°C for 12–16 hours, followed by washing. Other conditions, such as physiologically relevant conditions that may be encountered within the organism, can also be applied. Fundamental complementarity allows the relevant functions of nucleic acids to proceed, for example, the binding of oligonucleotides to polynucleotides during PCR-based assays, sequencing, etc. Technicians will be able to determine the most suitable set of conditions for testing the complementarity of the two sequences based on the final application of the hybridized nucleotides.

[0148] As used herein, the term "corresponding" refers to the relationship between two nucleotides present at the same or equivalent positions within different polynucleotide molecules having the same or complementary nucleobases. Corresponding nucleotides may have different modifications. Furthermore, as those skilled in the art will understand, uracil in an RNA molecule may correspond to thymine in a DNA molecule. Therefore, when the nucleotide sequence of a capped primer corresponds to the nucleotide sequence of a DNA template, the capped primer typically binds to the 3' strand of the DNA template (i.e., the template strand) at the relevant site of interest using complementary binding rules.

[0149] As used herein, the term "co-transcription" refers to an in vitro transcription reaction in which RNA molecules are transcribed from a DNA template, for example, by RNA polymerase, where the capping of the RNA molecules occurs in the same transcription reaction. In one form of co-transcription, transcription is initiated with a capping primer such that the capping primer is incorporated into the 5' end of the RNA molecule, which both caps the RNA molecule with a 5' cap analogue and initiates transcription.

[0150] As used herein, the term "DNA template" refers to a double-stranded DNA molecule that contains at least a promoter sequence and encodes a target RNA molecule of interest from which the RNA molecule can be transcribed. A DNA template can be double-stranded linear DNA, partially double-stranded linear DNA, circular double-stranded DNA, a DNA plasmid, a PCR amplicon, or a modified nucleic acid template compatible with RNA polymerase.

[0151] As used in this article, the term "effectively" when used with respect to a particular parameter or result is intended to refer to a sufficient percentage of the parameter or result in order to achieve the desired result.

[0152] As used herein, “expression” of a nucleic acid sequence refers to post-translational modifications that translate mRNA into a polypeptide, assemble multiple polypeptides into a complete protein (e.g., an enzyme), and / or a polypeptide or a fully assembled protein (e.g., an enzyme). In this application, the terms “expression” and “production”, as well as grammatical equivalents, are used interchangeably.

[0153] As used herein, the term "isolated" means material that is substantially or substantially free of the components that normally accompany it in its natural state. For example, "isolated polynucleotide" as used herein refers to a polynucleotide whose sequence has been purified from its natural state, such as a DNA fragment that has been removed from a sequence normally adjacent to the fragment. Alternatively, as used herein, "isolated peptide" or "isolated polypeptide," etc., refers to peptide or polypeptide molecules that have been isolated and / or purified in vitro from their natural cellular environment and association with other components of the cell, i.e., that they do not associate with substances in vivo.

[0154] As used herein, the term "messenger RNA" or "mRNA" refers to an RNA polynucleotide molecule that encodes at least one polypeptide. As used herein, mRNA encompasses both modified and unmodified RNA. mRNA may contain one or more coding and non-coding regions.

[0155] As used herein, the term "nucleotide" refers in its broadest sense to compounds and / or substances incorporated into or potentially incorporated into polynucleotide chains. Those skilled in the art will understand that nucleotides typically consist of three distinct chemical subunits: a pentose sugar molecule (i.e., the pentose sugar ring, deoxyribose in DNA, or ribose in RNA), a nucleobase (e.g., adenine (A), cytosine (C), guanine (G), thymine (T), or uracil (U)), and a phosphate group (e.g., a monophosphate, diphosphate, or triphosphate group). Chemical convention designates the carbon atoms in sugar molecules as 1' through 5', and this convention also determines that polynucleotide molecules have a 5' end and a 3' end. In a polynucleotide molecule, the 3' carbon of the first nucleotide is linked to the 5' carbon of the next nucleotide. Those skilled in the art will understand that, unless otherwise specifically stated, polynucleotide sequences are generally read in a 5' through 3' orientation.

[0156] In some embodiments, a nucleotide is a compound and / or substance incorporated into or potentially incorporated into a polynucleotide chain via a phosphodiester bond. In some embodiments, "nucleotide" refers to an individual nucleic acid residue (e.g., a nucleotide and / or a nucleoside). The term "nucleotide" may be used interchangeably with "nucleic acid." "Nucleotide" encompasses RNA (ribonucleotides) as well as single-stranded and / or double-stranded DNA and / or cDNA.

[0157] Nucleotides consist of a nucleoside and a phosphate group. However, those skilled in the art will understand that, in common use, the term "nucleotide" includes nucleosides. An example of a nucleotide is the nucleoside 5' triphosphate (NTP). Table 1 provides other common triphosphate nucleotides. Nucleotides include RNA nucleotides (i.e., ribonucleotides) and DNA nucleotides (i.e., deoxyribonucleotides).

[0158] As used herein, the term "nucleoside" refers to a molecule containing a pentose sugar (ribose for RNA, deoxyribose for DNA) linked to a nucleobase (a nitrogenous base), and includes any suitable natural, modified, or non-natural nucleoside. Examples of common natural nucleosides include adenosine, guanosine, cytidine, 5-methyluridine, uridine, and thymidine, as shown in Table 1. Naturally occurring nucleobases include purine rings (e.g., adenine, guanine, and N-nucleotides). 6 (-methyladenine) and pyrimidine rings (e.g., cytosine, thymine, 5-methylcytosine, pseudouridine). For example, naturally occurring nucleosides include, but are not limited to, the ribose, 2'-O-methyl, or 2'-deoxyribose derivatives of adenosine, guanosine, cytidine, thymidine, uridine, inosine, 7-methylguanosine, or pseudouridine.

[0159] Table 1: Examples of nucleosides and nucleotides

[0160]

[0161] As used herein, the terms “nucleoside analog,” “modified nucleoside,” or “nucleoside derivative” include synthetic nucleosides as described herein. Nucleoside derivatives also comprise nucleosides having modified bases and / or sugar moieties, with or without protecting groups, and comprising, for example, 2'-deoxy-2'-fluorouridine, 5-fluorouridine, etc. The compounds and methods provided herein comprise such base rings and their synthetic analogs, as well as non-natural heterocyclic substituted basic sugars and acyclic substituted basic sugars. Other nucleoside derivatives that can be used in this disclosure include, for example, LNA nucleosides, halogenated purines (e.g., 6-fluoropurines), halogenated pyrimidines, N… 6 -Ethyl adenine, N 4 -(alkyl)-cytosine, 5-ethylcytosine, etc.

[0162] As used herein, “oligonucleotide,” “oligonucleotide molecule,” “oligonucleotide primer,” or “primer” is a single-stranded polynucleotide molecule that may be naturally occurring or synthetically produced and have a user-specified sequence of interest. It may be an RNA, DNA, or chimeric RNA / DNA polynucleotide molecule and may contain modified nucleosides. Typically, at least a portion of the oligonucleotide molecule is specifically or perfectly complementary to the polynucleotide sequence of interest and hybridizes to a specifically complementary single-stranded polynucleotide molecule. Oligonucleotide molecules are generally considered to be short polynucleotide molecules; however, their lengths may vary. Their lengths are suitable for use in at least one application in a range of applications, including polymerase chain reaction (PCR) based applications, sequencing applications, molecular cloning, and molecular probes. Oligonucleotide primers may contain one or more modifying groups. Oligonucleotide primers may include RNA, DNA, and / or other modified nucleosides. Those skilled in the art are able to design and prepare oligonucleotide primers suitable for transcription of DNA template sequences.

[0163] As used herein, the term "operationally linked" or "operationally connected" refers to a functional relationship between two or more nucleic acid segments, such as genes and regulatory elements (including but not limited to promoters), which in turn regulate gene expression.

[0164] As used herein, the term "pharmaceuticalally acceptable" means a substance that, when administered to a subject, will not cause a significantly adverse allergic or immune response. "Pharmaceuticalally acceptable carriers" include, but are not limited to, solvents, coatings, dispersants, wetting agents, isotonic agents, absorption delay agents, and disintegrants.

[0165] As used herein, the term "polynucleotide molecule" refers to a DNA or RNA nucleic acid molecule containing a nucleotide chain, and may include oligonucleotide molecules or target nucleic acids of interest. The terms "RNA nucleic acid molecule," "RNA molecule," and "RNA polynucleotide molecule" are used interchangeably. Similarly, the terms "DNA nucleic acid molecule," "DNA molecule," and "DNA polynucleotide molecule" are used interchangeably.

[0166] The term "polynucleotide variant" refers to a polynucleotide that exhibits substantially sequence identity with a reference polynucleotide sequence or, under stringent conditions, hybridizes with a reference sequence. This term also covers polynucleotides that are distinguished from a reference polynucleotide by the addition, deletion, or substitution of at least one nucleotide. Thus, the term "polynucleotide variant" includes polynucleotides in which one or more nucleotides have been added or deleted, or substituted with different nucleotides. In this regard, it is well known in the art that certain alterations can be made to a reference polynucleotide, including mutations, additions, deletions, and substitutions, such altered polynucleotides retaining the biological function or activity of the reference polynucleotide. The term "polynucleotide variant" also includes naturally occurring allelic variants. The terms "peptide variant" and "polypeptide variant," etc., refer to peptides and polypeptides that are distinguished from a reference peptide or polypeptide by the addition, deletion, or substitution of at least one amino acid residue. In some instances, peptide or polypeptide variants are distinguished from reference peptides or polypeptides by one or more substitutions, which may be conserved or non-conserved. In some instances, peptide or polypeptide variants contain conserved substitutions, and in this regard, it is well known in the art that some amino acids may be changed to other amino acids having substantially similar properties without altering the active properties of the peptide or polypeptide. Peptide and polypeptide variants also include peptides and polypeptides in which one or more amino acids have been added or missing, or which have been replaced with different amino acid residues.

[0167] As used herein, the terms “preferred initiating nucleoside” or “preferred initiating nucleotide” refer to the nucleoside or nucleotide that a particular RNA polymerase typically uses to initiate transcription, such as a nucleoside triphosphate.

[0168] As used herein, the term "promoter" refers to a region in a DNA template that directs and controls the initiation of transcription of a specific DNA sequence to produce RNA molecules. Promoters are located on the same strand of the DNA and upstream of it (towards the 5' region of the sense strand). Promoters are typically adjacent to (or partially overlap with) the DNA sequence to be transcribed. The nucleotide positions within a promoter are specified relative to the transcription start site (where transcription from DNA to RNA begins (position + 1)).

[0169] As used herein, the term "specificity," when referring to the initiating capped oligonucleotide primer sequence and its ability to hybridize with a DNA template, means a sequence that has at least 50% sequence identity with a portion of the DNA template when the initiating capped oligonucleotide primer is aligned with the DNA strand. Possibly preferred higher levels of sequence identity include at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, and most preferably 100% sequence identity.

[0170] When applied to polynucleotide molecules, the term "synthetic" is intended to mean that polynucleotide molecules are produced through an in vitro transcriptional reaction and do not occur naturally within cells.

[0171] As used herein, the term "synthetic RNA" or "synthetic mRNA" refers to RNA polynucleotide molecules produced using non-natural processes such as in vitro transcription, wherein the synthetic RNA is synthesized in a chemical reaction using polymerases, DNA templates, and ribonucleotides. Synthetic mRNA, as used herein, may include non-natural nucleosides, non-natural phosphate backbones, and non-natural 5' caps.

[0172] For example, synthetic RNA may contain non-natural nucleoside analogues, such as analogues with chemically modified bases or sugars, backbone modifications, etc. In some embodiments, mRNA is or contains natural nucleosides (e.g., adenosine, guanosine, cytidine, uridine); nucleoside analogues (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C-5-propynyl-cytidine, C-5-propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methyluridine, C5-methyluridine, C5-propynyl-cyt ... Cytidine, 2-aminoadenosine, 7-deadenosine, 7-deadenosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine; chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., thiophosphates and 5'-N-phosphoramide bonds). In some embodiments, the synthetic mRNA molecule incorporates modified, artificial or non-natural nucleotides, nucleoside analogs, nucleotide derivatives, modified uridine, N1-methyl-pseudouridine, etc.

[0173] "Therapeutic effective dose" is the minimum concentration or amount required to produce at least a measurable improvement in a particular disease or symptom. The therapeutic effective dose in this context may vary depending on factors such as the patient's disease state, age, sex, and weight. Therapeutic effective dose is also the amount at which the beneficial therapeutic outcome outweighs any toxic or harmful effects.

[0174] As used herein, the term “transcription” or “transcriptional reaction” refers to a method known in the art for the enzymatic production of RNA molecules complementary to a DNA template, thereby producing an RNA copy of the DNA sequence. The RNA molecule synthesized in a transcriptional reaction is called an “RNA transcript” or “transcript.”

[0175] As used in this article, the term "transcription initiation site" refers to a specific nucleotide within or after the promoter region of a DNA template from which transcription begins.

[0176] As used herein, the term "uncapped RNA molecule" refers to an RNA molecule that does not have a 5' cap. In one embodiment, an uncapped mRNA molecule contains a 5' triphosphate group, which is a triphosphate group that is bound to the 5' carbon of the 5' terminal nucleotide.

[0177] As used herein, the terms “unsubstituted” or “unmodified” in the context of capped primers and NTPs refer to unmodified capped primers and NTPs.

[0178] As used in this article, the term "upstream" when used in conjunction with polynucleotide molecules refers to the 5' direction. The term "downstream" refers to the 3' direction.

[0179] As used herein, the term "5' end" of a polynucleotide refers to the 5' carbon of the first nucleotide in the chain (i.e., the 5' terminal nucleotide) and any chemical group attached to the 5' carbon. As used herein, the term "5' end" of a polynucleotide molecule refers to the terminal portion of a molecule having a 5' carbon at its end. In one embodiment, the 5' end is a 5' terminal. In one embodiment, the 5' terminal portion includes a 5' end, a 5' cap (if present), a 5' terminal nucleotide, and a plurality of nucleotides immediately adjacent to the 5' terminal portion, for example, fewer than 50 nucleotides, fewer than 40 nucleotides, fewer than 30 nucleotides, fewer than 20 nucleotides, or fewer than 10 nucleotides.

[0180] The group on the 5' carbon can be a phosphate group (e.g., monophosphate, diphosphate, triphosphate, etc.); however, it should be understood that other groups can be present at the 5' end. The capped mRNA molecule of this disclosure has a cap that is attached to a 5' triphosphate bridge at the 5' carbon.

[0181] As used herein, the “3’ end” of a polynucleotide molecule refers to the 3’ carbon of the last nucleotide in the polynucleotide chain (i.e., the 3’ terminal nucleotide) and any chemical groups attached to the 3’ carbon. The 3’ carbon usually has a hydroxyl group; however, it should be understood that other groups may be attached to the 3’ carbon.

[0182] As used herein, the term "+1 position" refers to the first nucleotide in a capped RNA molecule, and the term "+2 position" refers to the second nucleotide, and so on. The +1 position of a capped RNA molecule corresponds to the first transcription start site in the DNA template, and the +2 position of a capped RNA molecule corresponds to the second transcription start site, such as... Figure 1 As shown in E. For the uncapped RNA molecule of this disclosure, transcription may begin from a second transcription start site, which corresponds to the +2 position relative to the capped RNA molecule, as shown in E. Figure 1 The F in parentheses indicates the position. The +1 position on the DNA template is the nucleotide that initiates transcription with the capped primer, i.e., the first transcription start site (TSS). The 3' nucleotide on the DNA template immediately adjacent to the first TSS is at the +2 position.

[0183] Synthetic RNA molecules

[0184] The success of mRNA vaccines and therapies is partly due to advancements in manufacturing technologies that have enabled the safe production of billions of doses during the COVID-19 pandemic. The synthesis of mRNA vaccines and therapies typically involves in vitro transcription. The resulting RNA, including messenger RNA, transfer RNA, small nucleolar RNA (snoRNA), guide RNA, etc., can be generated by in vitro transcription of a DNA template using an RNA polymerase in a reaction solution containing nucleoside-5'-triphosphates (NTPs) under suitable conditions. A 5' cap (such as a 5' m7G cap or similar) can be incorporated during in vitro transcription (i.e., co-transcription) or added post-transcriptionally using enzymes such as Forstovirus capping enzymes, vaccinia virus capping enzymes, etc. Target RNA molecules can be purified to remove residual contaminants and formulated for delivery.

[0185] In one embodiment, the RNA molecule can be an mRNA molecule, a transfer RNA molecule, a small nucleolar RNA (snoRNA) molecule, a guide RNA molecule, etc. In one embodiment, the RNA molecule can be an mRNA molecule.

[0186] The synthetic messenger (mRNA) molecules disclosed herein may include many features found in naturally occurring mRNA molecules, such as:

[0187] a. 5' cap, such as 5'-7-methyl-guanosine cap or cap analogues, which promote ribosome recognition and improve translation and mRNA stability;

[0188] b.5' Untranslated region (UTR), which helps or regulates ribosome translation, and its sequence can be natural or artificial;

[0189] c. Coding region, which includes the open reading frame encoding the protein of interest;

[0190] d.3' untranslated region, which is beneficial for improving translation and stability, can have a sequence that is either natural or artificial; and

[0191] The e.polyA tail, the length of which is related to mRNA stability and expression.

[0192] Figure 2 This demonstrates the features that may exist in synthetic mRNA molecules.

[0193] To improve translation efficiency, open reading frame sequences can be codon-optimized to avoid rare codons and contain fewer uracil residues. Synthetic mRNA vaccines and therapies also frequently incorporate non-naturally modified nucleotides, such as N1-methyl-pseudouridine and methoxyuridine, for example, to minimize recognition by innate immune responses and / or reduce susceptibility to nuclease digestion.

[0194] The mRNA sequence can be modified to reduce RNA secondary structure. The synthesized mRNA molecule may have an internal ribosome entry site (IRES), which facilitates translation initiation in the 5' cap-independent process. In one embodiment, the mRNA molecule may have a polyA tail, which may be longer than 10 nucleotides and may be encoded and transcribed within a DNA template. Alternatively, a polyA tail may be added after transcription using an enzyme such as polyadenylate.

[0195] Synthetic mRNA may, for example, encode a vaccine or antigen sequence, or a therapeutic protein, such as an enzyme, antibody (or fragment thereof), or peptide. In one embodiment, the synthetic mRNA molecule is a vaccine and / or therapeutic agent. In one embodiment, this disclosure provides a pharmaceutical composition comprising a synthetic mRNA molecule that has been tested using the methods of this disclosure.

[0196] The presence of a 5' cap and other RNA quality characteristics on the mRNA molecules of vaccines or therapies is typically rigorously analyzed to ensure their quality, safety, and efficacy (mRNA Vaccine Quality Analysis Methods - Draft Guideline, USP, 2023). The USP currently recommends using reversed-phase ion-pair chromatography (IP-RP-HPLC) to verify the presence of the 5' cap. Alternatively, liquid chromatography-mass spectrometry (LC-MS) can also be used for this purpose (Galloway et al., 2020; Beverly et al., 2016). This technique first uses Antarctic phosphatase to convert uncapped monophosphates, diphosphates, and triphosphates to 5'OH for analysis. The technique uses complementary DNA oligonucleotides hybridized to the mRNA to guide RNase H cleavage, followed by LC-MS analysis of the products. LC-MS trace analysis shows peaks corresponding to uncapped (phosphatase-treated 5'OH) and capped peaks, expressed in mass.

[0197] The performance of different capping techniques (including their effects on mRNA translation and stability) is typically assessed using reporter gene mRNAs such as luciferases. These assessments measure the effect of the 5' cap on mRNA expression, providing an indirect inference of the capping state rather than a direct detection of the 5' cap. Furthermore, methods based on both HPLC and MS are both time-consuming and expensive, and provide poor quantitative measurements of capped and uncapped mRNA molecules in mixed samples.

[0198] 5' cap

[0199] While not wanting to be bound by theory, the 5' cap is generally considered to facilitate the translation of open-read forms encoded by mRNA molecules. It should be understood that uncapped mRNA molecules, even if otherwise intact, are generally not translatable. Furthermore, uncapped mRNA can activate RIG-I (a cytoplasmic antiviral innate immune sensor), leading to an interferon response (Hornung et al. 2006). The presence of the internal ribosome entry site (IRES) may allow for translation initiation in a cap-independent pathway. The 5' cap may also have additional functions, such as regulating nuclear export of mRNA and / or preventing mRNA degradation by exonucleases.

[0200] Naturally occurring cap structures typically contain a 7-methylguanosine (m7G) cap, which is linked to the 5' end of the first transcribed nucleotide via a triphosphate bridge, forming a dinucleotide cap of m7G(5')ppp(5')N, where N is any nucleoside. This can also be represented interchangeably as m7GpppN. In vivo, the 5' terminal nucleoside (i.e., the cap-adjacent nucleotide) is usually guanosine. In vivo, the cap is enzymatically added to the cell nucleus immediately after transcription, a reaction typically catalyzed by guanylate transferases.

[0201] The 5' cap is beneficial for the effectiveness of synthetic mRNA vaccines and therapies. For synthetic mRNAs produced using in vitro transcription, the 5' cap is typically added to the mRNA molecule during the initiation of in vitro transcription (co-transcription), where a synthetic 5' cap analog is incorporated at the transcription initiation site; or after mRNA synthesis (i.e., post-transcriptional), where enzymes perform transferase and methylation reactions to convert the 5' mRNA terminus to a Cap0 or Cap1 structure.

[0202] While not wanting to be bound by theory, the 5' cap is generally considered to facilitate the translation of open-read forms encoded by mRNA molecules. It should be understood that uncapped mRNA molecules, even if otherwise intact, are generally untranslatable. Furthermore, uncapped mRNA can activate RIG-I (a cytoplasmic antiviral innate immune sensor), leading to an interferon response (Hornung et al. 2006). Additionally, the 5' cap is involved in the nuclear export, stability, and degradation of mRNA molecules.

[0203] In eukaryotic cells, shortly after transcription is initiated by guanylate transferase, a cap is added to the nascent RNA post-transcriptionally. During the capping reaction, a first intermediate, called Cap0(m), is generated. 7 GpppN). The Cap0 structure lacks a 2'-O-methyl residue at the first and second 5' terminal (i.e., cap-adjacent) nucleotides. The formation of Cap0 may subsequently involve methylation at the 2'-O position of the first 5' terminal nucleotide, resulting in Cap1 (m 7 The Cap1 structure can be further methylated at the 2'-O position of the second 5' terminal nucleotide to form the Cap2 (m7GpppNmNm) structure. The presence of Cap0, Cap1, and Cap2 is crucial for the cellular innate immune system's self / non-self recognition of mRNA. When the first 5' nucleotide is adenosine, adenosine can be further methylated by N... 6 Methylation to form m 6 Am cap.

[0204] In one embodiment, the capping primer of this disclosure comprises a Cap0 structure. In one embodiment, the capping primer of this disclosure comprises a Cap1 structure. The Cap1 structure has a 2'-O-methyl residue at the nucleotide adjacent to the first cap (5' end). In one embodiment, the capping primer of this disclosure comprises a Cap2 structure. The Cap2 structure has a 2'-O-methyl residue attached to the nucleotides adjacent to the first and second caps (5' ends). In one embodiment, the capping primer of this disclosure comprises a capM6 structure. The CapM6 structure has adenosine as the nucleotide adjacent to the first cap (5' end), which is methylated at the 2'O and N6 positions of the ribose.

[0205] Table 2: Methylation of Cap0, Cap1, Cap2, Cap M6 and TMG Cap structures

[0206]

[0207] in:

[0208] m indicates methylation

[0209] G-indicator guanosine

[0210] p indicates phosphate groups

[0211] N 1 and N 2 Indicates any nucleoside

[0212] A indicates adenosine.

[0213] The 5' cap of a synthetic mRNA molecule may contain the same 5' cap as that of an in vivo-produced mRNA molecule (e.g., m7G(5')ppp(5')N, where N is any nucleoside). Alternatively, the 5' cap of a synthetic mRNA molecule may contain a synthetic 5' cap analog. A variety of m7G cap analogs are known in the art. Cap structures may include ARCA 3'-OCH3, ARCA 2'-OCH3 cap analogs (Jemielity, J. et al., 2003), N7-benzylated dinucleotide tetraphosphate analogs (described in Grudzien, E. et al., 2004), phosphate thioester cap analogs (described in Grudzien-Nogalska, E. et al., 2007), and cap analogs (including biotinylated cap analogs), described in U.S. Patent Nos. 8,093,367, 8,304,529, and 11,377,642.

[0214] Capped primer

[0215] The capped primers disclosed herein may contain the universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n ;where N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8. In one embodiment, N 1 Not a preferred starting nucleoside for RNA polymerase. The capping primer may contain a 5' cap or cap analog linked to at least one nucleoside, which can be co-transcribed into the RNA molecule to provide a 5' cap analog to the RNA molecule.

[0216] The capping primer disclosed herein may contain 7-methylguanosine (m 7 G) A cap or the like, or an alternative cap structure as detailed herein.

[0217] The cap or similar can be linked to at least one nucleoside via any suitable linker. In one form, the linker is a phosphate group. In one embodiment, the linker is a triphosphate. In another form, the linker is a reverse 5' to 5' triphosphate bond.

[0218] At least one nucleoside may comprise any nucleoside suitable for incorporation into the 5' end of an RNA molecule. It may include any suitable natural, modified, or non-natural nucleoside, such as naturally occurring RNA and DNA nucleosides, synthetic nucleosides, or any suitable modified nucleoside, nucleoside analog, nucleoside derivative, or modified nucleobase. Examples include, but are not limited to, base- and sugar-modified nucleosides, nucleotides, and nucleic acids such as inosine, 7-deoxyguanosine, 2'-O-methylguanosine, 2'-fluoro-2'-deoxycytidine, pseudouridine, locked nucleic acid (LNA), and peptide nucleic acid (PNA).

[0219] Modified nucleosides may include modified, synthetic, or natural nucleobases, such as deoxythymidine (dT), 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine; 2-thiouracil, 2-thiothymidine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyluracil and cytosine, 6-azauracil... Pyrimidines, cytosines and thymines, 5-uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halogenated (especially 5-bromine), 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deadenine and 7-deadenine, and 3-deadenine and 3-deadenine.

[0220] The modified nucleoside may have one or more substituted sugar moieties, for example, it may include one of the following at the 2'-position: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-ynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and ynyl groups may be substituted or unsubstituted C1 to C2 groups. 10 Alkyl or C2 to C 10 Alkenyl and ynyl groups.

[0221] Capped primers may contain natural phosphodiester bonds between nucleotides or modifications thereof, or combinations thereof. Examples of oligonucleotide internucleotide bond modifications include thiophosphates, phosphate triesters, and methylphosphonate derivatives.

[0222] In one respect, capped primers have a universal form: [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n ; wherein the cap includes any suitable cap as disclosed herein; the connector includes any suitable connector as disclosed herein; N1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8.

[0223] In one respect, the capped primer has the universal form m7GpppN 1 [N 2 ] m [N 3 ] n Where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; and m is 0 or 1; and n is any integer from 0 to 10.

[0224] In one embodiment, the capping primer can be a dinucleotide, for example, containing a nucleotide linked to the first nucleoside (N-nucleotide) via a triphosphate bridge. 1 7-methylguanosine at the 5' end of ) 7 G), producing m 7 G(5')ppp(5')N 1 The dinucleotide cap, in which N 1 It can be any nucleoside. Commercially available examples of dinucleotide capping primers are shown in Table 3.

[0225] Table 3: Examples of synthesized dinucleotide capping primers

[0226]

[0227] In one embodiment, m = 1 and n = 0.

[0228] In one embodiment, the capping primer can be a trinucleotide, for example, containing a nucleotide linked to the first nucleoside (N-nucleotide) via a triphosphate bridge. 1 7-methylguanosine at the 5' end of ) 7 G), which is linked to the second nucleotide (N). 2 ), producing m 7 G(5')ppp(5')N 1 N 2 The trinucleotide cap, in which N 1 and N 2 It can be any nucleoside independently. Commercially available examples of trinucleotide capping primers are shown in Table 4.

[0229] Table 4: Examples of synthesized trinucleotide capping primers

[0230]

[0231] in:

[0232] m or Me indicates methylation.

[0233] G-indicator guanosine,

[0234] p indicates a phosphate group.

[0235] N indicates any nucleoside, and

[0236] A indicates adenosine.

[0237] CleanCap reagent is commercially available from TriLink Biotechnologies (San Diego, USA).

[0238] In one embodiment, the capping primer is selected from m7GpppAmG, m7G3'OmpppAmG, m7GpppAmU, and m7G3'Ompppm6AmG.

[0239] Other synthetic trinucleotide capping primers applicable to the methods of this disclosure are selected from the group consisting of: m7GpppApA, m7GpppApC, m7GpppApG, m7GpppApU, m7GpppCpA, m7GpppCpC, m7GpppCpG, m7GpppCpU, m7GpppGpA, m7GpppGpC, m7GpppGpG, m7GpppGpU, m7GpppUpA, m7GpppUpC, m7GpppUpG, m7GpppUpU, m7G3'OmepppApA, m7G3'OmepppApC, m7G3'OmepppApG, m7G3'OmepppApG, m7G3'OmepppApA ... pppApU, m7G3'OmepppCpA, m7G3'OmepppCpC, m7G3'OmepppCpG, m7G3'OmepppCpU, m7G3'OmepppGpA, m7G3'OmepppGpC, m7G3'OmepppGpG, m7G3'OmepppG pU, m7G3'OmepppUpA, m7G3'OmeppUpC, m7G3'OmepppUpG, m7G3'OmepppUpU, m7G3'OmepppA2'OmepA, m7G3'OmepppA2'OmepC, m7G3'OmepppA2'OmepG, m7 G3'OmepppA2'OmepU, m7G3'OmepppC2'OmepA, m7G3'OmepppC2'OmepC, m7G3'OmepppC2'OmepG, m7G3'OmepppC2'OmepU, m7G3'OmepppG2'OmepA, m7G3'O mepppG2'OmepC, m7G3'OmepppG2'OmepG, m7G3'OmepppG2'OmepU, m7G3'OmepppU2'OmepA, m7G3'OmepppU2'OmepC, m7G3'OmepppU2'OmepG, m7G3'Omepp U2'OmepU, m7GpppA2'OmepA, m7GpppA2'OmepC, m7GpppA2'OmepG, m7GpppA2'OmepU, m7GpppC2'OmepA, m7GpppC2'OmepC, m7GpppC2'OmepG, m7GpppC2'O mepU, m7GpppG2'OmepA, m7GpppG2'OmepC, m7GpppG2'OmepG, m7GpppG2'OmepU, m7GpppU2'OmepA, m7GpppU2'OmepC, m7GpppU2'OmepG and m7GpppU2'OmepU.

[0240] In one embodiment, the capped primers of this disclosure comprise structures selected from Cap0, Cap1, Cap2, CapM6, or TMG Cap structures.

[0241] In one embodiment, the capping primer comprises a structure selected from m7GpppG, m7GpppN, m7G3'OmpppG, m7G3'OmpppA, modified ARCA, and β-S-ARCA. In another embodiment, the capping primer comprises a structure selected from m7GpppNN, m7G3'OpppNN, m7GpppNmN, m7G3'OpppNmN, m7GpppAN, m7G3'OpppAN, m7G3'OmpppAmN, m7Gpppm6AmN, m7G3'Opppm6AmN, m7GpppAmN, m7G3'OmpppAmN, m7GpppAG, m7G3' The structures of OpppAG, m7G3'OmpppAmG, m7Gpppm6AmG, m7G3'Opppm6AmG, m7GpppAmG, m7G3'OmpppAmG, m7GpppAU, m7G3'OpppAU, m7G3'OmpppAmU, m7Gpppm6AmU, m7G3'Opppm6AmU, m7GpppAmU, and m7G3'OmpppAmU. In one embodiment, the capping primer comprises a structure selected from the preceding claims, wherein the capping primer is selected from the group consisting of: 5'm7GpppAN, m7Gpppm2AN, m7Gpppm2Am2N, m7Gpppm6AN, and m7Gpppm6Am2N, m7G(5')ppp(5')G, m7(3')-O-Me-G(5')ppp(5')G), m7G(5')ppp(5')(2'OMeA)pG and m7G(5')ppp(5')(2'OMeA)pU), m7(3'OMeG)(5')ppp('5)m6(2'OMeA)pG), m7Gpppm6A mN, m7GpppN mN and 7mG3OMe5'ppp5'N.

[0242] In one embodiment, N is any nucleoside. In one embodiment, the capping primer may further contain 1 to 8 additional nucleosides.

[0243] In one embodiment, the capping primer may contain a natural RNA or DNA nucleoside, or a modified nucleoside analog, a natural or immobilized phosphodiester bond, or one or more modified sugars.

[0244] In some embodiments, the capping primer comprises at least one modified nucleoside. In one embodiment, the modified nucleoside comprises a 2'-O-methyl modification (2'Ome, 2'Om). In one embodiment, the modified nucleoside comprises a methyl group at position 6 (m6A).

[0245] In one embodiment, N 1 Nucleosides containing a 2'-O-methyl modification can provide a Cap1 structure when incorporated during transcription initiation. In one embodiment, N 1 The nucleoside contains a methyl group at position 6 (m6A). In one embodiment, N 1 The nucleoside comprises adenosine (m6(2'OMeA)) having a 2'-O-methyl modification and a methyl group at position 6. In one embodiment, N 1 The nucleoside is adenosine (m6A) with a methyl group at position 6. Compared to CleanCap AG or CleanCap AG (3'OMe), m6A modification can further increase protein expression. It is speculated that m6A modification adjacent to the 7-methylguanosine cap can positively influence mRNA stability by preventing enzyme-mediated uncapping.

[0246] In one embodiment, N 2 The nucleoside contains a 2'-O-methyl modification. In one embodiment, N 2 The nucleoside comprises adenosine (m6A) having a methyl group at position 6. In one embodiment, N 1 Nucleosides contain 2'-O-methyl modification and N 2 Nucleosides contain 2'-O-methyl modifications.

[0247] In one embodiment, N 1 The nucleoside contains a methyl group (m6A) at position 6 and N 2 The nucleoside contains a methyl group (m6A) at position 6.

[0248] In one embodiment, 7-methylguanosine is m7-methyl-3'-O-methyl-guanosine (7mG3'Ome).

[0249] In one embodiment, the capping primer comprises an ARCA analog carrying a modified 7MeG residue, wherein the 3' and / or 2' positions on the ribose are blocked to facilitate the directed initiation of in vitro transcription.

[0250] Using dinucleotides (m7GpppN), trinucleotides (m7GpppNN), tetranucleotides (m7GpppNNN), etc., capping primers can increase the frequency of capping primer incorporation into RNA molecules. In one embodiment, the dinucleotide, trinucleotide, or tetranucleotide capping primer of this disclosure may contain 1 to 8 additional nucleosides. The capping primer may be, for example, 2 to 10 nucleotides long. In one embodiment, the capping primer has a form selected from m7GpppN, m7GpppNN, m7GpppNNN, m7GpppNNNNNN, m7GpppNNNNNN, m7GpppNNNNNNNN, m7GpppNNNNNNNN, m7GpppNNNNNNNN, m7GpppNNNNNNNN, m7GpppNNNNNNNN, m7GpppNNNNNNNNNN, etc., wherein each instance of N is independently any natural, modified, or non-natural nucleoside. The m7G cap may also be modified as described herein.

[0251] In one embodiment, the capped primer contains nucleotides corresponding to the DNA template at the transcription start site, which may overlap with the promoter sequence. In one embodiment, N 1 This corresponds to the first transcription start site in the DNA template. In one embodiment, N 2 or N 3 At least one instance corresponds to a second transcription start site in the DNA template. In one embodiment, N 1 It is any natural, modified, or non-natural nucleoside, in which N 1 Not a preferred starting nucleotide for RNA polymerase. In one embodiment, N 2 or N 3 At least one instance corresponds to a preferred starting nucleotide of RNA polymerase. In one embodiment, N 1 Not with N 2 or N 3 At least one instance of the same nucleoside.

[0252] In some embodiments, the capping primer contains nucleotides corresponding to the DNA template sequence at the 5' transcription site of the transcribed RNA (i.e., the target RNA). In one embodiment, the capping primer has a 3' terminal OH group, which is an effective substrate for RNA polymerase (RNAP) and enables the transcribed RNA molecule to elongate.

[0253] DNA template

[0254] The DNA template can be double-stranded linear DNA, partially double-stranded linear DNA, circular double-stranded DNA, DNA plasmid, PCR amplicons, or modified nucleic acid templates compatible with RNA polymerase. The DNA template disclosed herein is typically a double-stranded DNA molecule that contains at least a promoter sequence and a sequence capable of transcribing an RNA molecule of interest (e.g., an mRNA molecule; see [link to relevant documentation]). Figure 1 ).

[0255] Those skilled in the art will understand that during transcription, RNA polymerase binds to a specific promoter region of a DNA molecule and uses the 3' strand (i.e., the template strand) of the double-stranded DNA molecule as a template for transcription from the transcription start site to transcribe the RNA molecule. Nucleosides are typically added to the RNA strand using the rules of complementary base pairing. The promoter sequence of the DNA template contains or is immediately followed by one or more transcription start sites (TSS), from which transcription can be initiated (see [link to relevant documentation]). Figure 1 and Figure 3-5 The transcription start site of a DNA template can also be called the start nucleotide and is the nucleotide from which transcription begins. For example... Figure 1 and 3 As shown in Figure 5, the corresponding nucleotide in the RNA molecule is the 5' terminal nucleotide, which can also be called the +1 transcription nucleotide. The next transcription nucleotide in the RNA molecule is called the +2 transcription nucleotide, and so on. The +1 transcription nucleotide corresponds to the +1 nucleotide in the DNA template, and the +2 transcription nucleotide corresponds to the +2 nucleotide in the DNA template (see Figure 5). Figure 1 and 3 -5).

[0256] In one embodiment, the nucleoside at the +1 position of the DNA template comprises adenosine, and N 1 Contains adenosine. In one embodiment, the nucleoside at the +1 position of the DNA template comprises cytidine, and N 1 Contains cytidine. In one embodiment, the nucleoside at the +1 position of the DNA template comprises thymidine, and N 1 Urate. In one embodiment, the nucleoside at the +1 position of the DNA template is adenosine and N... 1 The sample is selected from the group consisting of adenosine, N2-methyladenosine, or N6-methyladenosine. In one embodiment, the nucleotide at the +2 position of the DNA template is guanosine and N... 2 It is guanosine. In one embodiment, the nucleoside at the +2 position of the DNA template comprises thymidine and N 2 Contains guanosine. In one embodiment, the nucleoside at the +2 position of the DNA template comprises thymidine, and N 2 Urate. In one embodiment, the nucleoside at the +2 position of the DNA template comprises cytidine, and N 2 It contains cytidine.

[0257] The promoter sequence of the DNA template can be modified depending on the RNA polymerase to be used and / or the capping primer to be used. The sequence of the DNA template at the transcription start site can be modified to improve the effectiveness of initiating transcription using a capped primer that produces capped RNA molecules, rather than using a nucleoside triphosphate (such as guanosine triphosphate) that produces uncapped RNA molecules. Specifically, the DNA template sequence can be varied at the transcription start site to optimize hybridization between the capped primer and the nucleotides of the 3' strand (i.e., the template strand) of the DNA template at the transcription start site, thereby enabling transcription to be initiated using the capped primer.

[0258] Therefore, in one embodiment, the sequence of the DNA template may be modified at the transcription start site to correspond to the nucleotides in the capping primer (see...). Figure 1 and 3 -5). For example, when the capping primer has the form m7GpppAG, the 5' to 3' sequence of the 5' strand of the DNA template transcription start site can be modified to AG to optimize hybridization of the capping primer with the DNA template during transcription, thereby promoting the incorporation of the capping primer into the generated RNA molecule.

[0259] In another example, when the capped primer has the form m7GpppAU, the 5' to 3' sequence of the 5' strand of the transcription start site of the DNA template can be modified to AU to optimize hybridization of the capped primer with the DNA template during transcription, thereby promoting the incorporation of the capped primer into the resulting RNA molecule. Those skilled in the art will understand that the transcription start site can be modified to correspond to a nucleoside present in other capped primers.

[0260] In one embodiment, the DNA template may contain a natural or modified promoter sequence. Examples of natural and modified promoter sequences for T7, T3, and SP6 RNA polymerases and self-amplified RNA are shown in Table 6. The natural T7 promoter sequence is TAATACGACTCACTATAGG (SEQ ID NO: 13). In one embodiment, the modified T7 (HG) promoter sequence is TAATACGACTCACTATA. HG (SEQ ID NO: 14), wherein H = A, C, or T (A is most commonly used), and preferably the starting nucleoside is G. The natural T3 promoter sequence is AATTAACCCTCACTAAAG (SEQ ID NO: 16). In one embodiment, the modified T3 promoter sequence is TAATACGACTCACTAAA. H(SEQ ID NO: 17), wherein H = A, C, or T, and preferably the starting nucleoside is G. The natural SP6 promoter sequence is ATTTAGGTGACACTATAGA (SEQ ID NO: 18). In one embodiment, the modified SP6 promoter sequence is ATTTAGGTGACACTATA. H (SEQ ID NO: 20), wherein H = A, C, or T, and preferably the starting nucleotide is G. The natural self-amplifying RNA promoter is TAATACGACTCACTATAGG (SEQ ID NO: 21). In one embodiment, the modified self-amplifying RNA promoter sequence is TAATACGACTCACTATA. HT (SEQ ID NO:22), wherein H = A, C, or T, and preferably the starting nucleotide is G. In one embodiment, the (+) strand genome of the self-amplified virus begins at 5'-AU. Underlined indicates an alternative transcription start site. Bold indicates when combined with the form m7GpppN 1 [N 2 ] m [N 3 During co-transcription with capped primers, the transcription start site of the RNA molecule, where N... 1 It is X.

[0261] In one embodiment, the DNA template comprises a sequence selected from the group consisting of TAATACGACTCACTATA (SEQ ID NO: 23), TAATACGACTCACTATAAG (SEQ ID NO: 15), and TAATACGACTCACTATAAT (SEQ ID NO: 24), wherein the RNA molecule has been transcribed by T7 RNA polymerase; a sequence selected from AATTAACCCTCACTAAA (SEQ ID NO: 25), AATTAACCCTCACTAAAG (SEQ ID NO: 26), and AATTAACCCTCACTATA (SEQ ID NO: 27), wherein the RNA molecule has been transcribed by T3 RNA polymerase; or the sequence ATTTAGGTGACACTATA (SEQ ID NO: 28), wherein the RNA molecule has been transcribed by Sp6 RNA polymerase. In one embodiment, the preferred starting nucleotide for T7 RNA polymerase is guanosine.

[0262] Those skilled in the art will understand that the methods disclosed herein can be used with any suitable RNA polymerase by incorporating an appropriate promoter sequence at a suitable location within the DNA template.

[0263] In one embodiment, the nucleotides at the transcription start sites of the promoter sequences of at least T7, T3, and SP6 RNA polymerases may be modified to correspond to the sequences of the capping primers to optimize the incorporation of the capping primers into the RNA molecule. Similarly, the nucleotides at the transcription start sites of the promoter sequences of other polymerases or self-amplifying RNAs may be modified to correspond to the nucleosides present in a particular capping primer, which can enhance the incorporation of the capping primers into the RNA molecule.

[0264] In one embodiment, the DNA template contains a promoter having the sequence [N]yTATA, where N is any nucleoside and y is any suitable integer.

[0265] Those skilled in the art will understand that, by optimizing the DNA template sequence and transcription reaction, the methods disclosed herein can be applied to a range of capped primers, polymerases, and target RNA molecules.

[0266] In vitro transcription

[0267] The synthesized RNA, including messenger RNA, transfer RNA, small nucleolar RNA (snoRNA), guide RNA, etc., can be generated by in vitro transcription of a DNA template by an RNA polymerase in a suitable reaction solution containing nucleoside-5'-triphosphate (NTP). Phage RNA polymerases, such as T3 polymerase, T7 polymerase, SP6 polymerase, and other polymerases, are typically used to drive transcription from a sequence-specific promoter upstream of the template mRNA sequence of interest in the DNA template. For example, a plasmid DNA template can be synthesized containing an RNA polymerase promoter sequence (e.g., a T7 promoter sequence), followed by a sequence encoding the target RNA sequence of interest, and optionally a restriction enzyme site. In one embodiment, the DNA template can then be amplified using in vitro or bacterial amplification methods and linearized by restriction enzyme digestion prior to purification. The DNA template can be transcribed in vitro using an RNA polymerase to synthesize a target RNA molecule. In one embodiment, suitable polymerases include those derived from T7, T3, SP6, K1-5, K1E, K1F, or K11 phages. In one embodiment, the RNA polymerase is selected from T7 RNA polymerase, T3 RNA polymerase, and SP6 RNA polymerase. In one embodiment, the RNA molecule has been transcribed by T7 RNA polymerase. In one embodiment, the RNA molecule is an mRNA molecule.

[0268] In one embodiment, RNA molecules are post-transcribed and capped using a capping enzyme (such as vaccinia virus capping enzyme) after in vitro transcription, which produces Cap0 or Cap1 RNA molecules.

[0269] Synthetic RNA molecules found in mRNA vaccines and therapies can be transcribed in vitro using capping primers as described herein to initiate transcription and co-transcribe a 5' cap analogue into the RNA molecule. In one embodiment, the RNA molecule is co-transcribed and capped, for example, by using an excess of a capping primer (e.g., having the universal form m7GpppN as described herein) during the transcription reaction. 1 [N 2 ] m [N 3 ] n In this embodiment, a capped primer can initiate transcription. While not wishing to be bound by theory, it should be understood in one embodiment that, in the absence of a capped primer, T7 RNA polymerase has a strong preference for initiating transcription using guanosine triphosphate (GTP), resulting in the incorporation of guanosine as a 5' terminal nucleotide, producing uncapped RNA molecules. However, in the presence of a universal form m7GpppN... 1 [N 2 ] m [N 3 ] n In the case of capped primers, T7 polymerase may preferentially initiate transcription with capped primers, allowing transcription to be initiated with capping of various 5' sequences and RNA molecules.

[0270] In one embodiment, transcription begins at a different TSS when a capped primer is used, compared to the case without a capped primer. For example, when a capped primer containing the form m7GpppAG is incorporated and the TSS is modified to AG, T7 RNA polymerase initiates transcription at the A nucleotide of AG in the TSS region (see [link to documentation]). Figure 1 and 4 However, when transcription is initiated with a nucleotide triphosphate, T7 may preferentially initiate transcription with guanosine triphosphate. Therefore, transcription may be initiated from the next guanosine nucleotide downstream of the promoter region. In this example, transcription may be initiated from the G nucleotide of the AG in the TSS region (see [link to TSS]). Figure 1 and 5 ).

[0271] In one instance, when the capping primer is m7GmpppAmG (e.g., CleanCap reagent AG) and the polymerase is T7, the DNA template can be designed to include a promoter sequence of T7 polymerase with AG at the transcription start site to facilitate the binding of the capping primer to the transcription start site, i.e., TAATACGACTCACTATA. A Specific binding to G (SEQ ID NO: 15).

[0272] In this example, when transcription is initiated (i.e., co-transcription) using a capped primer, the resulting RNA molecule is a capped RNA molecule (i.e., capped with a capped primer). The 5' terminal nucleotide of the capped RNA molecule is adenosine, which is methylated. The 5' terminal nucleotide of the RNA molecule corresponds to the adenosine at the first transcription start site of the DNA template (i.e., +1 A in the AG TSS region). The RNA molecule is longer than the RNA molecule that begins at the second (i.e., downstream) transcription start site.

[0273] In the same example, when the capped primer does not initiate transcription, T7 preferentially uses guanosine triphosphate (GTP) to initiate transcription. The resulting RNA molecule is uncapped. The 5' terminal nucleotide of the uncapped RNA molecule is guanosine, which is not expected to be methylated. The 5' terminal nucleotide of the RNA molecule corresponds to the guanosine at the second transcription start site of the DNA template (i.e., +2 G in the AG TSS region in this example). The RNA molecule is shorter than the RNA molecule that begins at the first (i.e., upstream) transcription start site.

[0274] RNA sample

[0275] In one embodiment, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules. In one embodiment, the sample comprises a mixture of capped and uncapped RNA molecules representing a single population of RNA molecules having an RNA sequence of interest transcribed from a corresponding DNA template.

[0276] In one embodiment, the sample contains more than one population of RNA molecules, each transcribed from a different / distinct DNA template, such that each RNA molecule population has a different RNA sequence of interest. Those skilled in the art will understand that an RNA molecule sample containing a population of RNA molecules, after in vitro transcription from a DNA template, will contain a mixture of capped and uncapped RNA molecules with different RNA sequences of the relevant population.

[0277] In one embodiment, the sample comprises a mixture of capped and uncapped RNA molecules from two RNA molecule populations, each population having the RNA sequence of interest, wherein each population is transcribed in vitro from its corresponding DNA template. In one embodiment, the sample comprises a mixture of capped and uncapped RNA molecules from three RNA molecule populations, each population having the RNA sequence of interest, wherein each population is transcribed in vitro from its corresponding DNA template. In one embodiment, the sample comprises a mixture of capped and uncapped RNA molecules from four RNA molecule populations, each population having the RNA sequence of interest, wherein each population is transcribed in vitro from its corresponding DNA template. In one embodiment, the sample comprises a mixture of capped and uncapped RNA molecules from multiple (e.g., two, three, four, five, six, seven, eight, nine, ten, etc.) RNA molecule populations, each population having the RNA sequence of interest, wherein each population is transcribed in vitro from its corresponding DNA template.

[0278] Identification of capped and uncapped RNA molecules

[0279] In one aspect, this disclosure relates to a method for distinguishing capped RNA molecules from uncapped RNA molecules after in vitro transcription of DNA, the method being performed by identifying at least one difference between capped and uncapped RNA molecules associated with a differential transcription initiator (i.e., a capped primer or its nucleotide or analogue) used to initiate transcription, as well as the associated differential transcription initiation site.

[0280] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules;

[0281] This capped RNA molecule has been used in a universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n Capped primers; where N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes identifying capped RNA molecules and / or uncapped RNA molecules from a sample, wherein the identification includes steps selected from the group consisting of:

[0282] Identify N 1 The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N...1 The absence of the 5' terminal nucleotide identified an uncapped RNA molecule.

[0283] Identifying the 5' terminal nucleotide as the first transcription start site corresponding to the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleotide as the second transcription start site corresponding to the DNA template identifies uncapped RNA molecules.

[0284] Identifying the 5' terminal nucleoside as the nucleoside corresponding to the +1 position of the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template identifies uncapped RNA molecules.

[0285] Where N 1 N 2 and N 3 At least one of each instance independently identifies any modified or non-natural nucleoside, thus identifying a capped RNA molecule; and / or identifies the absence of a modified or non-natural nucleoside, thus identifying an uncapped RNA molecule; and

[0286] The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; RNA molecules with a length of at least x+1 nucleotides are identified, which identifies capped RNA molecules; and RNA molecules with a length of x nucleotides are identified, which identifies uncapped RNA molecules.

[0287] The identification step identifies capped RNA molecules and / or uncapped RNA molecules.

[0288] In one embodiment, the capped primer may have the universal form m7GpppN 1 [N 2 ] m [N 3 ] n Where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8.

[0289] Structural differences can be identified using any suitable method known to those skilled in the art capable of distinguishing or identifying structural differences. In one embodiment, the method is selected from the group consisting of: polymerase chain reaction (PCR) based methods (including reverse transcription PCR (RT-PCR), quantitative PCR (qPCR), digital PCR (dPCR), nested PCR, multiplex PCR, touchdown PCR, hot-start PCR, high-fidelity PCR); next-generation sequencing (NGS) methods (Illumina sequencing, IonTorrent sequencing, PacBio sequencing, Nanopore sequencing, Sanger sequencing, and RNA-seq sequencing); direct RNA sequencing; sequencing-by-synthesis; ligation or probe hybridization methods; RNA sequencing, RNA ligation, cDNA sequencing, isothermal amplification methods (such as LAMP), and cDNA synthesis or hybridization with specific oligonucleotides using primers switched with specific templates. Those skilled in the art will understand how to combine these techniques to design appropriate primers and / or probes and / or optimize reaction conditions.

[0290] In one aspect, the method further includes quantifying the identified capped RNA molecules and / or uncapped RNA molecules. In one embodiment, quantification includes determining the absolute amount of capped RNA molecules in the sample; determining the relative abundance of capped RNA molecules and / or uncapped RNA molecules in the sample; and / or determining the relative abundance of capped RNA molecules in the sample relative to total RNA molecules.

[0291] Identification of capped RNA molecules by identifying the 5' terminal nucleotide.

[0292] By identifying N 1 The presence or absence of 5'-terminal nucleotides is used to identify capped RNA molecules.

[0293] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been transcribed in vitro from a DNA template by RNA polymerase. 1 [N 2 ] m [N 3 ] n Capped primers; where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes identifying N 1The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N... 1 The absence of the 5'-terminal nucleotide identified an uncapped RNA molecule.

[0294] For example, in capped primers with the universal form m7GpppAN 2 In the case of N 1 The group consisting of free adenosine, N2-methyladenosine, or N6-methyladenosine, or adenosine with alternative modifications, can be selected, and identifying adenosine as the 5' terminal nucleotide of an RNA molecule will identify the capped RNA molecule. Therefore, in one embodiment, in N... 1 In cases involving modified nucleosides, the presence of base nucleosides (e.g., via sequencing technology) can identify N. 1 It exists as a 5'-terminal nucleoside.

[0295] N, as a 5' terminal nucleotide 1 The presence or absence of can be determined using methods known to those skilled in the art, such as those described elsewhere herein.

[0296] In one embodiment, identification includes identifying at least one to three 5' nucleotides of the RNA molecule. In one embodiment, identification includes identifying at least one to ten 5' nucleotides of the RNA molecule. In one embodiment, identification includes identifying at least one to 20 5' nucleotides of the RNA molecule. In one embodiment, identification includes identifying at least one to 50 5' nucleotides of the RNA molecule. In one embodiment, identification includes identifying at least one to fifty 5' nucleotides of the RNA molecule. In one embodiment, identification includes identifying at least one to 100 5' nucleotides of the RNA molecule. In one embodiment, the identification step includes identifying substantially all nucleotides in the RNA molecule that determine the RNA molecule.

[0297] Capped RNA molecules are identified by recognizing the +1 or +2 nucleotides corresponding to the DNA template as 5' terminal nucleotides.

[0298] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been transcribed in vitro from a DNA template by RNA polymerase. 1 [N 2 ] m [N 3 ] n Capped primers; where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes identifying a 5' end nucleoside as the nucleoside corresponding to the +1 position of the DNA template, which identifies capped RNA molecules; and / or identifying a 5' end nucleoside as the nucleoside corresponding to the +2 position of the DNA template, which identifies uncapped RNA molecules.

[0299] For example, in a DNA template containing the sequence TAATACGACTCACTATA A When the T7 promoter is G (SEQ ID NO: 15) and the capped primer has the universal form m7GmpppAG, the T7 polymerase is understood to preferentially initiate transcription using the capped primer. The 5' terminal nucleotide of the resulting capped RNA molecule is adenosine, and the nucleotide at the +1 position of the DNA template is adenosine, while the nucleotide at the +2 position of the DNA template is guanosine. When transcription is initiated using a nucleotide instead of the capped primer, the T7 polymerase is understood to preferentially initiate transcription using the guanosine at the G(+2) position (underlined) of the DNA template. In this embodiment, identifying the 5' terminal nucleotide as adenosine corresponding to the nucleotide at the +1 position of the DNA template identifies the capped RNA molecule; and / or identifying the 5' terminal nucleotide as guanosine corresponding to the nucleotide at the +2 position of the DNA template identifies the uncapped RNA molecule.

[0300] In one embodiment, the nucleoside at the +1 position of the DNA template comprises adenosine, and N 1 Contains adenosine. In one embodiment, the nucleoside at the +1 position of the DNA template comprises cytidine, and N 1 Contains cytidine. In one embodiment, the nucleoside at the +1 position of the DNA template comprises thymidine, and N 1 Urate. In one embodiment, the nucleoside at the +1 position of the DNA template is adenosine and N... 1 Select the group consisting of free adenosine, N2-methyladenosine, or N6-methyladenosine.

[0301] Capped RNA molecules are identified by identifying the nucleotide corresponding to the first or second transcription start site of the DNA template as the 5' terminal nucleotide.

[0302] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been transcribed in vitro from a DNA template by RNA polymerase. 1 [N 2 ] m [N 3 ] nCapped primers; where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes identifying a 5'-terminal nucleoside as a first transcription start site corresponding to a DNA template, which identifies capped RNA molecules; and / or identifying a 5'-terminal nucleoside as a second transcription start site corresponding to a DNA template, which identifies uncapped RNA molecules.

[0303] For example, in a DNA template containing the sequence TAATACGACTCACTATA AG When the T7 promoter (SEQ ID NO: 15) is used, and the capping primer has the universal form m7GmpppAG, the T7 polymerase is understood to preferentially initiate transcription using the capping primer at the first transcription start site (underlined, bold). The 5' terminal nucleotide of the resulting capped RNA molecule is adenosine, corresponding to the first transcription start site of the DNA template. When transcription is initiated using a nucleotide instead of a capping primer, the T7 polymerase is understood to preferentially initiate transcription using guanosine at the second transcription start site (underlined) of the DNA template. In this embodiment, identifying the 5' terminal nucleotide as adenosine corresponding to the first transcription start site of the DNA template identifies the capped RNA molecule; and / or identifying the 5' terminal nucleotide as guanosine corresponding to the nucleotide at the second transcription start site of the DNA template identifies the uncapped RNA molecule.

[0304] Capped RNA molecules were identified by length.

[0305] In one aspect, this disclosure provides a method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been transcribed in vitro from a DNA template by RNA polymerase. 1 [N 2 ] m [N 3 ] n Capped primers; where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3Each instance is independently any natural, modified, or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes: wherein the uncapped RNA molecule is x nucleotides long, and x is any suitable integer; identifying RNA molecules at least x+1 nucleotides long, which identifies capped RNA molecules; and identifying RNA molecules x nucleotides long, which identifies uncapped RNA molecules.

[0306] For example, in a DNA template containing the sequence TAATACGACTCACTATA AG When using the T7 promoter of (SEQ ID NO:15) and the capped primer having the universal form m7GmpppAG, the T7 polymerase is understood to preferentially initiate transcription using the capped primer at the first transcription start site (underlined, bold). When transcription is initiated using a nucleotide instead of a capped primer, the T7 polymerase is understood to preferentially initiate transcription using guanosine at the second transcription start site (underlined) of the DNA template. In this embodiment, if the length of the uncapped RNA molecule is 1000 nucleotides, the length of the capped RNA molecule will be 1001 nucleotides because transcription starts one nucleotide upstream compared to the uncapped molecule. In this embodiment, identifying an RNA molecule 1001 nucleotides long is identifying a capped RNA molecule; and identifying an RNA molecule 1000 nucleotides long is identifying an uncapped RNA molecule. Those skilled in the art will understand that this embodiment can be applied to RNA molecules of different lengths. Additionally, in one embodiment, it can be applied to discontinuous transcription start sites. For example, if the second transcription start site is two nucleotides downstream of the first transcription start site, the capped primer will be two nucleotides longer than the uncapped RNA molecule, and so on.

[0307] Capped RNA molecules were identified through nucleotide modification.

[0308] In one form, this disclosure relates to a method for identifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules; wherein the capped RNA molecules have been used with a universal form m7GpppN 1 [N 2 ] m [N 3 ] n Capped primers; where m7G is N7-methylated guanosine or a guanosine analogue; ppp is triphosphate; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside, wherein the N present in the capped primer is... 1 N2 and N 3 At least one of each instance includes any modified or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; the method includes: identifying the presence of a modified or non-natural nucleoside to identify a capped RNA molecule; and identifying the absence of a modified or non-natural nucleoside to identify an uncapped RNA molecule. In one embodiment, the modified or non-natural nucleoside can be detected by the method of this disclosure.

[0309] For example, when the capped primer contains one or more modified nucleosides (e.g., methylated adenosine such as that found in 7GmpppAmG), the capped RNA molecule produced according to this disclosure will contain the modified nucleosides, while the uncapped RNA molecule will contain the unmodified nucleosides from the transcription reaction. In this embodiment, the presence of a modified or non-natural nucleoside as a 5' terminal nucleoside is used to identify the capped RNA molecule; while the absence of a modified or non-natural nucleoside as a 5' terminal nucleoside is used to identify the uncapped RNA molecule. Those skilled in the art will understand that this embodiment of the present disclosure can identify modified nucleosides at different positions within the capped primer.

[0310] In one embodiment, the presence of a modified or non-natural nucleoside at the 5' end of an RNA molecule identifies a capped RNA molecule; and the absence of a modified or non-natural nucleoside at the 5' end of an RNA molecule identifies an uncapped RNA molecule. In one embodiment, the presence of a modified or non-natural nucleoside as the 5' terminal nucleoside of an RNA molecule identifies a capped RNA molecule; and the absence of a modified or non-natural nucleoside as the 5' terminal nucleoside of an RNA molecule identifies an uncapped RNA molecule. In one embodiment, N 1 Capped RNA molecules are identified by identifying the presence of modified or non-natural nucleosides as 5'-terminal nucleosides; uncapped RNA molecules are identified by identifying the absence of modified or non-natural nucleosides as 5'-terminal nucleosides.

[0311] In one embodiment, the modified or non-natural nucleoside is as described elsewhere herein. In one embodiment, the modified or non-natural nucleoside forms part of a 5' cap structure as described elsewhere herein.

[0312] In one embodiment, the modified or non-natural nucleoside contains a methyl group at position 6 (m6A). In another embodiment, the modified or non-natural nucleoside contains adenosine (N6-methyladenosine, m6A) having a methyl group at position 6.

[0313] In one embodiment, the modified or non-natural nucleoside includes N2-methyladenosine.

[0314] In one embodiment, the modified or non-natural nucleoside contains a 2'-O-methyl modification (2'Ome, 2'Om).

[0315] In one embodiment, the modified or non-natural nucleoside comprises m7-methyl-3'-O-methyl-guanosine (7mG3'Ome). In another embodiment, the modified or non-natural nucleoside comprises a modified 7MeG residue in which the 3' and / or 2' positions on the ribose are blocked to facilitate directed initiation of in vitro transcription.

[0316] In one embodiment, the modified or non-natural nucleoside is selected from N2-methyladenosine or N6-methyladenosine.

[0317] In one embodiment, the method can detect multiple modified or non-natural nucleosides to identify the presence of capped RNA molecules. In one embodiment, multiple 2'-O-methyl modifications can be detected. In one embodiment, multiple m6A modifications can be detected.

[0318] Methods for detecting capped and / or uncapped RNA molecules

[0319] In one aspect of this disclosure, RNA molecules can be identified as capped or uncapped RNA molecules using methods known to those skilled in the art. In one embodiment, identification includes methods selected from the group consisting of: polymerase chain reaction (PCR) based methods (including reverse transcription PCR (RT-PCR), quantitative PCR (qPCR), digital PCR (dPCR), nested PCR, multiplex PCR, touchdown PCR, hot-start PCR, high-fidelity PCR); next-generation sequencing methods (Illumina sequencing, Ion Torrent sequencing, PacBio sequencing, Nanopore sequencing, Sanger sequencing, and RNA-seq sequencing); direct RNA sequencing; sequencing-by-synthesis; ligation or probe hybridization methods; RNA sequencing, RNA ligation, cDNA sequencing, isothermal amplification methods (such as LAMP), and cDNA synthesis or hybridization with specific oligonucleotides using primers with a specific template.

[0320] Transformed into complementary DNA molecules

[0321] In some embodiments, reverse transcriptases (such as MMLV reverse transcriptase, Omniscript reverse transcriptase (Qiagen), EnzScript reverse transcriptase (Qiagen), StableScript reverse transcriptase (Qiagen), SuperScript III reverse transcriptase (Invitrogen), SuperScript IV reverse transcriptase (Invitrogen), ProtoScript II reverse transcriptase (NEB), Induro reverse transcriptase (NEB), WarmStart reverse transcriptase (NEB), AMV reverse transcriptase (NEB)) can be used to convert RNA molecules into complementary DNA (cDNA) molecules. Therefore, in embodiments of the method disclosed herein, the method further includes using a reverse transcriptase to convert RNA molecules into cDNA molecules. In embodiments of the method disclosed herein, the identification step further includes reverse transcribing RNA molecules from a sample into complementary DNA (cDNA) and analyzing the cDNA molecules to identify capped or uncapped RNA molecules.

[0322] Transformation of cDNA molecules can be used for further downstream analysis of RNA molecules using capillary electrophoresis, PCR, or NGS-based methods.

[0323] In one embodiment, the identification step further includes reverse transcribing RNA molecules from a sample into complementary DNA (cDNA) and analyzing the cDNA molecules to determine the 5' terminal nucleotide of the RNA molecule, thereby identifying capped or uncapped RNA molecules according to the method of this disclosure. Those skilled in the art will understand that the 5' terminal nucleotide of an RNA molecule can be determined by analyzing the sequence of the cDNA molecule. In one embodiment, the identification step further includes reverse transcribing RNA molecules from a sample into complementary DNA (cDNA) molecules and identifying at least the nucleotide of the cDNA molecule corresponding to the 5' terminal nucleotide of the RNA molecule, wherein identifying at least the nucleotide of the cDNA molecule corresponding to the 5' terminal nucleotide of the RNA molecule at least identifies the 5' terminal nucleotide of the RNA molecule. In one embodiment, the identification includes identifying at least one terminal nucleotide of the cDNA molecule corresponding to the 5' terminal nucleotide of the RNA molecule, wherein identifying at least the terminal nucleotide of the cDNA molecule at least identifies the 5' terminal nucleotide of the RNA molecule.

[0324] In one embodiment, the identification step further includes reverse transcribing RNA molecules from the sample into complementary DNA (cDNA) and analyzing the cDNA molecules to determine the length of the RNA molecules, thereby identifying capped or uncapped RNA molecules according to the methods of this disclosure. Those skilled in the art will understand that the length of the RNA molecule can be determined by analyzing the sequence of the cDNA molecule.

[0325] Next-generation sequencing

[0326] In some embodiments, capped or uncapped RNA molecules can be identified using next-generation sequencing (NGS) methods. In one embodiment, the output RNA molecule is reverse transcribed into cDNA, which can then be used for library preparation and sequencing according to the manufacturer's instructions for the specific NGS platform used. These methods include, but are not limited to, Illumina sequencing, Ion Torrent sequencing, PacBio sequencing, Nanopore sequencing, Sanger sequencing, direct RNA sequencing, RNA sequencing, and cDNA sequencing.

[0327] In some embodiments, direct RNA nanopore sequencing, such as that offered by Oxford Nanopore Technologies, can be used to determine the presence of a cap. This method allows for direct analysis of RNA molecules (capped or uncapped) without the need for reverse transcription to cDNA.

[0328] Next-generation sequencing (NGS) technologies, such as those from Illumina, Pacific Biosciences (PacBio), and Oxford Nanopore Technologies (ONT), can be used to determine the sequence of RNA molecules. This facilitates the identification of nucleotides at the 5' end of RNA molecules. After sequencing heterogeneous samples of capped and uncapped RNA molecules, the read dataset can be bio-aligned with a reference DNA template sequence. TSSs can be identified by analyzing the nucleotides at the 5' end of the sequenced reads aligned to the RNA molecules (e.g., +1, +2, +3, +4, +5, etc.). This allows for the identification of whether capped primers are incorporated into the RNA molecule, and consequently, whether the RNA molecule is capped or uncapped, according to the methods described herein.

[0329] In one embodiment, NGS can detect the presence of nucleotide modifications at the 5' end of an RNA molecule (e.g., proximal cap positions, including +1, +2, and +3 nucleotides, etc.). Detection of nucleotide modifications can be performed using various methods, including direct RNA sequencing, cDNA sequencing, and mass spectrometry. For example, direct RNA sequencing can be used to detect the m7GTP cap at the 5' end of an mRNA molecule. Direct RNA sequencing can also be used to detect N6-methylation of adenosine at the proximal cap nucleotide (i.e., the 5' terminal nucleoside), thereby indicating the incorporation and presence of a capping primer (such as m7Gpppm6A). In one embodiment, direct RNA sequencing can identify the presence of N6-methyladenosine at transcription position 1, the presence of a 2'O-methyl nucleoside residue (Cap1) at transcription position 1, the presence of a 2'O-methyl nucleoside residue (Cap2) at transcription position 2, the presence of a 6'O-methyl nucleoside residue (Cap1) at transcription position 1, and / or the presence of a 6'O-methyl nucleoside residue (Cap2) at transcription position 2.

[0330] In one embodiment, NGS analysis can be used to detect the length of RNA molecules, which can identify whether the RNA molecules are capped or uncapped, as disclosed herein.

[0331] Quantitative Polymerase Chain Reaction

[0332] In one embodiment, a polymerase chain reaction (PCR)-based method can be used to analyze the 5' terminal nucleotides and capping state of RNA molecules. The PCR-based method can be selected from the group consisting of: reverse transcription polymerase chain reaction (RT-PCR), real-time quantitative RT-PCR (qRT-PCR), TaqMan qPCR, multiplex RT-PCR, digital RT-PCR (dRT-PCR), nested RT-PCR, RT loop-mediated isothermal amplification (RT-LAMP), RT-droplet digital PCR (RT-ddPCR), RT-PCR combined with high-resolution melting (HRM) analysis, etc.

[0333] These methods can identify the 5' nucleotide of RNA molecules (e.g., in synthetic mRNA vaccines), thereby identifying N... 1 The presence or absence of a 5' terminal nucleotide determines whether the 5' terminal nucleotide corresponds to a first or second transcription start site, and / or whether the 5' terminal nucleotide corresponds to a nucleotide at position +1 or position +2 on the DNA template and / or another nucleotide at another position on the DNA template. In one embodiment, this further determines whether the RNA molecule is capped or uncapped. This method is generally applicable to any PCR or probe hybridization technique capable of adequately measuring differences in terminal nucleotide or mRNA molecule length.

[0334] In some embodiments, polymerase chain reaction (PCR)-based methods, including reverse transcription PCR (RT-PCR), quantitative PCR (qPCR), digital PCR (dPCR), nested PCR, multiplex PCR, touchdown PCR, hot-start PCR, and high-fidelity PCR, can be used to measure the relative abundance of RNA molecules containing the presence or absence of adaptor oligonucleotides. PCR requires short oligonucleotide primers that anneal and amplify the sequence of interest.

[0335] In one embodiment, the 5' terminal nucleotide and capping state can be determined by Taqman probe assay. In one embodiment, the Taqman assay includes using probes and / or primers selected from the group consisting of: a Taqman forward primer according to SEQ ID NO: 1, a Taqman probe for identifying capped RNA molecules according to SEQ ID NO: 2, a Taqman reverse primer according to SEQ ID NO: 3, and a Taqman probe for identifying uncapped RNA molecules according to SEQ ID NO: 4. Those skilled in the art will understand that other primers and probes may be suitable for use in the Taqman probe assay.

[0336] In one embodiment, the 5' terminal nucleotide and capping state can be determined by the SYBR Green Assay. In one embodiment, the SYBR Green Assay includes using primers selected from the group consisting of: a SYBR forward primer for detecting capped molecules according to SED ID NO: 5, a forward primer for detecting uncapped molecules according to SED ID NO: 6, a first reverse primer according to SEQ ID NO: 7, a second reverse primer according to SEQ ID NO: 8, a third reverse primer according to SEQ ID NO: 9, and a fourth reverse primer according to SEQ ID NO: 10. Those skilled in the art will understand that other primers and probes may be suitable for use in the SYBR Green Assay.

[0337] Quantitative analysis of capped and / or uncapped RNA molecules

[0338] In one aspect, this disclosure provides methods for quantifying capped or uncapped RNA molecules as described herein.

[0339] In one embodiment, quantification includes determining the absolute amount of capped RNA molecules in the sample. In one embodiment, quantification includes determining the relative abundance of capped RNA molecules and / or uncapped RNA molecules in the sample. In one embodiment, quantification includes determining the relative abundance of capped RNA molecules in the sample relative to total RNA molecules.

[0340] In one embodiment, quantification includes comparing the number of RNA molecules identified as capped RNA molecules with the number of RNA molecules identified as uncapped RNA molecules to determine the relative abundance of capped RNA molecules in the sample.

[0341] In one embodiment, the method for quantifying the identified capped and / or uncapped RNA molecules utilizes suitable methods for quantifying the identified RNA molecules as known to those skilled in the art. In one embodiment, the method for quantifying the identified RNA molecules includes detection methods as described herein. In one embodiment, quantification includes methods selected from the group consisting of: polymerase chain reaction (PCR) based methods (including reverse transcription PCR (RT-PCR), quantitative PCR (qPCR), digital PCR (dPCR), nested PCR, multiplex PCR, touchdown PCR, hot-start PCR, high-fidelity PCR); next-generation sequencing methods (Illumina sequencing, Ion Torrent sequencing, PacBio sequencing, Nanopore sequencing, Sanger sequencing, and RNA-seq sequencing); direct RNA sequencing; sequencing-by-synthesis; oligonucleotide probe hybridization methods; RNA sequencing, RNA ligation, cDNA sequencing, isothermal amplification methods (such as LAMP), restriction enzyme digestion, and cDNA synthesis or hybridization with specific oligonucleotides using primers with a specific template. In one embodiment, the method for quantifying the identified RNA molecules includes next-generation sequencing methods, TaqMan probe assays, SYBR Green assays, and / or qRT-PCR assays.

[0342] Bioinformatics Analysis

[0343] Analysis of NGS datasets can determine the amount of capped and uncapped mRNA molecules in a sample. The proportion of capped and uncapped RNA molecules in an RNA sample can be quantified by measuring the count of identified capped or uncapped nucleotides. The first step may involve analyzing the output dataset, for example in FASTQ format, indicating the sequenced reads. These reads can be aligned to a reference sequence using software such as minimap2, bwa Bowtie2, HISAT2, or STAR. This reference sequence may include plasmid DNA, a DNA template sequence, an RNA molecule sequence, and / or an adaptor oligonucleotide sequence. The comparison between the aligned reads and the reference sequence can indicate differences caused by mutations and errors. Alternatively, k-mer counting methods (such as Kallisto or Salmon) can be used to count the presence of k-mer sequences (short DNA or RNA sequences of length k) corresponding to capped or uncapped mRNA within the sequencing data. For example, measuring the count of k-mer sequences starting at +1 or +2 nucleotides from the 5' end of an RNA molecule can indicate the count of capped or uncapped mRNA molecules, respectively. The count of k-mer sequences that match or complement the identified RNA molecules can indicate the abundance of capped and / or uncapped RNA molecules.

[0344] Measuring other mRNA quality characteristics

[0345] In one embodiment, the analysis of mRNA capping status using the method described above can be performed simultaneously to measure other mRNA quality characteristics using a single assay. For example, the analysis of mRNA capping status can be performed simultaneously in addition to one or more other analyses, including mRNA integrity, sequence identity, polyA tails, and contamination. This method enables the analysis of many different mRNA quality characteristics using a single, simplified approach. When used in conjunction with single-molecule sequence methods, this method allows for comparison of capping status with other quality characteristics of individual molecules. In one embodiment, this method can analyze the capping status of different mRNA sequences that together constitute a multivalent pharmaceutical composition.

[0346] Identification of capped RNA molecules in multivalent RNA samples

[0347] The method disclosed herein is applicable to the detection of capped RNA molecules from multiple RNA populations using an RNA sample. For example, multiple different primers may be required to detect multiple different RNA molecules within a single sample via PCR-based methods (e.g., multiplex PCR amplification), and this method can be used during library preparation for next-generation sequencing or for PCR-based detection methods (such as qRT-PCR, ddPCR, etc.).

[0348] Pharmaceutical Composition

[0349] In embodiments of this disclosure, RNA molecules identified using the methods disclosed herein are intended for use as vaccine molecules, scientific or research reagents, or therapeutic reagents in pharmaceutical compositions.

[0350] In one embodiment, this disclosure provides a pharmaceutical composition comprising an RNA molecule identified by the methods of this disclosure and optionally a pharmaceutically acceptable carrier. In one embodiment, the pharmaceutical composition comprises a capped RNA molecule identified by the methods of this disclosure. In one embodiment, the pharmaceutical composition comprises an uncapped RNA molecule identified by the methods of this disclosure.

[0351] In some instances, this disclosure provides a multivalent pharmaceutical composition comprising a plurality of capped RNA molecule populations identified by the methods described herein. It should be understood that the plurality of RNA molecules in the multivalent pharmaceutical composition have different sequences and can therefore serve as templates for translating different polypeptides. In one embodiment, the pharmaceutical composition of this disclosure comprises different RNA molecule populations identified by the methods of this disclosure, wherein each population is encoded by a different RNA sequence of interest. In one embodiment, the pharmaceutical composition comprises two or more different RNA molecule populations. In one embodiment, the pharmaceutical composition is a bivalent formulation comprising two RNA molecule populations. In one embodiment, the pharmaceutical composition is a trivalent composition comprising three RNA molecule populations. In one embodiment, the pharmaceutical composition is a tetravalent composition comprising four RNA molecule populations. In one embodiment, the pharmaceutical composition is a multivalent composition comprising multiple RNA molecule populations. For example, the composition may comprise one, two, three, four, five, six, seven, eight, nine, or ten RNA molecule populations.

[0352] In one embodiment, the method of this disclosure can determine the abundance of capped and uncapped mRNA molecules in each population that commonly comprises a multivalent composition. In another embodiment, the method of this disclosure can be used to individually quantify capped RNA molecules in each RNA population of interest, and then optionally, different RNA molecule populations can be combined into a single multivalent pharmaceutical composition.

[0353] The composition may contain pharmaceutically acceptable carriers, excipients, diluents, and / or adjuvants. Pharmaceutically acceptable carriers, excipients, diluents, and / or adjuvants, as considered herein, are substances that do not produce adverse reactions when administered to a particular recipient, such as a human or non-human animal. Pharmaceutically acceptable carriers, excipients, diluents, and adjuvants are generally also compatible with the other components of the composition.

[0354] In one embodiment, the pharmaceutical composition further comprises lipid nanoparticles.

[0355] In one embodiment, this disclosure provides a method for preventing or treating a disease, the method comprising administering to a subject requiring treatment a therapeutically effective amount of an RNA molecule identified by the method of this disclosure.

[0356] In one embodiment, this disclosure provides the use of an RNA molecule identified by the methods of this disclosure in the manufacture of a medicament for the preventive or therapeutic treatment of a disease.

[0357] In one embodiment, this disclosure provides an RNA molecule identified by the methods of this disclosure for the preventive or therapeutic treatment of a disease.

[0358] Example

[0359] The nucleic acid sequences related to this disclosure are listed in Table 5.

[0360] Table 5: Polynucleotide sequences – primers and probes

[0361]

[0362] Table 6: Polynucleotide Sequence - Promoter Sequence

[0363]

[0364] Example 1: Identifying Capped Molecules Using Next-Generation Sequencing

[0365] The DNA template with an artificial T7 promoter used in Example 1 is provided in SEQ ID NO: 31. The DNA template with a natural T7 promoter used in Example 1 is provided in SEQ ID NO: 32. The expected RNA transcript of the capped RNA transcribed from the artificial T7 promoter in Example 1 is provided in SEQ ID NO: 33. The expected RNA transcript of the uncapped RNA transcribed from the artificial T7 promoter in Example 1 is provided in SEQ ID NO: 34. The expected RNA transcript of the uncapped RNA transcribed from the natural T7 promoter in Example 1 is provided in SEQ ID NO: 35.

[0366] In vitro transcription of capped and uncapped mRNA

[0367] mRNA with modified nucleotides was generated via in vitro transcription (IVT) using T7 RNA polymerase, following the protocol described in Henderson et al. (2021) and according to the manufacturer's instructions (NEB, E2080S). First, plasmid DNA encoding the following components was prepared: a modified T7 promoter (i.e., A+1; G+2), a modified α-globin 5' UTR, an open reading frame encoding green fluorescent protein (GFP), a modified α-globin 3' UTR, and a segmented Poly(A) tail (Trepotec et al. 2019), followed by a restriction digestion site (BsaI). As a control, matching plasmid DNA encoding the native T7 promoter (i.e., G+1; G+2) was also generated, differing by a single nucleotide at the transcription start site (from A+1 to G+1).

[0368] The plasmid DNA template was linearized and purified and used as a template for an IVT reaction at 32°C for 3 hours. The IVT reaction contained 16 µg / mL T7 RNA polymerase (NEB M0251), ribonucleotides (6 mM ATP, 5 mM CTP, 5 mM MTGTP; NEB N0450), 5 mM N1-methylpseudouridine-5'-triphosphate (TriLink BioTechnologies, TRN108110) or 5 mM UTP for a matched unmodified control, transcription buffer (40 mM Tris·HCl pH 8.0, 16.5 mM magnesium acetate, 10 mM dithiothreitol (DTT), 20 mM spermidine, 0.002% (v / v) Triton X-100), 2 U / mL yeast inorganic pyrophosphatase (NEB, M2403), and 1000 U / mL mouse RNase inhibitor (NEB, M0314).

[0369] For AG-capped mRNAs, Cap1 analogs were co-transcribed into the 5' end of the mRNA by adding 4 mM CleanCap AG reagent (TriLink, TRN711310). Uncapped mRNAs were transcribed in vitro in the absence of any synthetic cap analogs. The mRNA IVT reaction was terminated by adding 200 units of Dnase I (NEB, M0303) per mL of IVT reaction and incubating at 37°C for 15 min.

[0370] For enzymatic capping, mRNA was prepared from purified, uncapped eGFP mRNA. Briefly, 10 µg of uncapped mRNA was combined with 1x FCE capping buffer, 0.2 mM SAM, 0.5 mM GTP, 25 U Foster virus capping enzyme (FCE, NEBM2081), and 100 U Cap2'-O-me (2'-O-me, NEB M0366). The reaction was incubated at 37°C for 60 min. The capped mRNA was then cleaned using a Zymo clean and concentrator – 5kit (Zymo R1013) and eluted in 15 µL of nuclease-free water. The prepared capped mRNA sample (derived from co-transcription or enzymatic capping) was then used for library preparation and sequencing as described above.

[0371] The following describes the enzymatic decapping of purified co-transcribed capped mRNA or enzymatically capped eGFP mRNA. Briefly, we used the mRNA decapping enzyme (NEB M0608S) according to the manufacturer's instructions. In short, the mRNA decapping enzyme was combined with reaction buffer (10x), the mRNA vaccine sample, and nuclease-free water, and incubated at 37°C for 30 minutes. The prepared decapped sample (derived from co-transcribed or enzymatically capped mRNA) was then used for library preparation and sequencing as described above.

[0372] Following the manufacturer's instructions, the obtained mRNA was purified using the Monarch RNA Purification Kit (NEB, T2050) and finally eluted in distilled ultrapure water (ThermoFisher Scientific, 10977015). The yield, length, and purity of IVT mRNA were assessed using a range of different analytical methods. mRNA was quantified by UV spectrophotometry using a NanoPhotometer N120 (Implen), and size distribution was assessed using TapeStation electrophoresis and RNA ScreenTapes (Agilent Technologies, USA, 5067-5576).

[0373] Nanopore library preparation and detection using next-generation sequencing

[0374] cDNA-PCR sequencing was used to determine the accuracy and purity of in vitro transcribed mRNA. First, the mRNA concentration was calculated using the Qubit RNABR kit (ThermoFisher Scientific). The mRNA was diluted in nuclease-free water to an appropriate concentration for library preparation (approximately 1 ng / μL), and the concentration was confirmed using the Qubit RNA HS kit (ThermoFisher). A barcoded ONT cDNA-PCR library (SQK-PCS111.24) was prepared according to the manufacturer's instructions (Oxford Nanopore Technologies) (see [link to ONT cDNA-PCR kit]). Figure 6A and 6B However, there are exceptions. Evaporation during the cDNA synthesis step is assessed by measuring the reaction volume, and the tube is filled with nuclease-free water if appropriate. cDNA amplification is performed for 14–16 cycles (14–18 cycles recommended). Finally, the library is eluted in 8 μL of elution buffer (instead of the recommended 12 μL volume) to increase the final concentration of the library. This is beneficial because libraries prepared with templates containing modified bases appear to produce lower output libraries than those prepared with unmodified bases.

[0375] Libraries were quantified using a Qubit instrument (Invitrogen) and a dsDNA HS kit, and qualitative analysis of fragment length distribution was performed using a D5000 ScreenTapes (Agilent Technologies, USA). The results of the quantitative and qualitative analyses were used to adjust the concentration of the merged and loaded libraries. Barcoded libraries were sequenced on an R9.4.1 (FLO-MIN106D) flow cell using high-precision real-time base detection (Guppy v5.1.13 and MinKNOW Core 4.5.4). All nanowell readings with a quality score greater than 9 were assigned as pass and proceeded to further analysis.

[0376] Bioinformatics Analysis

[0377] Oxford Nanopore pDNA sequencing data were plotted and analyzed using a custom workflow. High-quality filtered tandem FASTQ reads were aligned with plasmid references using Minimap2 (Release 2.20-r1064), and Nanopore34 was mapped using -ax map-ont. The resulting SAM alignment files were processed using SAMtools v1.15 to generate sorted and indexed BAM files, as well as various other mapping analysis files. The generated BAM files were viewed and analyzed using Integrative Genomics Viewer (IGV v2.12.3). Further run and sample quality statistics were obtained using NanoPlot v1.38.1 and pycoQC vv2.5.2.

[0378] Results and discussion

[0379] Next-generation sequencing was used to measure the 5' terminal nucleotide of mRNA to determine the 5' capping state. In the presence of the CleanCap reagent AG cap analogue, mRNA vaccine molecules were transcribed from an artificial T7 promoter, with transcription initiating at the first TSS nucleotide (i.e., A(+1)). Figure 6A (see above), or in the absence of any synthetic Cap analogues, where transcription begins at the second TSS nucleotide (i.e., the G (+2) nucleotide). Figure 6A (See figure below). For comparison, transcription of the mRNA vaccine molecule begins with the first TSS nucleotide (i.e., the G (+1) nucleotide from the natural T7 promoter).

[0380] like Figure 6C As shown, next-generation sequencing analysis indicates that the 5' terminal nucleotide of the capped mRNA molecule is adenosine, corresponding to the transcription initiation at the first TSS nucleotide (i.e., +1A), and the 5' terminal nucleotide of the uncapped mRNA molecule is guanosine, corresponding to the transcription initiation at the second TSS nucleotide (i.e., +2G).

[0381] This confirms that the 5' terminal nucleotides of both capped and uncapped RNA molecules can be detected and quantified. It also confirms that transcription of capped RNA molecules begins at the first TSS (A+1), while transcription of uncapped RNA molecules begins at the second TSS (G+2), as shown below. Figure 9 As shown.

[0382] like Figure 7As shown, the genome browser view displays alignments from sequenced mRNA molecules compared to a reference plasmid sequence. It shows that mRNAs incorporating a co-transcribed 5' cap begin at +1A nucleotide (top), while mRNAs without a co-transcribed 5' cap begin at +2G nucleotide (middle). Simultaneously, mRNA molecules incorporating a co-transcribed 5' cap (which is subsequently enzymatically uncapped) also begin at +1A nucleotide (bottom). While not wanting to be bound by theory, this may indicate that +1A nucleotides represent different transcription start sites, rather than the 5' cap itself being directly detected by next-generation sequencing methods.

[0383] like Figure 8 As shown, the genome browser view displays alignments derived from sequenced mRNA molecules aligned with a reference plasmid sequence. It shows that mRNA lacking a co-transcriptional 5' cap begins transcription from +2 G nucleotides (top). However, if the 5' cap is enzymatically added to the mRNA molecule post-transcriptionally, no effect on sequencing alignment is observed. While not wanting to be bound by theory, this may indicate that the sequencing method does not directly detect the 5' cap. Furthermore, if this enzymatically added 5' cap is subsequently removed with an enzyme, no effect on alignment is observed. This may also indicate that the sequencing method does not directly detect the 5' cap. In summary, this supports the interpretation that the changes in 5' alignment initiation observed during sequencing are not due to direct detection of the 5' cap, but rather reflect differences in transcription initiation preferences of RNA polymerase in the presence of a co-transcriptional 5' cap.

[0384] like Figure 10 As shown, most capped RNA molecules are 1187 nucleotides long, while most uncapped RNA molecules are 1186 nucleotides long. This confirms that differences in transcription start sites can be detected and quantified by analyzing the length of RNA molecules.

[0385] Figure 11 The study showed a direct correlation between the proportion of RNA molecules with adenosine as a 5' terminal nucleotide in the sample and the proportion of capped molecules, providing further evidence that capped RNA molecules are identified by identifying nucleotides corresponding to the first transcription start site.

[0386] The incorporation of a capped primer during transcription initiation affects the TSS selected by RNA polymerase. For example, when the m7GpppAG capped primer is incorporated, T7 RNA polymerase initiates at the A (+1) nucleotide (see [link to original text]). Figure 1 and 4 This preference stems from the fact that the promoter sequence and RNA polymerase favor the incorporation of dinucleotide m7GpppAG analogs with greater efficiency than other adenine, thymine, uracil, and guanine nucleotides present in other ways in the reaction mixture.

[0387] Conversely, during in vitro transcription in the absence of synthetic capping primers, RNA polymerases tend to initiate transcription from different nucleotides. For example, in the absence of any synthetic capping primers, T7 RNA polymerase tends to initiate transcription at the G (+2) nucleotide (see [link to T7 RNA polymerase]). Figure 1 and 5 It is noteworthy that this preferred transcription start site is more similar to the natural promoter sequence of T7 RNA polymerase (GG) (see [link to original text]). Figure 1 and 3 ).

[0388] Therefore, the 5' terminal nucleotide of synthetic mRNA molecules (e.g., RNA vaccines or therapies) indicates the transcription start site. The selection of the transcription start site by RNA polymerase indicates whether a synthetic capping primer has been incorporated during transcription initiation. Therefore, by measuring the 5' terminal nucleotide of the mRNA molecule, it is possible to determine whether a synthetic capping primer has been incorporated during transcription initiation. For example, when using a capping primer with the form m7GpppAmG, it is possible to determine whether the mRNA has a 5' terminal A(+1) ribonucleotide or G(+2) ribonucleotide (corresponding to...). Figure 1 The transcript number of the capped RNA molecule shown can potentially determine whether transcription started from A (i.e., the first TSS) or G (i.e., the second TSS), thereby determining whether the synthetic m7GpppAG cap was incorporated during in vitro transcription by T7 RNA polymerase. Given that G is typically the preferred TSS nucleotide for RNA polymerase, this method is optimized where guanosine is not the first near-capped nucleotide within the capping primer, and where the first TSS in the modified T7 promoter is not guanosine. This method can be applied to detect the incorporation of other suitable cap analogues, where the +1 nucleotide is a non-preferred starting nucleotide.

[0389] Additionally, synthetic capping primers may contain nucleotide modifications incorporated into the 5' end of the mRNA. For example, the proximal nucleotides of the first and second caps may be methylated to improve performance and reduce innate immune recognition. These include m7Gpppm2AN, m7Gpppm2Am2N, m7Gpppm6AN, and m7Gpppm6Am2N. Detection of these nucleotide modifications at the 5' end of the mRNA molecule indicates the correct incorporation of the synthetic capping primer during co-transcription initiation. Conversely, the absence of nucleotide modifications at the 5' end of the mRNA molecule indicates the absence of incorporation of the synthetic capping primer.

[0390] Determining the presence of a 5' cap is a key characteristic for mRNA manufacturing and quality control. By measuring the 5' terminal nucleotide of a synthetic mRNA vaccine or therapy, it is possible to determine whether the mRNA molecule has been successfully incorporated with a capped primer. By determining the number of mRNAs starting from different TSSs, it is possible to determine the fraction of 5'-capped synthetic mRNA vaccines in a mixed sample and the proportion of uncapped mRNA. This provides an accurate quantitative method to indicate the fraction of capped mRNA molecules in a sample and to measure the efficiency of 5' capping during mRNA manufacturing.

[0391] There are various methods to measure the terminal nucleotides of mRNA vaccines or therapies, including but not limited to next-generation sequencing, PCR amplification, ligation, or probe hybridization methods. The identity of the first nucleotide can be detected using a variety of methods, including RNA sequencing, RNA ligation, cDNA sequencing, PCR-based techniques, isothermal amplification methods (such as LAMP), cDNA synthesis using specific templates with switched primers, or hybridization with specific oligonucleotides.

[0392] Next-generation sequencing (NGS) technologies, such as those from Illumina, Pacific Biosciences (PacBio), and Oxford Nanopore Technologies (ONT), can be used to determine the sequence of mRNA vaccine samples. This allows for the identification of ribonucleotides located at the 5' end of the mRNA molecule (Figure 6). After sequencing heterogeneous samples, the read dataset can be bio-aligned with a reference DNA template sequence. By analyzing the +1, +2, and +3 nucleotides aligned to the sequenced reads, TSSs are identified, thereby determining whether the capped primer (a synthetic 5' cap analog) was correctly incorporated during transcription initiation.

[0393] mRNA vaccine samples were analyzed using full-length complementary DNA (cDNA) sequencing. In this case, mRNA molecules were first converted to cDNA using reverse transcription, then library adaptors were ligated, and sequencing was performed using Oxford Nanopore Technology (GridION; Figure 6). The sequenced read dataset was then processed (including trimming of the library adaptors) and aligned with a reference DNA template. Analysis of the sequenced read alignments allowed us to determine the 5' (+1) nucleotide of the mRNA molecule in the sample, indicating the presence or absence of a 5' cap analogue on the RNA molecule. Figure 3 ) Detection of the starting nucleotide requires a kit containing NGS reagent.

[0394] NGS can provide quantitative measurements of the fractions of capped and uncapped mRNA molecules in a mixed sample. For example, the fractions of capped and uncapped mRNA in a mixture can be obtained by comparing the relative fractions of mRNA molecules starting with A (+1) and those starting with G (+2). Figure 11 The fraction of mRNA molecules can indicate the efficiency of the capping step during mRNA production.

[0395] Example 2: Identification of capped molecules using Taqman probes and the SYBR green detection method

[0396] Preparation of cDNA from mRNA molecules

[0397] The sequences of the primers and probes used in this example are shown in Table 5. Double-stranded cDNA was synthesized from capped and uncapped mRNA encoding the GFP coding region, which was then transcribed in vitro as described in Example 1. Briefly, 3.5 µg of RNA was combined with 1 µM Moligo(dT)20 (Invitrogen, 18418020) and 1 mM dNTPs and incubated at 70 °C for 5 min, then maintained at 4 °C until the next step. Next, 1x template-switching RT buffer, 3.75 µM template-switching oligonucleotide (TSO), and 1x template-switching RT enzyme mixture (NEB M0466) were added for first-strand cDNA synthesis. The template-switching oligonucleotide added a unique 39 nucleotide (nt) adaptor to the 5' end of each cDNA molecule. The reaction was incubated at 42 °C for 90 min, at 85 °C for 5 min, and then maintained at 4 °C until the next step. The single-stranded cDNA was then combined with 1x Q5Hot Start High Fidelity Master Mix (NEB M0494) and 25 U E. coli RNase H (NEB #M0523). The RNA was then hydrolyzed by incubation at 37°C for 15 min, and the second strand was synthesized by incubation at 95°C for 1 min and then at 65°C for 10 min. The double-stranded cDNA was stored at -20°C.

[0398] qPCR primer design

[0399] Primers and probes were designed to quantify capped and uncapped mRNA using (1) TaqMan probes and (2) a SYBR Green-based detection method.

[0400] 1) TaqMan probe assay

[0401] TaqMan probe assays were designed to quantify the relative proportions of capped and uncapped mRNA in two singlet qRT-PCR reactions. A single primer pair was designed to amplify all eGFP mRNAs (regardless of capping status), while two TaqMan probes were designed with different fluorophores to differentially label capped and uncapped mRNAs. The forward (Fwd) primer was designed to bind to the 5' adaptor sequence introduced by the template-switching oligonucleotide, located upstream of the 5' end of the mRNA (Figure 12). The reverse primer was template-specific, flanking the human α-globin 5' UTR, the KOZAK region, and the eGFP coding region. When used together in PCR, they were expected to generate an 87-nucleotide PCR product. The two TaqMan probes were designed to flank the TSS of both capped (i.e., A (+1)) and uncapped (i.e., G (+2)) mRNAs, the 5' adaptor introduced by the template-switching oligonucleotide, and the human α-globin 5' UTR. Capped mRNA detection probes were labeled with HEX fluorophores and designed to overlap with the TSS (i.e., A(+1)) introduced during the 5' capping reaction. The alternative TSS for uncapped mRNA (i.e., G(+2)) implies a mismatch between the capped probe and the uncapped mRNA and is not expected to anneal stably during qRT-PCR. Uncapped mRNA detection probes labeled with FAM fluorophores lack the capped mRNA TSS (i.e., A(+1)). Since the uncapped probes do not contain the capped mRNA TSS (A(+1)), they are not expected to bind stably to capped mRNA during qRT-PCR.

[0402] (2) SYBR Green Measurement

[0403] The SYBR green assay was designed to quantify the relative proportions of capped and uncapped mRNA in two singlet qRT-PCR reactions. Two Fwd primers were designed to bind differentially to capped and uncapped mRNA (Figure 13). These primers differed only in their 3' terminal nucleotides: the capped detection Fwd primer (Fwd1_SYBR_cap) terminated with GGGA, and the uncapped detection Fwd primer (Fwd2_SYBR_uncap) lacked the A and terminated with GGGG. These primers were primarily complementary to the TSO adaptor sequence, except for the TSS of capped mRNA (i.e., A (+1)) or the TSS of uncapped mRNA (i.e., G (+2)). The two Fwd primers were designed for use with a single Rev primer, which was designed to encode the human α-globin 5' UTR, KOZAK, or eGFP region. After testing four reverse primers, based on melting and efficiency curve analysis, Rev3_SYBR was identified as producing the most reliable results when used with Fwd1_SYBR_cap and Fwd2_SYBR_uncap primers. When used for PCR, Fwd1_SYBR_cap and Rev3_SYBR produced a PCR product of 156 nucleotides, while PCR, Fwd2_SYBR_uncap, and Rev3_SYBR produced a PCR product of 155 nucleotides. All provided SYBR capping assay results included both primer combinations. Expression of the cap-dependent SYBR green primer was compared with the cap-independent positive control eGFP primer pair (Fwd1_eGFP_poscon and Rev1_eGFP_poscon (Leveque-Serve et al., 2007)). The control eGFP primers served as internal controls, similar to housekeeping control genes.

[0404] qRT-PCR assay of samples with known capping status

[0405] A mixture of capped and uncapped eGFP cDNA was subjected to SYBR Green qRT-PCR assays to test whether the calculated ratio of capped and uncapped cDNA matched the cDNA input. The cDNA mixture contained: 1) 100% capped eGFP cDNA, 2) 95% capped eGFP cDNA, 3) 75% capped eGFP cDNA, 4) 50% capped eGFP cDNA, and 5) 0% capped eGFP cDNA. Capped and uncapped PCR primers were amplified in separate reactions along with a positive control primer pair, independent of capping status. All three primer pairs were amplified using the following reaction conditions: 40 cycles of 95°C for 3 min, followed by 95°C for 15 sec and 68.5°C for 30 sec. Melting curves were then performed to confirm that a single product was amplified in each well. The Ct of each reaction was compared to the Ct of the positive control primer using the δCt method. The δCt method involves normalization equivalent to the control sample. The δCt value is expressed as a ratio to the 100% capped sample, as 100% capped sample is used.

[0406] Next, using the cDNA input described above, a mixture of capped and uncapped cDNA was subjected to TaqMan qRT-PCR assays. Both capped and uncapped PCR probes were assayed using a single primer pair. Capped and uncapped probes were used in two separate reactions and amplified under the following conditions: 40 cycles of 95°C for 3 min, followed by 95°C for 15 sec and 60°C for 30 sec. The Ct of each reaction was compared to the Ct of the positive control primer (described above) using the δCt method. The δCt method involves normalization equivalent to the control sample. 100% capped samples were used, therefore the δCt values ​​are expressed as a ratio to 100% capped samples.

[0407] Polymerase chain reaction (PCR)-based methods can be used to analyze the 5' nucleotide terminus and capping status of mRNA. These methods include reverse transcription polymerase chain reaction (RT-PCR), quantitative real-time RT-PCR (qRT-PCR), TaqMan qPCR, multiplex RT-PCR, digital RT-PCR (dRT-PCR), nested RT-PCR, RT-Loop-mediated isothermal amplification (RT-LAMP), RT-droplet digital PCR (RT-ddPCR), and RT-PCR combined with high-resolution melting (HRM) analysis. This method measures the 5' nucleotide terminus of synthetic mRNA vaccines, thereby indicating TSS and whether a 5' cap analogue has been incorporated during transcription initiation. This method is generally applicable to any PCR or probe hybridization technique that can adequately measure terminal nucleotide or mRNA molecule length differences. Detection of the initiating nucleotide requires a kit containing qPCR reagents.

[0408] TaqMan probe assays were designed to quantify the relative proportions of capped and uncapped mRNA in multiplex qRT-PCR reactions. In this case, the presence of m7Gpppm2A in the synthesized mRNA vaccine was investigated to determine whether it was incorporated into the mRNA vaccine by starting with A (+1) or G (+2). Probes designed to bind to the mRNA vaccine starting with either A (+1) or G (+2) nucleotides were used to detect capped or uncapped mRNA molecules, respectively. Figure 4 ).

[0409] When plotting the expected and calculated capping ratios for the SYBR green qRT-PCR assay, R was observed. 2 A value of 0.8 indicates that the calculated ratio of capped to uncapped cDNA reflects the capped ratio in the input sample. Figure 14A ).

[0410] When plotting the expected and calculated capping ratios for TaqMan qRT-PCR assays, R was observed. 2 The value is 0.7711, confirming that the calculated ratio of capped to uncapped cDNA reflects the capped ratio in the input sample. Figure 14B ).

[0411] These results confirm that the presence of additional adenosine due to the transcriptional initiation of the first TSS found in capped RNA molecules can be used to distinguish and quantify capped and uncapped RNA molecules using qRT-PCR and TaqManqRT-PCR assays.

[0412] The fraction of capped or uncapped mRNA molecules in a mixed sample can be determined by using a PCR-based method to identify the 5' terminal nucleotide. The fractions of capped and uncapped mRNA molecules in the mixed sample can be determined separately by measuring the fraction of mRNA molecules beginning with A (+1) or G (+2) nucleotides. The absolute fraction of capped or uncapped mRNA molecules in the test sample can be determined by comparing the fraction of mRNA molecules beginning with A (+1) or G (+2) nucleotides to a reference standard mixture of capped and uncapped mRNAs with known proportions.

[0413] This disclosure is in the form of

[0414] Form 1: A method for distinguishing capped RNA molecules from uncapped RNA molecules in a sample comprising a mixture of capped RNA molecules and uncapped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample;

[0415] This capped RNA molecule has been used in a universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0416] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0417] m is 0 or 1; and

[0418] n is any integer from 0 to 8;

[0419] The method includes:

[0420] a) Identifying capped and / or uncapped RNA molecules from the sample, wherein the identification includes steps selected from the group consisting of:

[0421] Identify N 1 The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N... 1 The absence of the 5' terminal nucleotide identified an uncapped RNA molecule.

[0422] Identifying the 5' terminal nucleotide as the first transcription start site corresponding to the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleotide as the second transcription start site corresponding to the DNA template identifies uncapped RNA molecules.

[0423] Identifying the 5' terminal nucleoside as the nucleoside corresponding to the +1 position of the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template identifies uncapped RNA molecules.

[0424] Where N 1 N 2 and N 3 At least one of the following in each instance is independently any modified or non-natural nucleoside: the presence of a modified or non-natural nucleoside at the 5' end of an RNA molecule identifies a capped RNA molecule; and the absence of a modified or non-natural nucleoside at the 5' end of an RNA molecule identifies an uncapped RNA molecule; and

[0425] The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; RNA molecules with a length of at least x+1 nucleotides are identified, which identifies capped RNA molecules; and RNA molecules with a length of x nucleotides are identified, which identifies uncapped RNA molecules.

[0426] This identification step distinguishes between capped and uncapped RNA molecules.

[0427] Form 2: A method for identifying capped RNA molecules and / or uncapped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in an RNA molecule sample containing a mixture of capped and uncapped RNA molecules;

[0428] This capped RNA molecule has been used in a universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0429] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0430] m is 0 or 1; and

[0431] n is any integer from 0 to 8;

[0432] The method includes:

[0433] a) Identifying the capped RNA molecule and / or the uncapped RNA molecule from the sample, wherein the identification includes steps selected from the group consisting of:

[0434] Identify N 1 The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N... 1 The absence of the 5' terminal nucleotide identified an uncapped RNA molecule.

[0435] Identifying the 5' terminal nucleotide as the first transcription start site corresponding to the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleotide as the second transcription start site corresponding to the DNA template identifies uncapped RNA molecules.

[0436] Identifying the 5' terminal nucleoside as the nucleoside corresponding to the +1 position of the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template identifies uncapped RNA molecules.

[0437] Identifying modified or non-natural nucleosides identifies capped RNA molecules; and / or identifying the absence of modified or non-natural nucleosides identifies uncapped RNA molecules; where N 1 N 2 and N 3 At least one of each instance is independently any modified or non-natural nucleoside; and

[0438] The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; RNA molecules with a length of at least x+1 nucleotides are identified, which identifies capped RNA molecules; and RNA molecules with a length of x nucleotides are identified, which identifies uncapped RNA molecules.

[0439] in:

[0440] This identification step identifies capped RNA molecules and / or uncapped RNA molecules.

[0441] Form 3: A method for identifying capped and / or uncapped RNA molecules in an RNA molecule sample comprising a mixture of capped and uncapped RNA molecules from multiple RNA molecule populations, each of which is transcribed in vitro from a DNA template by an RNA polymerase;

[0442] Capped RNA molecules have been used in a universal form [cap]-[connector]-N 1 [N 2 ] m [N 3 ] n Capped primers; among which

[0443] N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside;

[0444] m is 0 or 1; and

[0445] n is any integer from 0 to 8;

[0446] The method includes:

[0447] a) Identifying the capped RNA molecule and / or the uncapped RNA molecule from the sample, wherein the identification includes steps selected from the group consisting of:

[0448] Identify N 1 The presence of the 5' terminal nucleotide in RNA molecules identifies capped RNA molecules; and / or identifies N... 1The absence of the 5' terminal nucleotide identified an uncapped RNA molecule.

[0449] Identifying the 5' terminal nucleotide as the first transcription start site corresponding to the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleotide as the second transcription start site corresponding to the DNA template identifies uncapped RNA molecules.

[0450] Identifying the 5' terminal nucleoside as the nucleoside corresponding to the +1 position of the DNA template identifies capped RNA molecules; and / or identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template identifies uncapped RNA molecules.

[0451] Identifying modified or non-natural nucleosides identifies capped RNA molecules; and / or identifying the absence of modified or non-natural nucleosides identifies uncapped RNA molecules; where N 1 N 2 and N 3 At least one of each instance is independently any modified or non-natural nucleoside; and

[0452] The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; RNA molecules with a length of at least x+1 nucleotides are identified, which identifies capped RNA molecules; and RNA molecules with a length of x nucleotides are identified, which identifies uncapped RNA molecules.

[0453] This identification step identifies capped RNA molecules and / or uncapped RNA molecules.

[0454] Form 4: The method according to any one of the foregoing forms, further comprising:

[0455] b) Quantitatively analyze the identified capped and / or uncapped RNA molecules.

[0456] Form 5: The method according to any one of the preceding forms, wherein the quantification in step b) includes:

[0457] Determine the absolute amount of the capped RNA molecule in the sample;

[0458] Determine the relative abundance of capped and / or uncapped RNA molecules in the sample;

[0459] Determine the relative abundance of capped RNA molecules relative to total RNA molecules in the sample.

[0460] Form 6: The method according to Form 3 further includes:

[0461] b) Quantify the identified capped RNA molecules and / or uncapped RNA molecules;

[0462] The quantification in step b) includes:

[0463] Determine the absolute amount of the capped RNA molecule in the sample;

[0464] Determine the relative abundance of capped and / or uncapped RNA molecules in the sample;

[0465] Determine the relative abundance of capped RNA molecules relative to total RNA molecules in the sample;

[0466] Determine the absolute amount of capped and / or uncapped RNA molecules in each of the multiple RNA molecule populations in the sample;

[0467] Determine the relative abundance of capped and / or uncapped RNA molecules in each of multiple RNA molecule populations in a sample;

[0468] Determine the relative abundance of capped RNA molecules relative to the total RNA molecules in each of the multiple RNA molecule populations in the sample; and / or

[0469] Determine the relative abundance of capped and / or uncapped RNA molecules for each of multiple RNA molecule populations.

[0470] References

[0471] Beverly, Michael, Amy Dell, Parul Parmar, and Leslie Houghton. 2016. "Label-Free Analysis of MRNA Capping Efficiency Using RNase H Probes and LC-MS." Analytical and Bioanalytical Chemistry 408 (18): 5021-30.

[0472] Galloway, Alison, Abdelmadjid Atrih, Renata Grzela, Edward Darzynkiewicz, Michael A. J. Ferguson, and Victoria H. Cowling. 2020. “CAP-MAP: Cap Analysis Protocol with Minimal Analyte Processing, a Rapid and Sensitive Approach to Analysing MRNA Cap Structures." Open Biology 10 (2): 190306.

[0473] Grudzien, E. et al. RNA, 10: 1479-1487 (2004).

[0474] Grudzien-Nogalska, E., et al., RNA, 13: 1745-1755 (2007).

[0475] Henderson, Jordana M., Andrew Ujita, Elizabeth Hill, Sally Yousif-Rosales, Cory Smith, Nicholas Ko, Taylor McReynolds, Charles R. Cabral, Julienne R. Escamilla-Powers, and Michael E. Houston. 2021. "Cap1 Messenger RNASynthesis with Co-Transcriptional CleanCap® Analogue by in Vitro Transcription." Current Protocols 1 (2): e39.

[0476] Hornung, Veit, Jana Ellegast, Sarah Kim, Krzysztof Brzozka, Andreas Jung, Hiroki Kato, Hendrik Poeck, et al. 2006. "5′-Triphosphate RNA Is the Ligand for RIG-I." Science (New York, N.Y.) 314 (5801): 994-97.

[0477] Jemielity, J. et al., “Novel ‘anti-reverse’ cap analogues with superior translational properties”, RNA, 9: 1108-1122 (2003)

[0478] Levesque-Sergerie, JP., Duquette, M., Thibault, C. et al Detection limits of several commercial reverse transcriptase enzymes: impact on the low- and high-abundance transcript levels assessed by quantitative RT-PCR. BMC Molecular Biol 8, 93 (2007).

[0479] Trepotec, Z. et al., "Segmented poly(A) tails significantly reduce recombination of plasmid DNA without affecting mRNA translation efficiency or half-life." RNA 25(4):507-518 (2019)

[0480] Analytical Procedures for mRNA Vaccine Quality - Draft Guidelines, USP, 2023

Claims

1. A method for identifying capped RNA molecules transcribed in vitro from a DNA template by RNA polymerase in a sample containing a mixture of capped and uncapped RNA molecules; The capped RNA molecule described above has been used with the universal form m7GpppN 1 [N 2 ] m [N 3 ] n Capped primers; in m7G is an N7-methylated guanosine or a guanosine analogue; PPP stands for triphosphate; N 1 , N 2 , and N 3 each instance independently is any natural, modified or non-natural nucleoside; m is 0 or 1; and n is any integer from 0 to 8; The method includes: a) Identify at least the 5' terminal nucleotide of multiple RNA molecules from the sample by: Identifying N 1 the presence of a 5' terminal nucleoside, and then identifying the 5' terminal nucleoside as corresponding to the nucleoside at the +1 position of the DNA template, thereby identifying a capped RNA molecule; and / or Identifying N 1 the absence of a 5' terminal nucleoside, and then identifying the 5' terminal nucleoside as the nucleoside corresponding to the +2 position of the DNA template, thereby identifying an uncapped RNA molecule, in: the nucleoside at the +1 position of the DNA template comprises adenosine, and N 1 comprises adenosine; the nucleoside at the +1 position of the DNA template comprises cytidine, and N 1 comprises cytidine; or The nucleotide at the +1 position of the DNA template comprises thymidine, and N 1 It is uridine; and b) Quantitative N 1 The presence of the aforementioned substance allows for the quantification of the capped RNA molecules in the sample.

2. The method according to claim 1, wherein the nucleoside at the +1 position of the DNA template is adenosine, and N 1 The nucleoside at position +2 of the DNA template is selected from the group consisting of adenosine, N2-methyladenosine, or N6-methyladenosine, and the nucleoside at position +2 of the DNA template comprises guanosine, and N... 2 It contains guanosine.

3. The method according to claim 1 or claim 2, wherein the quantification in step b) comprises: Determine the absolute amount of the capped RNA molecules in the sample; Determine the abundance of the capped RNA molecules relative to the uncapped RNA molecules in the sample; and / or Determine the abundance of the capped RNA molecules relative to the total RNA molecules in the sample.

4. The method according to any one of the preceding claims, wherein N 1 It is not the preferred starting nucleoside for the RNA polymerase.

5. The method according to any one of the preceding claims, wherein step a) further comprises identifying the length of the RNA molecule, wherein: The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; and in: Identifying an RNA molecule that is at least x+1 nucleotides long constitutes identification of a capped RNA molecule; and Identifying an RNA molecule that is x nucleotides long means identifying an uncapped RNA molecule.

6. The method according to any one of the preceding claims, wherein N 1 The nucleosides are modified or non-natural, and step a) further includes identifying the modified or non-natural nucleosides, wherein: The identification of capped RNA molecules is achieved by recognizing the presence of the modified or non-natural nucleoside as a 5' terminal nucleoside; and The absence of the modified or non-natural nucleoside as a 5'-terminal nucleoside is used to identify uncapped RNA molecules.

7. The method according to any one of the preceding claims, wherein m = 1 and n = 0.

8. The method according to any one of the preceding claims, wherein the capping primer is selected from the group consisting of: 5'm7GpppAN, m7Gpppm2AN, m7Gpppm2Am2N, m7Gpppm6AN, and m7Gpppm6Am2N, m7G(5')ppp(5')G, m7(3')-O-Me-G(5')ppp(5')G), m7G(5')ppp(5')(2'OMeA)pG and m7G(5')ppp(5')(2'OMeA)pU), m7(3'OMeG)(5')ppp('5)m6(2'OMeA)pG), m7Gpppm6A mN, m7GpppN mN and 7mG3OMe5'ppp5'N.

9. The method according to any one of the preceding claims, wherein the RNA molecule has been transcribed by T7 RNA polymerase.

10. The method of claim 9, wherein the DNA template comprises a sequence selected from the group consisting of: TAATACGACTCACTATA (SEQ ID NO: 23), TAATACGACTCACTATAAG (SEQ ID NO: 15), and TAATACGACTCACTATAAT (SEQ ID NO: 24).

11. The method according to any one of the preceding claims, wherein the identification in step a) further comprises: RNA molecules from the sample were reverse transcribed to form complementary DNA (cDNA) molecules, and Identifying at least the terminal nucleoside of the cDNA molecule, wherein identifying the terminal nucleoside of the cDNA molecule is equivalent to identifying at least the 5' terminal nucleoside of the RNA molecule.

12. The method according to any one of the preceding claims, wherein the identification in step a) comprises identifying at least one to three 5' nucleotides of the RNA molecule, identifying at least one to ten 5' nucleotides of the RNA molecule, or identifying substantially all nucleotides of the RNA molecule.

13. A method for quantifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample, said RNA molecule sample comprising a mixture of capped RNA molecules and uncapped RNA molecules; The capped RNA molecule described above has been used with the universal form m7GpppN 1 [N 2 ] m [N 3 ] n Capped primers; in m7G is an N7-methylated guanosine or a guanosine analogue; PPP stands for triphosphate; N1 is a modified or non-natural nucleoside, N2 is a modified or non-natural nucleoside, and N3 is any natural, modified, or non-natural nucleoside. m is 0 or 1; and n is any integer from 0 to 8; The method includes: To identify whether at least the 5' terminal nucleosides of multiple RNA molecules from the sample are natural, modified, or non-natural nucleosides; in: The identification of capped RNA molecules is defined as the presence of modified or non-natural nucleosides as 5'-terminal nucleosides; and The absence of modified or non-natural nucleosides as 5'-terminal nucleosides is used to identify uncapped RNA molecules; and The presence of modified or non-natural nucleosides as 5'-terminal nucleosides is quantified, thereby quantifying the capped RNA molecules in the sample.

14. A method for identifying capped RNA molecules transcribed in vitro from a DNA template in an RNA molecule sample, said RNA molecule sample comprising a mixture of capped RNA molecules and uncapped RNA molecules; The capped RNA molecule described above has been used with the universal form m7GpppN 1 [N 2 ] m [N 3 ] n Capped primers; in m7G is an N7-methylated guanosine or a guanosine analogue; PPP stands for triphosphate; N 1 Any natural, modified, or non-natural nucleoside, in which N 1 Not the preferred starting nucleotide for RNA polymerase; N 1 N 2 and N 3 Each instance is independently any natural, modified, or non-natural nucleoside; and m is 0 or 1; and n is any integer from 0 to 8; The method includes: The length of the RNA molecule was determined, wherein: The uncapped RNA molecule is x nucleotides long, and x is any suitable integer; and in: Identifying an RNA molecule that is at least x+1 nucleotides long constitutes identification of a capped RNA molecule; and Identifying an RNA molecule that is x nucleotides long means identifying an uncapped RNA molecule.

15. A pharmaceutical composition comprising a capped RNA molecule identified by the method according to any one of the preceding claims.