Nucleic acid untranslated regions

Optimized 5’UTR and 3’UTR combinations in mRNA constructs improve production yield and purity, addressing impurity issues and enhancing therapeutic efficacy.

WO2026061971A1PCT designated stage Publication Date: 2026-03-26LONZA AG
View PDF 12 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing mRNA constructs face challenges in achieving high yield and purity during production, with impurities such as RNA fragments and double-stranded RNA reducing therapeutic efficacy.

Method used

The use of specific 5’UTR and 3’UTR combinations derived from human alpha-globin, Glucose 6-phosphate isomerase 1 (GPI1), nascent polypeptide associated complex subunit alpha (NACA), and magnesium transporter (MRS2) nucleic acid sequences, optimized for improved RNA integrity, stability, and reduced dsRNA impurities.

Benefits of technology

Enhances mRNA production yield and purity, ensuring improved stability and translational efficiency, thereby increasing the effectiveness of mRNA-based therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000042_0001
    Figure IMGF000042_0001
  • Figure IMGF000045_0001
    Figure IMGF000045_0001
  • Figure IMGF000046_0001
    Figure IMGF000046_0001
Patent Text Reader

Abstract

1. An RNA molecule comprising in the 5'3' direction, a 5'-terminal transcription start site, a 5' untranslated region (5'UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3' untranslated region (3'UTR), wherein a) the 5'UTR comprises a transcript of the 5'UTR of the human alpha-globin nucleic acid sequence SEQ ID NO:1, a fragment or variant thereof which is at least 90% identical to SEQ ID NO:1; and b) the 3'UTR comprises a transcript of: i) the human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid sequence SEQ ID NO:19, or a fragment of SEQ ID NO:19 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO:19 which is at least 90% identical to SEQ ID NO:19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21, or a fragment of SEQ ID NO:21 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO:21 which is at least 90% identical to SEQ ID NO:21; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, or a fragment of SEQ ID NO:23 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO:23 which is at least 90% identical to SEQ ID NO:23; or iv) a combination of any two or more of i), ii), or iii).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] LO017P

[0002] -1-

[0003] NUCLEIC ACID UNTRANSLATED REGIONS

[0004] FIELD OF THE INVENTION

[0005] The invention refers to certain untranslated regions of nucleic acid expression constructs suitable for producing RNA molecules. The RNA molecule comprises in the 5’->3’ direction, a 5’ untranslated region (5’UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3’ untranslated region (3’UTR), and can be produced at high yield. A selection of 5’UTR and 3’UTR sequences is provided herein. The invention further refers to DNA templates comprising a nucleotide sequence suitable for transcription of the respective RNA molecules.

[0006] BACKGROUND mRNA-based therapeutics are promising treatment strategy in immunotherapy, gene therapy, and cancer treatments. Many of these applications require prolonged intracellular persistence of mRNA to improve bioavailability of the encoded expression product. mRNA molecules are intrinsically unstable and their intracellular kinetics depend on the UTRs embracing the coding sequence. Sequences of 5' and 3' untranslated regions (UTRs) are responsible for translational efficiency and stability of mRNA. An optimal combination of the UTR sequences allows for a significant increase of the protein’s expression. Chatterjee and Pal (Biol. Cell 2009, 101 :251-262) describe the role of 5’- and 3’UTR of RNAs in human disease.

[0007] RNA molecules can be produced by in vitro transcription of DNA molecules e.g., a vector, that serve as a template for transcription of the sequence.

[0008] A 5’UTR serves as the entry point for the ribosome during translation, can adopt elaborate RNA secondary and tertiary structures that may regulate translation initiation in a cap-dependent or cap-independent manner.

[0009] A 3’UTR participates in processes of cleavage and polyadenylation, translation, and localization of mRNA, affecting mRNA stability.

[0010] Among the best-known untranslated sequences that ensure high translational efficiency are globin gene sequences. In particular, sequences of human or rabbit a- and P-globin or of rabbit P-globin have been used to improve translational efficiency of mRNAs. LO017P

[0011] -2-

[0012] Certain 5’UTR and 3’UTR sequences are used in COVID-19 mRNA vaccines from Moderna (Cambridge, MA, USA) and Pfizer / BioNTech (Mainz, Germany). In particular, the 5’UTR and the 3’UTR sequences as used in the COVID-19 mRNA vaccine of Pfizer / BioNTech are considered as a reference standard when it comes to developing improved constructs. The respective 5’UTR is derived from the human alpha-globin nucleic acid sequence with some modifications, and the 3’UTR sequence comprises mitochondrially encoded 12S ribosomal RNA (mtRNRI) and amino-terminal enhancer of split (AES) motives. The 5’UTR and 3’UTR sequences are provided in World Health Organization MedNet, Sept. 2020 document 11889.

[0013] AES-mtRNR1 3'UTRs were found superior compared to other 3’UTRs (Orlandini von Niessen et al., Molecular Therapy 2019, 27(4):824).

[0014] Kozak (J. Mol. Biol. 1994, 235:95-110) discloses features in the 5’ non coding sequences of alpha globin mRNAs that effect translational efficiency. The so-called Kozak sequence is located near the initiation site of translation and significantly influences the performance of ribosomes in correctly identifying the initiation codon and starting protein synthesis. Different Kozak variants give rise to different amounts of protein synthesis, and thus modifications of Kozak sequences may have regulatory effects on translation efficiencies.

[0015] There is patent literature which discloses mRNA constructs comprising certain 5’UTR and 3’UTR combinations (WO2023062556A1 , WO2024 / 153324A1 , WO2023146230A1 , WO2023242817A2, WO2021214204A1 , WO2017060314A2, W02017059902A1). mRNA molecules for therapeutic and preventive medical applications are generally synthesized in vitro in cell-free systems. The synthesized mRNA molecules need to be purified using conventional laboratory methods. Impurities such as RNA fragments and double-stranded RNA may give rise to unwanted effects and reduce the therapeutic or prevention efficacy.

[0016] WO2020198697A1 discloses compositions and methods for editing, e.g., introducing double-stranded breaks, within the TTR gene, and a nucleic acid comprising an open reading frame encoding an RNA-guided DNA binding agent. A series of 5’UTR and 3’UTR combinations is disclosed.

[0017] WO2020185632A1 discloses modular and tunable protein expression systems, and a large set of coding nucleic acids. LO017P

[0018] -3-

[0019] WO2013151663A1 discloses polynucleotides comprising the native 5' UTR of any of the nucleic acids that encode any of SEQ ID NOs 8144-16131 , SEQ ID NOs: 1- 4, the native 3' UTR of any of the nucleic acids that encode any of SEQ ID NOs 8144- 16131 , SEQ ID NOs 5-21 , and a polypeptide of interest selected from the group consisting of SEQ ID NOs 8144-16131.

[0020] There is a need for improved RNA constructs and respective DNA templates which produce the respective RNA transcripts at high yield and purity.

[0021] SUMMARY OF THE INVENTION

[0022] It is the object of the invention to improve mRNA production from DNA templates in particular for higher yield and purity. It is a further object of the invention to provide specific 5’UTR and 3’UTR combinations which provide for improved mRNA production, in particular where the transcribed mRNA has a sufficient half-life suitable for various mRNA applications such as mRNA-based therapy.

[0023] The object of the invention is solved by the subject matter as claimed, and as further described herein.

[0024] The invention is based on the surprising finding that certain 5’UTRs and particular 5’UTR and 3’UTR combinations as used in a DNA template provide for improved RNA integrity during manufacturing, improved stability for purified RNA, improved protein yield in the cell, enhanced cellular mRNA stability, and a lower amount of dsRNA impurities as tested for specific mRNA constructs.

[0025] The invention provides for an RNA molecule comprising in the 5’->3’ direction, a 5’ untranslated region (5’UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3’ untranslated region (3’UTR), wherein a) the 5’UTR comprises a transcript of the 5’UTR derived from the human alphaglobin nucleic acid sequence SEQ ID NO:1 , a fragment of SEQ ID NO:1 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:1 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:1 ; and b) the 3’UTR comprises a transcript of: i) the human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid sequence SEQ ID NO: 19, a fragment of SEQ ID NO: 19 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any LO017P

[0026] -4- one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO: 19 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO: 19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21 , a fragment of SEQ ID NO:21 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:21 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:21 ; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, or a fragment of SEQ ID NO:23 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:23 which is at least any one of 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:23; or iv) a combination of any two or more of i), ii), or iii).

[0027] Specifically, any fragment or variant described herein is characterized by its functional activity as a respective 5’UTR or 3’UTR, and is herein understood as a functional fragment or functional variant. Functional activity, as used herein, denotes the role of the respective untranslated region (UTR) in influencing RNA properties, including but not limited to stabilization.

[0028] According to a specific aspect, the invention provides for an RNA molecule comprising in the 5’->3’ direction, a 5’-terminal transcription start site, a 5’ untranslated region (5’UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3’ untranslated region (3’UTR), wherein a) the 5’UTR comprises a transcript of the 5’UTR of the human alpha-globin nucleic acid sequence SEQ ID NO:1 , or a fragment of SEQ ID NO:1 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ LO017P

[0029] -5-

[0030] ID NO:1 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:1 ; and b) the 3’UTR comprises a transcript of: i) the human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid sequence SEQ ID NO: 19, or a fragment of SEQ ID NO: 19 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO: 19 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO: 19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21 , or a fragment of SEQ ID NO:21 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:21 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:21 ; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, or a fragment of SEQ ID NO:23 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:23 which is at least any one of 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:23; or iv) a combination of any two or more of i), ii), or iii).

[0031] Preferably, either one, or more or all of the 5’UTR and 3’UTR nucleic acid sequences comprise or consist of the full length-sequence further described herein, in particular SEQ ID NO:1 , 19, 21 , or 23, or a functionally active fragment of any of the foregoing, preferably wherein the functionally active fragment comprises at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full- length sequence. Functional activity is herein defined as the function as the respective UTR. LO017P

[0032] -6-

[0033] Specifically, the full-length or a fragment of any one of SEQ ID NO:1 , 19, 21 , or 23, can be used, preferably a fragment comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence. Specifically, the fragment of the respective SEQ ID NO:1 , 19, 21 , or 23 consists of a contiguous sequence of the respective SEQ ID NO:1 , 19, 21 , or 23 which is truncated at its 3’-end or 5’end.

[0034] Specifically, the variant or fragment of any one of SEQ ID NO:1 , 19, 21 , or 23 as described herein is functional, in particular to comprise functional activity as a respective 5’UTR or 3’UTR.

[0035] Specifically, the elements of the DNA constructs and RNA constructs described herein are operably linked such as to allow transcription of the DNA into an RNA molecule, in particular an mRNA or saRNA, and to further allow translation of the RNA molecule to produce an amino acid sequence encoded by the coding region of the RNA molecule.

[0036] Specifically, the 5’UTR DNA sequence comprises or consists of SEQ ID NO:1 , or a fragment, or variant thereof.

[0037] Specifically, the transcript of human alpha-globin nucleic acid 5’UTR DNA sequence comprising or consisting of SEQ ID NO:1 , or a fragment, or variant thereof is preferably the 5’UTR RNA sequence which comprises or consists of SEQ ID NO:2, or a fragment, or variant thereof.

[0038] Specifically, the 3’UTR DNA sequence comprises or consists of SEQ ID NO:19, or a fragment, or variant thereof.

[0039] Specifically, the transcript of human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid 3’UTR DNA sequence comprising or consisting of SEQ ID NO: 19, or a fragment, or variant thereof is preferably the 3’UTR RNA sequence which comprises or consists of SEQ ID NO:20, or a fragment, or variant thereof.

[0040] Specifically, the 3’UTR DNA sequence comprises or consists of SEQ ID NO:21 , or a fragment, or variant thereof.

[0041] Specifically, the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid 3’UTR DNA sequence comprising or consisting of SEQ ID NO:21 , or a fragment, or variant thereof is preferably the 3’UTR RNA sequence which comprises or consists of SEQ ID NO:22, or a fragment, or variant thereof. LO017P

[0042] -7-

[0043] Specifically, the 3’UTR DNA sequence comprises or consists of SEQ ID NO:23, or a fragment, or variant thereof.

[0044] Specifically, the human magnesium transporter (MRS2) nucleic acid sequence 3’UTR DNA sequence comprising or consisting of SEQ ID NO:23, or a fragment, or variant thereof is preferably the 3’UTR RNA sequence which comprises or consists of SEQ ID NO:24, or a fragment, or variant thereof.

[0045] Specifically, the variant of any of the UTR sequences provided herein comprises at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity.

[0046] Specifically, the variant of any one of SEQ ID NO:1 , 19, 21 , or 23 as described herein is at least any one of 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% up to 100% identical to the respective sequence.

[0047] Specifically, the 5’UTR DNA variant may comprise a substitution of one or more nucleotides e.g., within the 5’-terminal sequence, such as at position 2 of SEQ ID NO:1 or 2.

[0048] Specifically, the difference of the 5’UTR variant to SEQ ID NO:1 or 2 comprises a number of point mutations (in particular, a substitution, insertion, or deletion of one or more nucleotides) within the 5’-terminal and / or 3’-terminal sequence.

[0049] Specifically, the 5’UTR DNA variant differs from SEQ ID NO:1 in a number of point mutations which is 1 , 2, or 3.

[0050] Specifically, the 5’UTR RNA variant differs from SEQ ID NO:2 in a number of point mutations which is 1 , 2, or 3.

[0051] Specifically, the 5’UTR variant comprises a sequence of the consecutive nucleotides from position 4-31 of SEQ ID NO:1 or 2.

[0052] Specifically, the 5’UTR variant comprises a point mutation at position 2 of SEQ ID NO:1 or 2, in particular a substitution such as C2T in the DNA sequence SEQ ID NO:1 , or C2U in the RNA sequence SEQ ID NO:2.

[0053] Specifically, the 5’UTR variant comprises a 5’- and / or 3’- extension of the respective 5’UTR sequence.

[0054] Specifically, the 5’-extension of the 5’UTR sequence may comprise or consist of an insertion of one or more nucleotides, in particular a sequence of consecutive nucleotides with a length of 1-20 nucleotides, preferably at least any one of 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, preferably between 3 LO017P

[0055] -8- and 15 nucleotides, or between 3 and 11 nucleotides, more preferably about 11 nucleotides.

[0056] Specifically, the 5’-extension of the 5’UTR sequence may comprise or consist of a transcription start site, such as “AGG” or “GGG”, preferably AGG.

[0057] Specifically, the 5’UTR comprises a 5’-terminal transcription start site, preferably “AGG” or “GGG”, preferably AGG.

[0058] Specifically, the 5’-extension of the 5’UTR may comprise at its 5’-end a transcription start site, such as “AGG” or “GGG”, preferably AGG.

[0059] Specifically, the 5’-extension of the 5’UTR DNA sequence may comprise or consist of SEQ ID NO:51 or 53.

[0060] Specifically, the 5’-extension of the 5’UTR RNA sequence may comprise or consist of SEQ ID NO:52 or 54.

[0061] Specifically, the 3’-extension of the 5’UTR sequence comprises or consists of a sequence of consecutive nucleotides with a length of 1 -5 nt, preferably at least any one of 1 , 2, 3, 4, or 5 nt.

[0062] Specifically, the 3’-extension of the 5’UTR sequence comprises or consists of a spacer between the 5’UTR and the Kozak sequence, which may as well be understood as a 5’-terminal extension of the Kozak sequence.

[0063] Specifically, the 3’-extension of the 5’UTR sequence comprises or consists of nucleotides selected from C, G, A, or any of T (DNA sequence) and U (RNA sequence), respectively, preferably wherein at least 50% of the 3’-extension consists of a number of C which is any one of 1 , 2, 3, 4, or 5 nt, or wherein the 3’-extension consists of C, preferably consisting of 1 , 2, or 3 C, more preferably “CC”.

[0064] According to a specific aspect, the 5’UTR is a wild-type (herein also referred to as “native”) UTR, in particular a UTR that is natural-occurring as a UTR of a certain coding sequence or gene, such as of human origin. Specifically, the 5’UTR is a native UTR of alpha-globin or beta-globin.

[0065] According to another specific aspect, the 5’UTR is a synthetic UTR, preferably a UTR sequence that is derived from a native UTR by one or more modifications, such as described herein.

[0066] According to specific examples, the 5’UTR DNA comprises or consists of, or is comprised in or composed of, any one of SEQ ID NO:1 , 3, 5, 7, 9, 11 , 13, 15, or 17, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any LO017P

[0067] -9- one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0068] Specifically, the 5’UTR RNA comprises or consists of, or is comprised in or composed of, a transcript of any one of SEQ ID NO:1 , 3, 5, 7, 9, 11 , 13, 15, or 17, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0069] According to specific examples, the 5’UTR DNA is a variant of SEQ ID NO:1 which comprises or consists of, or is comprised in or composed of, any one of SEQ ID NO:5, 7, 9, 11 , 13, 15, or 17, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full- length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0070] Specifically, the variant of SEQ ID NO:1 that comprises or consists of, or is comprised in or composed of, SEQ ID NO:5 or 9 comprises the transcription start site AGG within the 5’-terminal sequence, and one or more (preferably up to 2 or 3), point mutations in the consecutive nt sequence following the transcription start site.

[0071] Specifically, the variant of SEQ ID NO:1 that comprises or consists of, or is comprised in or composed of, SEQ ID NO:7 or 11 comprises the transcription start site GGG within the 5’-terminal sequence, and one or more (preferably up to 2 or 3), point mutations in the consecutive nt sequence following the transcription start site.

[0072] Specifically, the variant of SEQ ID NO:1 that comprises or consists of, or is comprised in or composed of, SEQ ID NO:13 comprises the 5’-terminal motif “ATT” positioned within the 5’-terminal sequence or at the 5’-terminus, a 3’-terminal motif LO017P

[0073] -10-

[0074] “ACC”, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’-terminal motif and the 3’-terminal motif.

[0075] Specifically, the variant of SEQ ID NO:1 that comprises or consists of, or is comprised in or composed of, SEQ ID NO: 15 comprises the 5’-terminal motif SEQ ID NO:51 or 53 within the 5’-terminal sequence or at the 5’-terminus, a 3’-terminal motif “ACC” positioned within the 3’-terminal sequence that is linking the 5’UTR to the Kozak sequence, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’-terminal motif and the 3’-terminal motif.

[0076] Specifically, the variant of SEQ ID NO:1 that comprises or consists of, or is comprised in or composed of, SEQ ID NO: 17 comprises the 5’-terminal motif SEQ ID NO:51 or 53 within the 5’-terminal sequence or at the 5’-terminus, an additional “AGG” transcription start site at the 5’-terminus, a 3’-terminal motif “ACC” positioned within the 3’-terminal sequence that is linking the 5’UTR to the Kozak sequence, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’- terminal motif and the 3’-terminal motif.

[0077] According to specific examples, the 5’UTR RNA comprises or consists of, or is comprised in or composed of, any one of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, or 18, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0078] According to specific examples, the 5’UTR RNA is a variant of SEQ ID NO:2 which comprises or consists, or is comprised in or composed of, of any one of SEQ ID NO: 6, 8, 10, 12, 14, 16, or 18, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full- length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence. LO017P

[0079] -11-

[0080] Specifically, the variant of SEQ ID NO:2 that comprises or consists of, or is comprised in or composed of, SEQ ID NO:6 or 10 comprises the transcription start site AGG within the 5’-terminal sequence, and one or more (preferably up to 2 or 3), point mutations in the consecutive nt sequence following the transcription start site.

[0081] Specifically, the variant of SEQ ID NO:2 that comprises or consists of, or is comprised in or composed of, SEQ ID NO:8 or 12 comprises the transcription start site GGG within the 5’-terminal sequence, and one or more (preferably up to 2 or 3), point mutations in the consecutive nt sequence following the transcription start site.

[0082] Specifically, the variant of SEQ ID NO:2 that comprises or consists of, or is comprised in or composed of, SEQ ID NO: 14 comprises the 5’-terminal motif “AULT positioned within the 5’-terminal sequence or at the 5’-terminus, a 3’-terminal motif “ACC”, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’-terminal motif and the 3’-terminal motif.

[0083] Specifically, the variant of SEQ ID NO:2 that comprises or consists of, or is comprised in or composed of, SEQ ID NO: 16 comprises the 5’-terminal motif SEQ ID NO:52 or 54 within the 5’-terminal sequence or at the 5’-terminus, a 3’-terminal motif “ACC” positioned within the 3’-terminal sequence that is linking the 5’UTR to the Kozak sequence, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’-terminal motif and the 3’-terminal motif.

[0084] Specifically, the variant of SEQ ID NO:2 that comprises or consists of, or is comprised in or composed of, SEQ ID NO: 18 comprises the 5’-terminal motif SEQ ID NO:52 or 54 within the 5’-terminal sequence or at the 5’-terminus, an additional “AGG” transcription start site at the 5’-terminus, a 3’-terminal motif “ACC” positioned within the 3’-terminal sequence that is linking the 5’UTR to the Kozak sequence, and one or more (preferably up to 2 or 3) point mutations in the consecutive nt sequence between the 5’- terminal motif and the 3’-terminal motif.

[0085] According to a specific aspect, a UTR variant may comprise one or more fragments of the respective UTR sequence provided herein. Specifically, the variant comprises one or two sequences of consecutive nucleotides of the respective full length UTR sequence, and a heterologous sequence of one or more consecutive nucleotides.

[0086] Specifically, the fragment of any of the UTR sequences provided herein comprises at least any one of 50%, or 60%, or 70%, or 80%, or 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length LO017P

[0087] -12- sequence, most preferred at least 95%, up to 100% of the respective full-length sequence.

[0088] According to a specific aspect, the 3’UTR is a wild-type (herein also referred to as “native”) nucleotide sequence, in particular a UTR that is natural-occurring as a UTR of a certain coding sequence or gene, such as of human origin.

[0089] Specifically, the 3’UTR comprises a native sequence of GPI1 , NACA, or MRS2, or a combination of any two or more of the GPI1 , NACA, or MRS2 sequences, specifically comprising a native UTR sequence of any of the foregoing, or a combination of two or more UTR sequences derived from one or more of GPU , NACA, or MRS2.

[0090] Specifically, a combination of any two or more of the GPU , NACA, or MRS2 sequences, in particular a combination of any two or more UTR sequences, comprises the respective sequences linked to each other, with or without a spacer, with or without a modification such as a modification at an intersection between the two or more UTRs, such as to remove a stop codon. A modification is e.g., a point mutation for a A->T substitution to delete a start codon (“ATG”).

[0091] A preferred 3’UTR comprises at least an MRS2 sequence and one or both of GPU or NACA.

[0092] According to a specific aspect, the 3’UTR DNA comprises or consists of a combination of two or more sequences of i), ii), or iii), and / or the 3’UTR RNA comprises or consists of a transcript of a combination of two or more sequences of i), ii), or iii), wherein i), ii) and iii) are specified as follows: i) the human Glucose 6-phosphate isomerase 1 (GPU) nucleic acid sequence SEQ ID NO: 19, a fragment of SEQ ID NO: 19 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO: 19 which is at least 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO: 19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21 , a fragment of SEQ ID NO:21 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ LO017P

[0093] -13-

[0094] ID NO:21 which is at least 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:21 ; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, a fragment of SEQ ID NO:23 comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full-length sequence, or a variant of SEQ ID NO:23 which is at least 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:23;.

[0095] Specifically, a combination of two or more 3’UTRs described herein can be used, preferably arranged in a head-to-tail orientation. The combination of two or more 3’UTRs may comprise at least two copies of a 3’UTR, or at least two different 3’UTRs.

[0096] Specifically, the 3’UTR DNA combination comprises or consists of SEQ ID NO:25,

[0097] 27, 29, or 31 , or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0098] Specifically, the 3’UTR RNA combination comprises or consists of SEQ ID NO:26,

[0099] 28, 30, or 32, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0100] An exemplary 3’UTR DNA comprises at least the MRS2 sequence comprising or consisting of SEQ ID NO:23, and either one of both of the GPI1 sequence comprising or consisting of SEQ ID NO: 19, or the NACA sequence comprising or consisting of SEQ ID NO:21.

[0101] The respective exemplary 3’UTR RNA comprises the MRS2 sequence comprising or consisting of SEQ ID NO:24, and either one of both of the GPI1 sequence LO017P

[0102] -14- comprising or consisting of SEQ ID NO:20, or the NACA sequence comprising or consisting of SEQ ID NO:22.

[0103] According to a specific aspect, the 3’UTR DNA comprises or consists of any one of SEQ ID NO: 19, 21 , 23, 25, 27, 29, or 31 , or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0104] According to a specific aspect, the 3’UTR RNA comprises or consists of any one of SEQ ID NO: 20, 22, 24, 26, 28, 30, or 32, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91%, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0105] Specifically, the 3’UTR comprises or consists of the native human GPI1 3’UTR, such as SEQ ID NO: 19 (3’UTR DNA) and SEQ ID NO:20 (the corresponding 3’UTR RNA), respectively.

[0106] Specifically, the 3’UTR comprises or consists of the native human NACA 3’UTR, such as SEQ ID NO:21 (3’UTR DNA) and SEQ ID NO:22 (the corresponding 3’UTR RNA), respectively.

[0107] Specifically, the 3’UTR is derived from the native human MRS2 sequence, and comprises a combination of two or more MRS2 sequences.

[0108] Specifically, the 3’UTR derived from the native human MRS2 sequence comprises in the 5’->3’ direction: the 27 nt 3’-terminal sequence of the native MRS2 5’UTR, followed by the Kozak sequence “GCACC” (SEQ ID NO:59), followed by “TTG” which is a motif that results from A->T substitution (at position 1 of the motif) to delete the start codon (“ATG”), which motif is followed by a 112 nt part of the MRS2 coding sequence.

[0109] Specifically, the 3’UTR derived from MRS2 comprises or consists of SEQ ID NO:23 (3’UTR DNA) and SEQ ID NO:24 (the corresponding 3’UTR RNA), respectively. LO017P

[0110] -15-

[0111] According to a specific aspect, the RNA construct (and respective DNA construct) described herein comprises a translation initiation sequence which is a Kozak sequence. For example, a Kozak sequence is operably linked to the 3’-end of the 5’UTR with or without a spacer nt sequence, such as further described herein.

[0112] Specifically, an ATG start codon (in the DNA molecule, or the corresponding AUG in the RNA molecule) is positioned immediately downstream of the Kozak sequence.

[0113] Specifically, the Kozak sequence comprises a sequence motif comprising “CACC” or “CGCC” immediately upstream of the start codon, preferably wherein the Kozak sequence may further comprise a number of C or G nucleotides upstream of the “CACC” motif. Based on sequence alignments, a Kozak consensus sequence was identified, for example, a consensus sequence which includes any of SEQ ID NO:88-92 (Kozak, 1987, Nucleic Acid Research, 1987 Volume, 15, Number 20:8125-8148).

[0114] According to specific examples, the Kozak sequence comprises or consists of any one of SEQ ID NO:57-92. A preferred Kozak sequence is any one of SEQ ID NO:88- 92, or any one of SEQ ID NO:57, 62, 69, 72, 77, 80, 84, 85, 86, or 87, preferably SEQ ID NO:57.

[0115] According to a specific aspect, the RNA molecule comprises: a) a transcription start site which consists of the nucleotide sequence AGG or GGG; and / or b) a 5’UTR which comprises the 5’-terminal transcription start site, preferably AGG or GGG, more preferably AGG; and / or c) a Kozak sequence which comprises or consists of any one of SEQ ID NO:57- 92, more preferably SEQ ID NO:57.

[0116] A preferred RNA molecule comprises a) SEQ ID NO:10; and b) SEQ ID NO:20.

[0117] Specifically, SEQ ID NO: 10 comprises the transcription start site “AGG”, the 5’UTR comprising a transcript of SEQ ID NO:1 , and the Kozak sequence SEQ ID NO:57.

[0118] Specifically, SEQ ID NO:20 comprises the 3’UTR of the native human GPI1 nucleic acid sequence.

[0119] Specifically, the respective preferred DNA molecule that can be used for transcribing the RNA molecule comprises: a) SEQ ID NO:9, and b) SEQ ID NO:19. LO017P

[0120] -16-

[0121] Specifically, the respective preferred RNA molecule comprises: a) SEQ ID NO:10, and b) SEQ ID NO:20.

[0122] Specifically, a section of the DNA molecule which serves as a template for transcribing the RNA molecule described herein, comprises a section ( / .e., a sequence of consecutive nucleotides) that comprises the 5’UTR and the Kozak sequence, which is any one of SEQ ID NO:3, 9, 11 , 15 or 17, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0123] Specifically, a section of the RNA molecule which comprises the 5’UTR and the Kozak sequence, comprises a transcript of any one of SEQ ID NO:3, 9, 11 , 15, or 17, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0124] Specifically, a section of the RNA molecule, which comprises the 5’UTR and the Kozak sequence, comprises any one of SEQ ID NO:4, 10, 12, 16, or 18, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant of any of the foregoing comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to the respective sequence.

[0125] According to a specific aspect, the DNA molecule that serves as a template for transcribing the RNA molecule, further comprises one or more 3’-tailing sequences, in particular at the 3’-end of the 3’UTR. According to a specific aspect, the RNA molecule, such as the transcribed RNA molecule, further comprises one or more 3’-tailing sequences, in particular at the 3’-end of the 3’UTR.

[0126] Specifically, the one or more 3’ tailing sequences are selected from the group consisting of a poly-A or poly-Adenine sequence, polyadenylation signal, a G- quadruplex, a poly-C sequence, a stem loop and combinations thereof.

[0127] Specifically, the 3’-tailing sequence comprises or consists of one or more poly-A sequences (also referred to as “poly-A tail”) each comprising between 10 and 300 or between 80 and 130 consecutive adenosine nucleotides.

[0128] Specifically, the 3’-tailing sequence comprises a poly-A sequence of at least 80 adenosine (A) nucleotides, wherein the poly-A sequence preferably is an un-interrupted sequence of adenosine (A) nucleotides.

[0129] According to a specific aspect, there are between 10 and 20, or 20 and 30, or 30 and 40, or 40 and 50, or 50 and 60, or 60 and 70, or 70 and 80, or 80 and 90, or 90 and 100, or 100 and 130, or 130 and 150, or 150 and 175, or 175 and 200, or 200 and 225, or 225 and 250, or 250 and 275, or 275 and 300 consecutive adenosine nucleotides in the poly-A tail.

[0130] Specifically, the 3’-tailing sequence may or may not comprise a sequence of one or more consecutive nucleotides other than adenosine (A) nucleotides.

[0131] Specifically, said sequence of one or more consecutive nucleotides containing nucleotides other than adenosine (A) nucleotides has a length of at least any one of 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 nt e.g., up to 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nt, preferably up to 30 or up to 50.

[0132] Specifically, the 3’-tailing sequence may or may not comprise two or more poly- A sequences, which are separated by an interrupting linker.

[0133] According to a specific aspect, the DNA molecule that serves as a template for transcribing the RNA molecule, further comprises a transcription start site for co- transcriptional capping.

[0134] According to a specific aspect, the RNA molecule, such as the transcribed RNA molecule, further comprises a 5’-cap.

[0135] Specifically, the 5’-cap comprises or consists of a 5’-terminal cap structure.

[0136] Specific examples of the 5’-terminal cap structure are a conventional or an endogenous cap, or analogue thereof, such as a guanine or guanine analogue thereof. -18-

[0137] Specifically, the conventional 5’-terminal cap structure comprises a cap structure that is naturally-occurring on the 5'-end of an mRNA molecule and generally consists of a 7-Methyl-guanosine 5 '-triphosphate (Gppp) which is connected via its triphosphate moiety to the 5'- end of the next nucleotide of the mRNA (i.e., the guanosine is connected via a 5' to 5' triphosphate linkage to the rest of the mRNA).

[0138] Specifically, the analogue of the conventional 5’-terminal cap structure comprises an analogue structure, such as e.g., selected from a group consisting of antireverse cap analogue (ARCA), N7,2'-0-dimethyl -guanosine (mCAP), inosine, N1 -methylguanosine, 2'fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA- guanosine, 2-azido-guanosine, N6,2'-O-dimethyladenosine, 7-methylguanosine (m7G), Cap1 , and Cap2, preferably a Cap1 structure.

[0139] Specifically, the 5’terminal cap structure can be linked to the 5’-end of the RNA, in particular at the 5’-end of the 5’UTR, by a 5'-5'-triphosphate linkage or a 5’-5’ phosphorothioate linkage.

[0140] According to a specific aspect, the DNA molecule that serves as a template for transcribing the RNA molecule, comprises a coding sequence which is a nucleic acid sequence encoding a polypeptide of interest.

[0141] According to a specific aspect, the RNA molecule further comprises a coding sequence which is a nucleic acid sequence encoding a molecule comprising an amino acid sequence of interest.

[0142] Specifically, the coding sequence encodes a peptide, polypeptide or protein (herein also referred to as “peptide, polypeptide or protein of interest”, abbreviated POI).

[0143] Specifically, the POI is a eukaryotic protein, preferably a mammalian derived or related protein such as a human protein or a protein comprising a human protein sequence, or a bacterial protein or bacterial derived protein

[0144] Preferably, the POI is a therapeutic or diagnostic protein or product, such as functioning in mammals e.g., a human pharmaceutical (e.g., therapeutic or prophylactic), or diagnostic. According to a specific aspect, the POI is selected from the group consisting of an antigen-binding protein, an enzyme, a peptide, a protein antibiotic, a toxin fusion protein, a structural protein, a regulatory protein, a cell surface receptor, a vaccine antigen (in particular a peptide or epitope of a vaccine antigen), a hormone, a growth factor, a cytokine, and a blood clotting or coagulation factor.

[0145] Specifically, the POI is produced from the coding sequence ex vivo or in vivo, such as to allow manufacturing of pharmaceutical or diagnostic products, or in vivo LO017P

[0146] -19- applications. In particular, mRNA can be used for in vivo applications, e.g., mRNA encoding vaccine antigens; mRNA therapy, or mRNA for in vivo CAR-T engineering.

[0147] Specifically, the RNA molecule is coding for a vaccine antigen, such as suitably used for mRNA vaccines.

[0148] Specifically, the RNA molecule is coding for an antigen suitably used for engineering mRNA-based CAR-T cells ex vivo or in vivo. CAR-T can be engineered from the coding sequence ex vivo or in vivo. For example, in v / tro-transcribed mRNA encoding CAR or TCR can be delivered to circulating T cells, effectively reprogramming them in vivo to combat diseases.

[0149] In particular, one or more mRNA molecules can be used to express a POI such as a POI which is a complex of POI fragments. For examples, more than one mRNA molecules can be used which encode POI fragments.

[0150] In specific cases, the POI is a multimeric protein, specifically a dimer or tetramer.

[0151] Specifically, the antigen-binding protein is an antibody molecule, preferably an antibody molecule comprising one or more epitope binding fragments of a full-length antibody, preferably wherein said one or more epitope binding fragments are Fab, Fab', F(ab')2, Fv, scFv fragments, or single domain antibodies;

[0152] Specifically, the antibody molecule is a full-length antibody, an antibody comprising one or more epitope binding fragments of a full-length antibody, or a bispecific or multi-specific antibody comprising one or more of said fragments.

[0153] Specifically, said one or more epitope binding fragments of a full-length antibody are Fab, Fab', F(ab')2, Fv, or scFv fragments, or single domain antibodies.

[0154] Specifically, the antigen-binding protein is selected from the group consisting of: a) antibodies or antibody fragments, such as any of chimeric antibodies, humanized antibodies, bi-specific antibodies, Fab, Fd, scFv, diabodies, triabodies, Fv tetramers, minibodies, single-domain antibodies like VH, VHH, IgNARs, or V-NAR, in particular camelid VHH; b) antibody mimetics, such as Adnectins, Affibodies, Affilins, Affimers, Affitins, Alphabodies, Anticalins, Avimers, DARPins, Fynomers, Kunitz domain peptides, Monobodies, or NanoCI_AMPS; or c) fusion proteins comprising one or more immunoglobulin-fold domains, antibody domains or antibody mimetics.

[0155] A specific POI comprises or consists of an antibody fragment, in particular an antibody fragment comprising an antigen-binding domain. Among specific POIs are Fab, LO017P

[0156] -20-

[0157] Fd, single-chain variable fragment (scFv), or engineered variants thereof such as for example Fv dimers (diabodies), Fv trimers (triabodies), Fv tetramers, or minibodies and single-domain antibodies like VH, VHH, IgNARs, or V-NAR, or any protein comprising an immunoglobulin-fold domain. Further antigen-binding molecules may be selected from antibody mimetics, or (alternative) scaffold proteins such as e.g., engineered Kunitz domains, Adnectins, Affibodies, Affiline, Anticalins, or DARPins.

[0158] Specifically, the POI is a DNA processing and / or genome editing enzyme, process enzyme or metabolic enzyme.

[0159] Specific examples of a POI comprise or consist of antibodies or antibody fragments or antibody-related structures, such as heavy and / or light chains, single-chain antibodies, nanobodies, CRISPR / Cas nucleases, cancer epitopes (in particular neoepitopes), epitopes of pathogens causing infectious diseases, transcription factors, mammalian (in particular human) enzymes such as used for enzyme replacement.

[0160] Specifically, the POI is a pharmaceutically active peptide, polypeptide or protein.

[0161] For example, the POI is an active substance of RNA-based therapy such as e.g., gene therapy, replacement therapy, or comprises an epitope for inducing an immune response against an antigen in a subject. The antigen or epitope is optionally derived from or is a protein of a pathogen, an immunogenic variant of the protein, or an immunogenic fragment of the protein or the immunogenic variant thereof. Specifically, the pathogen is a pathogen causing an infectious disease.

[0162] According to a specific aspect, the mRNA comprises an ORF encoding a POI such as a pharmaceutically active POI, in particular a POI whose expression is active in preventing or treating a disease. Specifically, the POI is functional in treating a disease such as cancer, auto-immune disease, or a disease caused by lack of function mutations in enzymes and other proteins.

[0163] Specific POI examples are vaccine antigens (in particular peptides or epitopes) of a pathogen, or those which comprise or mimic a POI of a pathogen.

[0164] Specific POI examples are vaccine antigens (in particular peptides or epitopes) of a virus, or those which comprise or mimic a viral POI.

[0165] A specific POI example is a vaccine antigen comprising a POI originating from a virus such as a virus selected from: a) Coronaviridae (P-coronavirus, such as SARS-CoV-2, MERS-CoV, SARS-CoV- 1 , HCoV-OC43, HCoV-HKU1 ; or a-coronavirus, such as HCoV-NL63, HCoV-229E or PEDV, including naturally-occurring variants or mutants of any of the foregoing); LO017P

[0166] -21- b) Adenoviridae (such as Adenoviruses or human Adenoviruses e.g., HAdVB, HAdVC, or HAdVD); c) Paramyxoviridae (such as RSV or human RSV e.g., hRSV subtype A or B); or d) Orthomyxoviridae (such as influenza viruses or human influenza viruses, preferably influenza virus A (IVA), such as H1 N1 , H3N3, or H5N1 , or influenza virus B (IVB), or influenza virus C (IVC), or influenza virus D (IVD)).

[0167] According to a specific example, the POI may originate from a coronavirus and may comprise a spike protein or a domain thereof e.g., an RBS domain. A specific example refers to the spike protein of SARS-CoV-2 or one or more epitopes thereof.

[0168] According to another specific example, the POI may originate from an influenza virus and may comprise a hemagglutinin and / or a neuraminidase, or one or more domains thereof such as comprising one or more epitopes of hemagglutinin.

[0169] According to another specific example, the POI may comprise or originate from Cystic Fibrosis Transmembrane Conductance Regulator (CTFR).

[0170] Specifically, the coding sequence comprises a functional POI coding gene.

[0171] Specifically, the POI is a heterologous POI.

[0172] Specifically, the coding sequence is not naturally-occurring ( / .e., that is not found in nature) with any one or both of the 5’UTR and 3’UTR.

[0173] Specifically, the coding sequence is heterologous to the 3’UTR.

[0174] According to a specific aspect, a) where the 3’UTR is of GPI1 , in particular comprising or consisting of SEQ ID NO: 19, or a fragment of SEQ ID NO: 19 comprising at least 50% of the full-length sequence (in particular, at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full- length sequence), or a variant of SEQ ID NO: 19 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:19, the coding sequence is not a GPI1 coding sequence; b) where the 3’UTR is of NACA, in particular comprising or consisting of SEQ ID NO:21 , or a fragment of SEQ ID NO:21 comprising at least 50% of the full-length sequence (in particular, at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full- length sequence), or a variant of SEQ ID NO:21 which is at least any one of 90%, or LO017P

[0175] -22-

[0176] 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:21 , the coding sequence is not a NACA coding sequence; c) where the 3’UTR is of MRS2, in particular comprising or consisting of SEQ ID NO:23, or a fragment of SEQ ID NO:23 comprising at least 50% of the full-length sequence (in particular, at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the full-length sequence, most preferred at least 95%, up to 100% of the full- length sequence), or a variant of SEQ ID NO:23 which is at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% identical to SEQ ID NO:23, the coding sequence is not a MRS2 coding sequence.

[0177] Specifically, the coding sequence is not operably linked to any one or both of the 5’UTR and 3’UTR in a wild-type nucleic acid construct.

[0178] Specifically, any one or both of the 5’UTR and 3’UTR are heterologous UTRs, in particular UTRs which are not naturally-occurring with the coding sequence.

[0179] Specifically, any one or both of the 5’UTR and 3’UTR is not operably linked to the coding sequence in a wild-type nucleic acid construct.

[0180] According to a specific aspect, the RNA molecule comprises at least one modified nucleoside.

[0181] In one example, the one or more RNAs comprises at least one chemically modified nucleotide.

[0182] Specifically, the RNA molecule comprises one or more modified nucleosides which are independently selected from pseudouridine (y), N1-methyl-pseudouridine (m1 i ), and 5-methyl -uridine (m5U).

[0183] According to a specific aspect, the RNA molecule is an mRNA.

[0184] Specifically, the mRNA encodes one or more POL

[0185] Specifically, the RNA is a self-replicating RNA.

[0186] A specific example of the RNA molecule comprises the 5’UTR as described herein.

[0187] According to a specific aspect, the invention further provides for an RNA molecule (e.g., an mRNA) comprising the nucleotide sequence which is any one of SEQ ID NO:6, 8, 10, 12, 14, 16, or 18 as a 5’UTR, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full- length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the LO017P

[0188] -23- respective full-length sequence, or a variant thereof comprising at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to any of the foregoing, preferably wherein the variant comprises the stretch of position 4-34 of SEQ ID NO:6. Specifically, the mRNA molecule comprises said 5’UTR and any suitable 3’UTR, and optionally one or more of the elements of an mRNA described herein.

[0189] According to a specific aspect, the invention further provides for a DNA molecule or a DNA template for transcribing the RNA molecule (e.g., an mRNA) described herein. Specific examples refer to a DNA molecule or template which comprises the nucleotide sequence SEQ ID NO:5, 7, 9, 11 , 13, 15, or 17 as a 5’UTR, or a fragment of any of the foregoing comprising at least any one of 50%, 60%, 70%, 80%, 90%, 95%, up to 100% of the respective full-length sequence, preferably at least any one of 80%, 90%, or 95%, up to 100% of the respective full-length sequence, most preferred at least 95%, up to 100% of the respective full-length sequence, or a variant thereof comprising at least at least any one of 90%, or 91 %, or 92%, or 93%, or 94%, or 95%, or 96%, or 97%, or 98%, or 99% sequence identity to any of the foregoing. Specifically, the DNA molecule or template comprises said 5’UTR and any suitable 3’UTR, and optionally one or more of the elements of a DNA molecule described herein which is suitable for transcribing the RNA molecule (e.g., an mRNA).

[0190] According to a specific example, the RNA molecule comprises the 5’UTR and any 3’UTR sequence, such as e.g., a 3’UTR as described herein, in particular a 3’UTR listed in Figure 7. Specifically, the RNA molecule comprises any of the 5’UTR and 3’UTR combinations as described herein.

[0191] According to a specific example, the RNA molecule comprises the 3’-tailing sequence described herein and / or the 5’-cap as described herein.

[0192] Specifically, the RNA molecule (e.g., the mRNA) comprises in operable linkage (in particular in the 5’->3’ direction) at least: a 5’-cap; a 5’-terminal transcription start site (preferably “AGG” or “GGG”, more preferably “AGG”), the 5’UTR described herein; a Kozak sequence; a coding sequence (e.g., an ORF comprising the coding sequence); a 3’UTR; and a poly-A sequence.

[0193] The invention further provides for a DNA molecule which comprises a nucleotide sequence suitable for transcription of an RNA molecule described herein, preferably in the form of a plasmid.

[0194] Specifically, the DNA molecule comprises a template for transcribing the RNA. LO017P

[0195] -24-

[0196] Specifically, the DNA molecule comprises the 5’UTR as described herein.

[0197] According to a specific example, the DNA molecule comprises the 5’UTR and 3’UTR combination as described herein.

[0198] According to a specific example, the DNA molecule comprises the 3’-tailing sequence described herein and / or the 5’-cap as described herein.

[0199] Specifically, the DNA molecule comprises in operable linkage, a promoter operably linked to a nucleic acid comprising a 5’ transcription start site; the 5’UTR as described herein; a Kozak sequence; a coding sequence (e.g., an ORF comprising the coding sequence); a 3’UTR; and a poly-A sequence.

[0200] Specifically, the DNA molecule is provided for transcription of RNA, preferably in vitro transcription of RNA, in particular mRNA. The respective description provided herein, thus, refers to the use of the DNA molecule described herein for transcription of RNA, preferably in vitro transcription of RNA.

[0201] The DNA molecule described herein is preferably a closed circular molecule prior to cleavage and a linear molecule after cleavage. Preferably, cleavage is carried out using a restriction cleavage site which is preferably a restriction cleavage site for a type IIS restriction endonuclease.

[0202] Specifically, the DNA molecule is an expression vector or plasmid, such as an in vitro transcription (IVT) vector, in particular a circular IVT vector.

[0203] Prior to in vitro transcription, circular IVT vectors (or plasmids) can be linearized e.g., downstream of the poly-A sequence (e.g., the poly-A cassette) by a suitable enzyme such as a type II restriction enzyme. The poly-A cassette corresponds to the poly-A sequence in the transcript. The linearized vector (or plasmid) can then be used as template for in vitro transcription, and the resulting transcript comprises the RNA construct or molecule as further described herein.

[0204] The invention further provides for a method of producing an RNA molecule described herein, by in vitro transcription (IVT) of the DNA molecule described herein, optionally after linearization.

[0205] The invention also relates to methods of obtaining RNA by transcribing the DNA molecule described herein, and RNA obtainable by the transcription of the DNA molecule described herein.

[0206] Specifically, the method of in vitro transcription produces an RNA molecule (e.g., an mRNA) with sufficient mRNA integrity and which is sufficiently pure. LO017P

[0207] -25-

[0208] The integrity of an mRNA molecule can be determined by Capillary gel electrophoresis (CGE) or HPLC, such as described in “Analytical Procedures for Quality of mRNA Vaccines and Therapeutics” USP draft guideline, third edition (2024), Biologies Monograph 3 - Complex Biologies & Vaccines. Integrity is typically measured as the degree of full-length mRNA in an mRNA preparation (w / w). Specifically, the mRNA integrity as measured by capillary gel electrophoresis is at least 50%, or 60%, or at least any one of 70%, 80%, or 90%, and at least any one of 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40%, as measured by HPLC.

[0209] Specifically, the mRNA is a functional mRNA, in particular an mRNA which has a function to translate a coding sequence comprised in the mRNA.

[0210] The capping efficiency is a measure of determining functionality of mRNA and can be determined by RNaseH digestion and analysis of the 5’ fragments such as described in WO 2014 / 152659A1. Specifically, the degree of capping is at least any one of 80%, 85%, 90%, or 95%. Specifically, an isolated mRNA provided herein comprises a capping efficiency of least any one of 80%, 85%, 90%, or 95%.

[0211] The RNA molecule is conveniently purified by a suitable purification method such as to remove impurities resulting from the production method e.g., resulting from incorrect transcription.

[0212] Specifically, the purification method comprises a method step of purifying using magnetic beads, such as the binding of RNA to the surface of the beads in the presence of a salt concentration (in particular a high concentration) of salt and molecular crowding agents, followed by washing of the beads with a solution containing ethanol (in particular a high percentage of ethanol) and finally elution in water.

[0213] Specifically, the purification method comprises a method step of oligo(dT) affinity purification, such as described by Levison et al. (“Recent developments of magnetic beads for use in nucleic acid purification”, Journal of Chromatography A 1998, Volume 816, Issue 1 , Pages 107-111). The oligo(dT) affinity purification may increase the mRNA integrity and reduce the amount of dsRNA in an mRNA preparation. Specifically, the amount of dsRNA in a purified mRNA preparation is less than 2%, or less than any one of 1.9, 1.8, 1.7, 1.6, or 1.5% (%, w / w). Specifically, the amount of dsRNA in a magnetic beads purified mRNA preparation and / or oligo-dT purified mRNA preparation is less than 2%, or less than any one of 1 .9, 1 .8, 1 .7, 1 .6, or 1 .5% (%, w / w).

[0214] Specifically, an isolated mRNA provided herein comprises less than 2% dsRNA, or less than any one of 1 .9, 1 .8, 1 .7, 1 .6, or 1 .5%. LO017P

[0215] -26-

[0216] Specifically, the purification method comprises a method step of hydrophobic interaction chromatography (HIC) such as described in as described in Gagnon P. Purification of Nucleic Acids: A handbook for purification of plasmid DNA and mRNA for gene therapy and vaccines. BIA Separations, Ajdovscina, 2020; 136. The HIC purification may increase the mRNA integrity and may reduce the amount of dsRNA in an mRNA preparation.

[0217] According to specific aspect, an mRNAs are electroporated or transfected into either a host cell to translate mRNAs inside cells to produce a vaccine antigen or therapeutic POI, which can e.g., be used for treating or preventing a disease or disorder. A preferred approach for delivering the mRNA into the cells is transfection by lipid nanoparticles.

[0218] Specific examples of host cells are T cells, primary NK cells, NK-92, DC, macrophages or other types of immune cells or mammalian cells, such as liver and lung cells.

[0219] Specifically, the RNA molecule is provided for in vitro (or in vivo) transfecting a host cell, preferably cells for gene editing, chimeric antigen receptor expression in T cells, cancer cells, antigen-presenting cells, monocytes, dendritic cells, T cells, macrophages, lung or liver cells. The respective description provided herein, thus, refers to the use of the RNA molecule described herein for in vitro transfecting a host cell.

[0220] The invention may be utilized, for example, for increasing expression of recombinant proteins in cellular transcription and expression. Specifically, a DNA molecule described herein which is e.g., an expression vector, can be used to produce a recombinant protein. Suitable methods comprise transcribing a recombinant nucleic acid encoding the recombinant protein from the DNA molecule, and expressing the recombinant protein in a cell-based system.

[0221] Specifically, the nucleic acid molecules can also be used for gene therapy applications. A nucleic acid molecule described herein may be a gene therapy vector and used for expression of a transgene. Specifically, any nucleic acid (DNA / RNA) - based vector systems (for example plasmids, adenoviruses, poxvirus vectors, influenza virus vectors, alphavirus vectors, and the like) may be used. Cells can be transfected with these vectors in vitro, for example in lymphocytes or dendritic cells, or else in vivo by direct administration.

[0222] An RNA molecule described herein (e.g., obtained using a nucleic acid molecule described herein as a transcription template) may be employed, for example, for LO017P

[0223] -27- transient expression of genes, with possible fields of application being RNA-based vaccines which are transfected into cells in vitro or administered directly in vivo, transient expression of functional recombinant proteins in vitro, for example in order to initiate differentiation processes in cells or to study functions of proteins, and transient expression of functional recombinant proteins such as erythropoietin, hormones, coagulation inhibitors, etc., in vivo, in particular as pharmaceuticals.

[0224] According to a specific aspect, the RNA molecule described herein may be used for transfecting antigen-presenting cells, in particular where the RNA molecule is used as a tool for delivering the antigen to be presented and for loading antigen-presenting cells, wherein said antigen or an epitope thereof to be presented corresponds to a POI expressed from said RNA or being derived therefrom, in particular by way of intracellular processing such as cleavage. Specifically, the antigen or epitope to be presented is, for example, a fragment of the POI expressed from the RNA. Such antigen-presenting cells may be used for stimulating T cells, in particular CD4+and / or CD8+T cells.

[0225] According to a specific aspect, the RNA molecule described herein is used in a method of medical treatment wherein a host cell is transfected in vitro or in vivo.

[0226] Specifically, the method is conducted in vitro i.e., the cells do not form part of an organ, a tissue and / or an organism of a subject. Specifically, the cells are an ex vivo cell culture.

[0227] Specifically, the present disclosure provides an in vitro method of transfecting cells, comprising adding the RNA molecule described herein to cells; and incubating the mixture of the RNA molecule and cells for a sufficient amount of time. Specifically, where the RNA molecule encodes a pharmaceutically active POI, the mixture of the composition and cells is incubated for a time sufficient to allow the expression of the pharmaceutically active protein. Specifically, the sufficient amount of time is at least one hour e.g., at least any one of 1 , 2, 3, 4, 5 ,6 ,7 ,8, 9, 10, 11 , or 12 hours, or even longer up to e.g., 96, 72, 48, 36, 24, or 12 hours.

[0228] Specifically, the present disclosure provides an in vivo method of transfecting cells i.e., the cells form part of an organ, a tissue and / or an organism of a subject.

[0229] According to a specific aspect, the RNA molecule described herein is provided for medical use.

[0230] Specifically, the RNA molecule is provided for use in mRNA-based therapy.

[0231] Specifically, the RNA molecule is used for vaccination e.g., for vaccinating a mammal, preferably a human subject. LO017P

[0232] -28-

[0233] Further provided herein is a pharmaceutical composition comprising the RNA molecule described herein, which comprises one or more of pharmaceutically acceptable carriers, diluents and excipients.

[0234] Specifically, the RNA molecule is mixed with lipid components to produce lipid nanoparticles (LNPs) with RNA encapsulated.

[0235] Specifically, the pharmaceutical composition is an immunogenic composition.

[0236] According to a specific aspect, the pharmaceutical composition is provided for use in mRNA-based therapy, in particular for vaccination e.g., for vaccinating a subject such as a mammal, preferably a human subject.

[0237] Specifically, the pharmaceutical composition is provided for use as a vaccine.

[0238] According to a specific aspect, the medical use comprises treatment of a subject by delivering the RNA molecule described herein to cells of the subject, the method comprising administering to the subject the pharmaceutical composition described herein.

[0239] According to a specific aspect, the medical use comprises treatment of a subject by delivering a therapeutic POI to a subject, the method comprising administering to the subject the RNA molecule described herein, wherein the RNA molecule encodes the therapeutic POI.

[0240] According to a specific aspect, the medical use comprises treating or preventing a disease or disorder in a subject, the method comprising administering to the subject the RNA molecule described herein, wherein delivering the RNA molecule to cells of the subject is beneficial in treating or preventing the disease or disorder. In a related aspect, the present disclosure relates to the RNA molecule described herein for use in a method for treating or preventing a disease or disorder in a subject, wherein delivering the RNA molecule to cells of the subject is beneficial in treating or preventing the disease or disorder.

[0241] According to a specific aspect, the medical use comprises treating or preventing a disease or disorder in a subject, the method comprising administering to the subject the RNA molecule described herein, wherein the RNA molecule encodes a therapeutic POI and wherein delivering the therapeutic POI to the subject is beneficial in treating or preventing the disease or disorder. In a related aspect, the present disclosure relates to the RNA molecule described herein for use in a method for treating or preventing a disease or disorder in a subject, wherein the RNA molecule encodes a therapeutic POI LO017P

[0242] -29- and wherein delivering the therapeutic POI to the subject is beneficial in treating or preventing the disease or disorder.

[0243] Specifically, the subject is a mammal, such as a human.

[0244] FIGURES

[0245] Figure 1 : Schematic overview of DNA sequence elements in synthetic genes.

[0246] Figure 2: Schematic representation of firefly luciferase expression profile and parameter extracted from it.

[0247] Figure 3: Expression level of the firefly luciferase constructs after 24 h in different cell-lines was normalized to the UTR17 construct, ref. Example 1.

[0248] Figure 4: Expression levels of the firefly luciferase constructs after 24 h in different cell-lines were normalized to UTR1 construct, ref. Example 2.

[0249] Figure 5: HPLC chromatograms for mRNA integrity analysis that reveal the formation of a truncated species in a front peak compared to the full-length product in the main peak for firefly luciferase mRNAs linked to UTR1 , UTR14-6, or UTR14-10.

[0250] Figure 6: mRNA integrity analysis for firefly luciferase mRNAs linked to UTR1 , UTR14-6, or UTR14-10 during incubation at 2-8°C, 22°C, and 40°C.

[0251] Figure 7: Sequences described herein.

[0252] Figure 8: Structure of a typical human mRNA molecule; source: mRNA technologies Insight Report, European Patent Office October 2023.

[0253] 5’-cap: A sub-structure in the mRNA molecule which protects the molecule from decomposition by specific enzymes. It plays a fundamental role in putting the mRNA molecule in a position to be read in the ribosome and to have the genetic information translated during protein synthesis. The 5’-cap remains untranslated during the protein synthesis.

[0254] 5’ untranslated region (5’UTR): Sub-structure in the mRNA which precedes the coding region and remains untranslated in most cases during protein synthesis in the ribosome. It plays an important role in regulating the process of translating genetic information into the protein. It also supports the ribosome in recognizing the mRNA molecule and helps to modify the mRNA after the translation process is completed.

[0255] Coding region: Sub-structure that contains a representation of the genetic information for the production of a protein.

[0256] 3’ untranslated region (3'UTR): Sub-structure which follows the coding region and generally remains untranslated during protein synthesis in the ribosome. It has several regulatory functions for the copying of the genetic information in the cell nucleus, the LO017P

[0257] -30- transportation of the mRNA molecule from the cell nucleus to the ribosome and for the regulation of the translation of genetic information during protein synthesis.

[0258] Poly(A) tail: This part of the mRNA molecule provides for the export of the molecule from the cell nucleus, for the translation process and for the stability of the mRNA molecule. It also protects the mRNA molecule from degradation and has an important effect on the lifespan of the molecule: the poly(A) tail is gradually shortened over time and, below a certain threshold, the molecule may be degraded enzymatically.

[0259] DETAILED DESCRIPTION OF THE INVENTION

[0260] Unless indicated or defined otherwise, all terms used herein have their usual meaning in the art, which will be clear to the skilled person. Reference is for example made to the standard handbooks, such as Sambrook et al., 2012, Molecular Cloning: A Laboratory Manual, volumes 1-4, Cold Spring Harbor Press, NY); Lewin, "Genes IV", Oxford University Press, New York, (1990), and Janeway et al., "Immunobiology" (5th Ed., or more recent editions), Garland Science, New York, 2001 , Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Vega et al., Gene Targeting, CRC Press, Ann Arbor Mich. (1995), and Vectors: A Survey of Molecular Cloning Vectors and Their Uses, Butterworths, Boston Mass. (1988).

[0261] As used herein, the terms “a”, “an” and “the” are used herein to refer to one or more than one i.e., to at least one. The terms “comprise”, “contain”, “have” and “include” as used herein can be used synonymously and shall be understood as an open definition, allowing further members or parts or elements. “Consisting” is considered as a closest definition without further elements of the consisting definition feature. Thus “comprising” is broader and contains the “consisting” definition.

[0262] The term “about” as used herein refers to the same value or a value differing by + / -10% or + / -5% of the given value.

[0263] Specific terms as used throughout the specification have the following meaning.

[0264] The term “coding sequence” as used herein shall mean the region (herein also referred to as “section”, “part” or “portion”) of a nucleic acid (e.g., a DNA or RNA) that codes for a POL In the context of the present disclosure, the term “coding sequence” which is comprised in an RNA molecule shall mean the sequence which can direct the assembly of the encoded amino acid sequence to produce a POI during the process of translation in an appropriate environment, such as within target cells.

[0265] The term "expression" as used herein is defined as the transcription and / or translation of a particular nucleotide sequence. LO017P

[0266] -31-

[0267] The term “expression cassette” is herein understood to refer to a nucleic acid molecule, which contains a desired coding sequence, and control sequences in operable linkage, so that hosts transfected with these molecules incorporate the respective sequences and are capable of producing the encoded POL The term “expression” is used herein for expression of a nucleic acid sequence, or for expression of the respective amino acid sequence.

[0268] With respect to RNA, the term "expression" or "translation" relates to the process in the ribosomes of a cell by which a strand of mRNA directs the assembly of a sequence of amino acids to make a POL

[0269] The term “heterologous” as used herein with respect to a nucleotide sequence, construct such as an expression cassette, amino acid sequence or POI, in the context of a heterologous nucleotide sequence comprised in a nucleic acid construct shall refer to a nucleotide sequence that is not found in the same relationship to the nucleic acid construct in nature i.e., wild-type. Any recombinant or otherwise artificial nucleotide sequence is understood to be heterologous. A nucleic acid molecule which comprises a heterologous nucleic acid sequence is herein also referred to as “recombinant” or “synthetic”.

[0270] An example of a heterologous nucleotide sequence is not natively associated with other elements of the nucleic acid construct with which it is operably linked. A specific example of a heterologous nucleotide sequence is a POI encoding nucleotide sequence that is operably linked to a 5’UTR and / or a 3’UTR, to which a naturally-occurring POI coding sequence is not normally operably linked.

[0271] The term “in vitro transcription” or “IVT” as used herein shall mean the transcription (i.e., the generation of RNA from a DNA) is conducted in a cell-free manner. In particular, IVT does not use living or cultured cells but rather the transcription machinery extracted from cells e.g., cell lysates or the isolated components thereof, including an RNA polymerase. Particular examples of RNA polymerases are the T7, T3, and SP6 RNA polymerases.

[0272] A series of in vitro transcription kits is commercially available, e.g., from Thermo Fisher Scientific (such as TranscriptAid™ T7 kit, MEGAscript® T7 kit, MAXIscript®), New England BioLabs Inc. (such as HiScribe™ T7 kit, HiScribe™ T7 ARCA mRNA kit), Promega (such as RiboMAX™, HeLaScribe®, Riboprobe® systems), Jena Bioscience (such as SP6 or T7 transcription kits), and Epicentre (such as AmpliScribe™). LO017P

[0273] -32-

[0274] As described herein, an RNA molecule such as an mRNA can be produced by in vitro of an appropriate DNA template. Specifically, the in vitro transcription is controlled by a promoter which can be any promoter for any RNA polymerase e.g., a T7 or SP6 promoter. A DNA template for in vitro transcription may be obtained by cloning of a nucleic acid, in particular cDNA, and introducing it into an appropriate vector for in vitro transcription. The cDNA may be obtained by reverse transcription of RNA.

[0275] The term human Glucose 6-phosphate isomerase 1 (GPI1) as used herein shall refer to Homo sapiens glucose-6-phosphate isomerase (GPI) e.g., transcript variant 2 NM_000175 (NCBI Reference Sequence).

[0276] The term human nascent polypeptide associated complex subunit alpha (NACA) as used herein shall refer to Homo sapiens nascent polypeptide associated complex subunit alpha e.g., transcript variants NM_001113201.3, NM_001113202.2, NM_001 113203.3, NM_001320193.2, NM_001320194.2, NM_001365896.1 , NM-005594.6).

[0277] The term human magnesium transporter (MRS2) as used herein shall refer to the Homo sapiens magnesium transporter MRS2, e.g., transcript variant 2, >NM_020662.4.

[0278] The term human alpha-globin as used herein shall refer to the human alpha globin (HBA1), Homo sapiens HBA1 hemoglobin subunit alpha 1 e.g.,

[0279] >NC_000016.10: 176680-177522 Homo sapiens chromosome 16, GRCh38.p14 Primary Assembly. The native 5’ UTR of HBA1 is identical with the 5’UTR of the Homo sapiens beta globin subunit alpha 2 (HBA2) e.g., >NC_000016.10: 172876-173710 Homo sapiens chromosome 16, GRCh38.p14 Primary Assembly.

[0280] The term “linked” as used herein shall refer to the joining together of two or more elements or components or domains.

[0281] The term “Kozak sequence” as used herein shall refer to a nucleic acid sequence upstream of the start codon or including the start codon of a coding sequence which mediates translation of an RNA molecule. Typically, the Kozak sequence is a Kozak consensus sequence. The Kozak sequence can be characterized as a specific set of nucleotides that serves as the initiation site where protein translation begins in eukaryotic mRNA produced from transcription. The Kozak sequence ensures accurate translation of the protein in terms of ribosome assembly and translation initiation. (Kozak, 1987, Nucleic Acid Research, 1987 Volume, 15, Number 20:8125-8148).

[0282] A Kozak sequence typically extends from approximately position -10 to position +6, where +1 is assigned to the adenine of the start codon. The Kozak sequence can LO017P

[0283] -33- be an unmodified wild-type Kozak sequence such as naturally-occurring with a coding sequence that is endogenous to a cell. The Kozak sequence can be an optimized Kozak sequence such as e.g., to increase translational efficiency

[0284] The term “nucleic acid” as used herein shall refer to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), combinations thereof, and modified forms thereof. The term comprises genomic DNA, cDNA, mRNA, recombinantly produced and chemically synthesized molecules. A nucleic acid may be present as a single-stranded or doublestranded and linear or covalently circularly closed molecule. A nucleic acid can be isolated.

[0285] The term “isolated” or “isolation” as used herein with respect to a nucleic acid molecule shall refer to such compound that has been sufficiently separated from the environment with which it would naturally be associated, so as to exist in “purified” or “substantially pure” form. Yet, “isolated” does not necessarily mean the exclusion of artificial or synthetic mixtures with other compounds or materials, or the presence of impurities that do not interfere with the fundamental activity, and that may be present, for example, due to incomplete purification. Isolated compounds can be further formulated to produce preparations thereof, and still for practical purposes be isolated.

[0286] With reference to nucleic acids described herein, the term “isolated nucleic acid” is sometimes used. This term, when applied to DNA, refers to a DNA molecule that is separated from sequences with which it is immediately contiguous in the naturally occurring genome of the organism in which it originated. For example, an “isolated nucleic acid” may comprise a DNA molecule inserted into a vector, such as a plasmid or expression vector. When applied to RNA, the term “isolated nucleic acid” refers primarily to an RNA molecule encoded by an isolated DNA molecule as defined above. Alternatively, the term may refer to an RNA molecule that has been sufficiently separated from other nucleic acids with which it would be associated in its natural state ( / .e., in cells or tissues). An “isolated nucleic acid” (either DNA or RNA) may further represent a molecule produced directly by biological or synthetic means and separated from other components present during its production.

[0287] An isolated nucleic acid can be produced by amplification in vitro, for example via polymerase chain reaction (PCR) for DNA or in vitro transcription for RNA. Isolated nucleic molecules can be produced recombinantly by cloning, or by purifying, for example, by cleavage and separation by gel electrophoresis, or by synthesizing e.g., by chemical synthesis. LO017P

[0288] -34-

[0289] The term “DNA” as used herein shall refer to a nucleic acid molecule which includes nucleotides that contains deoxyribose. The term shall encompass double stranded DNA, single stranded DNA, isolated DNA, and in particular synthetic DNA, recombinantly produced DNA, and modified DNA that differs from naturally-occurring DNA by one or more modification such as substitution, insertion or deletion of one or more nucleotides. Such modifications may refer to addition of non-nucleotide material to internal DNA nucleotides or to the end(s) of DNA. Nucleotides in DNA may be standard nucleotides (A, T, C, G), and non-standard nucleotides, such as chemically synthesized nucleotides or ribonucleotides.

[0290] DNA may be recombinant DNA and may be obtained by cloning of a nucleic acid, in particular cDNA. The cDNA may be obtained by reverse transcription of RNA.

[0291] The term “DNA template” as used herein shall refer to a DNA molecule comprising a nucleic acid sequence encoding an RNA transcript (e.g., an mRNA transcript) which can be synthesized by in vitro transcription. The template DNA is used as template for in vitro transcription in order to produce the mRNA transcript encoded by the template DNA. The template DNA typically comprises all elements necessary for in vitro transcription, particularly a promoter element for binding of a DNA-dependent RNA polymerase, such as, e.g., T3, T7 and SP6 RNA polymerases, which is operably linked to the DNA sequence encoding a desired mRNA transcript. Furthermore, the template DNA may comprise primer binding sites 5' and / or 3' of the DNA sequence encoding the mRNA transcript to determine the identity of the DNA sequence encoding the mRNA transcript, e.g., by PCR or DNA sequencing. The “DNA template” in the context of the present disclosure may be a linear or a circular DNA molecule, in particular a DNA vector, such as a plasmid DNA, which comprises a nucleic acid sequence encoding the desired mRNA transcript.

[0292] The term “RNA” as used herein shall refer to a nucleic acid molecule which includes ribonucleotide residues. The term shall encompass double stranded RNA, single stranded RNA, isolated RNA, synthetic RNA, recombinantly produced RNA, and modified RNA that differs from naturally occurring RNA by one or more modification such as substitution, insertion or deletion of one or more nucleotides. Such alterations may refer to addition of non-nucleotide material to internal RNA nucleotides or to the end(s) of RNA. Such modifications may refer to addition of non-nucleotide material to internal RNA nucleotides or to the end(s) of RNA. Nucleotides in RNA may be standard LO017P

[0293] -35- nucleotides (A, U, C, G), and non-standard nucleotides, such as chemically synthesized nucleotides or ribonucleotides.

[0294] The term “RNA” particularly refers to RNA molecules such as mRNA, selfamplifying RNA (saRNA), tRNA, ribosomal RNA (rRNA), small nuclear RNA (snRNA), inhibitory RNA (such as antisense ssRNA, small interfering RNA (siRNA), or microRNA (miRNA)), activating RNA (such as small activating RNA) and immunostimulatory RNA (isRNA). A preferred RNA molecule described herein is an mRNA or self-amplifying RNA (saRNA).

[0295] The term “sequence identity” of a variant as compared to a reference (also referred to as “parent” nucleotide) indicates the degree of identity of two or more sequences. Two or more nucleotide sequences may have the same or conserved base pairs at a corresponding position, to a certain degree, up to 100%. The compared molecules may comprise sequences of different lengths.

[0296] The degree of sequence identity can be determined by comparing similar overlapping sequences, wherein the similar overlapping sequences may have a certain degree of sequence identity. The overlapping sequences may comprise a part or the full-length of one of the compared sequences e.g., the shorter one of the compared sequences, wherein the part is e.g., at least any one of 50%, 60%, 70%, 80%, 90%, or 95% of said one the compared sequences. In particular, sequence identity refers to comparing the full-length sequence of one of the compared sequences e.g., the shorter one of the compared sequences.

[0297] The term “comprising or consisting of’ with respect to a certain sequence identified herein, shall particularly mean the respective sequence identity to the part or the full-length of one of the compared sequences e.g., the shorter one of the compared sequences, wherein the part is e.g., at least any one of 50%, 60%, 70%, 80%, 90%, or 95% of said one the compared sequences.

[0298] In particular, where a molecule comprises a certain sequence identity to a compared molecule, the sequence identity is determined for at least part of said compared molecule e.g., at least any one of 50%, 60%, 70%, 80%, 90%, 95%, or 100% of said compared sequence. Where a molecule consists of a certain sequence identity to a compared molecule, the sequence identity is determined for the full-length of said compared molecule i.e., 100% of said compared sequence. LO017P

[0299] -36-

[0300] Sequence similarity searching is an effective and reliable strategy for identifying homologs with excess (e.g., at least 50%) sequence identity. Sequence similarity search tools frequently used are e.g., BLAST, FASTA, and HMMER.

[0301] Sequence similarity searches can identify such homologous nucleic acid sequences or genes by detecting excess similarity, and statistically significant similarity that reflects common ancestry.

[0302] "Percent (%) identity" with respect to a nucleotide sequence is defined as the percentage of nucleotides in a candidate nucleic acid sequence that is identical with the nucleotides in the reference or parent sequence, after aligning the sequence and introducing gaps, if necessary, to achieve the maximum percent sequence identity, and optionally not considering any conservative substitutions as part of the sequence identity. Alignment for purposes of determining percent nucleotide sequence identity can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software. Those skilled in the art can determine appropriate parameters for measuring alignment, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared.

[0303] For purposes described herein (unless indicated otherwise), the sequence identity between two amino acid sequences is determined using the NCBI BLAST program version BLASTN 2.8.1 with the following exemplary parameters: Program: blastn, Word size: 11 , Expect threshold: 10, Hitlist size: 100, Gap Costs: 5.2, Match / Mismatch Scores: 2,-3, Filter string: Low complexity regions, Mark for lookup table only.

[0304] The term “subject” as used herein is understood to comprise mammalian subjects, in particular human subjects, and may further include livestock animals, companion animals, and laboratory animals. A subject can be a patient suffering from a specific disease or disorder or a healthy subject. In particular, the treatment and medical use described herein applies to a subject in need of prophylaxis or therapy of a disease or disorder. Specifically, the treatment may be by interfering with the pathogenesis of a disease or disorder. The subject may be a subject at risk of such disease or disorder, or suffering from such disease or disorder.

[0305] The term “therapeutic POI” as used herein shall refer to a POI that has a positive or advantageous effect on a disease or disorder of a subject when provided to the subject in a therapeutically effective amount. Specifically, a therapeutic POI has curative or palliative properties and may be administered to ameliorate, relieve, alleviate, reverse, LO017P

[0306] -37- delay onset of or lessen the severity of one or more symptoms of a disease or disorder. A therapeutic POI may have prophylactic properties and may be used to delay the onset of a disease or to lessen the severity of such disease or pathological condition.

[0307] The term “treatment” as used herein shall always refer to treating a subject for prophylactic ( / .e., to prevent disease or disorder) or therapeutic ( / .e., to cure, ameliorate, relieve, alleviate, reverse, delay onset of or lessen the severity of one or more symptoms of a disease or disorder) purposes.

[0308] The term “transfection” as used herein shall refer to the introduction of a nucleic acid (e.g., RNA) into a cell or the update of such nucleic acid by the cell. The cell can be present ex vivo e.g., in a cell isolated from an organism, or in an in vitro cell culture. The cell can also be present in vivo i.e., in a subject where the cell can form part of an organ, a tissue and / or an organism. Transfection can be transient or stable. For some applications of transfection, it is sufficient if the transfected genetic material is only transiently expressed. RNA can be transfected into cells to transiently express its coded protein. Since the nucleic acid introduced in the transfection process is usually not integrated into the nuclear genome, the expressed nucleic acid will be diluted through mitosis or degraded. Cells allowing episomal amplification of nucleic acids greatly reduce the rate of dilution. A stable transfection can be achieved by using virus-based systems or transposon-based systems for transfection.

[0309] Specifically, the RNA molecules described herein are provided for transfection into cells to transiently express its coded protein.

[0310] The term “untranslated region” or “UTR” as used herein shall refer to a region in a DNA molecule which is transcribed but is not translated into an amino acid sequence, or to the corresponding region in an RNA molecule, such as an mRNA molecule. An untranslated region (UTR) can be present 5’ (upstream) of an open reading frame (5’UTR) or 3’ (downstream) of an open reading frame (3’UTR). A 5’UTR is typically located upstream of the start codon of a coding sequence and downstream of a 5’-cap (if present), e.g., directly adjacent to the 5’-cap. A 3’UTR is typically located downstream of the termination codon of a coding sequence, and upstream of a poly-A sequence, if present e.g., directly adjacent to the poly-A sequence. A UTR can be autologous or heterologous to other elements of the nucleic acid molecule (e.g., the mRNA elements) into which they are introduced.

[0311] As used herein, the term “wild-type” means that the sequence is naturally occurring and is not artificially modified, including naturally occurring mutants. LO017P

[0312] -38-

[0313] Therefore, the present invention provides for improved nucleic acid constructs, in particular improved DNA molecules which can serve as a template for transcribing corresponding RNA molecules with a high degree of integrity, stability, yield of POI expression and a low amount of process-related impurities such as dsRNA, as exemplified herein.

[0314] According to a specific example, a novel 5’UTR has proven to be superior than a reference 5’UTR that is used in a COVID-19 mRNA vaccine in particular as regards transcription of the DNA to produce an mRNA product with less process-related impurities such as dsRNA, higher mRNA integrity, higher stability of the purified mRNA, and higher yield for POI.

[0315] According to another specific example, the novel 5’UTR or specific 5’UTR and 3’UTR combinations have proven to be advantageous for producing the encoded POI with a high translation efficiency and yield.

[0316] The invention is further described by one or more of the following items.

[0317] 1. An RNA molecule comprising in the 5’->3’ direction, a 5’ untranslated region (5’UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3’ untranslated region (3’UTR), wherein a) the 5’UTR comprises a transcript of the 5’UTR derived from the human alphaglobin nucleic acid sequence SEQ ID NO:1 , a fragment or variant thereof which is at least 90% identical to SEQ ID NO:1 ; and b) the 3’UTR comprises a transcript of: i) the human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid sequence SEQ ID NO: 19, a fragment or variant thereof which is at least 90% identical to SEQ ID NO:19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21 , a fragment or variant thereof which is at least 90% identical to SEQ ID NO:21 ; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, a fragment or variant thereof which is at least 90% identical to SEQ ID NO:23; or iv) a combination of any two or more of i), ii), or iii).

[0318] 2. The RNA molecule of item 1 , wherein the 5’UTR comprises or is comprised in a transcript of any one of SEQ ID NOU , 3, 5, 7, 9, 11 , 13, 15, or 17, or a fragment or variant of any of the foregoing comprising at least 95% sequence identity thereto. LO017P

[0319] -39-

[0320] 3. The RNA molecule of item 1 or 2, wherein the 5’UTR comprises or is comprised in any one of SEQ ID NO:2, 4, 6, 8, 10, 12, 14, 16, or 18, or a fragment or variant of any of the foregoing comprising at least 95% sequence identity thereto.

[0321] 4. The RNA molecule of any one of items 1 to 3, wherein the 5’UTR comprises a 5’-terminal transcription start site, preferably AGG or GGG.

[0322] 5. The RNA molecule of any one of items 1 to 4, wherein the Kozak sequence comprises or consists of any one of SEQ ID NO:57-92.

[0323] 6. The RNA molecule of any one of items 1 to 5, wherein a section of the RNA molecule which comprises the 5’UTR and the Kozak sequence, comprises a transcript of any one of SEQ ID NO:3, 9, 11 , 15, or 17, or a fragment or variant of any of the foregoing comprising at least 95% sequence identity thereto.

[0324] 7. The RNA molecule of any one of items 1 to 6, wherein a section of the RNA molecule, which comprises the 5’UTR and the Kozak sequence, comprises any one of SEQ ID NO:4, 10, 12, 16 or 18, or a fragment or variant of any of the foregoing comprising at least 95% sequence identity thereto.

[0325] 8. The RNA molecule of any one of items 1 to 7, wherein the 3’UTR comprises a transcript of a combination of two or more sequences of i), ii), or iii), preferably comprising SEQ ID NO:25, 27, 29, or 31 , or a fragment or variant of any of the foregoing comprising at least 95% sequence identity thereto.

[0326] 9. The RNA molecule of any one of items 1 to 8, wherein the 3’UTR comprises any one of SEQ ID NO: 20, 22, 24, 26, 28, 30, or 32, or a fragment or a variant of any of the foregoing comprising at least 95% sequence identity thereto.

[0327] 10. The RNA molecule of any one of items 1 to 9, wherein the RNA molecule further comprises a polyadenyl sequence at the 3’-end, optionally comprising within the polyadenyl sequence a sequence of one or more consecutive nucleotides other than A nucleotides.

[0328] 11. The RNA molecule of any one of items 1 to 10, wherein the RNA molecule further comprises a 5’-cap.

[0329] 12. The RNA molecule of any one of items 1 to 11 , wherein the RNA molecule is an mRNA or self-amplifying RNA (saRNA).

[0330] 13. A DNA molecule comprising a nucleotide sequence suitable for transcription of an RNA molecule of any one of items 1 to 12, preferably in the form of a plasmid.

[0331] 14. Use of a DNA molecule of item 13 as a template for in vitro transcription of

[0332] RNA. LO017P

[0333] -40-

[0334] 15. A method of producing an RNA molecule of any one of items 1 to 12, by in vitro transcription of the DNA molecule of item 13, optionally after linearization.

[0335] 16. Use of the RNA molecule of any one of items 1 to 12, for in vitro transfecting a host cell, preferably cells for gene editing, cancer cells, antigen-presenting cells, monocytes, dendritic cells, T cells, macrophages, or lung or liver cells.

[0336] 17. The RNA molecule of any one of items 1 to 12, for medical use, preferably for use in mRNA-based therapy.

[0337] 18. A mRNA comprising a nucleotide sequence SEQ ID NO:6, 8, 10, or 12 as a 5’ untranslated region (5’UTR), or a variant thereof comprising at least 95% sequence identity to any of the foregoing, preferably wherein the variant comprises the stretch of position 4-34 of SEQ ID NO:6.

[0338] 19. A DNA template for transcribing the mRNA of item 18.

[0339] The foregoing description will be more fully understood with reference to the following examples. Such examples are, however, merely representative of methods of practicing one or more embodiments of the present invention and should not be read as limiting the scope of invention.

[0340] EXAMPLES

[0341] EXPERIMENTAL DESCRIPTION

[0342] DNA template generation for small-scale transcription

[0343] Synthetic genes up to 3000 bp were ordered as gBIocks HiFi Gene Fragments from Integrated DNA Technologies (IDT), Leuven Belgium. Longer synthetic genes were ordered from GeneArt, Regensburg Germany. Synthetic genes contained the following sequence elements in 5’ to 3’ direction: a conserved upstream sequence, a T7 promotor sequence inclusive AGG transcription start site, 5’UTR, coding sequence, 3’UTR, and a conserved spacer sequence (Figure 1 ). LO017P

[0344] -41-

[0345] PCR amplification was performed as described in WO2013151663A1. PCR amplicons were purified using 0.45 reaction volumes of Sera-Mag Select magnetic beads (29343052, Cytiva) in a KingFisher Apex instrument (ThermoFisher Scientific). PCR amplicon DNA was eluted with nuclease-free water (AM9916, ThermoFisher Scientific). DNA concentration was determined by absorbance at 260 nm using a NanoDropOne device (ThermoFisher Scientific). DNA integrity was by capillary gel electrophoresis using a Fragment Analyzer 5300 device (Agilent) with the DNF-930 kit (Agilent).

[0346] Small-scale mRNA construct production and analytics

[0347] Transcription was performed in a buffer consisting of 40 mM Tris-HCI pH8.0 (AM9855G, ThermoFisher), 40 mM DTT (1370-100 GM, BioVectra), 2 mM Spermidine (S2626, Sigma-Aldrich), and 28 mM MgCh (AM9530G, ThermoFisher Scientific). The transcription reaction further contained 15 mM NTP (ATP:CTP:GTP:UTP 1 :1 :1:1 , ThermoFisher Scientific), 4 mM CleanCap AG (N-7113-100, Trilink), 22 ng / pL PCR amplicon DNA, 0.002 U / pL Pyrophosphatase (Roche) and 0.1 U / pL RNase inhibitor (Roche) and 7.5 U / pL T7 polymerase (ThermoFisher Scientific). The transcription reaction was performed for 90 min at 37°C. DNAse digestion was performed with 800 U / mL DNasel (Roche) with 0.5 mM CaCh (C5670, Sigma-Aldrich) for 30 min at 37°C. The reaction was quenched with 56 mM EDTA (AM9260G, ThermoFisher Scientific). LO017P

[0348] -42-

[0349] The RNA was purified using 0.5 reaction volumes of Sera-Mag Select magnetic beads (Cytiva) in a KingFisher Apex instrument (ThermoFisher Scientific).

[0350] In specific examples, the RNA was purified also by an oligo(dT) affinity step using CIM Oligo dT18 1 mL Monolithic 24-well Plate C12 Linker, 2 pm channels (BIA- 124.1219-2, Sartorius). Binding of poly(A)-containing mRNA was performed in 10 mM Tris pH 7.0 (AM9851 , ThermoFisher Scientific), 5 mM EDTA (AM9260G, ThermoFisher Scientific), 600 mM KCI (60142, Sigma-Aldrich). Washing was performed in 10 mM Tris pH 7.0, 5 mM EDTA, 50 mM KCI and elution was performed with nuclease-free water (AM9916, ThermoFisher Scientific). mRNA integrity was determined by capillary gel electrophoresis (CGE) using a Fragment Analyzer 5300 device (Agilent) with the DNF-471 kit (Agilent) as wells as by a Vanquish HPLC-UV (ThermoFisher Scientific) device using a DNAPac RP Column (3.0x100 mm) (ThermoFisher Scientific) and a gradient from 0.1 M tetraethylammonium acetate (TEAA) to 0.1 M TEAA with 25% acetonitrile. mRNA capping efficiencies were analyzed by RNaseH digestion as described in the USP draft guideline Analytical Procedures for mRNA Vaccines Quality - 3rd Edition. dsRNA content was determined by ELISA as described in the USP draft guideline Analytical Procedures for mRNA Vaccines Quality - 3rd Edition.

[0351] Scale-up of firefly luciferase transcription and downstream processing

[0352] Transcription was performed at 10 mL scale in Ambr15 bioreactors (Sartorius) for 2h at 37°C. The reaction buffer was identical to the one used for small-scale mRNA manufacturing. The transcription reaction further contained 20 mM NTP (ATP:CTP:GTP:N1-Methyl-pseudo-UTP 1 :1 :1 :1 , ThermoFisher Scientific), 4 mM CleanCap AG (N-7113-100, Trilink), 40 ng / pL linearized plasmid DNA, 0.002 U / pL Pyrophosphatase (Roche) and 0.1 U / pL RNase inhibitor (Roche) and 7.5 U / pL T7 polymerase (ThermoFisher Scientific). DNase digestion and EDTA quenching was performed identical to the small-scale mRNA manufacturing. mRNA was purified first by oligo-dT affinity chromatography followed by Hydrophobic Interaction Chromatography (HIC) using an Azura HPLC device (Knauer) as described in Gagnon P. Purification of Nucleic Acids: A handbook for purification of plasmid DNA and mRNA for gene therapy and vaccines. BIA Separations, Ajdovscina, 2020; 136. Ultrafiltration and diafiltration into 10 mM Tris pH 7.0 was performed us pTFF device (Formulatrix) with a 100 kDA modified polyethersulfone cassette. The mRNA was LO017P

[0353] -43- filtered using a 0.2 m Supor EKV membrane in Mini Kleenpak syringe filter capsules (KM2EKVS, Cytiva).

[0354] CELL-BASED ASSAYS

[0355] Kinetic firefly luciferase and HiBiT detection assay

[0356] For the kinetic monitoring of firefly luciferase expression in HEK293, cells were detached by try pie select (12563011 , Gibco) and resuspended at 1 x107cells / mL in Opti- MEM medium (31985062, Gibco). Electroporation of 30 pg / mL mRNA to 200000 cells / well was performed using a 4D Nucleofactor device (Lonza) using the HEK program and an electroporation plate from the SF Cell Line 96-well Nucleofector kit (Lonza). After electroporation, cells were resuspended at 5x105cells / mL in DMEM (Glutamax) medium (10566016, Gibco) supplemented with 10% FBS (16000044, Gibo) in Nunc MicroWell 96-well plates (136101 , ThermoFisher Scientific) and incubated at 37°C and 5% CO2 in a CO2-lncubator (Sanyo). For detection of firefly luciferase activity at 1 , 2, 3, 4, 5, 6, 7, 8, 24, 27, 30, and 33 hours, an equal volume of DualGlo luciferase reagent (E2920, Promega) was added. After 15 min incubation at room temperature in the dark, luminescence was recorded using a Tristar 5 multimode reader (Berthold). For each sample, four replicates were of the firefly luciferase expression assay were performed

[0357] For the kinetic monitoring of the HiBiT detection tag in HEK293, cells were detached by try pie select (12563011 , Gibco) and resuspended at 1 x107cells / mL in Opti- MEM medium (31985062, Gibco). Electroporation of 10 pg / mL mRNA to 200000 cells / well was performed using a 4D Nucleofactor device (Lonza) with the HEK program and an electroporation plate from the SF Cell Line 96-well Nucleofector kit (Lonza). After electroporation, cells were resuspended at 5x105cells / mL in DMEM (Glutamax) medium (1056601 , Gibco) supplemented with 10% FBS (6000044, Gibo) in Nunc MicroWell 96- well plates (136101 , ThermoFisher Scientific) and incubated at 37°C and 5% CO2 in a CO2-lncubator (Sanyo). For detection of the HiBit Tag, 1 , 2, 3, 4, 5, 6, 7, 8, 24, 28, 32, and 48 hours, an equal volume of Nano-Gio HiBiT Lytic Detection System (N3050, Promega) was added per well. After 15 min incubation at room temperature in the dark, luminescence was recorded using a Tristar 5 multimode reader (Berthold). For each sample, four replicates were performed.

[0358] From both the firefly luciferase and HiBiT detection assay, the translation efficiency was determined from the initial slope, the protein yield was determined from LO017P

[0359] -44- the area under the curve. For firefly luciferase only, the mRNA half-life time was determined for timepoints starting with 24 h by transforming the data points with the natural logarithm and extraction of the decay rate by linear regression. The half-life time is obtained by dividing ln(2) by the slope value (Figure 2).

[0360] The following examples refer to certain nucleic acid constructs comprising 5’UTR and 3’UTR sequences. The respective UTR RNA sequences are listed in the table below.

[0361] Table 2: DNA constructs as used in the examples, 5’ and 3’ UTR sequences specified LO017P

[0362] -45-

[0363] Example 1 : 5’UTR and 3’UTR sequences with superior properties compared to a reference construct in terms of translation efficiency and protein yield as well as manufacturability

[0364] The UTR1 construct as used in the COVID-19 mRNA vaccine has served as a reference construct (Table 2).

[0365] For substitution of 5’UTR or 3’UTR of the UTR1 construct, firefly luciferase mRNA SEQ ID NO:94 was transcribed and only purified by a one-step purification method employing magnetic beads. No additional oligo(dT) affinity purification was performed. Table 3 shows the results of mRNA integrity testing, capping efficiency testing, and the degree of purity in terms of product-related dsRNA.

[0366] The results show that the novel 5’UTR SEQ ID NO:9 (UTR14) mediates improved transcription as compared to the reference 5’UTR in UTR1 (SEQ ID NO: 17). For UTR14 relative to UTR1 , the amount of dsRNA was reduced even following a purification process that employs only magnetic beads purification without an oligo(dT) affinity purification step. Similarly, all constructs containing a different 3’UTR from UTR1 showed reduced dsRNA levels.

[0367] The mRNA was evaluated in the firefly luciferase kinetic assay in HEK293 cells to extract the translation efficiency, mRNA half-life-time, and protein yield relative to the UTR1 firefly luciferase construct (Table 4).

[0368] Relative to UTR1 , all constructs showed a higher translation efficiency. For the evaluated 5’UTR, the firefly luciferase constructs with UTR14 both increased translation efficiency and mRNA half-life time. UTR6, UTR9, UTR10, UTR11 , and UTR19 firefly luciferase constructs exhibited increased relative mRNA stability compared to the UTR17 reference construct with the beta-globin 3’ UTR. On the other hand, the UTR11 construct showed a decreased mRNA integrity due to by-product formation and UTR19 showed the lowest translation efficiency. Hence, it can be concluded that constructs UTR6, UTR9, UTR10, and UTR14 with enhanced performance in the cell and high mRNA integrity only require a one-step purification method to achieve a lower dsRNA level compared to UTR1 , which is a significant improvement of manufacturability of the mRNA product as well as improved translation efficiencies and protein yields. LO017P

[0369] -46-

[0370] UTR6, UTR9, UTR10, UTR14 (and UTR17 as comparator) were evaluated for expression in different cell-lines (Figure 1). All construct showed a similar expression in different cell-lines as the UTR17 comparator and therefore show a limited cell-line specificity for protein expression that enables a wide range of therapeutic applications.

[0371] Table 3: Physico-chemical properties of firefly luciferase constructs containing 5’ or 3’ UTR substitutions in UTR1.

[0372] Table 4: Parameter extracted from kinetic luciferase assay. LO017P

[0373] -47-

[0374] Table 4 continued

[0375] Example 2: Selected 5’UTR and 3’UTR to express firefly luciferase mRNA containing pairs of UTRs identified in Example 1 were purified after transcription by a two-step purification method, including magnetic beads and oligo(dT) affinity chromatography. For the UTR1 reference construct, purification by oligo(dT) affinity increased the mRNA integrity and significantly improved the translation efficiency and protein yield (Table 5, Table 6). Therefore, demonstrating that the UTR1 construct requires an additional purification step to reach a translation efficiency and protein yield comparable to the UTR6, UTR9, UTR10, and UTR14 constructs tested in Example 1.

[0376] On the other hand, for the UTR1 construct, the dsRNA level remained comparable before and after oligo(dT) purification (Example 1). A similar observation was also made for the UTR14 construct, where the dsRNA level increased slightly relative to the magnetic bead-purified mRNA in Example 1. For the constructs comprising 5’UTR and 3’UTR sequences with superior cellular performance selected in Example 1 , there was also limited improvement due to the two-step purification method compared to the one-step purification method (as in Example 1) with the exception of the mRNA integrity increase. Hence, the oligo(dT) purification step was not suitable to further reduce the dsRNA level. LO017P

[0377] -48-

[0378] Table 5: Physico-chemical properties of firefly luciferase constructs containing 5’ and 3’ UTR combinations, n.a. stands for not analyzed.

[0379] Table 6: Comparison for parameter extracted from the luciferase kinetic assay for UTR1 constructs purified by mag. beads or mag. beads + oligo (dT).

[0380] In context with the 5’UTR SEQ ID 10, the firefly luciferase construct UTR14-6 led to the highest protein yield followed by UTR14-9, and lastly UTR14-10 (Table 7). Creation of a combined 3’UTR from GPI1 and MRS2 for the constructs UTR14-6-10 and UTR14-10-6 did not show a significant improvement over constructs UTR14-6 or UTR14-10. On the other hand, UTR14-9-10 as a 3’UTR composed from both the NACA and MRS2 sequences showed a higher protein yield compared to UTR14-9 and UTR14- 9-10. LO017P

[0381] -49-

[0382] Table 7: Parameter extracted from kinetic luciferase assay for 5’ and 3’ UTR combinations.

[0383] The expression of selected 5’ and 3’ UTR combinations was additionally evaluated in different cell-lines (Figure 4). For the examined cell-lines, all constructs showed a comparable expression when normalized to the UTR1 construct. Therefore, the expression of the UTR constructs was not cell line specific.

[0384] Example 3: Selected 5’UTR and 3’UTR to express an eGFP-PEST-HiBiT coding sequence eGFP-PEST-HiBiT mRNA constructs SEQ ID: 96 containing the identified 5’ and 3’ UTR pairs of Example 2 were prepared. The PEST domain was added to decrease the protein life-time of eGFP to approximately 2 h (“Generation of destabilized green fluorescent protein as a transcription reporter constructs”, Li. et.aL, 1998, Biochemistry, Volume 273, Issue 52, Pages 34970-34975). The HiBiT tag was added for detection of the eGFP expression (“mRNAid, an open-source platform for therapeutic mRNA design and optimization strategies”, Vostrosablin et. al, 2024, NAR Genomics and Bioinformatics Volume 6, Issue 1). mRNA constructs were purified after transcription both by magnetic beads and oligo(dT) affinity chromatography. A constant mRNA yield was obtained for all constructs (Table 8). Importantly, a lower dsRNA content was measured for all constructs compared to UTR1. LO017P

[0385] -50-

[0386] Table 8: Physico-chemical properties of eGFP constructs containing 5’ and 3’ UTR combinations

[0387] The eGFP expression kinetics was monitored in HEK293 cells using the HiBiT detection assay (Table 9). All constructs with 5’ and 3’ UTR combinations showed a slightly higher translation efficiency relative to the UTR1 construct (Table 9). For eGFP constructs, the protein yield followed a different trend of UTR14-6 > UTR1 > UTR14-9 > UTR14-10 > UTR14-9-10 as observed for the firefly luciferase constructs of Examples 1 and Example 2. Importantly, the protein yield for the UTR14-6 construct exceeded the yield of the UTR1 construct.

[0388] Table 9: Parameter extracted from kinetic HiBiT detection assay for eGFP constructs.

[0389] Example 4: Selected 5’UTR and 3’UTR to express the Covid spike protein-HiBiT coding sequence

[0390] Covid Spike-HiBiT mRNA constructs SEQ ID: 98 were prepared with the 5’ and 3’UTR combinations from examples 2 and 3. mRNA constructs were purified after transcription both by magnetic beads and oligo(dT) affinity chromatography. Comparable physico-chemical properties and purity levels were obtained (Table 8). LO017P

[0391] -51-

[0392] Table 10: Physico-chemical properties of Covid Spike constructs containing 5’ and 3’ UTR combinations The Covid Spike protein expression kinetics was monitored in HEK293 cells using the HiBiT detection assay (Table 11 ). UTR14-10 showed an increased translation efficiency compared to UTR1 that correlates a lower dsRNA level for UTR14-10 (Table 10). For the protein yield, UTR14-6 showed a strongly increased protein yield compared to UTR1. This finding is comparable to example 3 for eGFP where UTR14-6 also outperformed UTR1 in terms of protein yield. Similarly, UTR14-10 and UTR14-9-10 led to higher protein yields compared to UTR1.

[0393] Table 11 : Parameter extracted from kinetic HiBiT detection assay for Covid Spike protein constructs.

[0394] LO017P

[0395] -52-

[0396] Example 5: Scale-up of mRNA synthesis in bioreactor - manufacturability

[0397] Firefly luciferase mRNA linked to UTR1 , UTR14-6, and UTR14-6 were transcribed at 10 mL scale in a bioreactor. mRNA integrities were analyzed at the end of the transcription reaction (Figure 5). A distinct front peak corresponding to truncated mRNA was detected in addition to the full-length product that elutes in the main peak. The ratio of the truncated front peak was significantly lower for UTR14-6 and also lower for UTR14-10 when compared to UTR1 (Table 13). Hence, a significantly higher yield of full-length firefly luciferase can be achieved for UTR14-6 compared to UTR1 due to reduced transcription termination in the 3’UTR regions.Example 6: Stability study of purified mRNA

[0398] As part of an accelerated stability study, firefly luciferase mRNAs purified by 2- step chromatography as well as concentrated and buffer exchanged by tangential flow filtration were diluted to 50 ng / L in nuclease-free water and incubated for 0, 4, 8, 24, 32, and 48 h at 2-8°C, 22°C, and 40°C. The mRNA integrity was determined by capillary gel electrophoresis (CGE) and normalized relative to the starting value. At 40°C, the mRNA integrity decreased significantly faster for UTR1 compared to UTR14-6 and UTR14-10 (Figure 6). At 2-8°C and 22°C, the absence of any significant mRNA integrity decrease for UTR1 , UTR14-6, and UTR14-10 argues against the presence of a contamination, which could lead to an mRNA integrity decrease, in the mRNA batches for UTR1 , UTR14-6 and UTR14-10. Hence, the difference in stability at 40°C reflects the difference in stability of the mRNA sequences.

Claims

LO017P-53-CLAIMS1. An RNA molecule comprising in the 5’->3’ direction, a 5’-terminal transcription start site, a 5’ untranslated region (5’UTR), a Kozak sequence, an open reading frame (ORF) comprising a coding region, and a 3’ untranslated region (3’UTR), wherein a) the 5’UTR comprises a transcript of the 5’UTR of the human alpha-globin nucleic acid sequence SEQ ID NO:1 , a fragment or variant thereof which is at least 90% identical to SEQ ID NO:1 ; and b) the 3’UTR comprises a transcript of: i) the human Glucose 6-phosphate isomerase 1 (GPI1) nucleic acid sequence SEQ ID NO: 19, or a fragment of SEQ ID NO: 19 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO: 19 which is at least 90% identical to SEQ ID NO: 19; or ii) the human nascent polypeptide associated complex subunit alpha (NACA) nucleic acid sequence SEQ ID NO:21 , or a fragment of SEQ ID NO:21 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO:21 which is at least 90% identical to SEQ ID NO:21 ; or iii) the human magnesium transporter (MRS2) nucleic acid sequence SEQ ID NO:23, or a fragment of SEQ ID NO:23 comprising at least 50% of the full-length sequence, or a variant of SEQ ID NO:23 which is at least 90% identical to SEQ ID NO:23; or iv) a combination of any two or more of i), ii), or iii).

2. The RNA molecule of claim 1 , wherein the 5’UTR comprises or is comprised in a transcript of any one of SEQ ID NOU , 3, 5, 7, 9, 11 , 13, 15, or 17, or a fragment of any of the foregoing comprising at least 50% of the full-length sequence, or a variant of any of the foregoing comprising at least 90% sequence identity thereto.

3. The RNA molecule of claim 1 or 2, wherein the 5’UTR comprises or is comprised in any one of SEQ ID NO:2, 4, 6, 8, 10, 12, 14, 16, or 18, or a fragment of any of the foregoing comprising at least 50% of the full-length sequence, or variant of any of the foregoing comprising at least 90% sequence identity thereto.

4. The RNA molecule of any one of claims 1 to 3, whereinLO017P-54- a) the transcription start site consists of the nucleotide sequence AGG or GGG; and / or b) the 5’UTR comprises the 5’-terminal transcription start site, preferably AGG or GGG; and / or c) the Kozak sequence comprises or consists of any one of SEQ ID NO:57-92.

5. The RNA molecule of any one of claims 1 to 4, wherein the 3’UTR comprises a transcript of a combination of two or more sequences of i), ii), or iii), preferably comprising SEQ ID NO:25, 27, 29, or 31 , or a fragment of any of the foregoing comprising at least 50% of the full-length sequence, or variant of any of the foregoing comprising at least 90% sequence identity thereto.

6. The RNA molecule of any one of claims 1 to 5, wherein the 3’UTR comprises any one of SEQ ID NO: 20, 22, 24, 26, 28, 30, or 32, or a fragment of any of the foregoing comprising at least 50% of the full-length sequence, or a variant of any of the foregoing comprising at least 90% sequence identity thereto.

7. The RNA molecule of any one of claims 1 to 6, wherein the RNA molecule further comprises a) a polyadenyl sequence at the 3’-end; and / or b) a 5’-cap.

8. The RNA molecule of any one of claims 1 to 7, wherein the coding sequence is heterologous to the 3’UTR.

9. The RNA molecule of any one of claims 1 to 8, wherein the RNA molecule is an mRNA or self-amplifying RNA (saRNA).

10. A DNA molecule comprising a nucleotide sequence suitable for transcription of an RNA molecule of any one of claims 1 to 9, preferably in the form of a plasmid.11 . Use of a DNA molecule of claim 10 as a template for in vitro transcription ofRNA.LO017P-55-12. A method of producing an RNA molecule of any one of claims 1 to 9, by in vitro transcription of the DNA molecule of claim 11 , optionally after linearization.

13. Use of the RNA molecule of any one of claims 1 to 9, for in vitro transfecting a host cell, preferably cancer cells, antigen-presenting cells, monocytes, dendritic cells,T cells, macrophages, lung or liver cells.

14. The RNA molecule of any one of claims 1 to 9, for medical use, preferably for use in mRNA-based therapy.

15. A mRNA comprising a nucleotide sequence which is any one of SEQ ID NO:6, 8, 10, 12, 14, 16, or 18 as a 5’ untranslated region (5’UTR), or a variant thereof comprising at least 90% sequence identity to any of the foregoing, preferably wherein the variant comprises the stretch of position 4-34 of SEQ ID NO:6.

16. A DNA template for transcribing the mRNA of claim 15.

Citation Information

Patent Citations

  • Quantitative assessment for cap efficiency of messenger RNA

    WO2014152659A1

  • 3' UTR sequences for stabilization of RNA

    WO2017059902A1

  • 3' UTR sequences for stabilization of RNA

    WO2017060314A2

  • RNA molecules

    WO2023062556A1

  • mRNA for protein expression and template therefor

    WO2023146230A1