mRNA platform

The mRNA platform with H2AX 5' UTR, COL6A2 3' UTR, and cystatin C signal peptide addresses delivery challenges, achieving efficient and stable protein expression for therapeutic applications.

WO2026003269A1PCT designated stage Publication Date: 2026-01-02OXFORD UNIVERSITY INNOVATION LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/068258
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2025-06-27
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing therapeutic interventions face challenges in delivering exogenous proteins and genetic information to cells due to limited delivery systems, low cellular uptake, immunogenicity, and off-target side effects, particularly in the context of DNA and viral vector systems, which hinder the scalability and flexibility needed for personalized medicine.

Method used

An mRNA platform comprising specific 5' UTR, 3' UTR, and ORF sequences, including H2AX 5' UTR, COL6A2 3' UTR, and cystatin C signal peptide, enhances the translational capacity and stability of therapeutic mRNAs, enabling efficient expression of intracellular and secreted proteins.

Benefits of technology

The mRNA platform supports highly efficient translation and stability of therapeutic mRNAs, overcoming limitations of existing delivery methods by providing a versatile and scalable solution for protein expression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000026_0001
    Figure IMGF000026_0001
  • Figure IMGF000054_0001
    Figure IMGF000054_0001
  • Figure IMGF000055_0001
    Figure IMGF000055_0001
Patent Text Reader

Abstract

The invention relates to an RNA platform comprising a 5' UTR, a 3' UTR and / or a sequence encoding a signal peptide, each having advantageous properties.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] mRNA platform

[0002] Field of the invention

[0003] The present invention relates to an mRNA platform comprising a 5’ UTR, a 3’ UTR and / or a sequence encoding a signal peptide, each having advantageous properties. The invention also relates to a polynucleotide, vector or cell comprising or encoding the platform, and uses of the platform.

[0004] Background of the invention

[0005] Our genetic information is stored in the nucleus of our cells in the form of DeoxyriboNucleic Acid (DNA). The genes present in the DNA can broadly be categorised into protein-encoding genes and noncoding RNA genes. Gene activation results in copying (transcription) of the DNA sequence of a particular gene into a complementary RiboNucleic Acid (RNA) molecule. The protein-encoding genes are transcribed into messenger RNAs (mRNAs). mRNAs function as blueprints for ribosomes in the cytoplasm of cells to translate the information stored in nucleic acids into amino acids, the building blocks for proteins. In cells, the genetic information flows from DNA to mRNA to protein.

[0006] Accordingly, therapeutic interventions that aim to replace non-functional disease-causing proteins or express exogenous proteins with novel functions (e.g., antibodies) in patients can achieve this by inserting DNA, mRNA or protein into targeted tissues and cells. Whilst the administration of proteins is the most direct approach, it is also the most challenging. This is particularly true for therapeutic proteins that function intracellularly as replacement or regenerative proteins. The practicality of these therapeutic proteins is severely limited, chiefly due to the lack of suitable delivery systems and the low cellular uptake of exogenous proteins by target cells (Porello & Cellesi, 2023).

[0007] In contrast, proteins that act in patient serum, most notably therapeutic monoclonal antibodies (mAbs), are successfully employed in immunotherapies combatting a variety of pathologies, including cancers. However, clinical practice of exogenous functional antibodies is complicated by challenging and labour-intensive manufacturing processes, contamination risks, immunogenicity and high cost.

[0008] The delivery of exogenous genetic information into patient cells via DNA has long been the backbone of gene therapy. Administering exogenous DNA that encodes a therapeutic protein can be achieved in two principles ways. DNA can be delivered into cells in the form of naked DNA plasmid vectors (Glover et al., 2005). However, the expression and efficacy of the encoded genetic information from plasmid vectors is limited by the inefficient uptake of plasmid DNA into the nucleus of targeted cells (Guan et al., 2024). Viral vectors are an effective alternative for the delivery of DNA into patient cells and tissues (Lundstrom, 2023). Whilst viral vector systems have been successfully employed, they do have limitations (Butt et al., 2022). The expression of clinically effective levels of the encoded proteins can be compromised because of immunogenicity of the viral vehicle. Viral vector delivery systems also bear considerable risk of unforeseen and undesirable off-target side effects, potentially delaying regulatory approval. Another issue is the limited capacity regarding the length of the therapeutic cargo DNA fragment they can incorporate and carry. In addition, the production methods can be complicated, requiring considerable expertise that compromise scalability and flexibility of viral delivery systems. These constraints can significantly impede the transition of DNA viralbased gene therapy systems into clinical practice, particularly within the realm of personalized medicine that requires versatile and easily adaptable platforms.

[0009] Delivery of therapeutic proteins via mRNA has matured into an alternative approach that avoids many of the issues highlighted above. Therapeutic mRNAs are produced by in vitro transcription (IVT), which is a simple, cost-effective and scalable cell-free chemical reaction. The production of IVT plasmids that serve as templates to transcribe therapeutic mRNAs relies on well-established methodology that is highly versatile. In addition, proven in vivo delivery methods such as lipid nanoparticles (LNPs) promoting effective cellular uptake of therapeutic mRNAs, are available (Hou et al., 2021). The incorporation of naturally-occurring modified nucleosides into the mRNA provides an elegant solution to suppress the cellular innate immune response otherwise triggered by exogenous RNA (Andries et al., 2015; Kariko et al., 2005, 2008). Unlike DNA, mRNA is short-lived and thus much less prone to trigger long-lasting unwanted side effects. Thus, mRNA-encoded therapeutics provide an affordable highly versatile alternative to supplement patient tissues with replacement, regenerative or protective proteins.

[0010] Functional mRNAs are single-stranded nucleic acids with an inherent 5’ to 3’ directionality. They feature a cap structure consisting of a methylated Guanosine nucleotide at their 5’ end and a polyadenosine tail at their 3’ end. The coding region contains the nucleotides or the code, that is read by the ribosome and translated into protein. This coding region is flanked by the 5’ and 3’ untranslated regions (5’ UTR, 3’UTR, respectively). Whilst the nucleotides of the UTRs do not hold any coding information required for the ribosome to translate the mRNA into protein, they harbour critical regulatory sequences that determine the fate of the mRNA in terms of its cellular localisation, translation efficiency, tissue-specific expression and the stability of the mRNA molecule. The choice of UTRs in therapeutic mRNAs thus provides significant scope to maximise and tailor the expression of the encoded payload proteins (Fang et al., 2022).

[0011] It is an aim of the present invention to improve the translational capacity of mRNA molecules, e.g. for therapeutic use.

[0012] Summary of the invention

[0013] The present inventors surprisingly found that RNA molecules that comprise a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR), and an open reading frame (ORF) and / or a non-coding functional RNA sequence, wherein: the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof; the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof support highly efficient translation of payload proteins. Using a luciferase reporter, the inventors also identified that a cystatin C signal peptide or variants thereof, outperform the signal peptide that normally guides IgG antibodies to the ER-associated ribosomes for subsequent secretion. Therefore, the present inventions serves as a basis for a plasmid platform, that affords the production of highly translatable and stable therapeutic mRNAs to drive the expression of therapeutic intracellular and secreted proteins.

[0014] Accordingly, the invention provides an RNA molecule comprising a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR), and an open reading frame (ORF) and / or a non-coding functional RNA sequence, wherein: the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof; the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

[0015] The invention also provides a polynucleotide comprising a nucleic acid sequence encoding the RNA molecule of the invention. The invention further provides a polynucleotide comprising a nucleic acid sequence encoding a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR) and (i) an open reading frame (ORF) and / or a non-coding functional RNA sequence, or (ii) a sequence for introducing an ORF and / or a non-coding functional RNA sequence, wherein: the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof; the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

[0016] The invention additionally provides a vector comprising the polynucleotide of the invention.

[0017] The invention also provides a cell comprising the polynucleotide of the invention of the vector of the invention.

[0018] The invention further provides a pharmaceutical composition comprising the RNA molecule of the invention, the polynucleotide of the invention or the vector of the invention. The invention also provides the pharmaceutical composition of the invention for use in a method of treating or preventing a disease in a subject.

[0019] The invention also provides an in vitro method of producing an RNA molecule, comprising contacting the polynucleotide of the invention, the vector of the invention or the cell of the invention, with an RNA polymerase.

[0020] The invention also provides a pharmaceutical composition produced by the method of the invention for producing an RNA molecule, wherein said method further comprises formulating the RNA molecule into a pharmaceutical composition.

[0021] The invention also provides an in vitro or an in vivo method of producing a peptide, polypeptide or protein, the method comprising translating the RNA molecule of the invention, or the RNA molecule obtained by the method of the invention for producing an RNA molecule.

[0022] The invention also provides a method of increasing the translational capacity of a polynucleotide comprising a nucleic acid sequence encoding a 5’ UTR, a 3’ UTR and / or an ORF comprising a sequence encoding a signal peptide, the method comprising: replacing the sequence encoding the 5’ UTR with a sequence encoding an Histone 2 A family member X (H2AX) 5’ UTR or variant thereof; replacing the sequence encoding the 3’ UTR with a sequence encoding a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or replacing the sequence encoding a signal peptide with the sequence encoding a cystatin C signal peptide or variant thereof.

[0023] Brief description of the figures

[0024] Figure 1 - Schematic of the mRNA platform parental plasmid. The plasmid comprises a CleanCap-compatible T7 RNA polymerase promoter followed by a sequence encompassing the unique restriction enzyme (RE) sites EcoRI and Xbal. Downstream of the two RE recognition sites are the sequences of the human H2AX 5 ’UTR followed by the two unique RE sites for BamHI and Xhol. A Kozak sequence is positioned immediately after the Bam HI RE site and is followed by the Renilla luciferase ORF. The ORF is followed by Spel and Hindlll specific recognition sites. The 3’UTR of the COL6A2 gene is flanked by the upstream-positioned Hindlll site and Kpnl and AfUI RE recognition sites. Downstream of these Kpnl and AfUI recognition sites is a polyadenosine tract, or a modified polyadenosine tract of varied length, which serves to encode poly(A) tails on the resulting mRNAs. Immediately after the poly(A) tract are a number of unique RE recognition sites (SapI, BssHII and Avril) that serve to linearize the plasmid for the in vitro transcription of run-off mRNAs.

[0025] Figure 2 - RNA integrity control of mRNAs featuring different UTR combinations.

[0026] Agarose gel depicting Ipg of loaded Renilla luciferase-encoding capped and polyadenylated (30 adenosines) unmodified (uridines) mRNAs featuring the H2AX and COL6A2 UTRs (lane 2) and equivalent mRNA containing the 5’ and 3’UTRs of Modema’s mRNA-1273 (lane 5; SEQ ID NOs: 11 and 12 respectively). The bands show the integrity of the transcribed mRNA in each case. M represent lanes containing DNA size markers with selected sizes indicated in kilobases (kb) on the left. Lanes 1 and 6 are left empty and lane 2 contains mRNA featuring the H2AX and the Col6A2 UTRs. Lane 5 features the Modema 1273 UTRs. Lanes 3 and 4 feature mRNAs with different combinations of H2AX, Col6A2 and Modema UTRs.

[0027] Figure 3 - Benchmarking UTRs. Comparison of Renilla luciferase activity measured using cell lysates from HEK-293T cells transfected with lOOng, I pg or 2pg of mRNAs comprising the H2AX 5’UTR, COL6A2 3’UTR combination (black filled bars) or lOOng, Ipg or 2pg of equivalent mRNAs featuring the UTRs of Moderna’s mRNA-1273 (striped bars). The relative luciferase activity is represented in Relative Light Units (RLU). Data is represented as mean ± standard deviation. mRNAs were transcribed using unmodified uridine nucleotides and featured a 30-nucleotide long poly(A) tail.

[0028] Figure 4 - Expressing monoclonal antibodies using the H2AX / Col6A2 mRNA platform. Expression of the R5.016 monoclonal antibody encoded by two mRNAs featuring the H2AX and COL6A2 UTRs that flank the ORF encoding the light or heavy chains of R5.016 and a 114-nucleotide long segmented poly(A) tail. The concentration of R5.016 antibodies present in the culture media of HEK-293T cells transfected with Ipg of heavy chain- and light chain- encoding mRNAs was determined by standardised ELISA (r2of the standard curve = 0.998 and all positive controls were within 20% of expected value). For the 3 biological repeats of the transfected mAb-encoding mRNA, data is represented as mean ± standard deviation. For the negative control in the ELISA an unrelated antibody was used.

[0029] Figure 5 - Schematic of the mRNA platform plasmid featuring a signal peptide cassette. The parental plasmid was used to insert a signal peptide cassette that can be exchanged using RE sites as indicated. The SP sequence is in-frame with the downstream main ORF. Other features including the UTRs are as described in figure 1.

[0030] Figure 6 - Signal peptide plasmids. Schematic of the IVT plasmid templates featuring the Cystatin C or the IgG signal peptides upstream of the Renilla luciferase ORF. The amino acid sequence of the two signal peptides are depicted.

[0031] Figure 7 - RNA integrity control of mRNAs featuring different signal peptides.

[0032] 1 pg of Renilla luciferase-encoding mRNAs containing unmodified uridines (U), pseudouridines ( ) or N1 -methylpseudouridines (m I ) featuring either an IgG signal peptide (lanes 1-3) or a human cystatin C signal peptide (lanes 4-6). All mRNAs feature poly(A) tails of 30 nucleotides. M represents a lane containing DNA size markers with selected sizes indicated in kilobases (kb) on the left.

[0033] Figure 8 - Comparison of secreted Renilla luciferase activity. Measurements of Renilla luciferase (RL) translational output from the culture media of HEK-293T cells transfected with 1 pg of mRNA transcribed with uridine (U), pseudouridine ( ) or Nl- methylpseudourdine (m l ) present in the transcription mix. The luciferase activity present in the HEK-293T cell culture media resulting from mRNAs containing unmodified or modified nucleotides and featuring either the IgG (striped bars) or the human cystatin C signal peptide (black filled bars) is depicted. The white bar represents a control measuring luciferase activity in culture media from cells transfected with an unmodified mRNA encoding RL lacking a signal peptide (baseline activity resulting from ruptured cells). All mRNAs feature poly(A) tails of 30 nucleotides. The relative luciferase activity is represented in Relative Light Units (RLU). Data is represented as mean ± standard deviation.

[0034] Figure 9 - Benchmarking UTRs with unmodified and mlT-modified mRNA.

[0035] Comparison of Renilla luciferase activity measured using cell lysates from HEK-293T cells transfected with lOOng mRNA per well in a 96-well plate (48 hours). mRNAs were transcribed using uridine nucleotides (unmodified mRNA, striped bars) or Nl- m ethylpseudouridine nucleotides (ml'P-modified mRNA, black bars) and featured a 114 or 198 nucleotide poly(A) tail. The UTRs consisted of the 5’UTR and 3’UTR of Modema’s mRNA-1273, the H2AX 5’UTR with a shortened version of the COL6A2 3’UTR or the H2AX 5’UTR with the full-length COL6A2 3’UTR. The relative luciferase activity is represented in Relative Light Units (RLU). Data is represented as mean ± standard deviation (3 biological repeats).

[0036] Detailed description

[0037] General definitions

[0038] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person skilled in the art to which this invention belongs.

[0039] In general, the term “comprising” is intended to mean including but not limited to. For example, the phrase “a pharmaceutical composition comprising an RNA molecule ” should be interpreted to mean that the pharmaceutical composition comprises the RNA molecule, but that the pharmaceutical composition may comprise further components (for example an excipient, a lipid nanoparticle, etc. . .).

[0040] In some embodiments of the invention, the word “comprising” is replaced with the phrase “consisting of” . The term “consisting of” is intended to be limiting. For example, the phrase “a pharmaceutical composition consisting of an RNA molecule" should be understood to mean that the pharmaceutical composition contains the RNA molecule and no further components.

[0041] In some embodiments of the invention, the word “comprising” is replaced with the phrase “consisting essentially of” . The term “consisting essentially of” means that specific further components can be present, namely those not materially affecting the essential characteristics of the subject matter. For example, the phrase “a pharmaceutical composition consisting essentially of an RNA molecule ” indicates that the pharmaceutical composition may further comprise one or more excipients that have no particular function.

[0042] For the purpose of this invention, in order to determine the percent identity of two sequences (such as two polynucleotide or two polypeptide sequences), the sequences are aligned for optimal comparison purposes (e.g. gaps can be introduced in a first sequence for optimal alignment with a second sequence). The nucleotide or amino acid residues at each position are then compared. When a position in the first sequence is occupied by the same nucleotide or amino acid as the corresponding position in the second sequence, then the nucleotides or amino acids are identical at that position. The percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity = number of identical positions / total number of positions in the reference sequence x 100).

[0043] Typically, the sequence comparison is carried out over the length of the reference sequence. For example, if the user wished to determine whether a given (“test”) sequence is 95% identical to SEQ ID NO: 1, SEQ ID NO: 1 would be the reference sequence. To assess whether a sequence is at least 95% identical to SEQ ID NO: 1 (an example of a reference sequence), the skilled person would carry out an alignment over the length of SEQ ID NO: 1, and identify how many positions in the test sequence were identical to those of SEQ ID NO: 1. If at least 95% of the positions are identical, the test sequence is at least 95% identical to SEQ ID NO: 1. If the sequence is shorter than SEQ ID NO: 1, the gaps or missing positions should be considered to be non-identical positions. If the sequence is longer than SEQ ID NO: 1, the additional positions that are within the section aligned with SEQ ID NO: 1 should be considered to be non-identical positions.

[0044] The skilled person is aware of different computer programs that are available to determine the homology or identity between two sequences. For instance, a comparison of sequences and determination of percent identity between two sequences can be accomplished using a mathematical algorithm. In an embodiment, the percent identity between two amino acid or nucleic acid sequences is determined using the Needleman and Wunsch (1970) algorithm which has been incorporated into the GAP program in the Accelrys GCG software package (available at http: / / www.accelrys.com / products / gcg / ), using either a Blosum 62 matrix or a PAM250 matrix, and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6.

[0045] As used herein, reference to a number of nucleotide or amino acid substitutions, insertions or deletions should be calculated based on the number of individual nucleotide or amino acid residues that have been substituted, inserted, or deleted. For example, deletion or insertion of a contiguous stretch of five amino acids should be considered a deletion or insertion of five amino acids, rather than a single deletion or insertion. The minimum number of substitutions, insertions or deletions needed to arrive at the modified sequence should be used. For example, modification of the nucleotide sequence ACGCCGCA to ACTA should be considered a total of 5 nucleotide substitutions, insertions or deletions (e.g. deletion of GCCG and substitution of C to T) rather than 6 substitutions, insertions or deletions (e.g. deletion of GCCGC and insertion of T).

[0046] The term “about” or “around” when referring to a value refers to that value but within a reasonable degree of scientific error. Optionally, a value is “about X” or “around X” if it is within 10%, within 5% or within 1% of X.

[0047] The singular forms “a”, “an” and “the ” include plural references unless the context clearly dictates otherwise. All publications, patents and patent applications cited herein, whether Supra or Infra, are hereby incorporated by reference in their entirety.

[0048] RNA molecule

[0049] The present invention relates to an RNA molecule comprising a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR), and an open reading frame (ORF) and / or a non-coding functional RNA sequence. The RNA molecule may comprise a 5’ UTR, a 3’ UTR and an ORF. The RNA molecule may comprise a 5’ UTR, a 3’ UTR and a noncoding functional RNA sequence. The RNA molecule may comprise a 5’ UTR, a 3’ UTR, an ORF and a non-coding functional RNA sequence.

[0050] In the RNA molecule, the 5’ UTR comprises an Histone 2A family member X (H2AX, H2AFX, H2A.X, H2A / X, Gene ID 3014, a sequence that overlaps with the 3’ side of the DPAGT1 gene: ID: XM 047426508.1) 5’ UTR or variant thereof, the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2, also known as: CDM1, BTHLM1, PP3610, UCMD1B, BTHLM1B, Gene ID 1292) 3’ UTR or variant thereof, and / or the ORF (where present) comprises a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof. For example, the RNA molecule may comprise a H2AX 5’ UTR or variant thereof and a COL6A2 3 ’ UTR or variant thereof. The RNA molecule may comprise a H2AX 5’ UTR or variant thereof. The RNA molecule may comprise a COL6A2 3’ UTR or variant thereof. The RNA molecule may comprise a H2AX 5’ UTR or variant thereof and an ORF comprising a sequence encoding a cystatin C signal peptide or variant thereof. The RNA molecule may comprise a COL6A2 3’ UTR or variant thereof and an ORF comprising a sequence encoding a cystatin C signal peptide or variant thereof. The RNA molecule may comprise an H2AX 5’ UTR or variant thereof, a COL6A2 3’ UTR or variant thereof and an ORF comprises a sequence encoding a cystatin C signal peptide or variant thereof.

[0051] The H2AX 5’ UTR may be a human H2AX 5’ UTR (as set out in SEQ ID NO: 1) or variant thereof. The H2AX 5’ UTR may be a non-human H2AX 5’ UTR, or variant thereof. This may be useful for embodiments where the RNA molecule is expressed and / or translated in the non-human species or in a cell from the non-human species. For example, The H2AX 5’ UTR may be a non-human mammal H2AX 5’ UTR, such as a primate H2AX 5’ UTR (such as a chimpanzee or a green monkey), a rodent H2AX 5’ UTR (such as a mouse, rat or hamster), or a domesticated animal H2AX 5’ UTR (such as a horse, cow, sheep, pig, goat, dog or cat).

[0052] The COL6A2 3’ UTR may be a human COL6A2 3’ UTR (as set out in SEQ ID NO: 3) or variant thereof. The COL6A2 3’ UTR may be a human COL6A2 3’ UTR as set out in SEQ ID NO: 13, or variant thereof. The COL6A2 3’ UTR may be a non-human COL6A2 3’ UTR, or variant thereof. This may be useful for embodiments where the RNA molecule is expressed and / or translated in the non-human species or in a cell from the non-human species. For example, The COL6A2 3’ UTR may be a non-human mammal COL6 A2 3 ’ UTR, such as a primate COL6 A2 3 ’ UTR (such as a chimpanzee or a green monkey), a rodent COL6A2 3’ UTR (such as a mouse, rat or hamster), or a domesticated animal COL6A2 3 ’ UTR (such as a horse, cow, sheep, pig, goat, dog or cat).

[0053] The cystatin C signal peptide may be a human cystatin C signal peptide (as set out in SEQ ID NO: 6) or variant thereof. The cystatin C signal peptide may be a non-human cystatin C signal peptide, or variant thereof. This may be useful for embodiments where the RNA molecule is expressed and / or translated in the non-human species or in a cell from the non-human species. For example, the cystatin C signal peptide may be a non- human mammal cystatin C signal peptide, such as a primate cystatin C signal peptide (such as a chimpanzee or a green monkey), a rodent cystatin C signal peptide (such as a mouse, rat or hamster), or a domesticated animal cystatin C signal peptide (such as a horse, cow, sheep, pig, goat, dog or cat).

[0054] In the RNA molecule, the 5’ UTR may comprise a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, the 3’ UTR may comprise a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, and / or the ORF (where present) may comprise a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. The RNA molecule may comprise the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, and the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof. The RNA molecule may comprise the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, and an ORF comprising the sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. The RNA molecule may comprise the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, and an ORF comprising the sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. The RNA molecule may comprise the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, and an ORF comprising the sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. Preferably, the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or a variant thereof, the 3’ UTR comprises a sequence as set out in SEQ ID NO: 3 or a variant thereof, and / or the ORF (where present) comprises a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or a variant thereof.

[0055] In the RNA molecule, the 5’ UTR may comprise a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, the 3’ UTR may comprise a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, and / or the ORF (where present) may comprise a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. The RNA molecule may comprise the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, and the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof. The RNA molecule may comprise the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, and an ORF comprising the sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. The RNA molecule may comprise the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, and an ORF comprising the sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. Preferably, the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or a variant thereof, the 3’ UTR comprises a sequence as set out in SEQ ID NO: 13 or a variant thereof, and / or the ORF (where present) comprises a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or a variant thereof.

[0056] The term “ encoding a signal peptide" is intended to mean that the sequence comprises sequence information from which the signal peptide may be derived. A signal peptide is therefore "encoded' by the sequence when the sequence or its reverse complement comprises nucleotide triplets that can be translated (e.g. by a ribosome) to produce the amino acid sequence of the signal peptide.

[0057] In an RNA molecule, the T nucleotides found in a corresponding DNA molecule are typically replaced by uridine (U), which retains the ability to base pair with A. RNA polynucleotides therefore typically comprise a combination of A, C, G and U nucleotides. As used herein, references to a polynucleotide comprising a T is considered to be a thymidine nucleotide in DNA and a uridine nucleotide in RNA, unless explicitly defined otherwise. Other naturally-occurring, non-naturally occurring, modified and / or synthetic nucleotides may also be present in the RNA molecules described herein, particularly nucleotides that may reduce immunogenicity, improve stability, transcription or translation of the polynucleotide described herein. For example, the RNA molecule may include inosine (I), 5 ’methyl cytidine (5meC), N6-methyladenosine, pseudouridine, N1 -methylpseudouridine, deoxyuridine (dU), abasic nucleotides, threose nucleotides (TNA), glycerol nucleotides (GNA), locked nucleotides (LNA) and peptide nucleotides (PNA). In some cases, the RNA molecule may comprise 5 ’methyl cytidine (5meC), N6- methyladenosine, pseudouridine and / or N1 -methylpseudouridine. In some cases, the RNA molecule may comprise pseudouridine and / or N1 -methylpseudouridine. In some cases, the RNA molecule may comprise N1 -methylpseudouridine. As shown in the Examples, the translational capacity of an mRNA comprising a H2AX 5’UTR and / or a COL6A2 3’UTRmay be improved when incorporating N1 -methylpseudouridine. In some cases, the RNA molecule may comprise a combination of A, C, G and pseudouridine nucleotides, i.e. where pseudouridine has been incorporated instead of uridine. In some cases, the RNA molecule may comprise a combination of A, C, G and N1 -methylpseudouridine nucleotides, i.e. where N1 -methylpseudouridine has been incorporated instead of uridine. 5’ untranslated region

[0058] The RNA molecule may comprise a 5’ untranslated region (5’ UTR) of an Histone 2A family member X (H2AX) or variant thereof such as a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof. SEQ ID NO: 1 is the 5’ UTR sequence of the human H2A Histone family member X (H2AX, H2AFX, H2A.X, H2A / X, Gene ID 3014, a sequence that overlaps with the 3’ side of the DPAGT1 gene: ID: XM 047426508.1). SEQ ID NO: 2 is the 5’ UTR sequence of the Mus musculus H2A Histone family member X. Preferably, the RNA molecule comprises a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1, or a variant thereof.

[0059] The H2AX 5’ UTR variant may comprise at least 50% sequence identity to the H2AX 5’ UTR (e.g. a human or non-human H2AX 5’ UTR, as described above), such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to the H2AX 5’ UTR. In some embodiments, the H2AX 5’ UTR variant comprises a sequence at least 95% identical to the sequence of an H2AX 5’ UTR, such as a human or non-human H2AX 5’ UTR.

[0060] The variant may comprise up to 20 nucleotide substitutions, insertions or deletions of a H2AX 5’ UTR (e.g. a human or non-human H2AX 5’ UTR, as described above). In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 20 when compared to the H2AX 5’ UTR. The variant may comprise up to 15 nucleotide substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 nucleotide substitution(s), insertion(s) or deletion(s), of the H2AX 5’ UTR. The variant may comprise 1 to 20 nucleotide substitutions, insertions or deletions of the H2AX 5’ UTR, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitution(s), insertion(s) or deletion(s) of the H2AX 5’ UTR. In some embodiments, the H2AX 5’ UTR variant comprises 1 to 3 mutations relative to the sequence of an H2AX 5 ’UTR, such as a human or non-human H2AX 5’ UTR. The variant may comprise at least 50% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2.

[0061] The variant may comprise up to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2. In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 20 when compared to the sequence set out in SEQ ID NO: 1 or 2. The variant may comprise up to 15 nucleotide substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 nucleotide substitution(s), insertion(s) or deletion(s), of a sequence set out in SEQ ID NO: 1 or 2. The variant may comprise 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitution(s), insertion(s) or deletion(s) of a sequence set out in SEQ ID NO: 1 or 2.

[0062] The RNA molecule comprising an H2AX 5’ UTR or variant thereof, may have at least 70% of the translational capacity of a reference RNA molecule. For example, the RNA molecule may have at least 80%, at least 90% or at least 100% of the translational capacity of a reference RNA molecule. The reference RNA molecule is typically an equivalent RNA molecule comprising the 5’ UTR of SEQ ID NO: 1 or 2. In other words, the reference RNA molecule is identical to the RNA molecule except in the sequence of the 5’ UTR (and in some cases as described below, the reference sequence may also differ in the sequence of the 3’ UTR). If the H2AX 5’ UTR comprises a sequence as set out in SEQ ID NO: 1, then the reference sequence may comprise the sequence of SEQ ID NO: 1. If the H2AX 5’ UTR comprises a sequence as set out in SEQ ID NO: 2, then the reference sequence comprises the sequence of SEQ ID NO: 2.

[0063] In some cases where the RNA molecule comprises a 5’ UTR comprising an H2AX 5’ UTR or variant thereof, and a 3’ UTR comprising a COL6A2 3’ UTR or variant thereof, the RNA molecule may have at least 70% of the translational capacity of a reference RNA molecule. For example, the RNA molecule may have at least 80%, at least 90% or at least 100% of the translational capacity of a reference RNA molecule. In this case, the reference RNA molecule is typically an equivalent RNA molecule comprising the 5’ UTR of SEQ ID NO: 1 or 2 and the 3’ UTR of any one of SEQ ID NOs: 3 to 5. In other words, the reference RNA molecule is identical to the RNA molecule except in the sequence of the 5’ UTR and the 3’ UTR. If the RNA molecule comprises UTRs comprising the sequences as set out in SEQ ID NOs: 1 and 3, then the reference sequence comprises the sequence of SEQ ID NOs: 1 and 3. If the RNA molecule comprising UTRs comprising the sequences as set out in SEQ ID NOs: 2 and 5, then the reference sequence comprises the sequence of SEQ ID NOs: 2 and 5, and the like. The reference RNA molecule may be an equivalent RNA molecule comprising the 5’ UTR of SEQ ID NO: 1 or 2 and the 3’ UTR of any one of SEQ ID NOs: 3 to 5 and 13.

[0064] Translational capacity may be calculated by any means known to the skilled person, for example by measured levels of a protein encoded by the RNA molecule and reference RNA molecule, for example expression levels. Translation of the protein encoded by the RNA molecule is under the control of the 5’ UTR and / or the 3 ’UTR. Measured levels can be determined directly or indirectly, for example, by purifying and quantifying the amount of protein present in a sample or by performing an enzymatic reaction or fluorescence measurement that correlates to the level of protein. Typically, the protein is a reporter protein, i.e. a protein that has a characteristic that can be easily measured and quantified. For example, the reporter protein may be a fluorescent protein, such as GFP, or an enzyme, such as a luciferase. Translational capacity is measured under standardised conditions to allow for a comparison between the RNA molecule and the reference RNA molecule. For example, translational capacity may be calculated as the measured level of luciferase activity in a sample. The sample may be a whole cell lysate, as used in Figure 3, Example 1. The sample may be secreted protein, e.g. cell culture media, as used in Figure 8, Example 2. Translational capacity may be calculated after a set time period from transfection into a cell, such as 24, 48 or 72 hours after transfection. Translational capacity may be calculated based on a set level of RNA transfected into a cell, such as 0.1 pg, 1 pg or 2 pg of RNA. Translational capacity may be calculated based on transfection into a specific cell, such as a HEK293T cell. The RNA molecule and the reference molecule may each comprise a sequence encoding a luciferase, such as Renilla luciferase (e.g. GenBank: UniProtKB / Swiss-Prot: P27652.1). Whole cell lysate may be prepared by culturing cells (e.g. HEK-293T cells) in DMEM + 10 % foetal bovine serum in 6 well plates, removing the cell-culture media 48-hours posttransfection, lysing the resulting cells with lysis buffer, shaking the plate at 500 rpm for 40 mins, centrifuging the lysate at 2000 rpm for 5 mins in a bench centrifuge and retaining the supernatant (e.g. 800 pl), as explained in Example 1. Luciferase activity in each case may be measured using the Pierce Renilla Luciferase Glow Assay Kit (Catalog No 16167) using a Clariostar instrument (emission setting F:530-40; 1 second integration time; autoadjust carried out on the well with the highest raw value measured in the initial read). For a non-secreted protein (e.g. cell-associated by intracellular or membrane expression), translational capacity may be calculated using whole-cell lysate as described above, e.g. as the measured level of luciferase activity in whole cell lysate 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule in a HEK293T host cell, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase, as exemplified in Example 1 and Figure 3.

[0065] In some cases, the 5’ UTR consists of the sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof. In some cases, the 5’ UTR comprises the sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof, and further comprises one or more further sequences. In some cases, the 5’ UTR does not comprise the sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof. For example, the RNA molecule may comprise a 3’ UTR comprising the sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof, and a 5’ UTR typically used in the field. The RNA molecule may comprise a 3’ UTR comprising the sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof, and a 5’ UTR typically used in the field. The sequence of the 5’ UTR that is not the sequence as set out in SEQ ID NO: 1 or 2 (e.g. the one or more further sequences of the 5 ’UTR discussed above), or a variant thereof, may: comprise a ribosomal binding site, such as a Kozak sequence or a translation initiation site; comprise an internal ribosome entry site (IRES), as described further herein; comprise other regulatory elements such as one or more upstream ORFs (uORFs) and / or secondary structure elements, such as hairpin-loops; in some cases, not comprise a uORF (a nucleic acid sequence comprising an AUG start codon, and optionally comprising a stop codon prior to the start codon of the main ORF). The presence of a uORF in the 5’UTR may attenuate translation of the subsequent ORF); and / or comprise a 5’UTR or fragment thereof of a TOP gene, such as a 5’ terminal oligopyrimidine tract (TOP) gene encoding a ribosomal large protein, optionally lacking the 5’ TOP motif.

[0066] 3’ untranslated region

[0067] The RNA molecule may comprise a 3’ untranslated region (3’ UTR) of a Collagen type VI alpha 2 chain (COL6A2) or variant thereof, such as a sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof. The RNA molecule may comprise a 3’ UTR of COL6A2 as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof. SEQ ID NO: 3 is the 3’ UTR sequence of the human Collagen type VI alpha 2 chain (COL6A2, also known as: CDM1, BTHLM1, PP3610, UCMD1B, BTHLM1B, Gene ID 1292). SEQ ID NO: 4 is the 3’ UTR sequence of the Mus musculus Collagen type VI alpha 2 chain. SEQ ID NO: 5 is the 3’ UTR sequence of the Rattus novergicus Collagen type VI alpha 2 chain. SEQ ID NO: 13 is a short version of the 3’ UTR sequence of the human COL6A2, which has been shown in the examples to have comparable or improved functionality when compared to SEQ ID NO: 3, indicating that SEQ ID NO: 13 is a functional human COL6A2 3 ’UTR. Preferably, the RNA molecule comprises a 3’ UTR that comprises SEQ ID NO: 3 or a variant thereof. The RNA molecule may comprise a 3’ UTR that comprises SEQ ID NO: 13 or a variant thereof.

[0068] The COL6A2 3’ UTR variant may comprise at least 50% sequence identity to the COL6A2 3’ UTR (e.g. a human or non-human COL6A2 3’ UTR, as described above), such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to the COL6A2 3’ UTR. In some embodiments, the COL6A2 3’ UTR variant comprises a sequence at least 95% identical to the sequence of an COL6A2 3’ UTR, such as a human or non-human COL6A2 3’ UTR.

[0069] The variant may comprise up to 20 nucleotide substitutions, insertions or deletions of a COL6A2 3’ UTR (e.g. a human or non-human COL6A2 3’ UTR, as described above). In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 20 when compared to the COL6A2 3’ UTR. The variant may comprise up to 15 nucleotide substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 nucleotide substitution(s), insertion(s) or deletion(s), of the COL6A2 3’ UTR. The variant may comprise 1 to 20 nucleotide substitutions, insertions or deletions of the COL6A2 3’ UTR, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitution(s), insertion(s) or deletion(s) of the H2AX 5’ UTR. In some embodiments, the COL6A2 3’ UTR variant comprises 1 to 3 mutations relative to the sequence of an COL6 A2 3 ’ UTR, such as a human or non-human COL6 A2 3 ’ UTR.

[0070] The variant may comprise at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5. The variant may comprise at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13.

[0071] The variant may comprise up to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5. In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 20 when compared to the sequence set out in any one of SEQ ID NOs: 3 to 5. The variant may comprise up to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13. The variant may comprise up to 15 nucleotide substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 nucleotide substitution(s), insertion(s) or deletion(s), of a sequence set out in any one of SEQ ID NOs: 3 to 5. The variant may comprise up to 15 nucleotide substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 nucleotide substitution(s), insertion(s) or deletion(s), of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13. The variant may comprise 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitution(s), insertion(s) or deletion(s) of a sequence set out in any one of SEQ ID NOs: 3 to 5. The variant may comprise 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitution(s), insertion(s) or deletion(s) of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13.

[0072] The RNA molecule comprising a 3’ UTR comprising a COL6A2 3’ UTR or variant thereof, may have at least 70% of the translational capacity of a reference RNA molecule. For example, the RNA molecule may have at least 80%, at least 90% or at least 100% of the translational capacity of a reference RNA molecule. The reference RNA molecule is typically an equivalent RNA molecule comprising the 3’ UTR of any one of SEQ ID NOs: 3 to 5. The reference RNA molecule may be an equivalent RNA molecule comprising the 3’ UTR of any one of SEQ ID NOs: 3 to 5 and 13. In other words, the reference RNA molecule is identical to the RNA molecule except in the sequence of the 3’ UTR (and in some cases as described herein, the reference sequence may also differ in the sequence of the 5’ UTR). If the 3’ UTR comprises a sequence as set out in SEQ ID NO: 3, then the reference sequence may comprise the sequence of SEQ ID NO: 3. If the 3’ UTR comprises a sequence as set out in SEQ ID NO: 4, then the reference sequence may comprise the sequence of SEQ ID NO: 4, and the like.

[0073] Translational capacity may be calculated by any means known to the skilled person, for example by measured levels of a protein encoded by the RNA molecule and reference RNA molecule, for example expression levels. Translation of the protein encoded by the RNA molecule is under the control of the 5’ UTR and / or the 3 ’UTR. Measured levels can be determined directly or indirectly, for example, by purifying and quantifying the amount of protein present in a sample or by performing an enzymatic reaction or fluorescence measurement that correlates to the level of protein. Typically, the protein is a reporter protein, i.e. a protein that has a characteristic that can be easily measured and quantified. For example, the reporter protein may be a fluorescent protein, such as GFP, or an enzyme, such as a luciferase. Translational capacity is measured under standardised conditions to allow for a comparison between the RNA molecule and the reference RNA molecule. For example, translational capacity may be calculated as the measured level of luciferase activity in a sample. The sample may be a whole cell lysate, as used in Figure 3, Example 1. The sample may be secreted protein, e.g. cell culture media, as used in Figure 8, Example 2. Translational capacity may be calculated after a set time period from transfection into a cell, such as 24, 48 or 72 hours after transfection. Translational capacity may be calculated based on a set level of RNA transfected into a cell, such as 0.1 pg, 1 pg or 2 pg of RNA. Translational capacity may be calculated based on transfection into a specific cell, such as a HEK293T cell. The RNA molecule and the reference molecule may each comprise a sequence encoding a luciferase, such as Renilla luciferase (e.g. GenBank: UniProtKB / Swiss-Prot: P27652.1). Whole cell lysate may be prepared by culturing cells (e.g. HEK-293T cells) in DMEM + 10 % foetal bovine serum in 6 well plates, removing the cell-culture media 48-hours posttransfection, lysing the resulting cells with lysis buffer, shaking the plate at 500 rpm for 40 mins, centrifuging the lysate at 2000 rpm for 5 mins in a bench centrifuge and retaining the supernatant (e.g. 800 pl), as explained in Example 1. Luciferase activity in each case may be measured using the Pierce Renilla Luciferase Glow Assay Kit (Catalog No 16167) using a Clariostar instrument (emission setting F:530-40; 1 second integration time; autoadjust carried out on the well with the highest raw value measured in the initial read). For a non-secreted protein (e.g. cell-associated by intracellular or membrane expression), translational capacity may be calculated using whole-cell lysate as described above, e.g. as the measured level of luciferase activity in whole cell lysate 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule in a HEK293T host cell, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase, as exemplified in Example 1 and Figure 3. In some cases, the 3’ UTR consists of the sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof. In some cases, the 3’ UTR comprises the sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof, and further comprises one or more further sequences. In some cases, the 3’ UTR consists of the sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof. In some cases, the 3’ UTR comprises the sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof, and further comprises one or more further sequences. In some cases, the 3’ UTR does not comprise the sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof. In some cases, the 3’ UTR does not comprise the sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof. For example, the RNA molecule may comprise a 5’ UTR comprising the sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof, and a 3’ UTR typically used in the field. The sequence of the 3’ UTR that is not the sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 (e.g. the one or more further sequences of the 3’ UTR discussed above), or a variant thereof, may: comprise regulatory elements such as one or more inverted repeats that are capable of forming a stem-loop structure (stem-loop structures are known to act as a barrier for exoribonucleases and / or interact with proteins known to increase RNA stability); comprise a binding site for a regulatory protein and / or an miRNA; comprise an AU-rich element (ARE); function to enhance protein expression and / or to improve the half-life of the transcribed polynucleotide, for example, when compared to a control lacking a 3’ UTR; and / or comprise the 3 ’UTR of an albumin gene, an a-globin gene, a P-globin gene, a ribosomal protein gene, a tyrosine hydroxylase gene, a lipoxygenase gene, or a collagen alpha gene.

[0074] Signal peptide

[0075] The RNA molecule may comprise an open reading frame (ORF) comprising a sequence encoding a signal peptide of a Cystatin C, such as a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof. SEQ ID NO: 6 is the signal peptide of human cystatin C. SEQ ID NO: 7 is the signal peptide of Mus musculus cystatin C. Preferably, the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide as set out in SEQ ID NO: 6.

[0076] The cystatin C signal peptide variant may comprise at least 50% sequence identity to the cystatin C signal peptide (e.g. a human or non-human cystatin C signal peptide, as described above), such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to the cystatin C signal peptide. In some embodiments, the cystatin C signal peptide variant comprises a sequence at least 95% identical to the sequence of a cystatin C signal peptide, such as a human or non-human cystatin C signal peptide.

[0077] The cystatin C signal peptide variant may comprise up to 20 amino acid substitutions, insertions or deletions of a cystatin C signal peptide (e.g. a human or non-human cystatin C signal peptide, as described above). In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 20 when compared to the cystatin C signal peptide. The variant may comprise up to 15 amino acid substitutions, insertions or deletions, such as up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 amino acid substitution(s), insertion(s) or deletion(s), of the cystatin C signal peptide. The variant may comprise 1 to 20 amino acid substitutions, insertions or deletions of the cystatin C signal peptide, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 amino acid substitution(s), insertion(s) or deletion(s) of the cystatin C signal peptide. In some embodiments, the cystatin C signal peptide variant comprises 1 to 3 mutations relative to the sequence of an cystatin C signal peptide, such as a human or non-human cystatin C signal peptide.

[0078] The variant may comprise at least 50% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7. The variant may comprise up to 10 amino acid substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7. In other words, the variant may comprise a mix of modifications, such as substitutions, insertions or deletions, totalling up to 10 when compared to the sequence set out in SEQ ID NO: 6 or 7. The variant may comprise up to 9 amino acid substitutions, insertions or deletions, such as up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or 1 amino acid substitution(s), insertion(s) or deletion(s), of a sequence set out in SEQ ID NO: 6 or 7. The variant may comprise 1 to 10 amino acid substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 amino acid substitution(s), insertion(s) or deletion(s) of a sequence set out in SEQ ID NO: 6 or 7.

[0079] The variant may comprise one or more amino acids substituted with an alternative amino acid having similar properties. Some properties of the 20 main amino acids, which can be used to select suitable substituents, are as follows:

[0080] The variant may comprise a derivatised amino acid, for example, a labelled or nonnatural amino acid, providing the function of the signal peptide is not significantly adversely affected.

[0081] The signal peptide encoded by the variant retains its ability to function as a signal peptide, e.g. to target the cell for translocation into the endoplasmic reticulum for secretion or membrane expression. The signal peptide encoded by the variant retains its ability to be cleaved by signal peptidase during or after translocation into the endoplasmic reticulum of a cell. The signal peptide encoded by the variant typically forms a single alpha helix. Typically, the signal peptide encoded by the variant does not comprise any substitutions of a hydrophobic amino amid for a charged amino acid. In some cases, the signal peptide encoded by the variant does not comprise any substitutions of a hydrophobic amino acid for a charged amino acid or polar amino acid. In some cases, the signal peptide encoded by the variant does not comprise any additions or deletions of a charged amino acid, and / or does not comprise a substitution or deletion of the arginine in the signal peptide sequence.

[0082] The RNA molecule comprising an ORF comprising a sequence encoding a cystatin C signal peptide or variant thereof, may have the same or greater translational capacity (e.g. measured as amount of secreted protein) than a reference RNA molecule. For example, the RNA molecule may have at least 100%, at least 150%, at least 200%, at least 250% or at least 300% of the translational capacity of a reference RNA molecule. The reference RNA molecule is typically an equivalent RNA molecule comprising an ORF comprising a sequence encoding &Mus musculus IgG signal peptide, for example as set out in SEQ ID NO: 9. In other words, the reference RNA molecule is identical to the RNA molecule except in the sequence encoding the signal peptide within the ORF. SEQ ID NO: 9 is the sequence of the Mus musculus IgG signal peptide, and is identical to the sequence of a corresponding Homo sapiens IgG signal peptide.

[0083] Translational capacity may be calculated by any means known to the skilled person, for example by measured levels of a protein encoded by the RNA molecule and reference RNA molecule, for example expression levels. The protein comprises the signal peptide encoded by the RNA molecule or reference RNA molecule described above, i.e. at its N- terminus. The protein does not comprise an additional, for example endogenous, signal peptide sequence. Measured levels can be determined directly or indirectly, for example, by purifying and quantifying the amount of protein present in a sample or by performing an enzymatic reaction or fluorescence measurement that correlates to the level of protein. Typically, the protein is a reporter protein, i.e. a protein that has a characteristic that can be easily measured and quantified. For example, the reporter protein may be a fluorescent protein, such as GFP, or an enzyme, such as a luciferase. Translational capacity is measured under standardised conditions to allow for a comparison between the RNA molecule and the reference RNA molecule. For example, translational capacity may be calculated as the measured level of luciferase activity in a sample. The sample may be secreted protein, e.g. cell culture media, as used in Figure 8, Example 2. Translational capacity may be calculated after a set time period from transfection into a cell, such as 24, 48 or 72 hours after transfection. Translational capacity may be calculated based on a set level of RNA transfected into a cell, such as 0.1 pg, 1 pg or 2 pg of RNA. Translational capacity may be calculated based on transfection into a specific cell, such as a HEK293T cell. The RNA molecule and the reference molecule may each comprise a sequence encoding a luciferase, such as Renilla luciferase (e.g. GenBank: UniProtKB / Swiss-Prot: P27652.1). Secreted protein may be prepared by culturing cells (e.g. HEK-293T cells) in DMEM (e.g. FluoroBrite™ DMEM) + 10% foetal bovine serum in 6-well plates, and collecting the cell media 48 hours posttransfection, as explained in Example 2. Luciferase activity in each case may be measured using the Pierce Renilla Luciferase Glow Assay Kit (Catalog No 16167) using a Clariostar instrument (emission setting F:530-40; 1 second integration time; autoadjust carried out on the well with the highest raw value measured in the initial read). Translational capacity may be calculated as the measured level of secreted luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule into a HEK293T host cell, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase, as exemplified in Example 2 and Figure 8.

[0084] ORF

[0085] The RNA molecule may comprise an open reading frame (ORF), which may or may not comprise a sequence encoding a signal peptide as described herein. The ORF may be operably linked to the 5’ UTR and / or the 3’ UTR. The ORF typically encodes an amino acid sequence. The ORF may also encode nonprotein coding sequences, such as introns or other regulatory elements. The ORF typically begins with a start codon and terminates with a stop codon. The ORF may encode a peptide, polypeptide or protein. The ORF may encode a secreted protein. The ORF may encode an antibody, such as an antibody light chain, an antibody heavy chain, an scFv, a VHH, or the like. The ORF may encode a therapeutic peptide, polypeptide or protein. Therapeutic peptides, polypeptides and proteins are well-known to the skilled person, and may be useful for the prevention or treatment of an inherited or acquired disease. The ORF may encode an antigen. An antigen may be considered a therapeutic peptide, polypeptide or protein that achieves the therapeutic effect by eliciting an immune response in a host, such as a human host. The antigen may be a viral, bacterial, fungal, parasitic, allergenic, autoimmune or tumorigenic antigen. The viral, bacterial, fungal or parasitic antigen may be derived from a virus, bacteria, fungus or parasite that is a pathogen, e.g. a human pathogen. For example, the viral, bacterial, fungal or parasitic antigen may be derived from an infectious virus, bacteria, fungus or parasite, e.g. a virus, bacteria, fungus or parasite that is capable of infecting humans. The ORF may encode a reporter protein, i.e. a protein that has a characteristic that can be easily measured and quantified. For example, the reporter protein may be a fluorescent protein, such as GFP, or an enzyme, such as a luciferase.

[0086] The term “peptide” typically refers to a molecule comprising 2 or more consecutive amino acids linked by peptide bonds, more typically 2 to 50 consecutive amino acids. The term “polypeptide” typically refers to a molecule comprising 50 or more consecutive amino acids linked by peptide bonds. There is no practical upper limit to the length of the polypeptide. The term “protein” can typically be used interchangeably with the term “polypeptide” , but also encompass structures formed of more than one polypeptide chain, such as via cleavage and / or disulphide bond formation. The peptide, polypeptide or protein may comprise substances that are not amino acids, such as post-translational modifications, covalently- or non-covalently-bound co-factors and the like.

[0087] Where the ORF comprises a sequence encoding a signal peptide, the signal peptide is at the N-terminus of the amino acid sequence encoded by the ORF. Where the ORF comprises a sequence encoding a signal peptide, the ORF may encode a secreted protein, such as an antibody as described herein.

[0088] The RNA molecule may comprise more than one ORF, such as two ORFs or three ORFs. For example, the RNA molecule may encode a multi-chain protein, such as an antibody. In this case, one ORF may encode an antibody heavy chain and a second ORF may encode an antibody light chain.

[0089] The RNA molecule typically does not comprise an ORF encoding a histone H2AX, a COL6A2 or a cystatin C.

[0090] Non-coding functional RNA

[0091] The RNA molecule may comprise a non-coding functional RNA sequence. The noncoding functional RNA sequence may be in addition, or as an alternative, to an ORF.

[0092] The non-coding functional RNA sequence may comprise an RNA sequence forming part or whole of a ribonucleoprotein, a ribozyme, ribosomal RNA, transfer RNA, small nuclear RNA, small nucleolar RNA, Y RNA, microRNA, antisense RNA, and / or an RNA sequence that is capable of sequestering miRNAs or RNA binding proteins that can affect mRNA metabolism. The antisense RNA may be capable of hybridising to a target sense transcript, to thereby inhibit further transcription or translation of said sense transcript. The RNA sequence that is capable of sequestering miRNAs and / or RNA binding proteins, may comprise a sequence capable of hybridising to miRNAs and / or a sequence that acts as a binding site for RNA binding proteins. Without being limited to theory, it is believed that such non-coding functional RNA sequences may control expression of other RNA molecules.

[0093] Other features of the RNA molecule

[0094] The RNA molecule may comprise a poly(A) tail. The RNA molecule may comprise a poly(A) signal sequence, such as a bGH poly(A) signal sequence. The poly(A) tail may be comprised at or towards the 3’ end of the RNA molecule. The poly(A) tail typically comprises at least 30 adenosine nucleotides, such as at least 60, at least 100, at least 120 or at least 150 adenosine nucleotides. The poly(A) tail may comprise no more than 300 adenosine nucleotides, such as no more than 250, no more than 200 or no more than 150 adenosine nucleotides. The poly(A) tail may comprise 30-300 adenosine nucleotides, such as 60-250, 100-250, 100-200, about 120 or about 150 adenosine nucleotides. The poly(A) tail may consist of adenosine nucleotides. In some cases, the adenosine nucleotides in the poly(A) tail are separated into segments by one or more spacer sequences. Segmented poly(A) tails are described in WO 2016 / 091391 Al, WO 2020 / 074642 Al and WO 2016 / 005324 Al. The poly(A) tail may comprise at least 60 adenosine nucleotides, such as at least 100 adenosine nucleotides, and the RNA molecule may comprise N1 -methylpseudouridine.

[0095] As described above, the RNA molecule comprises: an H2AX 5’ UTR or variant thereof, a COL6A2 3’ UTR or variant thereof, and / or an ORF comprising a sequence encoding a cystatin C signal peptide or variant thereof. Variants of the 5’ UTR, the 3’ UTR and the signal peptide are described above. For example, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; and / or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 10 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7.

[0096] In some cases, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; and a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13.

[0097] In some cases, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; and a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 10 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7.

[0098] In some cases, the RNA molecule may comprise: a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13; and a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 10 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7.

[0099] In some cases, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97% or at least 98% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 20 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13, such as 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13; and a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 50% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, such as at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 10 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3 or 1 to 2 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7.

[0100] In some cases, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least 80% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 10, or 1 to 5 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least 80% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 10, or 1 to 5 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, or a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least 80% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 10, or 1 to 5 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 80% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 5 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as the 5’ UTR and the 3’ UTR, the 5’ UTR and the sequence encoding the signal peptide, the 3’ UTR and the sequence encoding the signal peptide, or the 5’ UTR, the 3’ UTR and the sequence encoding the signal peptide.

[0101] In some cases, the RNA molecule may comprise: a 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, wherein the variant comprises (i) at least at least 90% sequence identity to a sequence as set out in SEQ ID NO: 1 or 2; and / or (ii) 1 to 4 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 1 or 2; a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, wherein the variant comprises (i) at least at least 90% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5; and / or (ii) 1 to 4 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5, a 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof, wherein the variant comprises (i) at least at least 90% sequence identity to a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or (ii) 1 to 4 nucleotide substitutions, insertions or deletions of a sequence set out in any one of SEQ ID NOs: 3 to 5 and 13; and / or a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof, wherein the variant comprises (i) at least 90% sequence identity to a sequence as set out in SEQ ID NO: 6 or 7, and / or (ii) 1 to 4 nucleotide substitutions, insertions or deletions of a sequence set out in SEQ ID NO: 6 or 7, such as the 5’ UTR and the 3’ UTR, the 5’ UTR and the sequence encoding the signal peptide, the 3’ UTR and the sequence encoding the signal peptide, or the 5’ UTR, the 3’ UTR and the sequence encoding the signal peptide.

[0102] The RNA molecule may comprise, in order: the 5’ UTR, the ORF and / or the non-coding functional RNA sequence and the 3’ UTR. The RNA molecule may comprise, in order: 5’ UTR, the ORF and / or the non-coding functional RNA sequence, the 3’ UTR and the poly(A) tail. In this context, the term “ / / / order" refers to the arrangement of features in the sequence of the RNA molecule. If the RNA molecule comprises an ORF and may be directly translated to produce a peptide, polypeptide or protein, then the term “in order" refers to the direction 5’ to 3’ in the RNA molecule. If a complementary stand must first be generated before translation of an ORF to produce a peptide, polypeptide or protein, then the term “in order" refers to the direction 3’ to 5’ in the RNA molecule. Typically, the term “in order" for an RNA molecule refers to the 5’ to 3’ direction of reading.

[0103] The RNA molecule may comprise an internal ribosome entry site (IRES). For example, the RNA molecule may comprise an IRES and an ORF. The RNA molecule may comprise an IRES, a 5’ UTR, a 3’ UTR, and / or an ORF. The IRES is typically present in the RNA molecule 5’ or 3’ of an ORF. The RNA molecule may comprise two or more ORFs, and an IRES located between the two or more of the ORFs. In some cases, the IRES is present in a 5’ UTR, as described herein, i.e. 5’ or 3’ of the 5’ UTR sequence as described herein and 5’ of the ORF.

[0104] The IRES may, for example, be derived from an IRES of a viral genome or from a eukaryotic genome. Examples of IRESs derived from a viral genome include the picornavirus IRES, apthovirus IRES, Kaposi’s sarcoma-associated herpesvirus IRES, hepatitis A IRES, hepatitis C IRES, pestivirus IRES, Cripavirus RES, Rhopalosiphum padi virus IRES or MDV IRES. Examples of IRESs derived from a eukaryotic genome include IRESs found in the mRNA of FGF-1, FGF-2, PDGF, VEGF, IGF-II, Antennapedia, Ultrabithorax, MYT-2, NF-KB repressing factor NRF, AML1 / RUNX1, Gtx homeodomain protein, eIF4Ga eIF4Gia, eIF4G2, c-myc, L-myc, Pim-1, Protein kinase p58PITSLRE, p53, SLC7A1, Cat-1, Notch 2, Voltage-gated potassium channel, Apaf-1, XIAP, HIAP2, Bcl-xL, Bcl-2, ARC, a-subunit of calcium calmodulin dependent kinase II dendrin, MAP2, RC3, Amyloid precursor protein, BiP, HSP70, P-subunit of mitochondrial H+-ATP synthase, Ornithine decarboxylase, connexins 32 and 43, HIF-la and APC.

[0105] The RNA molecule is typically an mRNA molecule. An mRNA molecule typically comprises, in the 5’ to 3’ direction, a 5’ UTR, an ORF, and a 3’ UTR. The mRNA molecule may comprise, in the 5’ to 3’ direction, a 5’ UTR, an ORF, a 3’ UTR and a poly(A) tail.

[0106] The RNA molecule may comprise a 5’ cap. The 5’ cap may be suitable for binding to eukaryotic translation initiation factor 4E (eIF4E). The 5’ cap may be a 7- methylguanosine cap. In some cases, the 5’ cap may be a functional analogue of a 7- methylguanosine cap. The term “functional analogue” is intended to refer to a molecule at the 5’ end of the RNA molecule that retains the function of the 7-methyl guanosine cap, e.g. to bind eIF4E and enable circularisation of the mRNA molecule. In some cases, the functional analogue may enhance the properties of the mRNA molecule when compared to the 7-methyl guanosine cap, e.g. by enhancing the stability of the RNA molecule and / or recruitment of ribosomes. In some cases, the RNA molecule comprises a structure capable of initiating capindependent translation initiation. In some cases, the RNA molecule comprises the structure in addition to a 5’ cap. In some cases, the RNA molecule does not comprise a 5’ cap. The structure may be an IRES. The structure may be a cap-independent translation enhancer (CITE, otherwise known as a cap-independent translation element). A CITE typically binds mRNA-recruiting translational components, such as translation initiation factors and / or a 60S ribosomal subunit. The CITE may be comprised within the 5’ UTR and / or the 3’ UTR. The CITE may be a CITE comprised within a 3’UTR of an RNA virus, such as an RNA plant virus. The structure may be the presence of a N6- methyladenosine at the 5’ end of the RNA molecule and / or double stranded RNA regions at the 5’ end or 5’ UTR of the RNA molecule.

[0107] In some cases, the RNA molecule is a circular RNA molecule (CircRNA). In this case, the RNA molecule does not comprise a cap. The CirRNA molecule is typically covalently closed, and its non-linear form provides protection from exonuclease degradation. The circular RNA molecule may comprise a 5’ UTR, an ORF and / or a noncoding functional RNA sequence, and / or a 3’ UTR.

[0108] In some cases, the RNA molecule is a self-amplifying RNA molecule (saRNA). saRNA molecules are capable of replicating within a cell. In some cases, the saRNA molecule comprises a sequence encoding an RNA-dependent RNA polymerase. The RNA- dependent RNA polymerase may be an alphavirus RNA-dependent RNA polymerase, such as Semliki Forest virus RNA-dependent RNA polymerase or Venezuelan equine encephalitis virus RNA-dependent RNA polymerase.

[0109] Polynucleotide

[0110] The invention also relates to a polynucleotide comprising a nucleic acid sequence encoding the RNA molecule of the invention. For example, the polynucleotide may comprise a nucleic acid sequence encoding a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR) and (i) an open reading frame (ORF) and / or a non-coding functional RNA sequence, or (ii) a sequence for introducing an ORF and / or a non-coding functional RNA sequence. In the nucleic acid sequence, the encoded 5’ UTR may comprise a sequence as set out in SEQ ID NO: 1 or 2 or a variant thereof, the encoded 3’ UTR may comprise a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof, and / or the encoded ORF or the encoded sequence for introducing an ORF may comprise a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant thereof. Suitable variants and combinations thereof are described herein.

[0111] The term “encoding” is intended to mean that the nucleic acid sequence comprises sequence information from which the sequence of the RNA molecule may be derived. A feature (e.g. the 5 ’UTR, 3 ’UTR or signal peptide) is therefore “encoded" by the nucleic acid sequence when the nucleic acid sequence comprises a sequence corresponding to the feature or its reverse complement. In an embodiment wherein the feature encoded is an RNA molecule and the nucleic acid sequence is a DNA molecule, the DNA nucleotides A, C, G and T correspond to the RNA nucleotides A, C, G and U, respectively, or their reverse complement U, G, C and A, respectively. For example, where the encoded sequence is ACGUUACG, a nucleic acid sequence encoding the RNA molecule may comprise the sequence ACGTTACGon a sense strand, and / or may comprise the sequence CGTAACGT on an antisense strand.

[0112] The term nucleic acid sequence is typically a transcribable nucleic acid sequence. The term “transcribable” means that the nucleic acid sequence in the polynucleotide, or its reverse complement, is capable of being transcribed, for example, by contact by the polynucleotide with a polymerase, such as a DNA-dependent RNA polymerase or an RNA-dependent RNA polymerase, to produce a further polynucleotide, such as an RNA molecule.

[0113] The nucleic acid sequence is typically transcribed to produce a common transcript. It is, however, envisaged, that a single nucleic acid sequence may give rise to two or more difference transcribed molecules, for example, due to co- and / or post-transcription processing, such as alternative splicing. The polynucleotide may be a DNA polynucleotide or an RNA polynucleotide. The polynucleotide may comprise a sense strand (otherwise known as a non-template strand or coding strand) and / or an antisense strand (otherwise known as a template strand or non-coding strand upon which the polymerase acts). Sense and antisense strands are typically complementary. The DNA polynucleotide may be a single-stranded or doublestranded DNA polynucleotide. The RNA polynucleotide may be a single-stranded or double-stranded DNA polynucleotide. Polynucleotides have a chemical orientation defined by the position of the linking carbon in the five-carbon sugar of each consecutive nucleotide in the chain. Accordingly, sequence elements positioned sequentially along the length of a polynucleotide may be defined by the directionality of the chain of nucleotides that is either 5’ to 3’ or 3’ to 5’. DNA polynucleotides typically comprise a combination of adenosine (A), guanosine (G), cytidine (C) and thymidine (T) nucleotides. In an RNA polynucleotide, the T nucleotides are typically replaced by uridine (U), which retains the ability to base pair with A. RNA polynucleotides therefore typically comprise a combination of A, C, G and U nucleotides. As used herein, references to a polynucleotide comprising a T is considered to be a thymidine nucleotide in DNA and a uridine nucleotide in RNA, unless explicitly defined otherwise. Other naturally-occurring, non-naturally occurring, modified and / or synthetic nucleotides may also be present in the polynucleotides described herein, particularly nucleotides that may reduce immunogenicity, improve stability, transcription or translation of the polynucleotide described herein. For example, the polynucleotide may include inosine (I), 5 ’methyl cytidine (5meC), N6-methyladenosine, pseudouridine, Nl- methylpseudouridine, deoxyuridine (dU), abasic nucleotides, threose nucleotides (TNA), glycerol nucleotides (GNA), locked nucleotides (LNA) and peptide nucleotides (PNA). In some cases, the polynucleotide may comprise 5 ’methyl cytidine (5meC), N6- methyladenosine, pseudouridine and / or N1 -methylpseudouridine. In some cases, the polynucleotide may comprise pseudouridine and / or N1 -methylpseudouridine. In some cases, the polynucleotide may comprise a combination of A, C, G and pseudouridine nucleotides, i.e. where pseudouridine has been incorporated instead of uridine or thymine. In some cases, the polynucleotide may comprise a combination of A, C, G and N1 -methylpseudouridine nucleotides, i.e. where N1 -methylpseudouridine has been incorporated instead of uridine or thymine. The polynucleotide comprising a nucleic acid sequence typically comprises one or more further sequence elements.

[0114] The polynucleotide may further comprise a promoter. The promoter aids in the transcription of the nucleic acid sequence. The promoter may be the T7 promoter, the SP6 promoter or the T3 promoter. The promoter is typically upstream of the nucleic acid sequence on the polynucleotide. For example, the promoter is 5’ of the nucleic acid sequence on a sense strand, and 3’ of the nucleic acid sequence on an antisense strand.

[0115] The polynucleotide may further comprise a marker gene. The marker gene typically encodes a polynucleotide or an amino acid sequence that enables a cell comprising the polynucleotide to be distinguished from a cell that does not comprise the polynucleotide.

[0116] The marker gene may be a screenable marker gene or a selectable marker gene. The screenable marker gene may encode a fluorescent protein such as green fluorescent protein or a variant thereof. The screenable marker gene may be a bacterial LacZ gene. The selectable marker gene may be an antibiotic resistance gene, such as a gene encoding P-lactamase.

[0117] The polynucleotide may further comprise an origin of replication (ori). The origin of replication may be a single origin of chromosomal replication (oriC). The oriC is typically a bacterial oriC, such as E. coli oriC. The ori may be ColEl, pMBl, a derivative of PMBl such as that found in the pUC vector, pBR322, pSClOl, R6K, pl 5 A or Fl ori.

[0118] The polynucleotide may further comprise a multiple cloning site (MCS). An MCS typically comprises two or more unique restriction endonuclease recognition sites for allowing a nucleic acid sequence to be inserted into the polynucleotide.

[0119] The polynucleotide preferably comprises a restriction endonuclease recognition site (restriction site), typically at the terminal end of the poly(A) tail encoded by the nucleic acid sequence. In other words, cleavage at the restriction site does not separate the 5’ UTR, the 3’UTR, and the ORF and / or the non-coding functional RNA sequence encoded by the nucleic acid sequence, and preferably retains intact the entire nucleic acid sequence encoding an RNA molecule. For example, if the polynucleotide is a doublestranded DNA polynucleotide, the restriction site is typically 3’ of the sequence encoding the 3’ UTR, and where present the poly(A) tail, on the sense strand and 5’ of the sequence encoding the 3’ UTR, and where present the poly(A) tail, on the antisense strand. The purpose of the restriction site is typically to linearise the polynucleotide for in vitro transcription. The restriction site may partially or completely overlap the sequence encoding the 3’ UTR, or where present the poly(A) tail. In some cases, the restriction site is a SapI restriction site:

[0120] 5’ ...G C T C T T C (N)i v ... 3’

[0121] 3’ ...C GA GAA G (N)4A... 5’

[0122] The polynucleotide may therefore be designed such that cleavage at the restriction site does not leave any nucleotides that do not encode the 3’ UTR, or where present the poly(A) tail, at the terminal end of the linear polynucleotide, or minimises the number of nucleotides that do not encode the 3’ UTR, or where present the poly(A) tail, at the terminal end of the linear polynucleotide. In some cases, the restriction site is located -5 to 50 nucleotides, such as 0-50, 1-50, 0-26, 5-26 or 24-26 nucleotides from the terminal end of the 3’ UTR, or where present the poly(A) tail, encoded by the nucleic acid sequence. The restriction site is preferably a sequence that is capable of being recognised and cleaved by a restriction endonuclease that results in a 5’ overhang or a blunt end. In some cases, the restriction site is a sequence that is capable of being recognised and cleaved by the restriction endonuclease SapI, Stul, SspI or Xbal, or isoschizomers thereof. The polynucleotide is suitable, typically after linearization, for in vitro transcription of RNA, such as of mRNA.

[0123] The polynucleotide may comprise an inverted terminal repeat (ITR) element. An ITR element typically comprises a pair of symmetrical nucleotide sequences flanking the ends of the polynucleotide. The ITR may be an adeno-associated virus (AAV) ITR. ITRs have particular use in AAV vectors. The ITR may be an AAV2 ITR. Vector

[0124] The invention also provides a vector comprising a polynucleotide of the invention.

[0125] The vector may be a linear vector or a circular vector. The vector is typically a circular vector capable of being linearised by the action of a restriction endonuclease enzyme at a restriction site at the terminal end of the sequence encoding the poly(A) tail, as described herein.

[0126] The vector may be a viral vector. The viral vector may be an adeno-associated virus (AAV) vector. The AAV vector may be for use in AAV gene therapy applications.

[0127] The vector is typically suitable for in vitro transcription (IVT), i.e. to produce an RNA molecule encoded by the nucleic acid sequence.

[0128] Cell

[0129] The invention also provides a cell comprising the RNA molecule, the polynucleotide, or the vector, as a described herein.

[0130] The cell may be suitable for transcribing a non-coding functional RNA sequence and / or translating an ORF. For example, where the RNA molecule encodes a therapeutic peptide, polypeptide or protein, the cell may be suitable for translating the RNA molecule to produce the peptide, polypeptide or protein. The cell may be a eukaryotic cell. The eukaryotic cell may be a mammalian cell. The cell may be a human cell. The cell may be a non-human mammalian cell, such as a mouse, rat, cat, dog, pig, goat, sheep, horse, cow, camel or non-human primate cell. The cell may be isolated from a subject. The cell may be an immortal cell line. The cell may be a human embryonic kidney cell, such as a HEK293 cell. The HEK293 cell may be a HEK293T cell. The cell may be an immune cell, such as professional antigen presenting cell (for example a dendritic cell, a monocyte or a macrophage). The eukaryotic cell may be an insect cell. The insect cell may be a Drosophila S2 cell, a Spodoptera frugiperda cell (such as Sf9 and Sf21 ) or a Trichoplusia ni “Hi -Five” cell.

[0131] Where the cell comprises a polynucleotide encoding the RNA molecule of the invention, the cell is typically suitable for propagating the polynucleotide. For example, where the vector is a double-stranded circular DNA molecule comprising the polynucleotide, the cell may be suitable for producing more copies of the double-stranded circular DNA molecule. The cell may be a prokaryotic cell. The prokaryotic cell may be a bacterial cell. The bacterial cell may be an E. coli cell, such as an E. coli DH5 alpha cell. The cell may comprise a mutation that reduces recombination. For example, the cell may be an E. coli cell comprising a mutation in the recA, recAl and / or recA13 gene, such as deletion of the recA, recAl and / or recA13 gene, and / or a mutation in the recBCD gene to reduce or abolish exonuclease V activity.

[0132] Where the polynucleotide is an RNA polynucleotide encoding the RNA molecule of the invention, the cell is typically suitable for propagating the polynucleotide using an RNA- dependent RNA polymerase. The cell may endogenously express an RNA-dependent RNA polymerase capable of propagating the RNA polynucleotide. The cell may express a recombinant RNA-dependent RNA polymerase. The cell may be co-transfected with a virus or viral vector encoding an RNA-dependent RNA polymerase. The vector may encode an RNA-dependent RNA polymerase, as described herein. Any suitable cell may be used for propagating the RNA polynucleotide. The cell may be a prokaryotic cell. The prokaryotic cell may be a bacterial cell. The bacterial cell may be an E. coli cell, such as an E. coli DH5alpha cell. The cell may comprise a mutation that reduces recombination. For example, the cell may be an E. coli cell comprising a mutation in the recA, recAl and / or recA13 gene, such as deletion of the recA, recAl and / or recA13 gene, and / or a mutation in the recBCD gene to reduce or abolish exonuclease V activity. The cell may be a eukaryotic cell. The eukaryotic cell may be a mammalian cell. The cell may be a human cell. The cell may be a non-human mammalian cell, such as a mouse, rat, cat, dog, pig, goat, sheep, horse, cow, camel or non-human primate cell. The cell may be isolated from a subject. The cell may be an immortal cell line. The cell may be a human embryonic kidney cell, such as a HEK293 cell. The HEK293 cell may be a HEK293T cell. The cell may be an immune cell, such as professional antigen presenting cell (for example a dendritic cell, a monocyte or a macrophage). The eukaryotic cell may be an insect cell. The insect cell may be a Drosophila S2 cell, a Spodoptera frugiperda cell (such as Sf9 and Sf21) or a Trichoplusia ni “Hi-Five” cell.

[0133] Pharmaceutical composition

[0134] The invention further provides a pharmaceutical composition comprising the RNA molecule, the polynucleotide or the vector of the invention. The pharmaceutical composition may further comprise a pharmaceutically acceptable excipient. For the RNA molecule, polynucleotide or vector to have pharmaceutical use, the RNA molecule typically encodes a therapeutic RNA sequence, a therapeutic peptide, a therapeutic polypeptide, a therapeutic protein or an antigen, as described herein. In other words, the RNA molecule, polynucleotide or vector comprises or encodes an ORF, wherein the ORF encodes a therapeutic polypeptide, protein or antigen, as described herein. The pharmaceutical composition may thus be administered to a subject. The subject’s cells transcribe and / or translate the RNA molecule to produce a therapeutic molecule in vivo.

[0135] The pharmaceutical composition may comprise further components to aid uptake of the RNA molecule, polynucleotide or vector by a subject’s cells, e.g. by transfection. The component typically sequesters or encapsulates the RNA molecule, polynucleotide or vector. For example, the pharmaceutical composition may comprise a micelle, such as a lipid or polymer micelle. The pharmaceutical composition may comprise a liposome or a lipoplex. The pharmaceutical composition may comprise a viral vector, such as an adeno-associated virus (AAV) vector. The pharmaceutical composition may comprise a polymeric nanoparticle. The pharmaceutical composition may comprise a lipid nanoparticle. The RNA molecule, polynucleotide or vector is typically encapsulated by the lipid nanoparticle. Suitable lipid nanoparticles are described, for example, in Hou et al., 2021. Currently approved lipid nanoparticles typically comprise a cationic or ionisable lipid, cholesterol, a helper lipid and a PEG-lipid. The pharmaceutical composition may comprise a polymer or polymer-based nanoparticle, such as poly(beta- amino ester)s. The pharmaceutical composition may comprise an exosome, suitable for delivering mRNA into a cell. The exosome is typically an exosome derived from a cell of the same genus or species as the subject. For example, the exosome may be a human embryonic kidney (HEK) cell-derived exosome, a bone marrow stem cells (BMSC)- derived exosome, a milk-derived exosome or a red-blood cell derived exosome. Further transfection agents are described in WO 2016 / 091391 Al.

[0136] The RNA molecule, polynucleotide or vector is typically included in the pharmaceutical composition in an effective amount. The term “ effective amount' is intended to refer to a quantity sufficient to achieve a measurable therapeutic response in a subject to which the pharmaceutical composition has been administered.

[0137] Therapeutic and non-therapeutic uses

[0138] The invention also relates to the therapeutic use of the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention.

[0139] The invention provides a method of treating or preventing a disease, disorder or condition in a subject, the method comprising administering the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention to the subject. The invention also provides the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention for use in a method of treating or preventing disease in a subject. The invention also relates to the use of the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention in a method of treating or preventing a disease in a subject, or use of the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention for treating or preventing a disease in a subject. The invention also provides a use of the polynucleotide, vector, cell, RNA molecule or pharmaceutical composition of the invention in the manufacture of a medicament for use in a method of treating or preventing a disease in a subject.

[0140] The disease, disorder or condition to be treated depends on the therapeutic agent encoded by the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention. The disease, disorder or condition is typically one which can be treated by the therapeutic agent encoded by the polynucleotide, vector, RNA molecule or pharmaceutical composition, for example by the antigen, peptide, polypeptide or protein encoded by the ORF, or by the encoded non-coding functional RNA. The disclosure relates to an mRNA platform which can be used and adapted to express molecules related to existing or future therapies. Accordingly, the therapeutic use is not particularly limited. For example, the disease, disorder or condition may be a viral infection, a bacterial infection, a fungal infection, a parasitic infection, cancer, an autoimmune disease, allergy, a genetic disorder or the like. The therapeutic use may be as a vaccine. mRNA vaccines are well known in the art, particularly in view of the recent SARS-CoV- 2 pandemic. The therapeutic use may be as part of enzyme replacement therapy (ERT). The therapeutic use may be as part of a gene therapy. The therapeutic use may be for expression of a peptide, polypeptide or protein therapeutic molecule, for example an antibody. The quantities of peptide, polypeptide or protein for such therapeutic uses are typically much greater than that required for mRNA vaccines, and the RNA molecules of the present invention are particularly adapted to allow for robust expression of such molecules.

[0141] The invention also relates to the non-therapeutic use of the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention. For example, the invention provides a cosmetic use of the polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention.

[0142] In some cases, the invention also relates to a non-therapeutic and / or cosmetic use of the polynucleotide, vector or RNA molecule of the invention. For example, the invention provides a non-therapeutic and / or cosmetic method of administering the polynucleotide, vector or RNA molecule of the invention to a subject. The invention also provides a non-therapeutic and / or cosmetic use of the polynucleotide, vector or RNA molecule of the invention in a subject. The invention also provides the polynucleotide, vector or RNA molecule of the invention for a non-therapeutic and / or cosmetic use in a subject. The polynucleotide, vector or RNA molecule typically does not comprise or encode a therapeutic molecule, such as an antigen, or a therapeutic peptide, polypeptide or protein. The polynucleotide, vector or RNA molecule typically does not have a therapeutic use. For example, the polynucleotide, vector or RNA molecule does not have a use for prophylaxis or treatment of a disease. In some cases, the polynucleotide, vector or RNA molecule may have a therapeutic use, but the non-therapeutic and / or cosmetic use is restricted to use in subjects that would not have a therapeutic benefit. For example, the cosmetic use may be weight loss, and the use would be in subjects that would not benefit therapeutically from the use, for example in individuals that are not overweight.

[0143] The subject is typically a mammalian subject. Typically, the subject is a human. The subject may be a non-human mammal, such as a mouse, rat, cat, dog, pig, goat, sheep, horse, cow, camel or non-human primate.

[0144] Also described herein is a method of transforming a host cell with a polynucleotide, vector, RNA molecule or pharmaceutical composition of the invention. The polynucleotide, vector, RNA molecule or pharmaceutical composition may encode or comprise a sequence encoding a therapeutic RNA sequence, a therapeutic polynucleotide, peptide, polypeptide, protein or antigen as described herein. The host cell may be a cell for use in cellular therapy. For example, the host cell may be a stem cell, such a haematopoietic stem cell, a skeletal muscle stem cell, a mesenchymal stem cell, or a tissue-specific progenitor cell, such as a cardiac progenitor cell or a liver progenitor cell. The host cell may be a cell for use in adoptive cell therapy. For example, the host cell may be a lymphocyte, such as a T cell or a natural killer cell. The host cell may be a professional antigen presenting cell, for example a dendritic cell, a monocyte or a macrophage. The host cell may be a tissue specific cell, such as a pancreatic islet cell.

[0145] In vitro transcription

[0146] The invention also relates to an in vitro method of producing an RNA molecule from the polynucleotide, vector or cell described herein. The method may be a method of performing in vitro transcription (IVT). The method comprises contacting the polynucleotide, vector or cell of the invention with an RNA polymerase. The polynucleotide or vector is typically a DNA polynucleotide or DNA vector, and the RNA polymerase is typically a DNA-dependent RNA polymerase. The DNA-dependent RNA polymerase may be T7 RNA polymerase, SP6 RNA polymerase or T3 RNA polymerase. The polynucleotide or vector may be an RNA polynucleotide or RNA vector, and the RNA polymerase may be an RNA-dependent RNA polymerase. Where the polynucleotide or vector is an RNA polynucleotide or RNA vector, the RNA polynucleotide or RNA vector may be a self-amplifying RNA polynucleotide or RNA vector, e.g. encoding an RNA-dependent RNA polymerase.

[0147] The polynucleotide or vector may be a circular molecule. In this case, the method typically comprises linearising the molecule prior to contact with the RNA polymerase. Linearisation of the polynucleotide or vector may be performed by any means known to the skilled person. Typically, the method comprises linearising the circular molecule using a restriction enzyme. Linearisation may occur at the terminal end of the poly(A) tail encoded by the nucleic acid sequence. For example, the circular molecule may comprise a restriction site at the terminal end of the poly(A) tail encoded by the nucleic acid sequence, and the method may comprise using a restriction enzyme to linearise the circular molecule at the restriction site. The restriction site at the terminal end of the poly(A) tail encoded by the nucleic acid molecule may be a SapI restriction site, and the restriction enzyme may be SapI or an isoschizomer thereof, such as BspQI, Lgul, Nt.BspQI or PciSI.

[0148] The method may further comprise isolating the RNA molecule following polymerisation. As discussed above, the term “isolating is intended to be used interchangeably with the term “purifyin and refers to the separation of the RNA molecule from molecules or components used in the method, such as the polynucleotide or vector. For example, a DNase enzyme, such as DNasel may be added to the reaction mixture following polymerisation to degrade the template polynucleotide or vector. The RNA molecule may then be purified using standard methods known to the skilled person, for example using ethanol precipitation-, spin column-, phenol-chloroform extraction- or bead-based extraction methods.

[0149] The method may further comprise formulating the RNA molecule into a pharmaceutical composition. The pharmaceutical composition may further comprise a pharmaceutically acceptable excipient and / or a micelle, liposome, exosome or lipid nanoparticle, as described herein. The invention also provides a pharmaceutical composition produced by this method.

[0150] The method is performed under conditions suitable for the production of an RNA molecule, for example in the presence of ribonucleotide triphosphates. The method may be performed in the presence of a 7-m ethyl guanosine cap nucleotide or cap analogue under conditions suitable to achieve 5’ capping of the RNA molecule. The cap may be added co-transcriptionally or post-transcriptionally.

[0151] The method may further comprise adding a poly(A) tail to the RNA molecule, i.e. post- transcriptionally. This is useful for embodiments where the polynucleotide does not encode a poly(A) tail. The method may comprise contacting the RNA molecule with a poly(A) polymerase enzyme and ATP. The polynucleotide may encode a polyadenylation signal sequence in the RNA molecule, e.g. AAUAAA near the 3’ end of the encoded RNA molecule.

[0152] Other methods

[0153] In some cases, an in cellulo or in vivo method of producing an RNA molecule from the polynucleotide or vector or propagated polynucleotide is provided. The method may comprise transfecting the polynucleotide, vector or propagated polynucleotide into a cell, and transcribing the polynucleotide, vector or propagated polynucleotide to produce the RNA molecule. The cell typically provides the conditions to allow for the production of the RNA molecule by the cell, such as an RNA polymerase. The RNA polymerase may be endogenous to the cell or exogenous to the cell. The cell is typically a eukaryotic cell. This may allow RNA modifications, such as 5’ capping, to be performed within the cell. The eukaryotic cell may be a yeast cell. The eukaryotic cell may be an animal cell, such as a mammalian cell. The cell may be a human cell. The cell may be a non-human mammalian cell, such as a mouse, rat, cat, dog, pig, goat, sheep, horse, cow, camel or non-human primate cell. The cell may be isolated from a subject. The cell may be an immortal cell line. The cell may be a human embryonic kidney cell, such as a HEK293 cell. The HEK293 cell may be a HEK293T cell. The cell may be an immune cell, such as professional antigen presenting cell (for example a dendritic cell, a monocyte or a macrophage). The eukaryotic cell may be an insect cell. The insect cell may be a Drosophila S2 cell, a Spodoptera frugiperda cell (such as Sf9 and Sf21 ) or a Trichoplusia ni “Hi-Five” cell. The method typically further comprises purifying the RNA molecule from the cell. The purification may be performed by any means known to the skilled person. The purification may include a separating the RNA molecule produced by transcription of the polynucleotide or vector from other RNA molecules produced by the cell. The in cellulo or in vivo method may involve the use of a selfamplifying RNA molecule, polynucleotide or vector, as described herein.

[0154] The invention also provides a method of producing an RNA molecule, peptide, polypeptide, protein or antigen. The invention may be an in vitro method or an in vivo method. The method comprises transcribing or translating the RNA molecule of the invention or obtained according to the method of the invention. The RNA molecule, peptide, polypeptide, protein or antigen is typically a therapeutic RNA molecule, peptide, polypeptide, protein or antigen, as described herein. The method is typically performed in vivo following administration of the RNA molecule to a subject and incorporation of the RNA molecule into the subject’s cells.

[0155] The invention also provides a method of increasing the translational capacity of a polynucleotide comprising a nucleic acid sequence encoding a 5’ UTR, a 3’ UTR and / or an ORF comprising a sequence encoding a signal peptide. The method comprises: replacing the sequence comprising or encoding the 5’ UTR with a sequence comprising or encoding the 5’ UTR comprising a sequence as set out in SEQ ID NO: 1 or 2 or a variant of as defined herein; replacing the sequence comprising or encoding the 3’ UTR with a sequence comprising or encoding the 3’ UTR comprising a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant of as defined herein or a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant of as defined herein; and / or replacing the sequence encoding a signal peptide with a sequence encoding a signal peptide as set out in SEQ ID NO: 6 or 7 or a variant of as defined herein.

[0156] The method may comprise directly replacing the sequence comprising or encoding the poly(A) tail within a polynucleotide or vector. The method may comprise cloning an ORF from a vector comprising or encoding a poly(A) tail and inserting it into a polynucleotide or vector of the invention comprising a segmented poly(A) tail.

[0157] Informal Sequence Listing Examples

[0158] Example 1

[0159] Therapeutic mRNAs have been instrumental in combatting the recent coronavirus pandemic, COVID-19. Their success in combatting infectious diseases has reignited the interest in using mRNA as a vehicle to deliver proteins for replacement, regeneration, genome editing and in vivo antibody production. The transition of mRNA vectors from combatting infections to delivering a variety of therapeutic proteins is hampered by the high protein expression levels required to achieve their pharmacological activity thresholds. To address this issue and maximise the expression of mRNA-encoded protein payloads, we used non-naturally combined 5’UTRs and 3’UTRs to generate mRNAs that are highly stable and efficiently translated. To that end, the sequence from the human 5’UTR of the Histone H2AX was cloned into a plasmid featuring a CleanCap® (TriLink)-compatible T7 promoter followed by the Renilla luciferase open reading frame (ORF) and the 3’UTR sequence from the human Collagen Type VI Alpha 2 gene (COL6A2) gene (SEQ ID NOs: 1 and 3, respectively). The plasmid further featured a 30- nucleotide long template-encoded polyadenosine (poly[A] tail). As illustrated in figure 1, each of these features is flanked by unique type II restriction enzyme recognition sites that allow the straightforward exchange of the relevant sequence if desired (Figure 1).

[0160] To benchmark the expression of protein payloads encoded by in w / ra-transcribed mRNAs featuring the non-natural combination of the H2AX 5’UTR and the COL6A2 3’UTR, their sequences were precisely replaced by the 5’UTRs and the 3’UTRs of Modema’s FDA-approved COVID-19 vaccine, mRNA-1273 (SEQ ID NOs: 11 and 12, respectively).

[0161] Both plasmids were subjected to a BspQI (isoschizomer of SapI) restriction enzyme digest to linearise the plasmids and enable mRNAs to be transcribed by in vitro run-off transcription (IVT) using T7 RNA polymerase. The resulting Renilla luciferase-encoding mRNAs are capped and feature a 30-nucleotide long poly(A) tail. Plasmids featuring a 30-nucleotide long polyadenosine tract were chosen because plasmids with short template-encoded polyadenosine sequences are stable in bacterial hosts. In contrast, long poly(A) tracts (e.g., exceeding 100 adenosines) are highly prone to shortening during subcloning and amplification in bacteria (Grier et al., 2016; Trepotec et al., 2019). This approach ensured that the IVT mRNAs from both templates have uniform 30-nucleotide long poly(A) tails, allowing direct comparison of the translational output from mRNAs featuring the different UTR sequences.

[0162] The integrity of the transcribed mRNAs containing the H2AX / COL6A2 UTR combination or the Moderna mRNA-1273 UTR combination were verified by agarose gel electrophoresis (Figure 2, compare lanes 2 and 5, respectively). Equal amounts of both mRNAs were subsequently transfected into HEK-293T cells employing the Lipofectamine MessengerMax transfection reagent and protocol. 48h post-transfection, the cells were harvested, whole cell lysates were prepared and luciferase activity from the two different mRNAs were compared. Cells were cultured in 6-well plates in DMEM + 10 % foetal bovine serum. Cells were harvested by removing the cell culture media from the wells, and lysed by addition of 1 mL of lysis buffer (from Pierce Renilla Luciferase Glow Assay Kit: Catalog No 16167) to each well. The 6-well plates were then placed on a shaker platform at 500rpm for 40mins. The lysates were centrifuged at 2000rpm for 5mins using a bench centrifuge and 800 pL supernatant was removed into a fresh tube and the cell debris was discarded. For the luciferase assay, 150 pl of cell lysate from this 800 pl volume was added to a well in a white opaque 96-well plate (PerkinElmer). For each biological repeat, the luciferase activity was measured twice in the assay in two separate wells and the average of the two technical repeats was taken as the final value. After loading all samples, 50 pl of working solution (from Pierce Renilla Luciferase Glow Assay Kit: Catalog No 16167) was added to each well and the plate was covered with tinfoil for at least 15mins before measuring luciferase activity using a Clariostar instrument (emission setting F:530-40). 1 second integration time was used for the measurements and the autoadjust was carried out on the well with the highest raw value measured in the initial read.

[0163] At all tested amounts of transfected mRNA (lOOng, Ipg, 2pg), the mRNAs featuring the H2AX 5 ’UTR and the COL6A2 3 ’UTR outperformed the mRNAs that featured the mRNA-1273 UTRs (Figure 3). These results confirm that our mRNA platform featuring the H2AX and COL6A2 UTR combination affords the production of mRNAs with high translational capacity.

[0164] We next used the platform to transcribe mRNAs that encode a monoclonal antibody (mAb) to explore whether our UTR sequences are suitable to express secreted proteins. To that end, we cloned the ORFs of the R5.016 mAb (Alanine et al., 2019) light (L) and heavy (H) chains into a parental plasmid featuring a 114-nucleotide segmented poly(A) tail, replacing the Renilla luciferase ORF. mRNAs encoding the light and heavy chains were transcribed using non-modified uracil nucleotides and 1 pg each of H chain- and L chain-encoding mRNAs were subsequently transfected into HEK-293T cells. 48h posttransfection, the media of the cells were harvested, and the concentration of mAbs that bind the target antigen (Plasmodium falciparum Reticulocyte Binding homologue 5 [RH5]) was measured by standardised indirect enzyme-linked immunosorbent assay (ELISA). HEK-293T cells transfected with I pg of H chain- and L chain-encoding mRNAs featuring the H2AX and COL6A2 UTRs produced more than 16 pg / mL of mAbs capable of binding the RH5 antigen (Figure 4). These results demonstrate that our platform is suitable for the transcription of mRNAs that encode proteins that function intracellularly as well as proteins that are secreted from the host cell. Critically, our platform enables the secretion of mAbs at levels exceeding those achieved in similar settings with other mRNA designs (Deng et al., 2022).

[0165] Example 2

[0166] The success of mRNA therapeutics is highly dependent on reaching the pharmacological threshold of the encoded protein payload. Our UTR combinations support high levels of expression of both intracellular and secreted proteins. However, to make the platform more versatile for the production of secreted proteins, we adapted our original mRNA expression plasmid and inserted an additional signal peptide cassette downstream of the 5 ’UTR. This “SP cassette” is flanked by restriction enzyme sites (Figure 5) allowing the straightforward insertion, exchange and testing of different signal peptides (SP) that guide the mRNAs to the endoplasmic reticulum (Owji et al., 2018) and direct secretion of the downstream protein payload ORF.

[0167] We cloned two plasmids, one that featured the mouse IgG SP-encoding sequence (SEQ ID NO: 10) and a second plasmid that contained the human Cystatin C SP-encoding sequence (SEQ ID NO: 8) upstream of the Renilla ORF. The plasmids featuring the two different SPs (Figure 6) were linearised with BspQI and mRNAs were transcribed by T7 RNA polymerase. From each linearised plasmid, three different mRNAs were produced that incorporated either unmodified (uridine) or the modified nucleosides (pseudouridine [ ] or N1 -methylpseudouridine [ml'P]). We transcribed mRNAs with these different modifications to test whether uridine stretches in the SP coding sequences, as featured in the IgG SP, cause ribosomal frameshifting (Mulroney et al., 2023) and so compromise the overall protein output. All transcribed mRNAs were quality controlled by agarose gel electrophoresis (Figure 7) and Ipg of each mRNA was transfected into HEK-293T cells cultured in FluoroBrite™ DMEM (+ 10% foetal bovine serum) in 6-well plates. Media was collected 48h post-transfection and luciferase activity assays were performed to compare the level of functional luciferase secreted by the IgG or Cystatin C SPs. For the luciferase assay, 300 pl of cell media was added to a well in a white opaque 96-well plate (PerkinElmer). For each biological repeat, the luciferase activity was measured twice in the assay in two separate wells and the average of the two technical repeats was taken as the final value. After loading all samples, 50 pL of working solution (from Pierce Renilla Luciferase Glow Assay Kit: Catalog No 16167) was added to each well and the plate was covered with tinfoil for at least 15mins before measuring luciferase activity using a Clariostar instrument (emission setting F:530-40). 1 second integration time was used for the measurements and the autoadjust was carried out on the well with the highest raw value measured in the initial read.

[0168] Irrespective of the type of uracil nucleoside present, mRNAs containing the Cystatin C SP consistently outperformed equivalent transcripts containing the IgG SP (Figure 8). Thus, the combination of the H2AX 5’UTR and COL6A2 3’UTRs together with the Cystatin C SP represents a versatile therapeutic mRNA production platform that is highly suitable for the expression of secreted proteins in host cells. Discussion

[0169] The success of therapeutic mRNAs in replacement protein therapies and in vivo antibody production is heavily dependent on maximising the expression of the encoded payload proteins from in iv' / m-tran scribed mRNAs. Here we present a novel therapeutic mRNA platform that features non-natural combinations of untranslated regions and signal peptides that maximise the expression of payload proteins designed to function inside cells or to be secreted by targeted host cells.

[0170] We have demonstrated that the H2AX 5’UTR and the COL6A2 3’UTR are a powerful combination that equip in vitro transcribed mRNAs with features that maximise the expression of their encoded protein payload. Our platform is highly versatile and as effective in directing protein expression as the combination of UTRs featured in the design of one of the market leaders. We have shown that the Cystatin C signal peptide in combination with our UTRs creates a powerful tool for the production of mRNAs that encode proteins destined for secretion from transfected host cells. This Cystatin C SP design will be particularly useful for mRNA-directed in patient production of therapeutic secreted proteins.

[0171] Example 3

[0172] To optimise the 3’UTR, we investigated the use of a short version of the COL6A2 3’ UTR (SEQ ID NO: 13). This version of the COL6A2 3’UTR was prepared by restriction enzyme digest (PstI and Kpnl) of the COL6A2 gene sequence followed by blunt end ligation, leaving the first 213 nucleotides of the COL6A2 3’UTR as set out in SEQ ID NO: 3.

[0173] A construct comprising the Moderna mRNA-1273 5’UTR and 3’UTR and a 114 nucleotide poly(A) tail was screened against a construct comprising the H2AX 5’UTR, the short COL6A2 3’UTR (SEQ ID NO: 13) and a 114 nt poly(A) tail, a construct comprising the H2AX 5’UTR, the COL6A2 3’UTR (SEQ ID NO: 3) and a 114 nt poly(A) tail, and a construct comprising the H2AX 5’UTR, the COL6A2 3’UTR (SEQ ID NO: 3) and a 198 nt poly(A) tail. Each mRNA construct was tested incorporating either unmodified uridines or the modified nucleosides N1 -methylpseudouridine [m l ] but no signal peptide sequence (Figure 9). lOOng mRNA was transfected per well in a 96-well plate seeded with HEK-293T cells.

[0174] The data shows that the short COL6A2 3’UTR of SEQ ID NO: 13 is comparable to or better than the COL6A2 3’UTR set out in SEQ ID NO: 3. The data indicates that incorporation of m l may improve translational capacity of mRNAs featuring H2AX 5’UTR and COL6A2 3’UTR sequences, for example in the constructs comprising a 114nt and 198nt poly(A) tail tested. This may be important for reaching the higher protein thresholds required when using mRNA as a vector to express therapeutic / replacement proteins.

[0175] References

[0176] Alanine, D. G. W., Quinkert, D., Kumarasingha, R., Mehmood, S., Donnellan, F. R., Minkah, N. K., Dadonaite, B., Diouf, A., Galaway, F., Silk, S. E., Jamwal, A., Marshall, J. M., Miura, K., Foquet, L., Elias, S. C., Labbe, G. M., Douglas, A. D., Jin, J., Payne, R. O., ... Draper, S. J. (2019). Human Antibodies that Slow Erythrocyte Invasion Potentiate Malaria-Neutralizing Antibodies. Cell, 178(f), 216-228. e21. https: / / doi.Org / 10.1016 / j.cell.2019.05.025

[0177] Andries, O., Me Cafferty, S., De Smedt, S. C., Weiss, R., Sanders, N. N., & Kitada, T. (2015). Nl-methylpseudouridine-incorporated mRNA outperforms pseudouridine- incorporated mRNA by providing enhanced protein expression and reduced immunogenicity in mammalian cell lines and mice. Journal of Controlled Release, 217, 337-344. https: / / doi.Org / 10.1016 / J.JCONREL.2015.08.051

[0178] Deng, Y.-Q., Zhang, N.-N., Zhang, Y.-F., Zhong, X., Xu, S., Qiu, H.-Y., Wang, T.-C., Zhao, H., Zhou, C., Zu, S.-L., Chen, Q., Cao, T.-S., Ye, Q., Chi, H., Duan, X.-H., Lin, D.-D., Zhang, X.-J., Xie, L.-Z., Gao, Y.-W ., ... Qin, C.-F. (2022). Lipid nanoparticle- encapsulated mRNA antibody provides long-term protection against SARS-CoV-2 in mice and hamsters. Cell Research 2022 32:4, 32(4), 375-382. https : / / doi . org / 10.1038 / s41422-022-00630-0

[0179] Fang, E., Liu, X., Li, M., Zhang, Z., Song, L., Zhu, B., Wu, X., Liu, J., Zhao, D., & Li, Y. (2022). Advances in COVID-19 mRNA vaccine development. Signal Transduction and Targeted Therapy, https: / / doi.org / 10.1038 / s41392-022-00950-y

[0180] Glover, D. J., Lipps, H. I, & Jans, D. A. (2005). TOWARDS SAFE, NON- VIRAL THERAPEUTIC GENE EXPRESSION IN HUMANS. Nat Rev Genet , 6(4), 299-310. https: / / doi.org / 10.1038 / nrgl577

[0181] Grier, A. E., Burleigh, S., Sahni, J., Clough, C. A., Cardot, V., Choe, D. C., Krutein, M. C., Rawlings, D. J., Jensen, M. C., Scharenberg, A. M., & Jacoby, K. (2016). pEVL: A Linear Plasmid for Generating mRNA IVT Templates With Extended Encoded Poly(A) Sequences. Molecular Therapy - Nucleic Acids, 5(March), e306. https: / / doi.org / 10.1038 / mtna.2016.21

[0182] Guan, X., Pei, Y., & Song, J. (2024). DNA-Based Nonviral Gene Therapy— Challenging but Promising. Molecular Pharmaceutics, 21(2), 427-453. https : / / doi . org / 10.1021 / AC S MOLPHARM ACEUT.3 C00907 / AS SET / IM AGES / LARGE / MP3C00907_0009.JPEG

[0183] Hou, X., Zaks, T., Langer, R., & Dong, Y. (2021). Lipid nanoparticles for mRNA delivery. Nature Reviews Materials 2021 6:12, 6(12), 1078-1094. https: / / doi.org / 10.1038 / s41578-021-00358-0

[0184] Kariko, K., Buckstein, M., Ni, H., & Weissman, D. (2005). Suppression of RNA Recognition by Toll-like Receptors: The Impact of Nucleoside Modification and the Evolutionary Origin of RNA. Immunity, 23(2), 165-175. https: / / doi.Org / 10.1016 / J.IMMUNI.2005.06.008 Kariko, K., Muramatsu, H., Welsh, F. A., Ludwig, J., Kato, H., Akira, S., & Weissman,

[0185] D. (2008). Incorporation of pseudouridine into mRNA yields superior nonimmunogenic vector with increased translational capacity and biological stability. Molecular Therapy, 76(11), 1833-1840. https: / / doi.org / 10.1038 / mt.2008.200

[0186] Lundstrom, K. (2023). Citation: Lundstrom, K. Viral Vectors in Gene Therapy: Where Do We Viral Vectors in Gene Therapy: Where Do We Stand in 2023? https: / / doi.org / 10.3390 / vl5030698

[0187] Muhammad Hammad Butt, Muhammad Zaman, Abrar Ahmad, Rahima Khan, Tauqeer Hussain Mallhi, Mohammad Mehedi Hasan, Yusra Habib Khan, Sara Hafeez, Ehab El Sayed Massoud, Md. Habibur Rahman, & Simona Cavalu. (2022). Appraisal for the Potential of Viral and Nonviral Vectors in Gene Therapy: A Review. Genes 2, 73(1370). https: / / doi.org / 10.3390 / genesl3081370

[0188] Mulroney, T. E., Pbyry, T., Yam-Puc, J. C., Rust, M., Harvey, R. F., Kalmar, L., Homer,

[0189] E., Booth, L., Ferreira, A. P., Stoneley, M., Sawarkar, R., Mentzer, A. J., Lilley, K. S., Smales, C. M., von der Haar, T., Turtle, L., Dunachie, S., Klenerman, P., Thaventhiran, J. E. D., & Willis, A. E. (2023). N1 -methylpseudouridylation of mRNA causes +1 ribosomal frameshifting. Nature 2023, 1-6. https: / / doi.org / 10.1038 / s41586-023-06800-3

[0190] Owji, H., Nezafat, N., Negahdaripour, M., Hajiebrahimi, A., & Ghasemi, Y. (2018). A comprehensive review of signal peptides: Structure, roles, and applications. European Journal of Cell Biology, 97(6), 422-441. https: / / doi.Org / 10.1016 / J.EJCB.2018.06.003

[0191] Porello, I., & Cellesi, F. (2023). Intracellular delivery of therapeutic proteins. New advancements and future directions. In Frontiers in Bioengineering and Biotechnology (Vol. 11). Frontiers Media S.A. https: / / doi.org / 10.3389 / fbioe.2023.1211798

[0192] Trepotec, Z., Geiger, J., Plank, C., Aneja, M. K., & Rudolph, C. (2019). Segmented poly(A) tails significantly reduce recombination of plasmid DNA without affecting mRNA translation efficiency or half-life. Rna, 25(4), 507-518. https: / / doi.Org / 10.1261 / rna.069286. l 18

[0193] Further embodiments

[0194] 1. An RNA molecule comprising a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR), and an open reading frame (ORF) and / or a non-coding functional RNA sequence, wherein:

[0195] (a) the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof;

[0196] (b) the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or

[0197] (c) the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

[0198] 2. The RNA molecule of embodiment 1, wherein the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2.

[0199] 3. The RNA molecule of embodiment 1 or 2, wherein the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to three nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2.

[0200] 4. The RNA molecule of any one of the preceding embodiments, wherein the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof having at least 80% sequence identity to any one of SEQ ID NOs: 3 to 5.

[0201] 5. The RNA molecule of any one of the preceding embodiments, wherein the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof having at least 80% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13. 6. The RNA molecule of any one of the preceding embodiments, wherein the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 or a variant thereof having at least 90% sequence identity to any one of SEQ ID NOs: 3 to 5.

[0202] 7. The RNA molecule of any one of the preceding embodiments, wherein the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof having at least 90% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13.

[0203] 8. The RNA molecule of any one of the preceding embodiments, which has at least 70% of the translational capacity of a reference RNA molecule, wherein the RNA molecule comprises an H2AX 5’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 5’ UTR of SEQ ID NO: 1 or 2, and / or wherein the RNA molecule comprises a COL6A2 3’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 3’ UTR of any one of SEQ ID NOs: 3 to 5, wherein translational capacity is calculated as the measured level of luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule into HEK293T host cells, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase.

[0204] 9. The RNA molecule of any one of the preceding embodiments, which has at least 70% of the translational capacity of a reference RNA molecule, wherein the RNA molecule comprises an H2AX 5’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 5’ UTR of SEQ ID NO: 1 or 2, and / or wherein the RNA molecule comprises a COL6A2 3’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 3’ UTR of any one of SEQ ID NOs: 3 to 5 andl3, wherein translational capacity is calculated as the measured level of luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule into HEK293T host cells, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase. 10. The RNA molecule of any one of the preceding embodiments, wherein the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to five amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0205] 11. The RNA molecule of any one of the preceding embodiments, wherein the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to three amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0206] 12. The RNA molecule of any one of the preceding embodiments, which has a greater translational capacity than a reference RNA molecule, wherein the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof, and the reference RNA molecule is an equivalent RNA molecule comprising an ORF comprising a sequence encoding a signal peptide as set out in SEQ ID NO: 9, wherein translational capacity is calculated as the measured level of secreted luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule into HEK293T host cells, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase.

[0207] 13. The RNA molecule of any one of the preceding embodiments, wherein the ORF comprises a sequence encoding a secreted peptide, a polypeptide or a protein.

[0208] 14. The RNA molecule of any one of the preceding embodiments, wherein the ORF comprises a sequence encoding an antibody.

[0209] 15. The RNA molecule of any one of the preceding embodiments, wherein the ORF does not comprise a sequence encoding human H2A histone family member X or a sequence encoding collagen type VI alpha chain 2. 16. The RNA molecule of any one of the preceding embodiments, further comprising a poly(A) tail.

[0210] 17. The RNA molecule of any one of the preceding embodiments, wherein:

[0211] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; and

[0212] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof having a least 80% sequence identity to any one of SEQ ID NOs: 3 to 5.

[0213] 18. The RNA molecule of any one of the preceding embodiments, wherein:

[0214] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; and

[0215] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 an 13, or a variant thereof having a least 80% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13.

[0216] 19. The RNA molecule of any one of the preceding embodiments, wherein:

[0217] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to four nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; and

[0218] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof having a least 90% sequence identity to any one of SEQ ID NOs: 3 to 5.

[0219] 20. The RNA molecule of any one of the preceding embodiments, wherein:

[0220] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to four nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; and (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof having a least 90% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13.

[0221] 21. The RNA molecule of any one of the preceding embodiments, wherein:

[0222] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2;

[0223] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof having a least 80% sequence identity to any one of SEQ ID NOs: 3 to 5; and

[0224] (c) the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to five amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0225] 22. The RNA molecule of any one of the preceding embodiments, wherein:

[0226] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2;

[0227] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof having a least 80% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13; and

[0228] (c) the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to five amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0229] 23. The RNA molecule of any one of the preceding embodiments, wherein:

[0230] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to four nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5, or a variant thereof having a least 90% sequence identity to any one of SEQ ID NOs: 3 to 5; and / or

[0231] (c) the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to four amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0232] 24. The RNA molecule of any one of the preceding embodiments, wherein:

[0233] (a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to four nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2;

[0234] (b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof having a least 90% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13; and / or

[0235] (c) the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to four amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

[0236] 25. The RNA molecule of any one of the preceding embodiments, which comprises in order:

[0237] (i) the 5’ UTR;

[0238] (ii) the ORF and / or the non-coding functional RNA sequence;

[0239] (iii) the 3 ’ UTR; and

[0240] (iv) optionally, a poly(A) tail.

[0241] 26. The RNA molecule of any one of the preceding embodiments, further comprising an internal ribosome entry site (IRES).

[0242] 27. The RNA molecule of any one of the preceding embodiments, which is an mRNA molecule. 28. The RNA molecule of any one of the preceding embodiments, which comprises a 5’ cap, preferably wherein the 5’ cap comprises a 7-m ethyl guanosine cap.

[0243] 29. A polynucleotide comprising a nucleic acid sequence encoding the RNA molecule of any one of the preceding embodiments.

[0244] 30. A polynucleotide comprising a nucleic acid sequence encoding a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR) and (i) an open reading frame (ORF) and / or a non-coding functional RNA sequence, or (ii) a sequence for introducing an ORF and / or a non-coding functional RNA sequence, wherein:

[0245] (a) the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof;

[0246] (b) the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or

[0247] (c) the ORF, the sequence encoding an ORF, or the sequence encoding a peptide, polypeptide or protein comprises a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

[0248] 31. The polynucleotide of embodiment 30, wherein the sequence encoded by the nucleic acid sequence is as defined in any one of embodiments 1 to 28.

[0249] 32. The polynucleotide of any one of embodiments 29 to 31, wherein the polynucleotide further comprises a promoter.

[0250] 33. The polynucleotide of any one of embodiments 29 to 32, wherein the polynucleotide further comprises:

[0251] (i) a marker gene;

[0252] (ii) an origin of replication;

[0253] (iii) a multiple cloning site; and / or

[0254] (iv) a restriction site, preferably at the terminal end of the sequence encoded by the nucleic acid sequence, optionally wherein the restriction site is a SapI restriction site. 34. The polynucleotide of any one of embodiments 29 to 33, wherein the polynucleotide further comprises an inverted terminal repeat.

[0255] 35. The polynucleotide of any one of embodiments 29 to 34, which is a DNA polynucleotide.

[0256] 36. A vector comprising the polynucleotide of any one embodiments 29 to 35.

[0257] 37. The vector of embodiment 36, which is a linear vector or a circular vector.

[0258] 38. The vector of embodiment 36, which is a viral vector, such as an adeno- associated virus (AAV) vector.

[0259] 39. A cell comprising the polynucleotide of any one of embodiments 29 to 35, or the vector of any one of embodiments 36 to 38.

[0260] 40. The cell of embodiment 39, which is a prokaryotic cell, preferably a bacterial cell, more preferably an Escherichia coli cell.

[0261] 41. The cell of embodiment 39, which is a eukaryotic cell, preferably an insect cell or a mammalian cell, such as a human embryonic kidney cell.

[0262] 42. A pharmaceutical composition comprising the RNA molecule of any one of embodiments 1 to 28, the polynucleotide of any one of embodiments 29 to 35, or the vector of any one of embodiments 36 to 38.

[0263] 43. The pharmaceutical composition of embodiment 42, further comprising a pharmaceutically acceptable excipient.

[0264] 44. The pharmaceutical composition of embodiment 42 or 43, further comprising a micelle, liposome, exosome or lipid nanoparticle. 45. A method of treating or preventing a disease, disorder or condition in a subject, the method comprising administering the pharmaceutical composition of any one of embodiments 42 to 44 to the subject.

[0265] 46. The pharmaceutical composition of any one of embodiments 42 to 44 for use in a method of treating or preventing a disease in a subject.

[0266] 47. Use of the pharmaceutical composition of any one of embodiments 42 to 44 in a method of treating or preventing a disease in a subject.

[0267] 48. Use of the RNA molecule of any one of embodiments 1 to 28, the polynucleotide of any one of embodiments 29 to 33, the vector of any one of embodiments 36 to 38, the cell of any one of embodiments 39 to 41, or the pharmaceutical composition of any one of embodiments 42 to 44 in the manufacture of a medicament for use in a method of treating or preventing a disease in a subject.

[0268] 49. An in vitro method of producing an RNA molecule, comprising contacting the polynucleotide of any one of embodiments 29 to 35, the vector of any one of embodiments 36 to 38, or the cell of any one of embodiments 39 to 41, with an RNA polymerase.

[0269] 50. The method of embodiment 49, wherein the polynucleotide or vector is a circular molecule, and the method further comprises linearising the circular molecule, optionally wherein linearising the circular molecule is carried out using a restriction endonuclease at the terminal end of the sequence encoded by the nucleic acid sequence.

[0270] 51. The method of embodiment 49 or 50, wherein the RNA polymerase is a DNA- dependent RNA polymerase.

[0271] 52. The method of any one of embodiments 49 to 51, wherein the RNA polymerase is T7 RNA polymerase, SP6 RNA polymerase or T3 RNA polymerase. 53. The method of any one of embodiments 49 to 52, which further comprises isolating the RNA molecule following the step of contacting the polynucleotide, vector or cell with the RNA polymerase.

[0272] 54. The method of any one of embodiments 49 to 53, which further comprises formulating the RNA molecule into a pharmaceutical composition.

[0273] 55. The method of any one of embodiments 49 to 54, wherein the pharmaceutical composition further comprises a pharmaceutically acceptable excipient and / or a micelle, liposome, exosome or lipid nanoparticle.

[0274] 56. A pharmaceutical composition produced by the method of embodiment 54 or 54.

[0275] 57. An in vitro or in vivo method of producing a peptide, polypeptide or protein, the method comprising translating the RNA molecule of any one of embodiments 1 to 28 or obtained by the method of any one of embodiments 49 to 55, wherein the RNA molecule comprises an ORF.

[0276] 58. A method of increasing the translational capacity of a polynucleotide comprising a nucleic acid sequence encoding a 5’ UTR, a 3’ UTR and / or an ORF comprising a sequence encoding a signal peptide, the method comprising:

[0277] (a) replacing the sequence encoding the 5 ’ UTR with a sequence encoding a 5 ’ UTR as defined in any one of embodiments 1 to 28;

[0278] (b) replacing the sequence encoding the 3 ’ UTR with a sequence encoding a 3 ’ UTR as defined in any one of embodiments 1 to 28; and / or

[0279] (c) replacing the sequence encoding a signal peptide with the sequence encoding a signal peptide as defined in any one of embodiments 1 to 28.

Claims

CLAIMS1. An RNA molecule comprising a 3’ untranslated region (3’ UTR), a 5’ untranslated region (5’ UTR), and an open reading frame (ORF) and / or a non-coding functional RNA sequence, wherein:(a) the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof;(b) the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof; and / or(c) the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

2. The RNA molecule of claim 1, wherein the 5’ UTR comprises an H2AX 5’ UTR or variant thereof, and the 3’ UTR comprises a COL6A2 3’ UTR or variant thereof.

3. The RNA molecule of claim 1 or 2, wherein:(a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2; and / or(b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13 or a variant thereof having at least 80% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13.

4. The RNA molecule of any one of the preceding claims, which has at least 70% of the translational capacity of a reference RNA molecule, wherein the RNA molecule comprises an H2AX 5’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 5’ UTR of SEQ ID NO: 1 or 2, and / or wherein the RNA molecule comprises a COL6A2 3’ UTR or variant thereof and the reference RNA molecule is an equivalent RNA molecule comprising a 3’ UTR of any one of SEQ ID NOs: 3 to 5 and 13, wherein translational capacity is calculated as the measured level of luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNAmolecule into HEK293T host cells, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase.

5. The RNA molecule of any one of the preceding claims, wherein the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to five amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

6. The RNA molecule of any one of the preceding claims, which has a greater translational capacity than a reference RNA molecule, wherein the RNA molecule comprises an ORF comprising a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof, and the reference RNA molecule is an equivalent RNA molecule comprising an ORF comprising a sequence encoding a signal peptide as set out in SEQ ID NO: 9, wherein translational capacity is calculated as the measured level of secreted luciferase activity 48 hours after transfection of 1 pg of the RNA molecule or the reference RNA molecule into HEK293T host cells, wherein the RNA molecule and the reference RNA molecule each comprise a sequence encoding a luciferase.

7. The RNA molecule of any one of the preceding claims, wherein the ORF :(a) comprises a sequence encoding a secreted peptide, a polypeptide or a protein;(b) comprises a sequence encoding an antibody; and / or(c) does not comprise a sequence encoding human H2A histone family member X or a sequence encoding collagen type VI alpha chain 2.

8. The RNA molecule of any one of the preceding claims, further comprising:(a) a poly(A) tail;(b) an internal ribosome entry site (IRES); and / or(c) a 5’ cap, preferably wherein the 5’ cap comprises a 7-methyl guanosine cap.

9. The RNA molecule of any one of the preceding claims, wherein:(a) the 5’ UTR comprises a sequence as set out in SEQ ID NO: 1 or 2, or a variant thereof having one to five nucleotide substitutions, insertions or deletions relative to SEQ ID NO: 1 or 2;(b) the 3’ UTR comprises a sequence as set out in any one of SEQ ID NOs: 3 to 5 and 13, or a variant thereof having a least 80% sequence identity to any one of SEQ ID NOs: 3 to 5 and 13; and(c) optionally, the signal peptide is a signal peptide as set out in SEQ ID NO: 6 or 7, or a variant thereof having one to five amino acid substitutions, insertions or deletions relative to SEQ ID NO: 6 or 7.

10. The RNA molecule of any one of the preceding claims, which comprises in order:(i) the 5’ UTR;(ii) the ORF and / or the non-coding functional RNA sequence;(iii) the 3 ’ UTR; and(iv) optionally, a poly(A) tail.

11. The RNA molecule of any one of the preceding claims, which is an mRNA molecule.

12. A polynucleotide comprising a nucleic acid sequence encoding the RNA molecule of any one of the preceding claims.

13. A polynucleotide comprising a nucleic acid sequence encoding a 5’ untranslated region (5’ UTR), a 3’ untranslated region (3’ UTR) and (i) an open reading frame (ORF) and / or a non-coding functional RNA sequence, or (ii) a sequence for introducing an ORF and / or a non-coding functional RNA sequence, wherein:(a) the 5’ UTR comprises an Histone 2A family member X (H2AX) 5’ UTR or variant thereof;(b) the 3’ UTR comprises a Collagen type VI alpha 2 chain (COL6A2) 3’ UTR or variant thereof; and / or(c) the ORF, the sequence encoding an ORF, or the sequence encoding a peptide, polypeptide or protein comprises a sequence encoding a signal peptide which is a cystatin C signal peptide or variant thereof.

14. The polynucleotide of claim 13, wherein the sequence encoded by the nucleic acid sequence is as defined in any one of claims 1 to 11.

15. The polynucleotide of any one of claims 12 to 14, wherein the polynucleotide further comprises:(a) a promoter;(b) a marker gene;(c) an origin of replication;(d) a multiple cloning site;(e) a restriction site, preferably at the terminal end of the sequence encoded by the nucleic acid sequence, optionally wherein the restriction site is a SapI restriction site; and / or(f) an inverted terminal repeat.1615. The polynucleotide of any one of claims 12 to 15, which is a DNA polynucleotide.

17. A vector comprising the polynucleotide of any one claims 12 to 16, optionally which is:(a) a linear vector or a circular vector; or(b) a viral vector, such as an adeno-associated virus (AAV) vector.

18. A cell comprising the polynucleotide of any one of claims 12 to 16, or the vector of claim 17, optionally which is:(a) a prokaryotic cell, preferably a bacterial cell, more preferably an Escherichia coli cell; or(b) a eukaryotic cell, preferably an insect cell or a mammalian cell, such as a human embryonic kidney cell.

19. A pharmaceutical composition comprising the RNA molecule of any one of claims 1 to 11, the polynucleotide of any one of claims 12 to 16, or the vector of claim17.

20. The pharmaceutical composition of claim 19, further comprising:(a) a pharmaceutically acceptable excipient; and / or(b) a micelle, liposome, exosome or lipid nanoparticle.

21. The pharmaceutical composition of claim 19 or 20 for use in a method of treating or preventing a disease in a subject.

22. An in vitro method of producing an RNA molecule, comprising contacting the polynucleotide of any one of claims 12 to 16, the vector of claim 17, or the cell of claim18, with an RNA polymerase.

23. The method of claim 22:(a) wherein the polynucleotide or vector is a circular molecule, and the method further comprises linearising the circular molecule, optionally wherein linearising the circular molecule is carried out using a restriction endonuclease at the terminal end of the sequence encoded by the nucleic acid sequence;(b) wherein the RNA polymerase is a DNA-dependent RNA polymerase;(c) wherein the RNA polymerase is T7 RNA polymerase, SP6 RNA polymerase or T3 RNA polymerase;(d) which further comprises isolating the RNA molecule following the step of contacting the polynucleotide, vector or cell with the RNA polymerase; and / or(e) which further comprises formulating the RNA molecule into a pharmaceutical composition, optionally wherein the pharmaceutical composition further comprises a pharmaceutically acceptable excipient and / or a micelle, liposome, exosome or lipid nanoparticle.

24. A pharmaceutical composition produced by the method of claim 23.

25. An in vitro or in vivo method of producing a peptide, polypeptide or protein, the method comprising translating the RNA molecule of any one of claims 1 to 11 or obtained by the method of claim 23 or 23, wherein the RNA molecule comprises an ORF.

26. A method of increasing the translational capacity of a polynucleotide comprising a nucleic acid sequence encoding a 5’ UTR, a 3’ UTR and / or an ORF comprising a sequence encoding a signal peptide, the method comprising: (a) replacing the sequence encoding the 5’ UTR with a sequence encoding a5’ UTR as defined in any one of claims 1 to 11;(b) replacing the sequence encoding the 3’ UTR with a sequence encoding a 3’ UTR as defined in any one of claims 1 to 11; and / or(c) replacing the sequence encoding a signal peptide with the sequence encoding a signal peptide as defined in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Stabilization of poly(a) sequence encoding DNA sequences

    WO2016005324A1

  • Artificial nucleic acid molecules for improved protein expression

    WO2016091391A1

  • Plasmid containing a sequence encoding an mRNA with a segmented poly(a) tail

    WO2020074642A1

  • Reagents and methods for producing bioactive secreted peptides

    WO2010129310A1

  • Terminal modifications of polynucleotides

    WO2016100812A1