Novel regulatory element for increasing RNA stability or mRNA translation and use thereof
Regulatory elements derived from viral genomes enhance mRNA stability and translation efficiency, addressing the limitations of current mRNA therapeutics by stabilizing the poly(A) tail and improving protein expression while reducing immune responses.
Patent Information
- Application Number
- PCT/KR2025/012497
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-01-22
- Filing Date
- 2025-08-18
- Publication Date
- 2026-02-19
AI Technical Summary
Current mRNA-based therapeutics face challenges with short half-life and limited stability, leading to reduced protein expression and potential immune responses, with existing RNA platforms like circular RNAs and self-amplifying RNA having limitations such as low translation efficiency and safety concerns.
Identification of regulatory elements from viral genomes that enhance mRNA stability and translation efficiency, including specific base sequences and stem-loop structures, which can be integrated into the untranslated regions of mRNA to stabilize the poly(A) tail and improve translation.
The identified regulatory elements significantly enhance mRNA stability and translation efficiency, allowing for sustained protein expression and reduced immune responses, even with modified bases, thus improving the effectiveness of mRNA-based therapeutics.
Smart Images

Figure KR2025012497_19022026_PF_FP_ABST
Abstract
Description
Novel regulatory elements for increasing RNA stability or mRNA translation and their uses
[0001] The present invention relates to novel regulatory elements for increasing RNA stability or mRNA translation and their uses.
[0002] Messenger RNA (mRNA) technology has revolutionized vaccine and therapeutic development. However, mRNA-based therapeutics suffer from a short half-life, and improving mRNA stability for sustained protein expression remains a technical challenge. Currently, mRNA vaccines are synthesized using in vitro transcription (IVT) as linear molecules with a 5′ cap and a poly(A) tail. Modified bases such as Ψ (pseudouridine), m1Ψ (N1-methylpseudouridine), m5C (5-methylcytidine), and mo5U (5-methoxyuridine) are introduced into the IVT reaction to increase protein expression and reduce the innate immune response.
[0003] The poly(A) tail is crucial for the degradation and translation regulation of IVT mRNA. The poly(A) tail interacts with the cytoplasmic poly(A)-binding protein (PABPC), which binds to the EIF4F complex and promotes translation initiation. Furthermore, the poly(A) tail is crucial for mRNA stability, protecting the mRNA termini. When the tail is shortened to approximately 27 nucleotides (nt), PABPC dissociates, leading to decapping and RNA degradation. Therefore, deadenylation serves as the rate-determining step in mRNA turnover. Therefore, maintaining the integrity of the poly(A) tail is a key strategy for sustaining protein production from therapeutic mRNA.
[0004] Various RNA platforms have been developed to improve the persistence of mRNA therapeutics [Wesselhoeft, RA et al. Engineering circular RNA for potent and stable translation in eukaryotic cells. Nat. Commun. 9, 2629 (2018)], [Geall, AJ et al. Nonviral delivery of self-amplifying RNA vaccines. Proc. Natl. Acad. Sci. USA 109, 14604-14609 (2012)], but each platform has its own limitations. For example, circular RNAs (circRNAs) are resistant to exoribonucleases, but they have low translation efficiency, are vulnerable to base modifications, and require a complex circularization process. Self-amplifying RNA (saRNA) and transfer RNA (taRNA) enable RNA replication within cells, but this process generates double-stranded RNA (dsRNA) intermediates that can trigger innate immune responses, raising safety concerns. Chemically synthesized or modified mRNAs containing 3′-terminal blockers, branched caps, or multiple poly(A) tails offer increased stability, but can introduce complexity for large-scale manufacturing.
[0005] To overcome these limitations, attempts have been made to optimize the untranslated region (UTR) of standard linear mRNA to maintain the poly(A) tail, thereby reducing mRNA degradation and increasing translation efficiency. However, improvements in stability have been limited. Furthermore, while attempts have been made to identify RNA stabilizing elements from viruses, previous screening studies have been limited in scale. Furthermore, the elements identified so far have shown limited compatibility with base modifications, limiting their use in therapeutic mRNAs or vaccines that require base modifications. While modified bases can reduce the production of double-stranded RNA (dsRNA) byproducts during the IVT process and their associated side effects (overall translation inhibition, RNA degradation, interferon production, and inflammatory cytokine responses), these modified bases also inhibit the function of mRNA regulatory elements, hindering attempts to use UTR regulatory elements to control mRNA stability and translation.
[0006] Inspired by these challenges, the researchers systematically screened a total of 196,277 viral genome fragments representing 337 virus species and 326 genera to identify RNA elements that strongly enhance mRNA stability and translation efficiency. As a result, they identified elements that function in diverse coding sequences and cell types, and demonstrated that these elements also function in mRNAs with modified bases, leading to the completion of the present invention.
[0007] One aspect is to provide regulatory elements for increasing RNA stability and / or mRNA translation.
[0008] Another aspect is to provide regulatory elements derived from fragments of the viral genome.
[0009] Another aspect is to provide a regulatory element comprising the base sequence of SEQ ID NOs: 1, 38, 41 and 43 to 50; or a base sequence having at least 80% identity thereto.
[0010] Another aspect is to provide a construct, vector, or recombinant host cell comprising a gene of interest and said regulatory elements.
[0011] Another aspect provides a composition comprising the construct, vector, or recombinant host cell.
[0012] Another aspect provides a method of preparing the construct, vector, recombinant host cell, or composition.
[0013] Another aspect provides a method for increasing RNA stability and / or mRNA translation of a target gene, comprising the step of inserting the regulatory element into the UTR of the target gene.
[0014] Another aspect provides a use for increasing RNA stability and / or mRNA translation of the construct, vector, recombinant host cell, or composition.
[0015] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for producing an mRNA construct or a protein of interest.
[0016] Another aspect provides a method for preventing, ameliorating or treating a disease comprising administering the construct, vector, recombinant host cell or composition to a subject in need thereof.
[0017] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for the prevention or treatment of a disease.
[0018] Each description and embodiment disclosed in this application may also be applied to each other description and embodiment. That is, all combinations of the various elements disclosed in this application fall within the scope of this application. Furthermore, the scope of this application is not limited by the specific descriptions set forth below. Furthermore, those skilled in the art will recognize or be able to ascertain, through routine experimentation alone, numerous equivalents to the specific embodiments of this application described in this application. Furthermore, such equivalents are intended to be encompassed by this application.
[0019]
[0020] One aspect provides a regulatory element capable of increasing RNA stability and / or mRNA translation. The regulatory element, which can increase RNA stability and / or mRNA translation, may be suitable for increasing protein production in a construct comprising a gene of interest.
[0021] The regulatory elements of the present application may be derived from a viral genome or a fragment thereof.
[0022] The viral genome may be Cytomegalovirus, Herpetohepadnavirus, Avihepadnavirus, Orthohepadnavirus, Nacovirus, Hepatovirus, Megrivirus, Cardiovirus, Teschovirus, Hunnivirus or Gallivirus.
[0023] The regulatory element of the present application may comprise a fragment of a cytomegalovirus, a herpeterhepadnavirus, an avihepadnavirus, an orthohepadnavirus, a nacovirus, a hepatovirus, a melegrivirus, a cardiovirus, a tescovirus, a hoonivirus or a gallivirus; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0024] The base sequence of the virus used in this application can be obtained from a known database (e.g., NCBI, etc.).
[0025] The regulatory element of the present application may be a fragment consisting of 1 to 200 base sequences derived from the viral genome. Specifically, the regulatory elements are fragments of the viral genome, 200, 199, 198, 197, 196, 195, 194, 193, 192, 191, 190, 189, 188, 187, 186, 185, 184, 183, 182, 181, 180, 179, 178, 177, 176, 175, 174, 173, 172, 171, 170, 169, 168, 167, 166, 165, 164, 163, 162, 161, 160, 159, 158, 157, 156, 155, 154, 153, 152, 151, 150, 149, 148, 147, 146, 145, 144, 143, 142, 141, 140, 139, 138, 137, 136, 135, 134, 133, 132, 131, 130, 129, 128, 127, 126, 125, 124, 123, 122, 121, 120, 119, 118, 117, 116, 115, 114, 113, 112, 111, 110, 109, 108, 107, 106, 105, 104, 103, 102, 101, 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 53, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43,It may comprise or consist of a fragment consisting of a sequence of 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6 or 5 bases.
[0026] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 1; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0027] In one specific example, the regulatory element of the present application may include a base sequence in which one or more nucleotides are mutated from the base sequence of SEQ ID NO: 1. Specifically, the regulatory element of the present application may include a base sequence in which 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer nucleotides are mutated from the base sequence of SEQ ID NO: 1.
[0028] In one specific example, the regulatory element of the present application may include a base sequence in which six or fewer bases are substituted corresponding to any one or more positions selected from the group consisting of positions 7, 8, 9, 11, 12, 13, 18, 20, 22, 24, 25, 26, 27, 28, and 30 in the base sequence of SEQ ID NO: 1.
[0029] In one specific example, the mutation may include a substitution of at least one base corresponding to positions 8, 18, or 28 of the nucleotide sequence of SEQ ID NO: 1 with A, G, C, or T. The mutation may include a substitution of a base corresponding to positions 7, 13, or 26 of the nucleotide sequence of SEQ ID NO: 1 with T, or a substitution of a base corresponding to positions 20 or 25 with G, or a substitution of a base corresponding to positions 30 with C. The mutation may include a substitution of a base corresponding to position 11 of the nucleotide sequence of SEQ ID NO: 1 with C or G, a substitution of a base corresponding to position 12 with T or G, or a substitution of a base corresponding to position 27 with A or T.
[0030] The regulatory element of the present application may comprise or consist of any one of the base sequences of SEQ ID NOs: 38 to 40; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0031] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 41 or 42; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity thereto. The SEQ ID NO: 41 may be ggccggtcc.
[0032] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 43; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0033] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 44; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0034] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 45; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0035] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 46; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0036] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 47; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0037] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 48; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0038] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 49; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0039] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 50; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0040] In one specific example, the regulatory element is a fragment of a base sequence of Melegrivirus A (Melegrivirus_A, NC_023858.1, SEQ ID NO: 77), wherein the fragment comprises the base sequence of SEQ ID NO: 1, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity to the base sequence of SEQ ID NO: 1.
[0041] In one specific example, the regulatory element is a fragment of a base sequence of human betaherpesvirus 5 (NC_006273.2), wherein the fragment comprises a base sequence of any one of SEQ ID NOs: 38 to 40, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to a base sequence of any one of SEQ ID NOs: 38 to 40.
[0042] In one specific example, the regulatory element may be a fragment of a base sequence of Tibetan frog hepatitis B virus (NC_030446.1), wherein the fragment comprises a base sequence of SEQ ID NO: 41 or 42, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to a base sequence of SEQ ID NO: 4140 or 4241.
[0043] In one specific example, the regulatory element is a fragment of a base sequence of Heron hepatitis B virus (NC_001486.1), wherein the fragment comprises the base sequence of SEQ ID NO: 43, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 43.
[0044] In one specific example, the regulatory element is a fragment of a base sequence of a hepatitis B virus (NC_003977.2), wherein the fragment comprises the base sequence of SEQ ID NO: 44, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 44.
[0045] In one specific example, the regulatory element is a fragment of a nucleotide sequence of Turkey calicivirus (NC_043516.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 45, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 45.
[0046] In one specific example, the regulatory element is a fragment of a base sequence of hepatitis A virus (Hepatovirus A, NC_001489.1), wherein the fragment comprises the base sequence of SEQ ID NO: 46, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity to the base sequence of SEQ ID NO: 46.
[0047] In one specific example, the regulatory element is a fragment of a base sequence of Saffold virus (NC_009448.2), wherein the fragment comprises the base sequence of SEQ ID NO: 47, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 47.
[0048] In one specific example, the regulatory element is a fragment of a base sequence of Teschovirus A (NC_003985.1), wherein the fragment comprises the base sequence of SEQ ID NO: 48, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity to the base sequence of SEQ ID NO: 48.
[0049] In one specific example, the regulatory element is a fragment of a nucleotide sequence of Hunnivirus A1 (NC_018668.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 49, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 49.
[0050] In one specific example, the regulatory element is a fragment of a base sequence of Gallivirus A1 (Gallivirus A1, NC_018400.1), wherein the fragment comprises the base sequence of SEQ ID NO: 50, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 50.
[0051]
[0052] In one specific example, the regulatory element may include a base sequence further comprising at least 1 and not more than 200 nucleotides in the 5' direction from the 9650th nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77) at the end of the base sequence of SEQ ID NO: 1; or a base sequence having at least 80% identity thereto.
[0053] Specifically, the nucleotide added above is a base sequence that additionally includes 1 to 200 consecutive nucleotides, 1 to 190 consecutive nucleotides, 1 to 180 consecutive nucleotides, 1 to 176 consecutive nucleotides, 1 to 160 consecutive nucleotides, 1 to 155 consecutive nucleotides, 1 to 150 consecutive nucleotides, 1 to 149 consecutive nucleotides, 1 to 120 consecutive nucleotides, 1 to 107 consecutive nucleotides, 1 to 100 consecutive nucleotides, 1 to 50 consecutive nucleotides, 1 to 49 consecutive nucleotides, 1 to 33 consecutive nucleotides, or 1 to 30 consecutive nucleotides in the 5' direction from the 9650th nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77); Or it may comprise or consist of a base sequence having at least 80% identity thereto. The additional nucleotide may be linked to the 5' end and / or the 3' end, preferably the 5' end, of the base sequence of SEQ ID NO: 1.
[0054] Specifically, the nucleotide added above may include a base sequence additionally including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 33, 40, 45, 49, 50, 100, 107, 120, 149, 150, 155, 160, 176, 180, 190, or 200 nucleotides in the 5' direction from the 9650th nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77); or a base sequence having at least 80% identity thereto. The above-mentioned additional nucleotide may be linked to the 5' end and / or the 3' end of the base sequence of SEQ ID NO: 1, preferably the 5' end.
[0055] In one specific example, the regulatory element may include a base sequence further comprising at least 1 and no more than 30 nucleotides in the 3' direction from the 9681st nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77) at the end of the base sequence of SEQ ID NO: 1; or a base sequence having at least 80% identity thereto.
[0056] Specifically, the nucleotide added above may include a base sequence additionally including 1 to 30, 1 to 25, 1 to 20, 1 to 19, 1 to 18, 1 to 17, or 1 to 12 nucleotides in the 3' direction from the 9681st nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77); or a base sequence having at least 80% identity thereto. The added nucleotide may be linked to the 5' end and / or the 3' end, preferably the 3' end, of the base sequence of SEQ ID NO: 1.
[0057] Specifically, the nucleotide added above may include a base sequence additionally including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 nucleotides in the 5' direction from the 9681st nucleotide in the base sequence of Melegrivirus_A (NC_023858.1, SEQ ID NO: 77); or a base sequence having at least 80% identity thereto. The added nucleotide may be linked to the 5' end and / or the 3' end, preferably the 3' end, of the base sequence of SEQ ID NO: 1.
[0058] Even if the present application describes "a regulatory element comprising a base sequence of a specific sequence number" or "a regulatory element having a base sequence of a specific sequence number," it is clear that a regulatory element having a base sequence with some of the sequences mutated may also be used in the present application, as long as it has the same or corresponding function as a regulatory element composed of the base sequence of the corresponding sequence number. The "mutation" refers to, but is not limited to, substitution, deletion, and / or insertion of bases.
[0059] For example, if it has the same or corresponding function as the above regulatory element, it is obvious that a regulatory element in which a meaningless sequence is added to or at the end of the regulatory element sequence of the corresponding sequence number, or in which a part of the sequence in or at the end of the regulatory element sequence of the corresponding sequence number is deleted, is also within the scope of the present invention.
[0060] In one specific example, the regulatory element of the present application may include a base sequence in which one or more nucleotides are mutated from the base sequence of SEQ ID NO: 9. Specifically, the regulatory element of the present application may include a base sequence in which 20 or fewer, 19 or fewer, 18 or fewer, 17 or fewer, 16 or fewer, 15 or fewer, 14 or fewer, 13 or fewer, 12 or fewer, 11 or fewer, 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer nucleotides are mutated from the base sequence of SEQ ID NO: 1.
[0061] In one specific example, the regulatory element is a base that is mutated at any one or more positions selected from the group consisting of bases 4, 6, 7, 9, 11, 15, 16, 18, 19, 24, 26, 27, 29, 31, 32, 39, 40, 46, 48, 55, 56, 57, 59, 60, 61, 66, 68, 70, 72, 73, 74, 75, 76, 80, 81, 83, 84, 86, 88, 90, 91, and 93 in the base sequence of SEQ ID NO: 9. It may include a base sequence. Specifically, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 55, 56, 57, 59, 60, 61, 66, 68, 70, 72, 73, 74, 75, and 76 in the base sequence of SEQ ID NO: 9 is mutated.
[0062] In one specific example, the mutation may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 15, 18, 26, 29, 32, 39, 40, 46, 81, and 84 of the base sequence of SEQ ID NO: 9 is mutated. Specifically, the mutation may include a mutation of 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer nucleotides selected from the group consisting of bases 15, 18, 26, 29, 32, 39, 40, 46, 81, and 84 of the base sequence of SEQ ID NO: 9. More specifically, the mutation may include substitution of bases corresponding to positions 15, 18, 26, 29, 32, 39, 40, 46, 81 and 84 of the base sequence of SEQ ID NO: 9 with other bases.
[0063] In one specific example, the mutation may include a substitution of a base corresponding to bases 4, 6, 11, 15, 18, 26, 29, 32, 39, 40, 46, 57, 66, 70, 75, 81, 84, 86, 90, and 93 of the base sequence of SEQ ID NO: 9 with another base.
[0064] In one specific example, the mutation may include substitution of bases corresponding to positions 6, 9, 15, 18, 26, 29, 32, 39, 40, 46, 55, 60, 68, 72, 81, 84, 88 and 91 of the base sequence of SEQ ID NO: 9 with other bases.
[0065] In one specific example, the mutation may include a substitution of at least one base corresponding to positions 56, 66, or 76 of the nucleotide sequence of SEQ ID NO: 9 with A, G, C, or T. The mutation may include a substitution of a base corresponding to positions 55, 61, or 74 of the nucleotide sequence of SEQ ID NO: 9 with T, a substitution of a base corresponding to positions 68 or 73 with G, or a substitution of a base corresponding to positions 78 with C. The mutation may include a substitution of a base corresponding to positions 59 of the nucleotide sequence of SEQ ID NO: 9 with C or G, a substitution of a base corresponding to positions 60 with T or G, or a substitution of a base corresponding to positions 75 with A or T.
[0066] Homology and identity refer to the degree to which two given base sequences are related, and can be expressed as a percentage. The terms homology and identity are often used interchangeably.
[0067] Whether any two sequences are homologous, similar or identical can be determined using known computer algorithms such as the "FASTA" program using default parameters, for example as in Pearson et al (1988) [Proc. Natl. Acad. Sci. USA 85]: 2444. Alternatively, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needleman program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) (version 5.0.0 or later) can be determined using the GCG program package (Devereux, J., et al, Nucleic Acids Research 12: 387 (1984)), BLASTP, BLASTN, FASTA (Atschul, [S.] [F.,] [ET AL, J MOLEC BIOL 215]: 403 (1990); Guide to Huge Computers, Martin J. Bishop, [ED.,] Academic Press, San Diego, 1994, and [CARILLO ETA / .](1988) SIAM J Applied Math 48: 1073). For example, sequence homology, similarity, or identity can be determined using BLAST or ClustalW of the National Center for Biotechnology Information database.
[0068] In one embodiment, the regulatory element may comprise or consist of any one or more base sequences of SEQ ID NOs: 2 to 36; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity thereto.
[0069] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 37; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0070] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 51 or 61; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0071] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 52 or 62; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0072] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 53 or 63; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0073] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 54 or 64; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0074] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 55 or 65; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0075] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 56 or 66; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0076] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 57 or 67; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0077] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 58 or 68; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0078] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 59 or 69; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0079] In one embodiment, the regulatory element comprises or consists of a base sequence of SEQ ID NO: 60 or 70; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.
[0080] In one embodiment, the control element may comprise at least one stem-loop structure.
[0081] Specifically, the regulatory element may comprise a base sequence of SEQ ID NO: 76, i.e., a CNGG motif, within the loop. More specifically, the regulatory element may comprise a CNGG motif within a penta loop, a tetra loop, a hexa loop, or an internal loop.
[0082] Alternatively, the regulatory element may comprise at least one internal loop; and a first stem-loop and a second stem-loop. Specifically, the regulatory element may comprise an internal loop, and the internal loop may comprise a first stem-loop and a second stem-loop. More specifically, the internal loop may comprise a base sequence of SEQ ID NO: 74, i.e., an AGACCY motif, and the second stem-loop may comprise a base sequence of SEQ ID NO: 75, i.e., a GGGGNARUD motif. The first stem-loop and the second stem-loop may form independent stem-loop structures by the internal loops, respectively.
[0083] The above N may be independently selected from a nucleotide selected from A, U, T, G and C or a nucleotide analogue thereof. The above Y may be a pyrimidine or an analogue thereof, the above R may be a purine or an analogue thereof, and the above D may be a nucleotide other than C or a nucleotide analogue thereof.
[0084] Among the regulatory elements of the present application, elements containing a CNGG motif can interact with ZCCHC14, which interacts with TENT4. The regulatory element of the present application can interact with ZCCHC14, which interacts with TENT4, to induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail through mixed tailing, or both. Among the regulatory elements of the present application, elements not containing a CNGG motif can interact with TENT4 but not with ZCCHC14, and can induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail through mixed tailing, or both.
[0085] In one embodiment, the regulatory element can increase the stability or translation of an RNA or mRNA comprising one or more modified nucleosides.
[0086] The above modified nucleoside may include, but is not limited to, the structure of Ψ (pseudouridine), m1Ψ (N1-methylpseudouridine), m5C (5-methylcytidine) or mo5U (5-methoxyuridine) and may include various other chemically modified nucleosides.
[0087] The regulatory element of the present application can increase the stability of RNA or induce an improvement in the translation efficiency of mRNA even when applied to RNA or mRNA into which the above-mentioned modified nucleoside has been introduced.
[0088]
[0089] Another aspect provides a construct comprising a target gene and the regulatory element described above. The regulatory element is as described above.
[0090] The term "construct" as used herein may be understood as a non-naturally occurring DNA or RNA. That is, a construct may be understood as an artificial nucleic acid molecule or a non-natural nucleic acid molecule, and may be designed and / or generated by genetic engineering or chemical synthesis. A construct may include at least one of the above regulatory elements and at least one open reading frame. A construct may be a DNA molecule, an RNA molecule, or a hybrid molecule comprising DNA and RNA portions.
[0091] In one specific example, the construct may be a construct including a UTR of a target gene and the regulatory element. Within the construct, the regulatory element may be inserted into the UTR of the target gene or linked to the UTR of the target gene in the 5' or 3' direction, and the insertion or linking method is not limited. Specifically, the insertion may be, but is not limited to, positioning the regulatory element in the 5' UTR or 3' UTR region of the target gene. The linking may include, but is not limited to, linking the regulatory element directly to the UTR of the target gene or linking the regulatory element and the UTR with an additional base sequence between them.
[0092] The target gene and the regulatory element within the construct may be of heterologous origin. The target gene may be derived from a gene other than the regulatory element and may not be naturally combined.
[0093] In one specific example, the construct may further comprise, but is not limited to, one or more barcode sequences, forward adapter sequences, reverse adapter sequences, poly(A) tail sequences, or combinations thereof.
[0094] In one specific example, the construct may further comprise a promoter sequence, wherein the gene of interest may be operably linked to the promoter sequence, but is not limited thereto. The term "operably linked" as used herein means that the gene sequence is functionally linked to a promoter sequence that initiates and mediates transcription of the gene of interest.
[0095] In one specific example, the construct may comprise a 5' repeat sequence and a 3' repeat sequence of a virus selected from the group consisting of, but not limited to, adeno-associated virus, adenovirus, alphavirus, retrovirus (e.g., gamma retrovirus, and lentivirus), parvovirus, herpesvirus, and SV40.
[0096] In one specific example, the construct may be an mRNA construct. The mRNA construct may further include, but is not limited to, a 5' UTR, a 3' UTR, a poly(A) tail sequence, or a combination thereof.
[0097] In the present application, the target gene may be at least one selected from the group consisting of a reporter, a protein, a physiologically active peptide, an antigen, or an antibody or a fragment thereof; or an antisense oligonucleotide, mRNA, dsRNA, shRNA, miRNA, siRNA, gRNA, saRNA, lncRNA, taRNA, ribozyme, ncRNA, exosomal RNA, and an aptamer, but the type is not limited as long as RNA stability and / or mRNA translation can be increased by the regulatory element of the present application.
[0098] In one specific example, the reporter may be, but is not limited to, luciferase, a fluorescent protein, beta-galactosidase, chloramphenicol acetyltransferase, or aequorin.
[0099] In one specific example, the physiologically active polypeptide may be, but is not limited to, a hormone, a cytokine, a cytokine binding protein, an enzyme, a growth factor, or insulin.
[0100] In one specific example, the antigen may be, but is not limited to, a vaccine antigen, a cancer-related antigen, or an allergy antigen.
[0101]
[0102] Another aspect provides a vector comprising the construct or a pool of the vector.
[0103] The term "vector" as used herein refers to a genetic construct containing a base sequence encoding a target protein or a target gene operably linked to suitable regulatory sequences so as to enable expression of the target protein in a suitable host. The regulatory sequences may include, but are not limited to, a promoter capable of initiating transcription, an optional operator sequence for regulating such transcription, a sequence encoding a suitable mRNA ribosome binding site, and sequences regulating the termination of transcription and translation. After being introduced into a suitable host cell, the vector may replicate or function independently of the host genome, or may be integrated into the genome itself.
[0104] The vector used in this application is not particularly limited as long as it is capable of expression within a host cell, and any vector known in the art may be used to introduce the vector into the host cell. Examples of commonly used vectors include plasmids, cosmids, viruses, and bacteriophages, either in their natural or recombinant form.
[0105]
[0106] Another aspect provides a recombinant host cell comprising the construct or vector.
[0107] The term "host cell" as used herein encompasses any cell capable of expressing a target protein, including cells that have undergone natural or artificial genetic modification. Furthermore, the host cell includes both eukaryotic and prokaryotic cells, and may be, but is not limited to, eukaryotic cells or cells derived from mammals (e.g., humans).
[0108] In the present application, the method for introducing a construct or vector into a cell includes any method for introducing a nucleic acid into a cell (e.g., transfection or transformation), and depending on the cell, a suitable standard technique known in the art can be selected and performed. Examples thereof include, but are not limited to, electroporation, calcium phosphate (CaPO4) precipitation, calcium chloride (CaCl2) precipitation, microinjection, polyethylene glycol (PEG) method, DEAE-dextran method, cationic liposome method, lipid nanoparticle method, and lithium acetate-DMSO method.
[0109]
[0110] Another aspect provides a composition comprising the construct, vector, or recombinant host cell. The construct, vector, recombinant host cell, or composition comprising them of the present application can express a protein of interest in vitro, in vivo, or ex vivo.
[0111] In one specific example, when the composition is administered to a subject, the target protein can be provided to the subject via the construct, vector, or recombinant host cell, thereby exhibiting a preventive or therapeutic effect on a disease (e.g., infectious disease) depending on the intended use of the provided target protein. Accordingly, the composition may be, but is not limited to, a pharmaceutical composition.
[0112] In one specific example, the construct, vector, or recombinant host cell may be used to produce the construct or target protein of the present application in vitro or ex vivo. Accordingly, the composition may be a composition for producing the construct or target protein of the present application, but is not limited thereto.
[0113] For example, if the target protein is a vaccine antigen, the construct, vector, recombinant host cell or composition itself can be used as a vaccine, or these can be used to produce a vaccine antigen.
[0114] In one specific example, the composition may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof. Specifically, the construct or vector of the present application may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof, or the recombinant host cell or composition of the present application may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof. The construct, vector, recombinant host cell or composition may increase RNA stability or mRNA translation by inducing an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both through the interaction of ZCCHC14 and a regulatory element and / or the interaction of ZCCHC14, TENT4 and a regulatory element.
[0115] Another aspect provides a method for preventing, ameliorating or treating a disease comprising administering the construct, vector, recombinant host cell, or composition to a subject in need thereof.
[0116] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for the prevention or treatment of a disease.
[0117]
[0118] Another aspect provides a method for producing a target protein, comprising the steps of culturing the recombinant host cell; and recovering the target protein.
[0119] In the present application, the method for producing a target protein using the recombinant host cell can be performed using a method widely known in the art. Specifically, the culture can be continuously cultured in a batch process, a fed batch process, or a repeated fed batch process, but is not limited thereto. The medium used for culture can be appropriately selected by a person skilled in the art depending on the host cell. Specifically, the recombinant host cell of the present application can be cultured under aerobic or anaerobic conditions while controlling temperature, pH, etc. in a general medium containing an appropriate carbon source, nitrogen source, phosphorus source, inorganic compound, amino acid, and / or vitamin.
[0120] The method for producing the target protein may further include an additional process after the culturing step. The additional process may be appropriately selected depending on the intended use of the target protein.
[0121] Specifically, the method for producing the target protein may include a step of recovering the target protein from at least one material selected from among the recombinant host cell, a dried product of the recombinant host cell, an extract of the recombinant host cell, a culture of the recombinant host cell, a supernatant of the culture, and a lysate of the recombinant host cell after the culturing step.
[0122] The method may further include a step of lysing the combination host cells prior to or concurrently with the recovery step. Lysis of the combination host cells may be performed using methods commonly used in the technical field to which the present application pertains, such as a lysis buffer, a sonicator, heat treatment, and a French press. In addition, the lysis step may include, but is not limited to, an enzymatic reaction such as a cell wall / membrane degrading enzyme, a nuclease, a nucleic acid transferase, and / or a protease.
[0123] In the present application, the dried product of the recombinant host cell can be produced by drying the cell that has accumulated the target substance, but is not limited thereto.
[0124] In the present application, the extract of a recombinant host cell may refer to the material remaining after separating the cell wall / cell membrane from the cell. Specifically, it may refer to the remaining components obtained by lysing the cell, excluding the cell wall / cell membrane. The cell extract includes the target protein, and components other than the target protein may include, but are not limited to, one or more of the following: protein, carbohydrate, nucleic acid, and fiber of the cell.
[0125] In the present application, the recovery step can recover the target protein using a suitable method known in the art (e.g., centrifugation, filtration, anion exchange chromatography, crystallization, HPLC, etc.).
[0126] In the present application, the recovery step may include a purification process. The purification process may separate only the target protein from the cell and purify it. Through the purification process, a pure, purified target protein can be produced.
[0127] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for producing an mRNA construct or a protein of interest.
[0128] Another aspect provides a method for increasing RNA stability and / or mRNA translation of a gene of interest, comprising the step of inserting or linking the regulatory element into a UTR of the gene of interest.
[0129] Another aspect provides uses of the construct, vector, recombinant host cell, or composition for increasing RNA stability and / or mRNA translation.
[0130]
[0131] Another aspect is a method for producing an mRNA construct, comprising the steps of transcribing the construct or vector in vitro; and recovering the transcribed mRNA construct.
[0132] The above-mentioned transfer method and recovery method can utilize any suitable method known in the art.
[0133] In one specific example, the method may further include, but is not limited to, a step of removing DNA of the construct or vector used as a template by treating with DNase I after transcription; and / or a washing step.
[0134] [Sequence List]
[0135] Intentionally omitted sequence (DNA)
[0136] Sequence number 41: ggccggtcc
[0137] Sequence number 74: agaccy
[0138] Sequence number 75: gggnarud
[0139] Sequence number 76: cngg
[0140] A regulatory element according to one specific embodiment can increase RNA stability or mRNA translation, which is a transcription product of a target gene, thereby increasing the expression level of a target protein. Therefore, the regulatory element of the present disclosure can be usefully utilized in systems requiring precise control of gene expression, such as gene therapy, vaccine development, and protein therapeutic production. Furthermore, the regulatory element of the present disclosure exhibits excellent stability-increasing and translation-regulating properties not only in unmodified RNA but also in RNA containing modified bases, making it useful for therapeutic mRNA or vaccine platforms requiring base modification.
[0141] Figure 1 is a graph showing the luciferase expression level of IVT mRNA when unmodified (U), m1Ψ-modified, or m5C-modified IVT mRNA was transfected using Lipofectamine (left) or LNP (right).
[0142] Figure 2 is a graph showing the luciferase expression levels of the 3′UTR of mRNA-1273, the 3′UTR of BNT162b2, and the unmodified (U) or m1Ψ modified IVT mRNA into which the K element was introduced.
[0143] Figures 3 and 4 are schematic diagrams showing a screening method for selecting candidate regulatory elements.
[0144] Figure 5 is a graph showing luciferase activity (left) and average log2 enrichment values (right) measured in positive controls (K4, K5) and negative controls (K4m, K5m) after primary fluorescence-based cell sorting for candidate regulatory element screening.
[0145] Figure 6 is an image showing the Spearman correlation coefficient between samples calculated based on elements with UMI counts of 200 or more in the IVT mRNA sample.
[0146] Figure 7 is a graph showing the log2 fold change (36h / 0h) of stabilized candidate elements obtained from 11 virus species (left) and the results of 4 repetitions of positive controls K3 and K4 (right).
[0147] Figure 8 is a circular chart (left) showing the distribution number of tiles including stabilizing elements A1S to A11S and a graph (right) listing the log2 fold change (36h / 0h) of individual elements A1 to A11.
[0148] Figure 9 is a graph showing the luciferase activity of RNA constructs into which candidate elements A1 to A11 and A1S to A11S have been introduced.
[0149] Figures 10, 11, and 12 are diagrams showing the predicted RNA secondary structure and genomic location of A elements. Below each element, the virus type from which the element originated and the base position within the full-length viral sequence are indicated, with the first base indicated as 0.
[0150] Figure 13 is a schematic diagram showing the mutagenesis experiment workflow.
[0151] Figures 14 to 24 are graphs showing the effects of single base substitutions and 1-2 nt deletion mutations within the A element as log2 fold change (36h / 0h) and deltalog2 fold change (36h / 0h, ΔLog2FC) images of the substitutions and deletions on the A element secondary structure. Figures 14 to 16 are images for A7S, Figures 17 and 18 are images for A1S, Figures 19 and 20 are images for A2S, Figures 21 and 22 are images for A4S, and Figures 23 and 24 are images for A6S. The positions indicated in each figure represent base positions in the A7, A1S, A2, A4, and A6 sequences, respectively, with the first base indicated as 0.
[0152] Figure 25 is a graph showing the effect of loop length on log2 fold change (36h / 0h) by introducing insertion mutations in A1, A2, A4, and A6.
[0153] Figure 26 is a graph showing the luciferase expression levels of IVT mRNA into which the A7 core sequence has been introduced and IVT mRNA into which the A7t3 sequence with random mutations has been introduced.
[0154] Figure 27 is a graph showing the luciferase expression level of IVT mRNA into which A1 to A6 and A8 to A11 core sequences have been introduced.
[0155] Figure 28 is a graph showing the luciferase activity of m1Ψ modified firefly mRNA containing the A element or unmodified firefly mRNA evaluated in parental HCT116 and ZCCHC14 knockout (KO) cells under treatment with RG7834 or the inactive isomer RO0321 (100 nM).
[0156] Figures 29 and 30 are graphs showing the results of measuring the poly(A) tail length of m1Ψ modified firefly mRNA containing the A element using Hire-PAT. Figure 29 shows the results of electroporation transfection of mRNA into HCT116 cells and treatment with RG7834 (100 nM) for 12 hours, and Figure 30 shows the results of analysis after transfection of mRNA into HCT116 cells using LNP and treatment with RO0321 or RG7834 (100 nM) for 24 hours.
[0157] Figure 31 is a graph showing the luciferase activity of m1Ψ modified firefly mRNA containing the A element or unmodified firefly mRNA evaluated in parental HeLa and ZCCHC2 KO cells under treatment with RG7834 or the inactive isomer RO0321 (100 nM).
[0158] Figures 32 to 35 are graphs showing the results of measuring the poly(A) tail length of m1Ψ modified firefly mRNA containing the A element using Hire-PAT. Figure 32 shows the results of analysis 12 hours after electroporation transfection of mRNA into HCT116 and ZCCHC14 KO cells; Figure 33 shows the results of analysis 16 hours after transfection of mRNA into HCT116 and ZCCHC14 KO cells using LNP; Figure 34 shows the results of analysis 12 hours after electroporation transfection of mRNA into Parental HeLa and ZCCHC2 KO cells; and Figure 35 shows the results of analysis 24 hours after transfection of mRNA into Parental HeLa and ZCCHC2 KO cells using LNP.
[0159] Figure 36 is a schematic diagram of the mechanism by which the A element stabilizes mRNA by extending the poly(A) tail by TENT4.
[0160] Figure 37 is a graph showing the activity of element A in various cell lines (HepG2, A549, MCF7, THP1, Jurkat, HeLa) through the level of luciferase expression.
[0161] Figure 38 is a graph showing the activity of m5C modified mRNA upon introduction of the A element through the amount of luciferase expression.
[0162] Figure 39 is a graph showing the activity of the A element compared to the 3′UTR of commercial vaccine mRNA (mRNA-1273 and BNT162b2) through the level of luciferase expression.
[0163] Figure 40 is a schematic diagram (top) showing the d2EGFP mRNA reporter structure including the A7S element, and a graph (bottom) showing the results of measuring d2EGFP expression in HCT116 cells transfected with the reporter using flow cytometry for 1 to 5 days.
[0164] Figure 41 is a schematic diagram (top) showing the HA reporter structure including the A7S element and an image (bottom) showing the quantification of HA protein expression with various codon-optimized CDSs.
[0165] Figure 42 is an image showing quantification of HA protein expression in a similar manner to Figure 41 for a 3-fold diluted sample.
[0166] Figure 43 shows a schematic diagram (left) of linear mRNA (Lin-Ctrl, Lin-A7S) and circRNA constructs containing the firefly reporter, and the results of RNA Tapestation analysis (right) confirming the status of circRNA production. GTP promotes the circularization reaction of the IVT product, RNase R digests the linear RNA to remove it, and gel purification separates and purifies the circular RNA by size.
[0167] Figure 44 is a graph showing the levels of ISG15 and IFIT2 mRNA measured by RT-qPCR after transfection of HCT116 cells with poly(I:C), m1Ψ modified mRNA, or circRNA.
[0168] Figure 45 is a graph showing the results of measuring the level of firefly RNA expression in HepG2 cells at 12, 18, 24, and 30 hours after electroporation and in HCT116 cells at 12, 24, 36, and 48 hours by RT-qPCR (top) and the results of calculating the level of firefly RNA expression and its AUC in HepG2 and HCT116 cells for 7 days (bottom).
[0169] Hereinafter, the present invention will be described in more detail through experimental examples and examples. These examples are intended solely to illustrate the present invention, and the scope of the present invention is not to be construed as being limited by these experimental examples and examples.
[0170]
[0171] Comparative Example. Confirming the Gene Expression Enhancement Efficacy of Existing Elements.
[0172] (1) Confirmation of expression performance of modified IVT mRNA
[0173] First, we measured firefly luciferase expression over 5 days with unmodified (U) or modified (m1Ψ, m5C) IVT mRNA. Specifically, 10 ng of IVT mRNA was transfected using Lipofectamine MessengerMAX, and data were normalized to the 0-hour value of unmodified mRNA.
[0174] Modified mRNA, particularly mRNA containing m1Ψ, exhibited superior expression performance compared to unmodified mRNA (Fig. 1, left). TRIM25-mediated repression is activated during delivery through acidified endosomes [Kim et al. Exogenous RNA Surveillance by Proton-Sensing TRIM25. Science (2025).], and this performance improvement was more pronounced when delivered via LNP than via lipofectamine (Fig. 1, right).
[0175]
[0176] (2) Confirmation of the efficacy of existing 3' UTR-derived elements
[0177] The 3′UTR strongly influences mRNA stability by inducing regulatory proteins, and base modifications can alter these interactions. Therefore, we tested the performance of various 3′UTRs in the context of m1Ψ-modified mRNA. Specifically, 10 ng of m1Ψ-modified mRNA was transfected with Lipofectamine MessengerMAX and normalized to the 0-hour value of the control m1Ψ mRNA.
[0178] As a result of the luciferase firefly expression level verification, the 3′UTR of mRNA-1273, derived from human alpha-globin, the UTR used in the commercial COVID-19 vaccine, did not show superior performance to the control UTR containing the plasmid-derived region and the microRNA-1 binding site. Furthermore, the 3′UTR of BNT162b2, a fusion of the human AES UTR and the mtRNR gene, only increased expression by approximately 2-fold (Fig. 2).
[0179] In addition, K1, K2, K5, K6, K7, K8, and K9, identified in a previous study [Seo, JJ et al. Functional viromic screens uncover regulatory RNA elements. Cell 186, 3291-3306.e21 (2023)], lost activity in the m1Ψ-modified mRNA environment, and only K3 and K4 retained some function (Fig. 2).
[0180] That is, many existing UTR elements are incompatible with m1Ψ, and there is a need to discover new elements that can strongly enhance gene expression not only in unmodified but also in base-modified mRNA.
[0181]
[0182] Example 1. Screening for selection of candidate regulatory elements.
[0183] (1) Tiling oligo design
[0184] A screening was performed to identify viral-derived elements that could enhance mRNA performance.
[0185] Specifically, for library construction, we collected genomic sequences of vertebrate viruses with the potential to function in human cells. The final selection included all seven Baltimore taxonomic lineages and 337 viral species across 297 genera. To densely cover the viral genome, tiling oligos were designed by shifting 197-nt windows by 20-nt intervals, generating a library of 196,277 unique oligos. This new library significantly expanded phylogenetic diversity and genome coverage compared to the human virus-based screening (96 genera, 143 species, and approximately 30,000 oligos) performed in a previous study [Seo, JJ et al. Functional viromic screens uncover regulatory RNA elements. Cell 186, 3291-3306.e21 (2023)].
[0186] Controls included Saffold virus-derived K4 (SEQ ID NO: 71), which is active on m1Ψ-modified mRNA but more potently on unmodified mRNA, and Aichivirus-derived K5 (SEQ ID NO: 72), which is active only on unmodified mRNA, along with their inactive mutants K4m and K5m. To address technical issues due to library complexity, the oligos were divided into two subsets, VL.1 and VL.2, and each was cloned into the 3′UTR of the EGFP plasmid (Figs. 3 and 4).
[0187]
[0188] (2) Primary and secondary screening
[0189] In the initial screening to assess the effect on protein yield, each plasmid pool was integrated into HEK293T cells using the Pa01 integrase system, ensuring that only one viral element was inserted per cell. Fluorescence-based cell sorting was then performed twice to rigorously select for elements that enhance protein expression. Sequence analysis of the sorted cells revealed that, as expected, the positive controls (K4 and K5) were enriched, while their inactive mutant counterparts, the negative controls (K4m and K5m), were not enriched (Fig. 5).
[0190] After the primary screening, genomic DNA was extracted from the sorted cell population with high fluorescence signals. Equal amounts of DNA from VL.1 and VL.2 were pooled to construct a single integrated library, from which viral fragments were amplified and inserted into the 3′UTR of the firefly luciferase reporter plasmid. As an additional control, the K3 element (SEQ ID NO: 73) was included in this integrated library. Using this library as a template, IVT mRNA containing m1Ψ was synthesized and delivered to HCT116 cells using LNPs.
[0191] Sequencing results confirmed that the integrated library contained a total of 10,784 viral tiles. After transfection of HCT116 cells with LNPs, the medium was replaced after 3 hours of incubation, and cells were harvested at 0 and 36 hours thereafter.
[0192]
[0193] (3) Evaluation of RNA stabilization activity
[0194] RNA stabilization activity was evaluated by amplifying and sequencing the 3′UTR via RT-PCR and measuring the amount of RNA at 36 hours compared to 0 hours.
[0195] Data obtained from the second screening showed high reproducibility, with correlation coefficients between four replicates at each time point being greater than 0.97 (Fig. 6).
[0196] As expected, the positive controls, K3 and K4, were detected in higher amounts in the 36-h samples than in the 0-h samples in all replicate experiments, while K5, which is incompatible with m1Ψ, was not enriched as expected (Fig. 7). Among the viral sequences obtained in the first screening, 131 tiles, corresponding to 1.21%, significantly enhanced the stability of m1Ψ-modified mRNA (adjusted p < 0.001, log2 fold change [36h / 0h] > 0.4), which showed better performance than K3 and K4. Many of the 131 tiles overlapped with adjacent tiles to form clusters. Therefore, these were classified into clusters, and elements with mRNA stabilizing activity were identified from the clusters, which were named A1 to A11 and are shown in Table 1 below (Fig. 8). Start and End indicate the positions within the full-length viral sequence from which the corresponding sequences were derived.
[0197]
[0198] Type Origin STARTEND Length Sequence Number A1HUMAN_BETAHERPESVIRUS_545134709197nt61A2TIBETAN_FROG_HEPATITIS_B_VIRUS26672863 197nt62A3HERON_HEPATITIS_B_VIRUS21212317197nt63A4HEPATITIS_B_VIRUS12191415197nt64A5TURKEY_CALICIVIR US71027298197nt65A6HEPATOVIRUS_A582778197nt66A7MELEGRIVIRUS_A95029698197nt11A8SAFFOLD_VIRUS78027998 197nt67A9TESCHOVIRUS_A67426938197nt68A10HUNNIVIRUS_A171827378197nt69A11GALLIVIRUS_A166536849197nt70
[0199]
[0200] Example 2. Confirmation of inhibition of RNA expression of element A
[0201] To validate the screening results of Example 1, two luciferase expression constructs were constructed for each stabilized cluster. Specifically, one was designed as a long sequence (L) encompassing the entire 197-nt tile with the lowest p-value and excellent RNA structure prediction, and the other was designed as a short sequence (S) containing only sequences common to all tiles within the cluster, likely representing a minimal functional element (Fig. 9, top). For A1 derived from HCMV RNA2.7, SL2.7, a 50-nt sequence known to be the minimal functional element, was used for A1S. A1 to A11 and A1S to A11S were cloned into the 3′UTR of a luciferase plasmid and used for IVT mRNA synthesis. Experiments were performed with mRNAs containing uridine or m1Ψ modifications, and a reporter mRNA lacking the A element was used as a control.
[0202] To correct for transfection efficiency between wells, m1Ψ-mutated firefly mRNA was cotransfected with m1Ψ-mutated renilla mRNA, and unmodified firefly mRNA was cotransfected with unmodified renilla mRNA. Luciferase expression was measured 4 days after transfection and normalized to the expression level of the control mRNA. (Expressed as mean ± SEM (n = 3 for m1Ψ, 6 for U). Statistical significance was determined by a two-tailed Student's t-test (*p < 0.05, **p < 0.01).)
[0203] The length and sequence number of the short element are as shown in Table 2 below, and Start and End indicate the position within the full-length viral base sequence from which the sequence is derived.
[0204]
[0205] TypeSTARTENDLengthSection NumberA1S (=SL2.7)4539458850nt51A2S2767284377nt52A3S21412257117nt53A4S1279133557nt54A5S7202725877nt55A6S58271 8137nt56A7S95449698155nt10A8S78647998135nt57A9S6842693897nt58A10S72227378157nt59A11S66536809157nt60
[0206]
[0207] As a result, all elements used in the experiment (A1 to A11) strongly increased luciferase expression when inserted into unmodified mRNA, and the expression levels were generally similar. This allowed us to verify the functions of these elements.
[0208] Meanwhile, performance in m1Ψ-modified mRNA varied significantly depending on the element. A1, A2, A4, A6, and A7 maintained over 60% activity in m1Ψ-modified mRNA, demonstrating their m1Ψ compatibility. A5, A8, A9, A10, and A11 showed statistically significant activity, albeit with reduced efficiency, while A3 was inactive (Fig. 9).
[0209] Additionally, for the shorter versions of each element (A1S to A11S), many of them (A1S, A4S, A5S, A6S, A7S, A9S, A10S, A11S) showed luciferase expression equivalent to or higher than the original 197nt version, indicating that these short sequences contain the core functional motif (Fig. 9).
[0210]
[0211] Example 3. Confirmation of the core structure and sequence characteristics of element A.
[0212] As discussed in Example 2, the identified RNA elements were derived from various viral lineages. A1 was derived from lncRNA2.7 of HCMV (Herpesviridae), a linear double-stranded DNA virus. A2, A3, and A4 were derived from hepatitis B virus (Hepadnaviridae), which has a partially double-stranded DNA genome, and A5 was derived from Turkey calicivirus (Caliciviridae), a positive-stranded single-stranded RNA virus. The remaining six elements were all derived from various species of the positive-stranded single-stranded RNA virus family Picornaviridae, and corresponded to Hepatitis A virus (A6), Melegrivirus A (A7), Saffold virus (A8), Teschovirus A (A9), Hunnivirus A (A10), and Gallivirus A1 (A11), respectively.
[0213]
[0214] Additionally, we aimed to elucidate the core structural and sequence characteristics of the A elements. The structure of A3 was derived based on mutagenesis experiments, and the remaining elements were predicted using EternaFold.
[0215] Structural analysis revealed that A7 does not contain a CNGG motif at all, but instead forms two independent stem-loop structures separated by an internal loop.
[0216] In addition, we confirmed that A elements other than A7 contain a CNGG stem-loop structure that depends on TENT4. Among them, A1, A2, A4, A5, A6, and A9 contain a penta-loop in the form of a CNGGN, A8, A10, and A11 contain a CNGG tetra-loop, and A3 contains a CNGGN motif in its internal loop (Figs. 10, 11, and 12).
[0217]
[0218] To better understand the functional determinants of these elements, we performed additional sequencing-based deep mutagenesis analysis on five elements (A1, A2, A4, A6, and A7). Specifically, the following types of mutations were introduced into each element:
[0219] (1) Single base substitution across the element;
[0220] (2) single or double base deletions;
[0221] (3) Complementary mutations that substitute or remove predicted base pairs;
[0222] (4) Examining the importance of loop size by introducing insertion mutations within the CNGGN penta-loop of A2, A4, and A6; and
[0223] (5) Analysis of base sequence specificity through double substitution for two consecutive bases of the loop, bulge, or the second and fifth base pairs within the CNGGN loop.
[0224] A total of 4,542 mutant oligos were synthesized and cloned into the 3′UTR of the luciferase reporter. Using these constructs as templates, m1Ψ-mutant IVT mRNA was synthesized, transfected into HCT116 cells, and the effects of the mutations were measured by RNA sequencing at 0 and 36 h (Fig. 13).
[0225] As a result, A7 was found to have a branched structure that starts from a bulged stem (S1) to form an internal loop (IL), and then branches out into two small stem-loops (SL1 and SL2). Substitution mutagenesis analysis confirmed that the activity was significantly reduced by mutagenesis in the sequence between positions 116 and 190 of A7S, especially between positions 150 and 179. Specifically, the IL and SL2 regions were found to be essential regions for the function of A7, and the core motif 'AGACCY' (SEQ ID NO: 74, Y is pyrimidine) exists in IL, and 'GGGGNARUD' (SEQ ID NO: 75, R is purine, D is a base other than C) exists in SL2, and both of these were confirmed to be essential for the stabilizing function of A7. Deletion mutagenesis analysis also supported the importance of IL and SL2, and sequences within the stem region of SL1 and S1 were also shown to assist the activity of A7 (Figs. 14 to 16).
[0226] The A1, A2, A4, and A6 elements share a similar loop structure characterized by a CNGGN penta-loop structure, and it was confirmed that the C, G, and G bases at positions 1, 3, and 4 play a crucial role in the function of the elements (Figs. 17 to 24). The activity of A1 was significantly reduced by mutagenesis between positions 141 and 155 (positions 20 to 34 in A1S), A2 between positions 134 and 151, A4 between positions 81 and 102, and A6 between positions 66 and 80.
[0227] In addition, it was confirmed that not only elements having a penta loop (5 bases), but also elements having a tetra loop (4 bases) or a hexa loop (6 bases) showed significant activity (Fig. 25).
[0228]
[0229] Example 4. Confirming the minimum functional unit of element A
[0230] (1) Confirmation of the minimum functional sequence of element A
[0231] To identify the core sequence that determines the function of the A element, a small element was additionally produced based on the base sequence of the region where activity was significantly reduced upon base substitution and deletion in the deep mutagenesis analysis results. Next, IVT mRNA containing the small element was synthesized, introduced into cells, and firefly luciferase activity was measured. The small element sequence is shown in Table 3, and Start and End indicate the positions within the full-length viral base sequence from which the sequence was derived.
[0232]
[0233] 서열번호서열이름StartEnd서열1A7_30nt96519680CAGACCCTGGTCCGGGGCAATGGGACCACT2A7_32nt96509681TCAGACCCTGGTCCGGGGCAATGGGACCACTG3A7_34nt96499682GTCAGACCCTGGTCCGGGGCAATGGGACCACTGT4A7_36nt96489683GGTCAGACCCTGGTCCGGGGGCAATGGGACCACTGTT 5A7_40nt96469685TGGGTCAGACCCTGGTCCGGGGCAATGGGACCACTGTTTC6A7_43nt96469688TGGGTCAGACCCTGGTCCGGGGCAATGGGACCACT GTTTCGCG7A7_47nt96469692TGGGTCAGACCCTGGTCCGGGGCAATGGGACCACTGTTTCGCGTTTA38A1_15nt45554569CGTAGGCTGGTCCTG39A1_ 21nt45524572CCTCGTAGGCTGGTCCTGGGG40A1_35nt45454579ATCCATTCCTCGTAGGCTGGTCCTGGGGAACGGGT41A2_9nt28052813GGCCGG TCC42A2_18nt28002817GCAGAGGCCGGTCCATGT43A3_29nt21922220CGCAATATCCCATATCACCGGCGGGAGCG44A4_22nt12991320TGCTCGC AGCAGGTCTGGAGCA45A5_16nt72317246TTGCTGCTGGTCAGAA46A6_15nt647661GACAGCTGGACTGTT47A8_16nt78077822GCACTGCTGGCA GTGC48A9_18nt67946811AACATTGCAGGCCAAGTT49A10_17nt72987314GCACGTGCTGGCACTGT50A11_16nt66806695GTGTCGCTGGCGTTAC
[0234]
[0235] As a result, we confirmed a significant increase in luciferase expression activity of IVT mRNA containing the A7_30nt element, which consists of the 150th to 179th base sequence of A7S (9651st to 9680th base sequence of Melegrivirus A), whose activity was greatly reduced when mutagenesis was performed. In addition, A7_32nt, A7_34nt, A7_36nt, A7_40nt, A7_43nt, and A7_47nt, which contain the same region as the A7_30nt element, were all confirmed to have an increase in luciferase expression activity. In the case of not including the core sequence of the A element (A7S-A7min, SEQ ID NO: 37), the increase activity was lower than in the case of including the core sequence, but luciferase expression increased by about 2 times compared to the control group (Fig. 26).
[0236] In other A elements (A1, A2, A3, A4, A5, A6, A8, A9, A10, and A11), small elements composed of base sequences in the region where activity was significantly reduced in the mutagenesis analysis were also confirmed to exhibit luciferase expression-increasing activity (Fig. 27).
[0237]
[0238] (2) Verification of the efficacy of the A element with introduced mutations
[0239] To determine whether the original activity is maintained even when a mutation is introduced at a specific position in the A7 element, a mutant element was constructed by introducing mutations at arbitrary positions based on A7t3 (SEQ ID NO: 9), which consists of base sequences 9603 to 9698 of Melegrivirus A. The positions are indicated in bold and underlined in Table 4 below.
[0240]
[0241] 서열번호서열이름염기서열(돌연변이 위치는 볼드 및 밑줄 처리)9A7t3CTTGATTAAG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTGGTC CGGGGCAATG GGACCACTGT TTCGCGTTTA ATCCTC12A7t3m1CTTGATTAAG CGAACGAAAT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTGGTC CGGGGCAATG GGACCACTGT TTCGCGTTTA ATCCTC13A7t3m2CTTGATTAAG CGAACGAAAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTGGTC CGGGGCAATG GGACCACTGT TTCGCGTTTA ATCCTC14A7t3m3CTTGATTAAG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTAGTGCGGGGGAATAGCACTACTGT TTCGCGTTTA ATCCTC15A7t3m4CTTGATTAAG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCTTCGTGCGGGGCAGTG GCACGACTGT TTCGCGTTTA ATCCTC16A7t3m5CTTGATTAAG CGAAGGATAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTGGTC CGGGGCAATG GGACCACTGTATCCCGTTTA ATCCTC17A7t3m6CTTGATTAAG CGAACTAAGT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTGGTC CGGGGCAATG GGACCACTGCTTAGCGTTTA ATCCTC18A7t3m7CTTCATAAAGGGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTGGTC CGGGGCAATG GGACCACTGT TTCGCTTTTTATGCTC19A7t3m8CTTGACTATG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTGGTCCGGGGCAATG GGACCACTGT TTCGCGTATAGTCCTC20A7t3m9CTTCATAAAGGGAAGGATAT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTGGTC CGGGGCAATG GGACCACTGTATCCCTTTTTATGCTC21A7t3m10CTTGACTATG CGAAGGATAT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTGGTC CGGGGCAATG GGACCACTGTATCCCGTATAGTCCTC22A7t3m11CTTCATAAAGGGAACTAAGT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTGGTC CGGGGCAATG GGACCACTGCTTAGCTTTTTATGCTC23A7t3m12CTTGACTATG CGAACTAAGT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTGGTC CGGGGCAATG GGACCACTGCTTAGCGTATAGTCCTC24A7t3m13CTTCATAAAGGGAAGGATAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTGGTC CGGGGCAATG GGACCACTGTATCCCTTTTTATGCTC25A7t3m14CTTGACTATG CGAAGGATAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTGGTC CGGGGCAATG GGACCACTGTATCCCGTATAGTCCTC26A7t3m15CTTCATAAAGGGAACTAAGT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTGGTC CGGGGCAATG GGACCACTGCTTAGCTTTTTATGCTC27A7t3m16CTTGACTATG CGAACTAAGT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTGGTC CGGGGCAATG GGACCACTGCTTAGCGTATAGTCCTC28A7t3m17CTTGATTAAG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCCTAGTCCGGGGGAATAGGACTACTGT TTCGCGTTTA ATCCTC29A7t3m18CTTGATTAAG CGAACGAAAT GGGACATACG ATCCATCTTA GAGTGGGTCA GACCTTGGTGCGGGGCAGTG GCACCACTGT TTCGCGTTTA ATCCTC30A7t3m19CTTCATAAAGGGAACTAAGT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCCTAGTC CGGGGGAATAGGACTACTGCTTAGCTTTTTATGCTC31A7t3m20CTTCATAAAGGGAACTAAGT GGGGCACACGGTCCATCTCGGAGTGGGCCA GACCTTGGTGCGGGGCAGTG GCACCACTGCTTAGCTTTTTATGCTC32A7t3m21CTTCATAAAGGGAAGGATAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTAGTC CGGGGGAATAGGACTACTGTATCCCTTTTTATGCTC33A7t3m22CTTCATAAAGGGAAGGATAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCTTGGTGCGGGGCAGTG GCACCACTGTATCCCTTTTTATGCTC34A7t3m23CTTGACTATG CGAAGGATAT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCTTGGTGCGGGGCAGTG GCACCACTGTATCCCGTATAGTCCTC35A7t3m24CTTCATAAAGGGAACTAAGT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCCTAGTC CGGGGGAATG GGACCACTGCTTAGCTTTTTATGCTC36A7t3m25CTTCATAAAGGGAACTAAGT GGGACGTAAG ACCCATCTCGGAGTGAGTCA GACCTTGGTGCGGGGCAGTG GCACCACTGCTTAGCTTTTTATGCTC
[0242]
[0243] As a result, it was confirmed that the RNA expression-increasing activity of the A7 element was maintained even when approximately 10 to 20 base mutations were introduced (Fig. 27). This suggests that a mutant sequence with at least approximately 80% sequence identity with the A7 element can provide a level of IVT mRNA performance enhancement comparable to that of the A7 element.
[0244]
[0245] Example 5. Confirmation of the interaction of element A with TENT4, ZCCHC14, and ZCCHC2.
[0246] (1) Confirm interaction with TENT4
[0247] Since most of the viral elements identified in Example 3 contain a CNGG motif, which is known to be recognized by ZCCHC14, a cofactor of TENT4 [Kim, D. et al. Viral hijacking of the TENT4-ZCCHC14 complex protects viral RNAs via mixed tailing. Nat. Struct. Mol. Biol. 27, 581-588 (2020).], we sought to examine the role of TENT4 in their mechanisms of action. To this end, parental HCT116 and ZCCHC14 knockout (KO) cells were treated (100 nM) with the TENT4 inhibitor RG7834 or its inactive isomer RO0321, and transfected with m1Ψ-modified firefly mRNA containing the A element in the 3′UTR or unmodified mRNA using Lipofectamine MessengerMAX. Renilla m1Ψ-modified mRNA lacking the functional element was also transfected and used for normalization. Luciferase signals were measured 4 days after transfection and normalized to control mRNA. The results are expressed as the mean ± SEM (n = 3), and statistical significance was determined by a two-tailed Student's t-test (*p < 0.05, **p < 0.01).
[0248] As a result, we confirmed that the expression increase effect was lost upon RG7834 treatment in all 11 tested A elements. This indicates that A elements are dependent on TENT4. A7, which does not contain the CNGG motif and is structurally different, also had its activity suppressed upon RG7834 treatment, and its activity was also reduced in TENT4-deficient (KO) cells (Fig. 28).
[0249]
[0250] Additionally, to analyze the effect of A elements on the regulation of poly(A) tail, m1Ψ modified mRNA with a poly(A) length of 60 nt was introduced using electroporation (Fig. 29) or LNPlipofectamine (Fig. 30), and then the change in poly(A) tail length was observed using high-resolution poly(A) tail analysis (Hire-PAT).
[0251] HCT116 cells were transfected with RNA via electroporation or LNP and treated with RG7834 (100 nM) for 12 or 24 h, while the control group was left untreated. * indicates a nonspecific band, which was not observed in the poly(A) tail profile due to the short 3′UTR of the control RNA. Signal intensities were corrected by normalizing the area under the curve (AUC) between the compared conditions, and the units are arbitrary units (au).
[0252] As a result, the control mRNA ('Ctrl') underwent significant deadenylation within 12 hours after transfection, but the IVT mRNA containing the A element elongated the poly(A) tail, forming a tail much longer than the initial 60 nt. This result suggested that A elements induce poly(A) tail elongation and increase the lifespan of mRNA. On the other hand, when TENT4 activity was inhibited with RG7834, the poly(A) tail length of the mRNA containing the A element was shortened and showed a length distribution similar to that of the control mRNA. These results imply that poly(A) tail elongation is mediated through TENT4 activity.
[0253]
[0254] (2) Confirmation of interaction with auxiliary factors
[0255] TENT4 is known to have two cofactors, ZCCHC14 and ZCCHC2, that regulate substrate specificity in mRNA tailing. ZCCHC14 binds to the CNGG motif, whereas ZCCHC2 binds to the Kobuvirus-derived K5 element, which lacks the CNGG motif and forms a double hairpin structure. To determine whether these cofactors mediate the functions of these viral elements, experiments were performed using ZCCHC14 knockout (KO) cells, ZCCHC2 knockout cells, and their respective parental cell lines. Specifically, the luciferase activity of m1Ψ-mutated firefly mRNA containing the A element in the 3′ UTR or unmutated firefly mRNA was evaluated in parental HCT116 and ZCCHC14 knockout (KO) cells under treatment (100 nM) of RG7834 or the inactive isomer RO0321. Transfections were performed using Lipofectamine MessengerMAX, and cotransfection with renilla m1Ψ mutant mRNA lacking the functional element was used for normalization. Luciferase signals were measured 4 days after transfection and normalized to control mRNA. The results are expressed as the mean ± SEM (n = 3), and statistical significance was determined using a two-tailed Student's t-test (*p < 0.05, **p < 0.01).
[0256] As a result, all viral elements (A1, A2, A3, A4, A5, A6, A8, A9, A10, and A11) containing the CNGG loop lost activity in ZCCHC14 KO cells, but maintained activity in ZCCHC2 KO cells. This indicates that these elements are strictly dependent on ZCCHC14. In contrast, A7, which does not have the CNGG motif, maintained activity in both KO cell lines, suggesting that A7 utilizes an as-yet-unidentified adapter protein other than ZCCHC14 or ZCCHC2 to induce TENT4 (Fig. 31).
[0257]
[0258] These results were consistently confirmed by Hire-PAT analysis. Poly(A) tail extension was observed in the parental cell line at elements A1, A2, A4, and A6, but was lost in ZCCHC14 KO cells. RNAs containing A7 were unaffected by deletion of ZCCHC14 or ZCCHC2 (Figs. 32–35).
[0259]
[0260] Taken together, these results indicate that the RNA elements identified in this study act by promoting TENT4-mediated mixed tailing, and that A7 is a unique enhancer that acts independently of known TENT4 cofactors (Fig. 36).
[0261]
[0262] Example 6. Confirmation of the adaptability of element A to various molecular environments.
[0263] (1) Confirmation of activity in various cells
[0264] To confirm that element A functions in various molecular environments, the following experiments were performed.
[0265] First, the activity of the A element was confirmed in various cell lines (HepG2, A549, MCF7, THP1, Jurkat, and HeLa). Transfections were performed using Lipofectamine MessengerMAX, and m1Ψ-modified renilla mRNA was co-transfected and used for normalization. Luciferase signals were measured after 3 days in THP1 and Jurkat, and after 4 days in the other cells. The expression values were normalized to the expression values of the control mRNA, and are expressed as the mean ± SEM (n = 3). Statistical significance was calculated according to a two-tailed Student's t-test (*p < 0.05, **p < 0.01).
[0266] As a result, we confirmed that the A elements exhibited different activities in various cell lines, including HepG2 (hepatocyte-derived), A549 (lung epithelial cell-derived), MCF7 (breast cancer), THP1 (monocyte), Jurkat (T cell), and HeLa (cervical cancer). Specifically, A2 and A7 showed a wide range of activities. A2 increased luciferase expression in all cell lines except Jurkat, and A7S showed strong activity in all cell lines except THP1, demonstrating strong functionality across various cell lineages (Fig. 37). A1S enhanced expression in HCT116 (Fig. 1h), MCF7, and HeLa. A4S showed activity in HCT116, A549, MCF7, THP1, and HeLa, and A6S was confirmed to function in HCT116.
[0267] These results suggest that the efficacy of the identified elements may vary depending on cell type, with highly cell-specific elements being advantageous for targeted therapy, whereas elements that can act across a wide range of cells are suitable for general mRNA applications. For example, A2 showed potential for monocyte-targeted immunotherapy through its potent activity in THP1 cells, while A7 showed potential for applications in metabolic disease treatment and T cell-based immunotherapy through its activity in HepG2 and Jurkat cells.
[0268]
[0269] (2) Confirmation of activity in RNA with modified bases
[0270] We further evaluated the compatibility activity with 5-methylcytosine (m5C), another nucleotide modification known to attenuate the immune response to IVT mRNA. m5C-modified mRNA was transfected into HCT116 cells using Lipofectamine MessengerMAX, and luciferase expression was measured. m5C-modified renilla mRNA was cotransfected for normalization, and cells were harvested 4 days later. (Mean ± SEM (n = 3), **p < 0.01)
[0271] As a result, we confirmed that A7S retains its function even under m5C modification conditions (Fig. 38). This indicates that the A7 element has a wide range of compatibility with RNA modifications.
[0272]
[0273] (3) Confirmation of activity by comparison with commercial 3'UTR
[0274] Firefly luciferase signals from the 3′UTR of commercial vaccine mRNAs (mRNA-1273 and BNT162b2) were measured and compared with those of A7. Specifically, 10 ng of m1Ψ-modified mRNA was transfected, and the signals from untransfected samples were subtracted and normalized to the control mRNA values. (Mean ± SEM (n = 3), *p < 0.05, **p < 0.01)
[0275] As a result, A7S showed a 65-fold and 17-fold increase in protein production 3 days after transfection, and a 114-fold and 78-fold increase 5 days after transfection, and a 6-fold and 4-fold increase in total area under the curve (AUC) (Fig. 39).
[0276]
[0277] (4) Confirmation of compatibility with mRNA encoding a protein with a short half-life
[0278] One important property of an effective UTR is that it maintains its function in diverse molecular environments. To confirm this, we designed an mRNA expressing d2EGFP, a short-lived protein with a half-life of approximately 2 hours, and linked it to the 5′UTR (5′) of the COVID-19 vaccine mRNA-1273. d2EGFP expression in HCT116 cells transfected with the m1Ψ-modified d2EGFP mRNA containing the A7S and HBA 3′UTRs was measured by flow cytometry for 1–5 days. The human α-globin (HBA) 3′UTR used in mRNA-1273 was used as a control (Fig. 40, top). (Mean ± SEM (n = 3))
[0279] Fluorescent cell analysis results showed that m1Ψ variant mRNA including A7S produced GFP signals even 5 days after transfection, but the control m1Ψ variant mRNA signal disappeared after 2 days (Fig. 40, bottom).
[0280]
[0281] (5) Check compatibility with codon-optimized coding sequences
[0282] Additionally, we assessed whether A7 is compatible with membrane-associated proteins translated in the endoplasmic reticulum (ER) membrane. To this end, we used the hemagglutinin (HA) gene of influenza A virus (A / California / 07 / 2009, H1N1, pdm09), and incorporated an HBA-derived replacement sequence from BNT162b2 (mRNA vaccine) into the 5′UTR (Fig. 41, top).
[0283] Additionally, we applied several codon optimization strategies to experiment with various nucleotide configurations while maintaining the same protein sequence. Specifically, iCodon was used to generate HA-co1, HA-co2, and HA-co3, and LinearDesign was used to generate HA-co4, HA-co5, and HA-co6:
[0284] - 'HA-co1' (default options applied, sequence number 79);
[0285] - 'HA-co2' (added a term that lowers uridine content, SEQ ID NO: 80);
[0286] - 'HA-co3' (added term to lower minimum free energy (MFE), sequence number 81).
[0287] - 'HA-co4' (lambda value 0, sequence number 82);
[0288] - 'HA-co5' (optimizing codon adaptation index (CAI) and minimizing MFE by setting lambda value to 2.8, SEQ ID NO: 83); and
[0289] - 'HA-co6' (additionally reduces uridine content, SEQ ID NO: 84)
[0290] HA protein expression with various codon-optimized CDSs was quantified by western blot, and the intensity of HA bands was quantified relative to diluted controls and normalized to GAPDH (mean ± SEM (n = 3), *p < 0.05, **p < 0.01).
[0291] As a result, protein expression increased in all strategies, although the extent varied. HA-co2 showed a relatively low increase, while HA-co1, HA-co3, HA-co4, HA-co5, and HA-co6 showed approximately 5- to 10-fold increases in protein (Figures 41 and 42). A7 improved protein yield in all seven coding sequences (CDSs) tested. These results indicate that A7 can increase expression even in highly optimized CDSs.
[0292]
[0293] Taken together, these results suggest that A7 has high tolerance to various variables such as 5′UTR composition, coding sequence length, codon optimization level, RNA secondary structure, uridine content, and translation site (cytoplasmic free ribosomes or ER-associated ribosomes), and can be very universally utilized to improve the performance of IVT mRNA.
[0294]
[0295] Example 7. Comparison of the stability of linear mRNA and circRNA containing the A element.
[0296] To measure the stability of linear mRNA containing A7, we evaluated the effect of A7 on mRNA half-life (Fig. 43, left). Additionally, we compared this with a circRNA known to have a long half-life. The circRNA used was an optimized cirRNA reported in a recent study [Chen, R. et al. Engineering circular RNA for enhanced protein production. Nat. Biotechnol. 41, 262-272 (2023)] (Fig. 43, right). The optimized IRES derived from HRV-B3 and Apt-eIF4G were inserted at the 5′ end of the firefly CDS, and the human α-globin 3′ UTR sequence was ligated to the 3′ end. As a negative control, linear IVT mRNA containing a vector-derived 3′ UTR sequence was used.
[0297]
[0298] First, we measured the expression levels of ISG15 and IFIT2 using RT-qPCR and confirmed that the RNAs did not induce an innate immune response (Fig. 44).
[0299]
[0300] To assess mRNA stability, RNA half-life was measured by RT-qPCR at 12, 18, 24, and 30 h after transfection into HepG2 cells. Electroporation was used to exclude the influence of endosomal RNA or TRIM25.
[0301] As a result, circRNA exhibited a half-life of 15.1 hours, showing significantly higher stability than the control linear mRNA (6.0 hours). Notably, linear mRNA containing A7S exhibited a half-life of 15.9 hours, demonstrating a similar level of stability as circRNA. This trend was also observed in HCT116 cells, with the half-lives of A7S-containing RNA being 19.0 hours, circRNA being 19.5 hours, and control linear RNA being 6.2 hours (Figure 45, top).
[0302]
[0303] Additionally, we measured luciferase protein expression over time. For this, 10 ng of m1Ψ-modified linear mRNA and circRNA were transfected into HepG2 and HCT116 cells using Lipofectamine MessengerMAX, and the expression levels were normalized to the control group at 0 h (mean ± SEM (n = 3)). In addition, the area under the curve (AUC) for luciferase expression was calculated by integrating the baseline signal at 7 days in Lin-Ctrl, and expressed as the mean ± SEM (n = 3). Statistical significance was determined using a two-tailed Student's t-test (**p < 0.01).
[0304] As a result, although circRNAs showed low expression levels at early post-transfection times (3 hours and 1 day), their expression persisted for a long time, and their total protein yield (based on area under the curve) was similar to that of the rapidly degraded control linear RNA. In contrast, linear mRNAs containing A7S showed significantly higher protein expression levels than the control linear mRNAs and circRNAs at all time points, and maintained higher expression levels than circRNAs even 7 days after transfection. Based on area under the curve, A7S-linear mRNAs showed 5.3- and 5.8-fold higher protein yields than circRNAs in HepG2 and HCT116 cells, respectively (Figure 45, bottom).
[0305]
[0306] These results imply that m1Ψ-modified linear mRNAs containing A7 have comparable stability to circRNAs, while significantly outperforming circRNAs in terms of translation efficiency.
[0307]
[0308] From the above description, those skilled in the art will understand that the present invention can be implemented in other specific forms without altering its technical concept or essential characteristics. In this regard, it should be understood that the experimental examples and embodiments described above are illustrative in all respects and not restrictive. The scope of the present invention should be interpreted as encompassing all changes or modifications derived from the meaning and scope of the following claims and their equivalent concepts, rather than the detailed description above.
Claims
1. A regulatory element comprising a base sequence selected from any one of SEQ ID NOs: 1, 38, 41, and 43 to 50; or a base sequence having at least 80% identity therewith.
2. In claim 1, (a) a base sequence additionally including at least 1 and not more than 200 nucleotides in the 5' direction from the 9650th nucleotide in the base sequence of SEQ ID NO: 77 to the end of the base sequence of SEQ ID NO: 1; or a base sequence having at least 80% identity therewith; (b) a base sequence additionally including at least 1 and not more than 30 nucleotides in the 3' direction from the 9681st nucleotide in the base sequence of SEQ ID NO: 77 at the end of the base sequence of SEQ ID NO: 1; or a base sequence having at least 80% identity therewith; or (c) a base sequence further comprising, at each end of the base sequence of SEQ ID NO: 1, 1 to 200 nucleotides in the 5' direction from the 9650th nucleotide in the base sequence of SEQ ID NO: 77 and 1 to 30 nucleotides in the 3' direction from the 9681st nucleotide in the base sequence of SEQ ID NO: 77; or a base sequence having at least 80% identity thereto. A control element including:
3. In claim 1, (i) a fragment of a base sequence of human betaherpesvirus 5 (Human betaherpesvirus 5, NC_006273.2), wherein the fragment is a base sequence of SEQ ID NO: 38 or a base sequence having at least 80% identity therewith; (ii) a fragment of the base sequence of Tibetan frog hepatitis B virus (NC_030446.1), wherein the fragment is a base sequence of SEQ ID NO: 41 or a base sequence having at least 80% identity therewith; (iii) A fragment of the base sequence of Heron hepatitis B virus (NC_001486.1), wherein the fragment is the base sequence of SEQ ID NO: 43 or a base sequence having at least 80% identity therewith; (iv) A fragment of a base sequence of hepatitis B virus (NC_003977.2), wherein the fragment is a base sequence of SEQ ID NO: 44 or a base sequence having at least 80% identity therewith; (v) A fragment of the base sequence of Turkey calicivirus (NC_043516.1), wherein the fragment is the base sequence of SEQ ID NO: 45 or a base sequence having at least 80% identity therewith; (vi) A fragment of a base sequence of hepatitis A virus (Hepatovirus A, NC_001489.1), wherein the fragment is a base sequence of SEQ ID NO: 46 or a base sequence having at least 80% identity therewith; (vii) A fragment of a base sequence of Saffold virus (NC_009448.2), wherein the fragment is a base sequence of SEQ ID NO: 47 or a base sequence having at least 80% identity therewith; (viii) The above regulatory element is a fragment of a base sequence of Teschovirus A (NC_003985.1), wherein the fragment is a base sequence of SEQ ID NO: 48 or a base sequence having at least 80% identity therewith; (ix) a fragment of the base sequence of Hunnivirus A1 (NC_018668.1), wherein the fragment is the base sequence of SEQ ID NO: 49 or a base sequence having at least 80% identity therewith; or (x) A fragment of the base sequence of Gallivirus A1 (Gallivirus A1, NC_018400.1), wherein the fragment is the base sequence of SEQ ID NO: 50 or a base sequence having at least 80% identity therewith. A control element including:
4. In claim 1, the base sequence is a regulatory element including a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 4, 6, 7, 9, 11, 15, 16, 18, 19, 24, 26, 27, 29, 31, 32, 39, 40, 46, 48, 55, 57, 60, 66, 68, 70, 72, 75, 80, 81, 83, 84, 86, 88, 90, 91, and 93 in the base sequence of SEQ ID NO: 9 is mutated.
5. A regulatory element according to claim 1, wherein the base sequence comprises at least one base sequence selected from the group consisting of SEQ ID NOs: 2 to 36, or 51 to 70.
6. In claim 1, a control element including an internal loop, wherein the internal loop includes a first stem-loop structure and a second stem-loop structure, The above internal loop contains the base sequence of sequence number 74, A regulatory element wherein the second stem-loop comprises a base sequence of SEQ ID NO:
75.
7. In claim 1, comprising at least one stem-loop structure, The above stem-loop structure is a regulatory element comprising a base sequence of SEQ ID NO: 76 within the loop.
8. A regulatory element that increases RNA stability or mRNA translation in claim 1.
9. A regulatory element according to claim 8, wherein the RNA or mRNA is an RNA or mRNA containing one or more modified nucleosides.
10. In claim 1, the regulatory element interacts with TENT4, ZCCHC14, or both to induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both.
11. A construct comprising a target gene; and a regulatory element of any one of claims 1 to 10.
12. A construct according to claim 11, wherein the target gene is at least one selected from the group consisting of a reporter, a protein, a physiologically active peptide, an antigen, or an antibody or a fragment thereof; or an antisense oligonucleotide, mRNA, dsRNA, shRNA, miRNA, siRNA, gRNA, saRNA, lncRNA, taRNA, ribozyme, ncRNA, exosomal RNA, and an aptamer.
13. In claim 11, the construct is an mRNA construct.
14. A vector comprising the construct of claim 11.
15. A recombinant host cell comprising the construct of claim 11 or a vector comprising the construct.
16. A recombinant host cell according to claim 15, wherein the host cell further comprises TENT4 or a gene encoding it, or ZCCHC14 or a gene encoding it.
17. A composition comprising the construct of claim 11; a vector comprising the construct; or a recombinant host cell comprising the construct or vector.
18. A composition according to claim 17, wherein the composition is for preventing or treating a disease; or for producing an mRNA construct or a protein encoded by a target gene.
19. A composition according to claim 17, wherein the construct or vector further comprises a gene encoding TENT4 or ZCCHC14 or a gene encoding it, or the recombinant host cell or composition further comprises TENT4 or a gene encoding it or ZCCHC14 or a gene encoding it.
20. A method for increasing RNA stability or mRNA translation of a target gene, comprising the step of inserting or linking a regulatory element of any one of claims 1 to 11 into a UTR of the target gene.
Citation Information
Patent Citations
Method for generating educational scenario using generative artificial intelligence
KR1020240146874A
Gold deposition method using photodeposition through low temperature and photocatalytic sterilization filter on which gold is deposited by the deposition method
KR1020250109289A
Novel viral regulatory elements
WO2023062359A2
Lentiviral vectors
WO2023062367A1
Cited By
High-stability mRNA construct as well as preparation method and application thereof
CN121759475A