Novel regulatory element for increasing gene expression, screening method therefor, and use thereof

By constructing a high-density oligonucleotide library and using MPRA, RNA regulatory elements are systematically identified and characterized, addressing inefficiencies in current methods and enhancing gene expression and translation.

WO2026038930A1PCT designated stage Publication Date: 2026-02-19SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/012501
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-22
Filing Date
2025-08-18
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Current methods for characterizing viral RNA regulatory elements are inefficient and limited to a tiny fraction of the vast viral genome, hindering comprehensive understanding and application of their regulatory mechanisms.

Method used

A high-density oligonucleotide library is constructed from vertebrate virus genome sequences, and a massively parallel reporter assay (MPRA) is employed to identify and characterize RNA regulatory elements across various viruses, leveraging sequence conservation within viral genera.

Benefits of technology

This approach systematically identifies RNA regulatory elements that can enhance gene expression, stability, and translation, offering potential biotechnological applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025012501_19022026_PF_FP_ABST
    Figure KR2025012501_19022026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a novel regulatory element for increasing gene expression, a screening method therefor, and a use thereof. A regulatory element according to one embodiment can increase RNA stability or mRNA translation, which is a transcription product of a target gene, and thus can increase the expression level of a target protein, and can be effectively used in systems requiring a high level of gene expression, such as gene therapy, vaccine development, protein therapeutic agent production, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Novel regulatory elements that increase gene expression, methods for screening thereof, and uses thereof

[0001] The present invention relates to novel regulatory elements that increase gene expression, methods for screening the same, and uses thereof.

[0002] Viruses have evolved various mechanisms to hijack cellular gene expression machinery, and this process has contributed significantly to the advancement of RNA biology and biotechnology. For example, key discoveries such as the 7-methyl guanosine cap, the internal ribosome entry site (IRES), and the RNA triple helix were made through studies on reovirus, poliovirus, and Kaposi's sarcoma-associated herpesvirus, respectively (Non-patent literature 0001 (Furuichi et al., 1975)); (Non-patent literature 0002 (Pelletier & Sonenberg, 1988)); (Non-patent literature 0003 (Mitton-Fry et al., 2010)).

[0003] Despite this importance, major discoveries to date have been achieved through in-depth, low-throughput analyses of pathogenic viruses, representing only a tiny fraction of the total viral genome, or virome. Given the vast diversity of viral sequences revealed by recent metagenomics studies, efficient strategies for functionally characterizing them are essential (Non-Patent Document 0004 (Simmonds et al., 2017)).

[0004] Accordingly, in the present invention, we aimed to systematically identify and characterize regulatory RNA elements present in various vertebrate viruses. By constructing a high-density oligonucleotide library containing vertebrate virus genome sequences and performing a comprehensive massively parallel reporter assay (MPRA) screening based on this library, we aimed to derive new insights into the regulatory mechanisms of viral genomes. By leveraging the high sequence conservation within viral genera, we selected representative viruses for each genus, enabling us to effectively identify RNA regulatory elements in various viruses. The elements identified in this invention may offer potential future biotechnological applications.

[0005] [Prior Art Literature]

[0006] [Non-patent literature]

[0007] (Non-patent Document 0001) Furuichi, Y., Morgan, M., Muthukrishnan, S., & Shatkin, A. J. (1975). Reovirus messenger RNA contains a methylated, blocked 5'-terminal structure: m-7G(5')ppp(5')G-MpCp-. Proceedings of the National Academy of Sciences of the United States of America, 72(1), 362-366.

[0008] (Non-patent Document 0002) Pelletier, J., & Sonenberg, N. (1988). Internal initiation of translation of eukaryotic mRNA directed by a sequence derived from poliovirus RNA. Nature, 334(6180), 320-325.

[0009] (p. 0003) Mitton-Fry, RM, DeGregorio, SJ, Wang, J., Steitz, TA, & Steitz, JA (2010). Recognition of the poly(A) tail by a viral RNA element through assembly of a triple helix. Science, 330(6008), 1244-1247.

[0010] (MN 0004) Simmonds , P , Adams , MJ , Benko , M , Breitbart , M , Brister , JR , Carstens , EB , Davison , AJ , Delwart , E , Gorbalenya , AE , Harrach , B , Hull , R , King , AMQ , Koonin , EV , Krupovic , M , . Kuhn , JH , Lefkowitz , EJ , Nibert , ML , Orton , R. , Roossinck , MJ , ... Zerbini , FM (2017). Consensus statement: Virus taxonomy in the age of metagenomics. Nature Reviews. Microbiology, 15(3), 161-168.

[0011] (Paper 0005) Seo, J. J., Jung, S.-J., Yang, J., Choi, D.-E., & Kim, V. (2023). Functional viromic screens uncover regulatory RNA elements. Cell, 186(15), 3291–3306.e21.

[0012] (비특허문헌0006) Durrant, M. G., Fanton, A., Tycko, J., Hinks, M., Chandrasekaran, S. S., Perry, N. T., Schaepe, J., Du, P. P., Lotfy, P., Bassik, M. C., Bintu, L., Bhatt, A. S., & Hsu, P. D. (2022). Systematic discovery of recombinases for efficient integration of large DNA sequences into the human genome. Nature Biotechnology. https: / doi.org / 10.1038 / s41587-022-01494-w

[0013] (비특허문헌0007) Aviv, Tzvi, Zhen Lin, Stefanie Lau, Laura M. Rendl, Frank Sicheri, and Craig A. Smibert. 2003. "The RNA-Binding SAM Domain of Smaug Defines a New Family of Post-Transcriptional Regulators." Nature Structural Biology 10 (8): 614-21.

[0014] (비특허문헌 0008) Green, Justin B., Cary D. Gardner, Robin P. Wharton, and Aneel K. Aggarwal. 2003. "RNA Recognition via the SAM Domain of Smaug." Molecular Cell 11 (6): 1537-48.

[0015] (비특허문헌 0009) She, Richard, Anupam K. Chakravarty, Curtis J. Layton, Lauren M. Chircus, Johan O. L. Andreasson, Nandita Damaraju, Peter L. McMahon, Jason D. Buenrostro, Daniel F. Jarosz, and William J. Greenleaf. 2017. "Comprehensive and Quantitative Mapping of RNA-Protein Interactions across a Transcribed Eukaryotic Genome." Proceedings of the National Academy of Sciences of the United States of America 114 (14): 3619-24.

[0016] (비특허문헌 0010) Aviv, Tzvi, Zhen Lin, Giora Ben-Ari, Craig A. Smibert, and Frank Sicheri. 2006. "Sequence-Specific Recognition of RNA Hairpins by the SAM Domain of Vts1p." Nature Structural & Molecular Biology 13 (2): 168-76.

[0017] (비특허문헌 0011) Hyrina, Anastasia, Christopher Jones, Darlene Chen, Scott Clarkson, Nadire Cochran, Paul Feucht, Gregory Hoffman, et al. 2019. "A Genome-Wide CRISPR Screen Identifies ZCCHC14 as a Host Factor Required for Hepatitis B Surface Antigen Production." Cell Reports 29 (10): 2970-78.e6.

[0018] (Non-patent Document 0012) Wang, Y., Fan, X., Song, Y., Liu, Y., Liu, R., Wu, J., Li, SAMD4 family members suppress human hepatitis B virus by directly binding to the Smaug recognition region of viral RNA. Cellular & Molecular Immunology, 18(4), 1032-1044.

[0019] One aspect of the present invention is to provide a regulatory element that increases gene expression, a screening method therefor, and a use thereof.

[0020] Another aspect is to provide regulatory elements for increasing RNA stability and / or mRNA translation.

[0021] Another aspect is to provide regulatory elements derived from the viral genome or fragments thereof.

[0022] Another aspect is to provide a regulatory element comprising a base sequence of SEQ ID NO: 1 to 71; or a base sequence having at least 80% identity therewith.

[0023] Another aspect is to provide a construct, vector, or recombinant host cell comprising a gene of interest and said regulatory elements.

[0024] Another aspect provides a composition comprising the construct, vector, or recombinant host cell.

[0025] Another aspect provides a method of preparing the construct, vector, recombinant host cell, or composition.

[0026] Another aspect provides a method for increasing RNA stability and / or mRNA translation of a gene of interest, comprising the step of inserting or linking the regulatory element into a UTR of the gene of interest.

[0027] Another aspect provides a use for increasing RNA stability and / or mRNA translation of the construct, vector, recombinant host cell, or composition.

[0028] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for producing an mRNA construct or a protein of interest.

[0029] Another aspect provides a method for preventing, ameliorating or treating a disease comprising administering the construct, vector, recombinant host cell or composition to a subject in need thereof.

[0030] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for the prevention or treatment of a disease.

[0031] Each description and embodiment disclosed in this application may also be applied to each other description and embodiment. That is, all combinations of the various elements disclosed in this application fall within the scope of this application. Furthermore, the scope of this application is not limited by the specific descriptions set forth below. Furthermore, those skilled in the art will recognize or be able to ascertain, through routine experimentation alone, numerous equivalents to the specific embodiments of this application described in this application. Furthermore, such equivalents are intended to be encompassed by this application.

[0032]

[0033] One aspect of the present invention provides a regulatory element capable of increasing RNA stability and / or mRNA translation. The regulatory element, capable of increasing RNA stability and / or mRNA translation, may be suitable for increasing protein production in a construct comprising a gene of interest.

[0034] The regulatory elements of the present application may be derived from a viral genome or a fragment thereof. Specifically, the regulatory elements of the present application may be derived from various regions of the viral genome or a fragment thereof, including, but not limited to, regions such as a 5' UTR, a 3' UTR, a coding region, a non-coding region, an intron, an intergenic region, a packaging signal, an IRES (Internal Ribosome Entry Site), an RNA structural motif, or a fragment thereof.

[0035] The above viral genomes are West Nile virus, dengue virus type I, gallivirus A1, Canine picornavirus, Passerivirus A1, Sicinivirus A, Chicken astrovirus, Infectious bronchitis virus, Turkey calicivirus, Melegrivirus A, Eel picornavirus 1, anativirus A1, cardiovirus C1, Oscivirus A1, Tibetan frog hepatitis B virus, Heron hepatitis B virus, Equin lineitis B virus 1 (Equine rhinitis B virus 1), BtMr-AlphaCoV / SAX2011, Nebraska virus, Porcine sapelovirus 1, Lloviu cuevavirus, Paslahepevirus balayani, Chinook salmon bafinivirus, tremovirus A1, Human gammaherpesvirus 4, Rous sarcoma virus, Marburg marburgvirus, Human respirovirus 3,It may be Sunguru virus, Rabbit hemorrhagic disease virus, Hepatovirus A, Teschovirus A, Eel virus European X, Tupaia virus, Human betaherpesvirus 6B, Human papillomavirus 5, Snakehead retrovirus, Saffold virus, hunnivirus A1, Human betaherpesvirus 5, Atlantic salmon calicivirus, Astrovirus MLB1, or Eel picornavirus 1 strain F15 / 05.

[0036] The regulatory elements of the present application are West Nile virus, Dengue virus type I, Gallivirus A1, Canine picornavirus, Phaserivirus A1, Sisinivirus A, Chicken astrovirus, Infectious bronchitis virus, Turkish calicivirus, Melegrivirus A, Il picornavirus 1, Anathivirus A1, Cardiovirus C1, Ossivirus A1, Tibetan frog hepatitis B virus, Heron Tibetan frog hepatitis B virus, Equin lineitis B virus 1, BtMr-alphacoronavirus / SAX2011, Nebraska virus, Fosain sapelovirus 1, Yobiu cuevavirus, Phaslahepevirus Balayani, Chinook salmon baffinivirus, Tremovirus A1, Human gammaherpesvirus 4, Rhus sarcoma virus, Marburg Marburg virus, Human respirovirus 3, Sungguru virus, Rabbit hemorrhagic A fragment of a gene of a virus, Hepatovirus A, Tescovirus A, Ilvirus European X, Tupaia virus, human betaherpesvirus 6B, human papillomavirus 5, snakehead retrovirus, Sapfold virus, Hoonivirus A1, human betaherpesvirus 5, Atlantic salmon calicivirus, Astrovirus MLB1, or Ilpicornavirus 1 strain F15 / 05; or a sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity thereto.

[0037] The base sequence of the virus used in this application can be obtained from a known database (e.g., NCBI, etc.).

[0038] The regulatory element of the present application may be a fragment consisting of 1 to 200 base sequences derived from the viral genome. Specifically, the regulatory elements are fragments of the viral genome, 200, 199, 198, 197, 196, 195, 194, 193, 192, 191, 190, 189, 188, 187, 186, 185, 184, 183, 182, 181, 180, 179, 178, 177, 176, 175, 174, 173, 172, 171, 170, 169, 168, 167, 166, 165, 164, 163, 162, 161, 160, 159, 158, 157, 156, 155, 154, 153, 152, 151, 150, 149, 148, 147, 146, 145, 144, 143, 142, 141, 140, 139, 138, 137, 136, 135, 134, 133, 132, 131, 130, 129, 128, 127, 126, 125, 124, 123, 122, 121, 120, 119, 118, 117, 116, 115, 114, 113, 112, 111, 110, 109, 108, 107, 106, 105, 104, 103, 102, 101, 100, 99, 98, 97, 96, 95, 94, 93, 92, 91, 90, 89, 88, 87, 86, 85, 84, 83, 82, 81, 80, 79, 78, 77, 76, 75, 74, 73, 72, 71, 70, 69, 68, 67, 66, 65, 64, 63, 62, 61, 60, 59, 58, 57, 56, 55, 54, 53, 52, 51, 50, 49, 48, 47, 46, 45, 44, 43,It may comprise or consist of a fragment consisting of a sequence of 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6 or 5 bases.

[0039] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 1; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0040] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 2; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0041] In one specific example, the regulatory element is a fragment of a West Nile virus gene (NCBI Reference Sequence: NC_009942.1), wherein the fragment comprises the base sequence of SEQ ID NO: 2, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 2.

[0042] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 3; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0043] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 4; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0044] In one specific example, the regulatory element is a fragment of a Dengue virus type I gene (NCBI Reference Sequence: NC_001477.1), wherein the fragment comprises the base sequence of SEQ ID NO: 4, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 4.

[0045] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 5; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0046] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 6; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0047] In one specific example, the regulatory element is a fragment of a Gallivirus A1 gene (NCBI Reference Sequence: NC_018400.1), wherein the fragment comprises the base sequence of SEQ ID NO: 6, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 6.

[0048] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 7; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0049] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 8; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0050] In one specific example, the regulatory element is a fragment of a Canine picornavirus gene (NCBI Reference Sequence: NC_016964.1), wherein the fragment comprises the base sequence of SEQ ID NO: 8, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 8.

[0051] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 9; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0052] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 10; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0053] In one specific example, the regulatory element is a fragment of the Passerivirus A1 gene (NCBI Reference Sequence: NC_014411.1), wherein the fragment comprises the base sequence of SEQ ID NO: 10, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 10.

[0054] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 11; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0055] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 12; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0056] In one specific example, the regulatory element is a fragment of a Sicinivirus A gene (NCBI Reference Sequence: NC_023861.1), wherein the fragment comprises the base sequence of SEQ ID NO: 12, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 12.

[0057] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 13; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0058] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 14; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0059] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 15; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0060] In one specific example, the regulatory element is a fragment of an infectious bronchitis virus gene (NCBI Reference Sequence: NC_001451.1), wherein the fragment comprises the base sequence of SEQ ID NO: 15, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 15.

[0061] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 16; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0062] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 17; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0063] In one specific example, the regulatory element is a fragment of a Turkey calicivirus gene (NCBI Reference Sequence: NC_043516.1), wherein the fragment comprises the base sequence of SEQ ID NO: 17, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 17.

[0064] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 18; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0065] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 19; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0066] In one specific example, the regulatory element is a fragment of a Melegrivirus A gene (NCBI Reference Sequence: NC_023858.1), wherein the fragment comprises the base sequence of SEQ ID NO: 19, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 19.

[0067] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 20; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0068] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 21; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0069] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 22; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0070] In one specific example, the regulatory element is a fragment of a Cardiovirus C1 gene (NCBI Reference Sequence: NC_038305.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 22, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 22.

[0071] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 23; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0072] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 24; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0073] In one specific example, the regulatory element is a fragment of an Oscivirus A1 gene (NCBI Reference Sequence: NC_014412.1), wherein the fragment comprises the base sequence of SEQ ID NO: 24, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 24.

[0074] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 25; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0075] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 26; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0076] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 27; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0077] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 28; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0078] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 29; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0079] In one specific example, the regulatory element is a fragment of a Tibetan frog hepatitis B virus gene (NCBI Reference Sequence: NC_030446.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 29, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 29.

[0080] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 30; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0081] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 31; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0082] In one specific example, the regulatory element is a fragment of a Heron hepatitis B virus gene (NCBI Reference Sequence: NC_001486.1), wherein the fragment comprises the base sequence of SEQ ID NO: 31, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 31.

[0083] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 32; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0084] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 33; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0085] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 34; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0086] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 35; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0087] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 36; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0088] In one specific example, the regulatory element is a fragment of a Cardiovirus C1 gene (NCBI Reference Sequence: NC_038305.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 36, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 36.

[0089] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 37; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0090] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 38; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0091] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 39; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0092] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 40; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0093] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 41; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0094] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 42; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0095] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 43; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0096] In one specific example, the regulatory element is a fragment of a gallivirus A1 gene (NCBI Reference Sequence: NC_018400.1), wherein the fragment comprises the base sequence of SEQ ID NO: 43, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 43.

[0097] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 44; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0098] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 45; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0099] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 46; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0100] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 47; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0101] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 48; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0102] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 49; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0103] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 50; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0104] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 51; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0105] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 52; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0106] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 53; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0107] In one specific example, the regulatory element is a fragment of a Rabbit hemorrhagic disease virus gene (NCBI Reference Sequence: NC_001543.1), wherein the fragment comprises the base sequence of SEQ ID NO: 53, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 53.

[0108] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 54; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0109] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 55; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0110] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 56; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0111] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 57; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0112] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 58; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0113] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 59; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0114] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 60; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0115] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 61; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0116] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 62; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0117] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 63; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0118] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 64; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0119] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 65; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0120] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 66; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0121] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 67; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0122] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 68; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0123] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 69; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0124] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 70; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0125] In one specific example, the regulatory element is a fragment of an Eel picornavirus 1 gene (NCBI Reference Sequence: NC_022332.1), wherein the fragment comprises the nucleotide sequence of SEQ ID NO: 70, or comprises or consists of a nucleotide sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the nucleotide sequence of SEQ ID NO: 70.

[0126] The regulatory element of the present application may comprise or consist of a base sequence of SEQ ID NO: 71; or a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% identity thereto.

[0127] In one specific example, the regulatory element is a fragment of an Eel picornavirus 1 gene (NCBI Reference Sequence: NC_022332.1), wherein the fragment comprises the base sequence of SEQ ID NO: 71, or comprises or consists of a base sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity to the base sequence of SEQ ID NO: 71.

[0128]

[0129] Even if the present application describes a "regulatory element comprising a base sequence of a specific sequence number" or a "regulatory element having a base sequence of a specific sequence number," it is clear that a regulatory element having a base sequence in which some of the sequences are modified may also be used in the present application, as long as it has the same or corresponding function as a regulatory element composed of the base sequence of the corresponding sequence number. The "modification" refers to, but is not limited to, substitution, deletion, and / or insertion of bases.

[0130] For example, if it has the same or corresponding function as the above regulatory element, it is obvious that a regulatory element in which a meaningless sequence is added to or at the end of the regulatory element sequence of the corresponding sequence number, or in which a part of the sequence in or at the end of the regulatory element sequence of the corresponding sequence number is deleted, is also within the scope of the present invention.

[0131] In one specific example, the regulatory element of the present application may include a base sequence in which one or more nucleotides are modified in the base sequence of each sequence number. Specifically, the regulatory elements of the present application are 50 or less, 49 or less, 48 ​​or less, 47 or less, 46 or less, 45 or less, 44 or less, 43 or less, 42 or less, 41 or less, 40 or less, 39 or less, 38 or less, 37 or less, 36 or less, 35 or less, 34 or less, 33 or less, 32 or less, 31 or less, 30 or less, 29 or less, 28 or less, 27 or less, 26 or less, 25 or less, 24 or less, 23 or less, 22 or less, 21 or less, 20 or less, 19 or less, 18 or less, 17 or less, 16 or less, 15 or less, 14 or less, 13 or less It may include a base sequence in which 12 or fewer, 11 or fewer, 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, 2 or fewer, or 1 or fewer nucleotides are modified.

[0132] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1 to 80, 81 to 95, 117 to 119, 131 to 133, 135 to 141, 144 to 155, and 156 to 197 in the base sequence of SEQ ID NO: 5 is modified.

[0133] In one specific example, the substitution may include substitution of at least one base with A, G, C, U, or T among the bases corresponding to positions 1 to 80, 81 to 95, 117 to 119, 131 to 133, 135 to 141, 144 to 155, and 156 to 197 in the base sequence of SEQ ID NO: 5 (substitution with a different base).

[0134] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1 to 15, 37 to 39, 51 to 53, 55 to 61, 64 to 75, 76, and 77 in the base sequence of SEQ ID NO: 6 is modified.

[0135] In one specific example, the substitution may include substitution of at least one base with A, G, C, U, or T among the bases corresponding to positions 1 to 15, 37 to 39, 51 to 53, 55 to 61, 64 to 75, 76, and 77 in the base sequence of SEQ ID NO: 6 (substitution with a different base).

[0136] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1 to 41, 42 to 51, 55 to 58, 64 to 67, 69, 74, 77, 81 to 83, 86 to 97, and 98 to 197 in the base sequence of SEQ ID NO: 30 is modified.

[0137] In one specific example, the substitution may include substitution of at least one base with A, G, C, U, or T among the bases corresponding to positions 1 to 41, 42 to 51, 55 to 58, 64 to 67, 69, 74, 77, 81 to 83, 86 to 97, and 98 to 197 in the base sequence of SEQ ID NO: 30 (substitution with different bases).

[0138] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1, 2 to 11, 15 to 18, 24 to 27, 29, 34, 37, 41 to 43, and 46 to 57 in the base sequence of SEQ ID NO: 31 is modified.

[0139] In one specific example, the substitution may include substitution of at least one base with A, G, C, U, or T among the bases corresponding to positions 1, 2 to 11, 15 to 18, 24 to 27, 29, 34, 37, 41 to 43, and 46 to 57 in the base sequence of SEQ ID NO: 31 (substitution with a different base).

[0140] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1 to 75, 76 to 77, 90, 98, 101 to 102, 129 to 130, 135, 144 to 146, 148, 150, and 151 to 197 in the base sequence of SEQ ID NO: 69 is modified.

[0141] In one specific example, the substitution may include substitution of at least one base with A, G, C, U or T among the bases corresponding to positions 1 to 75, 76 to 77, 90, 98, 101 to 102, 129 to 130, 135, 144 to 146, 148, 150, and 151 to 197 in the base sequence of SEQ ID NO: 69 (substitution with a different base).

[0142] In one specific example, the regulatory element may include a base sequence in which a base corresponding to any one or more positions selected from the group consisting of bases 1 to 4, 5 to 6, 19, 27, 30 to 31, 58 to 59, 64, 73 to 75, 77, and 79 in the base sequence of SEQ ID NO: 71 is modified.

[0143] In one specific example, the substitution may include substitution of at least one base with A, G, C, U, or T among the bases corresponding to positions 1 to 4, 5 to 6, 19, 27, 30 to 31, 58 to 59, 64, 73 to 75, 77, and 79 in the base sequence of SEQ ID NO: 71 (substitution with a different base).

[0144] Homology and identity refer to the degree to which two given base sequences are related, and can be expressed as a percentage. The terms homology and identity are often used interchangeably.

[0145] Whether any two sequences are homologous, similar or identical can be determined using known computer algorithms such as the "FASTA" program using default parameters, for example as in Pearson et al (1988) [Proc. Natl. Acad. Sci. USA 85]: 2444. Alternatively, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needleman program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) (version 5.0.0 or later) can be determined using the GCG program package (Devereux, J., et al, Nucleic Acids Research 12: 387 (1984)), BLASTP, BLASTN, FASTA (Atschul, [S.] [F.,] [ET AL, J MOLEC BIOL 215]: 403 (1990); Guide to Huge Computers, Martin J. Bishop, [ED.,] Academic Press, San Diego, 1994, and [CARILLO ETA / .](1988) SIAM J Applied Math 48: 1073). For example, sequence homology, similarity, or identity can be determined using BLAST or ClustalW of the National Center for Biotechnology Information database.

[0146] In one embodiment, the control element may comprise at least one stem-loop structure.

[0147] Specifically, the regulatory element may comprise a CNGG motif (SEQ ID NO: 72 (CNGG motif): CNGG) within the loop. More specifically, the regulatory element may comprise a CNGG motif within a penta-loop, a tetra-loop, or an internal loop.

[0148] The above N can be independently selected from a nucleotide selected from A, U, T, G and C or a nucleotide analogue thereof.

[0149] Among the regulatory elements of the present application, an element containing a CNGG motif can interact with ZCCHC14, which interacts with TENT4A or TENT4B (hereinafter collectively referred to as TENT4).

[0150] The term 'TENT4' used in this application means TENT4A or TENT4B, and is hereinafter collectively referred to as TENT4.

[0151] Among the regulatory elements of the present application, elements containing a CNGG motif can interact with ZCCHC14, which interacts with TENT4, thereby inducing an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail by mixed tailing, or both.

[0152] Among the regulatory elements of the present application, elements that do not contain a CNGG motif may interact with TENT4 but not with ZCCHC14 and induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail by mixed tailing, or both, or may not interact with TENT4 and induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail by mixed tailing, or both in a TENT4-independent manner.

[0153]

[0154] Another aspect provides a construct comprising a target gene and the regulatory element described above. The regulatory element is as described above.

[0155] The term "construct" as used herein may be understood as a non-naturally occurring DNA or RNA. That is, a construct may be understood as an artificial nucleic acid molecule or a non-natural nucleic acid molecule, and may be designed and / or generated by genetic engineering or chemical synthesis. A construct may include at least one of the above regulatory elements and at least one open reading frame. A construct may be a DNA molecule, an RNA molecule, or a hybrid molecule comprising DNA and RNA portions.

[0156] In one specific example, the construct may be a construct including a UTR of a target gene and the regulatory element. Within the construct, the regulatory element may be inserted into the UTR of the target gene or linked to the UTR of the target gene in the 5' or 3' direction, and the insertion or linking method is not limited. Specifically, the insertion may be, but is not limited to, positioning the regulatory element in the 5' UTR or 3' UTR region of the target gene. The linking may include, but is not limited to, linking the regulatory element directly to the UTR of the target gene or linking the regulatory element and the UTR with an additional base sequence between them.

[0157] The target gene and the regulatory element within the construct may be of heterologous origin. The target gene may be derived from a gene other than the regulatory element and may not be naturally combined.

[0158] In one specific example, the construct may further comprise, but is not limited to, one or more barcode sequences, forward adapter sequences, reverse adapter sequences, poly(A) tail sequences, or combinations thereof.

[0159] In one specific example, the construct may further comprise a promoter sequence, wherein the gene of interest may be operably linked to the promoter sequence, but is not limited thereto. The term "operably linked" as used herein means that the gene sequence is functionally linked to a promoter sequence that initiates and mediates transcription of the gene of interest.

[0160] In one specific example, the construct may comprise a 5' repeat sequence and a 3' repeat sequence of a virus selected from the group consisting of, but not limited to, adeno-associated virus, adenovirus, alphavirus, retrovirus (e.g., gamma retrovirus, and lentivirus), parvovirus, herpesvirus, and SV40.

[0161] In one specific example, the construct may be an mRNA construct. The mRNA construct may further include, but is not limited to, a sequence of a 5' UTR, a 3' UTR, a poly(A) tail, or a combination thereof.

[0162] In the present application, the target gene may be at least one selected from the group consisting of a reporter, a protein, a physiologically active peptide, an antigen, or an antibody or a fragment thereof; or an antisense oligonucleotide, mRNA, dsRNA, shRNA, miRNA, siRNA, gRNA, saRNA, lncRNA, taRNA, ribozyme, ncRNA, exosomal RNA, and an aptamer, but the type is not limited as long as RNA stability and / or mRNA translation can be increased by the regulatory element of the present application.

[0163] In one specific example, the reporter may be, but is not limited to, luciferase, a fluorescent protein, beta-galactosidase, chloramphenicol acetyltransferase, or aequorin.

[0164] In one specific example, the physiologically active polypeptide may be, but is not limited to, a hormone, a cytokine, a cytokine binding protein, an enzyme, a growth factor, or insulin.

[0165] In one specific example, the antigen may be, but is not limited to, a vaccine antigen, a cancer-related antigen, or an allergy antigen.

[0166]

[0167] Another aspect provides a vector comprising the construct or a pool of the vector.

[0168] The term "vector" as used herein refers to a genetic construct containing a base sequence encoding a target protein or a target gene operably linked to suitable regulatory sequences so as to enable expression of the target protein in a suitable host. The regulatory sequences may include, but are not limited to, a promoter capable of initiating transcription, an optional operator sequence for regulating such transcription, a sequence encoding a suitable mRNA ribosome binding site, and sequences regulating the termination of transcription and translation. After being introduced into a suitable host cell, the vector may replicate or function independently of the host genome, or may be integrated into the genome itself.

[0169] The vector used in this application is not particularly limited as long as it is capable of expression within a host cell, and any vector known in the art may be used to introduce the vector into the host cell. Examples of commonly used vectors include plasmids, cosmids, viruses, and bacteriophages, either in their natural or recombinant form.

[0170]

[0171] Another aspect provides a recombinant host cell comprising the construct or vector.

[0172] The term "host cell" as used herein encompasses any cell capable of expressing a target protein, including cells that have undergone natural or artificial genetic modification. Furthermore, the host cell includes both eukaryotic and prokaryotic cells, and may be, but is not limited to, eukaryotic cells or cells derived from mammals (e.g., humans).

[0173] In the present application, the method for introducing a construct or vector into a cell includes any method for introducing a nucleic acid into a cell (e.g., transfection or transformation), and depending on the cell, a suitable standard technique known in the art can be selected and performed. Examples thereof include, but are not limited to, electroporation, calcium phosphate (CaPO4) precipitation, calcium chloride (CaCl2) precipitation, microinjection, polyethylene glycol (PEG) method, DEAE-dextran method, cationic liposome method, lipid nanoparticle method, and lithium acetate-DMSO method.

[0174]

[0175] Another aspect provides a composition comprising the construct, vector, or recombinant host cell. The construct, vector, recombinant host cell, or composition comprising them of the present application can express a protein of interest in vitro, in vivo, or ex vivo.

[0176] In one specific example, when the composition is administered to a subject, the target protein can be provided to the subject via the construct, vector, or recombinant host cell, thereby exhibiting a preventive or therapeutic effect on a disease (e.g., infectious disease) depending on the intended use of the provided target protein. Accordingly, the composition may be, but is not limited to, a pharmaceutical composition.

[0177] In one specific example, the construct, vector, or recombinant host cell may be used to produce the construct or target protein of the present application in vitro, in vivo, or ex vivo. Accordingly, the composition may be a composition for producing the construct or target protein of the present application, but is not limited thereto.

[0178] For example, if the target protein is a vaccine antigen, the construct, vector, recombinant host cell or composition itself can be used as a vaccine, or these can be used to produce a vaccine antigen.

[0179] In one specific example, the composition may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof. Specifically, the construct or vector of the present application may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof, or the recombinant host cell or composition of the present application may further comprise TENT4 or a gene encoding it, ZCCHC14 or a gene encoding it, or a combination thereof. The construct, vector, recombinant host cell or composition may increase RNA stability or mRNA translation by inducing an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both through the interaction of ZCCHC14 and a regulatory element and / or the interaction of ZCCHC14, TENT4 and a regulatory element.

[0180] Another aspect provides a method for preventing, ameliorating or treating a disease comprising administering the construct, vector, recombinant host cell, or composition to a subject in need thereof.

[0181] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for the prevention or treatment of a disease.

[0182]

[0183] Another aspect provides a method for producing a target protein, comprising the steps of culturing the recombinant host cell; and recovering the target protein.

[0184] In the present application, the method for producing a target protein using the recombinant host cell can be performed using a method widely known in the art. Specifically, the culture can be continuously cultured in a batch process, a fed batch process, or a repeated fed batch process, but is not limited thereto. The medium used for culture can be appropriately selected by a person skilled in the art depending on the host cell. Specifically, the recombinant host cell of the present application can be cultured under aerobic or anaerobic conditions while controlling temperature, pH, etc. in a general medium containing an appropriate carbon source, nitrogen source, phosphorus source, inorganic compound, amino acid, and / or vitamin.

[0185] The method for producing the target protein may further include an additional process after the culturing step. The additional process may be appropriately selected depending on the intended use of the target protein.

[0186] Specifically, the method for producing the target protein may include a step of recovering the target protein from at least one material selected from among the recombinant host cell, a dried product of the recombinant host cell, an extract of the recombinant host cell, a culture of the recombinant host cell, a supernatant of the culture, and a lysate of the recombinant host cell after the culturing step.

[0187] The method may further include a step of lysing the combination host cells prior to or concurrently with the recovery step. Lysis of the combination host cells may be performed using methods commonly used in the technical field to which the present application pertains, such as a lysis buffer, a sonicator, heat treatment, and a French press. In addition, the lysis step may include, but is not limited to, an enzymatic reaction such as a cell wall / membrane degrading enzyme, a nuclease, a nucleic acid transferase, and / or a protease.

[0188] In the present application, the dried product of the recombinant host cell can be produced by drying the cell that has accumulated the target substance, but is not limited thereto.

[0189] In the present application, the extract of a recombinant host cell may refer to the material remaining after separating the cell wall / cell membrane from the cell. Specifically, it may refer to the remaining components obtained by lysing the cell, excluding the cell wall / cell membrane. The cell extract includes the target protein, and components other than the target protein may include, but are not limited to, one or more of the following: protein, carbohydrate, nucleic acid, and fiber of the cell.

[0190] In the present application, the recovery step can recover the target protein using a suitable method known in the art (e.g., centrifugation, filtration, anion exchange chromatography, crystallization, HPLC, etc.).

[0191] In the present application, the recovery step may include a purification process. The purification process may separate only the target protein from the cell and purify it. Through the purification process, a pure, purified target protein can be produced.

[0192] Another aspect provides a use of the construct, vector, recombinant host cell, or composition for producing an mRNA construct or a protein of interest.

[0193] Another aspect provides a method for increasing RNA stability and / or mRNA translation of a gene of interest, comprising the step of inserting or linking the regulatory element into a UTR of the gene of interest.

[0194] Another aspect provides uses of the construct, vector, recombinant host cell, or composition for increasing RNA stability and / or mRNA translation.

[0195] Another aspect is a method for producing an mRNA construct, comprising the steps of transcribing the construct or vector in vitro; and recovering the transcribed mRNA construct.

[0196] The above-mentioned transfer method and recovery method can utilize any suitable method known in the art.

[0197] In one specific example, the method may further include, but is not limited to, a step of removing DNA of the construct or vector used as a template by treating with DNase I after transcription; and / or a washing step.

[0198] Regulatory elements according to one specific example can increase RNA stability or mRNA translation, which are transcription products of a target gene, thereby increasing the expression level of a target protein. Therefore, the regulatory elements of the present disclosure can be usefully utilized in systems requiring high levels of gene expression, such as gene therapy, vaccine development, and protein therapeutics production.

[0199] Figure 1 is a diagram of viruses known to infect vertebrates. The total number of species and average genome size for each virus family are indicated by white bars, while the number of species is indicated by shaded bars. To preserve oligonucleotide diversity, the viruses were divided into two libraries: VL.1 contains single-stranded RNA viruses, and VL.2 contains other viruses.

[0200] Figure 2 is a schematic diagram of the vector construction and experimental procedures. 196,277 197-nt segments were selected by tiling at 20-nt intervals. Each oligonucleotide was cloned into the 3' UTR region of the EGFP gene, and the plasmid pool was transfected into HCT116 cells (for RNA abundance and polysome analysis) or HEK293T cells (for total expression fractionation analysis). To quantify RNA stability effects, reporter DNA and RNA were extracted, amplified by PCR, and sequenced. For polysome profiling analysis, five fractions were collected from sucrose density gradient centrifugation, and the reporter RNA was sequenced. For total expression fractionation analysis, a HEK293T cell library integrating the EGFP reporter into the genome was sorted, and the top and bottom cell populations were enriched and sequenced.

[0201] Figure 3 is a graph ranking the RNA abundance scores of VL.1 (left) and VL.2 (right). The RNA abundance score is calculated as the log2 of the ratio of RNA reads to DNA reads. Positive controls (K5, K4), negative controls (K5m, K4m), self-cleaving ribozymes of Hepatitis delta virus, and viral miRNAs of Epstein-Barr virus (EBV) and Human Cytomegalovirus (HCMV) are shown.

[0202] Figure 4 is a diagram showing the results of total expression sorting analysis of VL.1 (left) and VL.2 (right). The top and bottom enrichment log2 values ​​were compared for each library. Positive controls (K5, K4) and negative controls (K5m, K4m, HDV) are indicated.

[0203] Figure 5 is a diagram showing the reporter analysis results for the control group used in this analysis. Results are included for K5, K5m, K4, and K4m (130 nt).

[0204] Figure 6 is a graph showing the Spearman correlation coefficient for the RNA abundance MPRA results.

[0205] Figure 7 is a diagram showing the RNA distribution of cluster 1 (top) and the enrichment pattern (top enrichment: Tpo, bottom enrichment) of the control group (K5, K4, HDV) in the total expression sorting library.

[0206] Figure 8 shows the distribution of BFP in Pa01 clone cells (left) and a graph verifying a single site of integration (right). After co-transfection of EGFP reporter plasmids and mCherry reporter plasmids at a 1:1 ratio, approximately 2.5% of cells should be double-positive if multiple integration sites are present (0.19 x 0.13 = 0.025).

[0207] Figure 9 shows the results of flow cytometry analysis showing GFP expression measured after cell sorting and subsequent cell proliferation.

[0208] Figure 10 is a diagram showing the enrichment results for the integrase library, comparing the T (top) and B (bottom) enrichment values ​​in replicates.

[0209] Figure 11 is a graph comparing the effects on RNA abundance (x-axis) and translation (y-axis). T-enriched and B-enriched elements are indicated in purple (◆ symbol).

[0210] Figure 12 shows a comparison of the effects on RNA abundance (x-axis) and translation (y-axis) in HCT116 cells (left) and a comparison of replicates of integrase library enrichment in HEK293T cells (right). Eighteen elements of VL.1 and two elements of VL.2 that were shown to increase expression in both HCT116 and HEK293T are highlighted in purple (◀ symbol). These elements were individually cloned into a dual-luciferase reporter.

[0211] Figure 13 is a graph showing the validation results for the 19 selected elements. The VL.1-derived elements begin with Q, and the VL.2-derived elements begin with V. The control group, a reporter without inserted Q or V elements, was used for normalization. Data are expressed as the mean ± standard error of the mean (SEM) (n = 3), and statistical analysis was performed using a two-sided Student's t-test, with significance indicated as *p < 0.05.

[0212] Figure 14 is a diagram depicting the overall genomic pattern of the library data, expressed by RNA / DNA ratio (HCT116), free mRNA to heavy polysome ratio (HP / Free mRNA ratio, HCT116), T-cell enrichment (T-enrichment, HEK293T), and B-cell enrichment (B-enrichment, HEK293T). For the HEK293T data, only elements with an FDR < 0.1 are displayed. Relative genomic positions were calculated only for single-genome viruses.

[0213] Figure 15 is a diagram showing the results of RT-qPCR analysis for the Q and V reporters presented in Figure 13.

[0214] Figure 16 is a diagram showing the luciferase activity of in vitro transcribed firefly mRNA constructs in the presence or absence of elements. Each value was normalized to co-transduced renilla mRNA, and the control reporter activity at 0 h was set to 1.

[0215] Figure 17 is a diagram comparing the effects on RNA abundance (x-axis) and translation (y-axis) in HCT116 cells (top), and the results of measuring luciferase activity by individually cloning 12 (VL. 1) elements and 6 (VL. 2) elements that were shown to increase expression in HCT116 cells among VL. 1 and VL. 2-derived elements not shown in Figure 12 into a dual luciferase reporter (bottom).

[0216] Figure 18 is a diagram showing the results of comparing the results of integrase library enrichment in HEK293T cells among replicates (top), and the results of individually cloning nine (VL. 1) elements and four (VL. 2) elements that were shown to have increased expression in HEK293T cells among VL. 1 and VL. 2-derived elements not shown in Figure 12 into a dual luciferase reporter and measuring luciferase activity (bottom).

[0217] In Figures 17 and 18, the control group was a reporter without Q or V elements, which was used for normalization. Data are expressed as the mean ± standard error of the mean (SEM) (n = 3, biological replicates). *p < 0.05, **p < 0.01 indicate significance by a two-sided Student's t test.

[0218] Figure 19 is a graph showing luciferase activity measured from the top element at Day 2 under treatment with RG7834 or its inactive R-isomer, RO0321 (150 nM). The luciferase activity of the control plasmid was used for normalization, and data are expressed as the mean ± standard error of the mean (SEM) (n = 3, biological replicates). Statistical significance was analyzed using a two-sided Student's t-test, and is indicated as *p < 0.05, **p < 0.01.

[0219] Figure 20 is a schematic diagram of the TENT4-dependent screening. Additional sorting was performed on the genome-integrated, top-enriched cell library to eliminate false positives from the initial FACS sorting. The genome-integrated cell library, which underwent two rounds of enrichment, was treated with RG7834 or RO0321, then sorted into three bins and sequenced. Relative protein expression levels were calculated by weighting the ratios of each bin.

[0220] Figure 21 is a diagram showing the results of the TENT4-dependent screening, with the results for Library 1 presented, and spike-normalized reads expressed as a percentage within each fraction. Positive controls K5 (red, bold line) and K4 (orange, dashed line) are also shown.

[0221] Figure 22 is a graph comparing relative protein expression levels calculated from sequencing data in the presence of RO0321 (control) or RG7834 (TENT4 inhibitor), as a result of the TENT4-dependent screening. Candidates with an FDR < 0.1 and higher protein expression levels in the control group than in the TENT4 inhibitor-treated group are highlighted in cyan (◀ symbol). Elements containing the CNGG motif are highlighted with a border (● symbol).

[0222] Figure 23 is a diagram showing the proportion of TENT4-dependent tiles related to Figure 22 and the proportion of tiles containing CNGG motifs within each library candidate.

[0223] Figure 24 is a diagram showing the validation results for elements confirmed as positive among elements containing the CNGG motif stem. Luciferase activity was measured on Day 2 after treating HCT116 parental and ZCCHC14 knockout (KO) cells with RO0321 or RG7834. Luciferase activity was normalized to the control plasmid, and data are expressed as the mean ± standard error of the mean (SEM) (n = 3, biological replicates). Statistical significance was analyzed using a two-tailed Student's t-test, and is indicated as *p < 0.05, **p < 0.01.

[0224] Figure 25 is a diagram showing the validation results for newly cloned elements among elements containing the CNGG motif stem. Luciferase activity was measured on Day 2 after treatment of HCT116 parental and ZCCHC14 knockout (KO) cells with RO0321 or RG7834. Luciferase activity was normalized to the control plasmid, and data are expressed as the mean ± standard error of the mean (SEM) (n = 3, biological replicates). Statistical significance was analyzed using a two-tailed Student's t-test, and *p < 0.05, **p < 0.01.

[0225] Figure 26 is a diagram showing luciferase activity measured under the same conditions as Figure 24 for an element lacking the CNGG motif to evaluate ZCCHC14 dependence.

[0226] Figure 27 is an individual plot showing the fractional ratio of each candidate group (only the verified elements) that does not contain CNGG stem-loops, with the dotted lines representing each replicate and the solid lines representing the average values.

[0227] Figure 28 shows the luciferase activity measured at Day 2 in HeLa parental cells and ZCCHC2 knockout (KO) cells treated with RO0321 or RG7834, targeting elements lacking the CNGG motif, to assess ZCCHC2 dependence. The luciferase activity of the control plasmid was used for normalization, and the data are expressed as the mean ± standard error of the mean (SEM) (n = 3, biological replicates). Statistical significance was analyzed using a two-sided Student's t-test, and *p < 0.05, **p < 0.01.

[0228] Figure 29 is a diagram evaluating the TENT4 cofactor dependence of non-CNGG stem-loop-containing candidates in ZCCHC2 knockout (ZCCHC2 KO) cells using a RO0321 / RG7834 luciferase reporter assay. Luciferase assays were performed in wild-type HeLa cells and ZCCHC2 knockout cells treated with RO0321 or RG7834.

[0229] Figure 30 is a diagram illustrating the dual recruitment of TENT4, where the CNGG stem-loop is located adjacent to K4 or Q7, which are non-CNGG TENT4-dependent elements but do not contain CNGG. m1 and m2 represent mutations in each region, and each mutation was introduced into a construct encompassing the entire range and analyzed in the same manner as in Figure 24.

[0230] Figure 31 is a schematic diagram illustrating the stabilization mechanism of regulatory elements that recruit TENT4 using different adaptor proteins and induce mixed tailing.

[0231] Figure 32 is a schematic diagram of the MPRA mutagenesis library, in which a library containing substitutions, deletions, and paired mutations for elements Q4, Q17, Q20, K4, V6, and V8 was designed. The element pool was integrated into the genome of HEK293T cells, sorted into four bins, and sequenced.

[0232] Figure 33 is a graph showing the flow cytometry distribution of the library measured together with the negative control (noEL, eK5m) and the wild-type of each element. Gating bins are indicated by the dotted lines on the vertical axis.

[0233] Figure 34 is a graph showing the correlation between replicates for the predicted expression level of mutant MPRA, visualized for data with a sequencing count of 40 or more.

[0234] Figure 35 is a diagram showing luciferase activity measured from truncated reporters. The minimal boundaries were established based on sequencing data from adjacent positive elements, and reporters with further narrowed boundaries, truncated Left (trL) or Truncated Right (trR), were also constructed and evaluated using luciferase assays.

[0235] Figure 36 is a diagram showing the Poly(A) length distribution measured by Hire-PAT (Hierarchical poly(A) tail) analysis. Normalized intensity (arbitrary unit (au)) is expressed as a percentile of reads, and the same criterion was applied to all subsequent Hire-PAT analyses. HCT116 cells were transfected with control, K5 reporter, or its mutant (K5m) plasmids and elements, and then analyzed.

[0236] Figures 37 and 38 are diagrams showing the results of the secondary screening, in which the expression level of each element was measured using substitution mutants. The horizontal dotted line indicates the wild-type expression level of each element, and the minimal elements are highlighted in red boxes (shaded boxes).

[0237] Figure 39 is a diagram illustrating the dual recruitment of TENT4, where the V6 element contains two key stem-loop structures. Mutations m1 and m2, respectively, represent a CNG"G→C"N substitution within the corresponding loop motif. Luciferase assays were performed in the presence of RO0321 or RG7834.

[0238] Figure 40 is a diagram summarizing the mutation results in the V6 element structure, showing the core elements and the average expression value of each nucleotide expressed in purple (shading). In addition, the bases corresponding to Δexpression (expression amount of paired nucleotides - expression amount of unpaired nucleotides) are indicated by line thickness at the connection between base pairs.

[0239] Figure 41 is a diagram showing luciferase activity measured from truncated reporters. The minimal boundaries were established based on sequencing data from adjacent positive elements, and reporters with further narrowed boundaries, truncated Left (trL) or Truncated Right (trR), were also constructed and evaluated using luciferase assays.

[0240] Figures 42 and 43 are diagrams showing the results of the secondary screening, in which the expression level of each element was measured using substitution mutants. The horizontal dotted line indicates the wild-type expression level of each element, and the minimal elements are highlighted in red boxes (shaded boxes).

[0241] Figure 44 is a diagram showing the results of luciferase expression measured after treating HCT116 cells with siNC or siSAMD4A / B.

[0242] Figure 45 is a diagram showing luciferase activity measured from truncated reporters. The minimal boundaries were established based on sequencing data from adjacent positive elements, and reporters with further narrowed boundaries, truncated Left (trL) or Truncated Right (trR), were also constructed and evaluated using luciferase assays.

[0243] Figure 46 is a diagram showing the distribution of Poly(A) lengths measured by Hire-PAT (Hierarchical poly(A) tail) analysis. Normalized intensity (arbitrary unit (au)) is expressed as a percentile of reads, and the same criteria were applied to all subsequent Hire-PAT analyses. HCT116 cells were transfected with control, K5 reporter, or its mutant (K5m) plasmids and elements, and then analyzed.

[0244] Figures 47 and 48 are diagrams showing the results of the secondary screening, in which the expression level of each element was measured using substitution mutants. The horizontal dotted line indicates the wild-type expression level of each element, and the minimal elements are highlighted in red boxes (shaded boxes).

[0245] Figure 49 is a diagram summarizing the mutation results in the Q4 element structure, showing the core elements and the average expression value of each nucleotide expressed in purple (shading). In addition, the bases corresponding to Δexpression (expression amount of paired nucleotides - expression amount of unpaired nucleotides) are indicated by line thickness at the connection between base pairs.

[0246] Figure 50 is a diagram showing luciferase activity measured from truncated reporters. The minimal boundaries were established based on sequencing data from adjacent positive elements. Truncated Left (trL) or Truncated Right (trR) reporters, which further narrowed these boundaries, were also constructed and evaluated using luciferase assays.

[0247] Figure 51 is a diagram showing the distribution of Poly(A) lengths measured by Hire-PAT (Hierarchical poly(A) tail) analysis. Normalized intensity (arbitrary unit (au)) is expressed as a percentile of reads, and the same criteria were applied to all subsequent Hire-PAT analyses. HCT116 cells were transfected with control, K5 reporter, or its mutant (K5m) plasmids and elements, and then analyzed.

[0248] Figures 52 to 55 illustrate the results of the secondary screening, in which the expression levels of each element were measured using substitution mutants. The horizontal dotted line indicates the wild-type expression level of each element, and the minimal elements are highlighted in red boxes (shaded boxes).

[0249] Figure 56 is a diagram summarizing the mutation results in the Q17 element structure, showing the core elements and the average expression value of each nucleotide expressed in purple (shading). In addition, the bases corresponding to Δexpression (expression amount of paired nucleotides - expression amount of unpaired nucleotides) are indicated by line thickness at the connection between base pairs.

[0250] Figure 57 is a diagram showing luciferase activity measured from truncated reporters. The minimal boundaries were established based on sequencing data from adjacent positive elements, and reporters with further narrowed boundaries, truncated Left (trL) or Truncated Right (trR), were also constructed and evaluated using luciferase assays.

[0251] Figures 58 to 61 are diagrams showing the results of the secondary screening, in which the expression level of each element was measured using substitution mutants. The horizontal dotted line indicates the wild-type expression level of each element, and the minimal elements are highlighted in red boxes (shaded boxes).

[0252] Figure 62 is a graph showing the results of confirming the gene expression increase activity of Q20min _79nt.

[0253] Figure 63 is a graph showing the results of confirming the gene expression increase activity of Q6min.

[0254] Figure 64 is a graph showing the results of confirming the gene expression increase activity of Q7min.

[0255] Figure 65 is a diagram showing increased luciferase activity in HCT116 cells using an in vitro transcription (IVT) construct containing Q4 and Q17.

[0256] Figure 66 is a diagram showing increased luciferase activity in HeLa and HCT116 cells using an IVT construct containing Q20.

[0257] Figure 67 is a drawing showing the results of increased luciferase activity of elements in HeLa cells using a RaPID construct containing Q3, Q5, Q11, Q13, Q14, Q18 and Q19.

[0258] Hereinafter, preferred embodiments are presented to aid understanding of the present invention. However, the following embodiments are provided solely to facilitate a better understanding of the present invention and are not intended to limit the scope of the present invention. The embodiments are susceptible to various modifications, and thus the embodiments are not limited to the embodiments disclosed below and may be implemented in various forms.

[0259] Terms or words used in the specification and claims of the present invention are not to be construed as limited to their usual or dictionary meanings, and should be interpreted as meanings and concepts that conform to the technical idea of ​​the present invention based on the principle that the inventor can appropriately define the concept of the term to explain his or her own invention in the best way.

[0260] Throughout the specification of the present invention, when a part is said to "include" a certain component, this does not mean that other components are excluded, but rather that other components may be included, unless specifically stated otherwise.

[0261] Throughout the specification of the present invention, “A and / or B” means A or B, or A and B.

[0262]

[0263] Example 1. Analysis method

[0264] Analysis of the experimental results related to the following examples was performed in the following manner.

[0265]

[0266] 1.1 Data Analysis

[0267] For all samples, sequenced reads were aligned to each oligonucleotide sequence using bowtie2.2.652 with the -local parameter. Aligned reads were filtered according to the following criteria:

[0268] - Perfect match ≥ 145;

[0269] - Insertion or deletion ≤ 1.

[0270] For polysome analysis, a variance-stabilizing transformation was performed using DESeq2, and then the mean value of the five fractions was subtracted from each fraction's value to calculate the relative distance between them. The calculated relative distances were used for hierarchical clustering analysis using the scipy module.

[0271] For TENT4 screening and mutation experiments, expression levels for each fraction were calculated as a weighted sum. Weights were derived using the relative median FITC value measured during the sorting process. Statistical tests were performed using limma.

[0272]

[0273] 1.2 Cutoff for element selection

[0274] Screening criteria for HCT116 cell line:

[0275] - Count_HP (z-score) > 0.2;

[0276] - Log₂(HP / Free) > 0.3 and FDR < 0.01;

[0277] - RNA log₂FC > 0.4 and FDR < 0.01.

[0278] Screening criteria for HEK293T cell line:

[0279] - FDR(T) < 0.01;

[0280] - FDR(TT) < 0.05

[0281] - Fold T > 50

[0282]

[0283] 1.3 Mutation screening analysis

[0284] Δexpression is defined as follows:

[0285] - The value obtained by subtracting the average expression level of unpaired nucleotides (excluding GU pairs) from the average expression level of paired nucleotides (AU / GC / CG / UA).

[0286]

[0287]

[0288] Example 2. MPRA screening (Massively Parallel Reporter Assay screening)

[0289] In this study, two comprehensive libraries, VL.1 and VL.2, were constructed, each containing 158 and 179 vertebrate viruses, respectively. The viruses in these libraries were selected from 37 classified and unclassified virus families known to infect humans, with a focus on representative viruses and pathogenic pathogens within each genus (Fig. 1). A total of 2,290 viruses belonging to 37 families were collected from the NCBI database. Excluding some pathogenic viruses, one virus per genus was selected to ensure complete coverage of the entire database without sequence redundancy. As a result, this catalog encompasses all seven groups of the Baltimore classification system and all virus genera infecting vertebrates. For RNA viruses, complete genome sequences were used, and for DNA viruses, non-coding RNA, introns, and predicted 3' UTR regions were tiled. As a result, 99,922 oligonucleotides were generated for the VL.1 library and 96,438 oligonucleotides for the VL.2 library, which were achieved using a sliding window of 197 nt and a step size of 20 nt (Fig. 2).

[0290] And as positive controls, K4 elements and K5 elements derived from Saffold virus and Aichi virus, and their corresponding mutations (K4m: mutation that disrupts the first pair of stem-loop, K5m) were included, and as negative controls, a ribozyme sequence derived from hepatitis B virus was included (Fig. 5) (Non-patent literature 0005 (Seo et al., 2023)).

[0291] The K4 element (SEQ ID NO: 73) from Saffold virus and the K5 element (SEQ ID NO: 74) from Aichi virus are regulatory elements that were previously shown to have RNA stability and / or mRNA translation enhancing activities in a previous study [Seo, JJ et al. Functional viromic screens uncover regulatory RNA elements. Cell 186, 3291-3306.e21 (2023)] (Non-patent literature 0005). K4m and K5m are inactive mutants of K4 and K5.

[0292] The synthesized oligonucleotide was amplified by PCR and then inserted into the 3'UTR of the EGFP gene.

[0293] In the first plasmid pool, oligonucleotides were raised and sequences were cloned into the 3' UTR of an EGFP reporter with the mPGK promoter, allowing direct expression after transfection of the plasmid pool into a human colon cancer cell line (HCT116). Forty-eight hours after transfection, RNA abundance and translation efficiency were calculated through polysome profiling to quantify the effect of each element on gene expression (Fig. 2). The RNA abundance score for each segment was calculated as the log2 ratio of the read fraction of DNA to RNA, and revealed a distinct regulatory effect between viral elements with high reproducibility (Figs. 3 and 6). The K5 and K4 elements showed higher RNA abundance compared to their corresponding mutants (K5m and K4m). In addition, tiles containing the negative control (HDV ribozyme) and some miRNA regions showed reduced stability (Fig. 3).

[0294] Furthermore, the translational effects of these segments were also evaluated. A total of 49 (VL. 1) and 46 (VL. 2) mRNA clusters were identified, which were classified based on the read ratio between heavy polysomes (H) and free mRNA, and showed differences in polysome binding behavior. A smaller cluster number indicates a higher proportion of heavy polysomes (H), indicating a group with relatively high translational activity. K5 and K4 were classified as Cluster 1 in both libraries (VL. 1 and VL. 2), which is consistent with the results that they positively influence translation (Fig. 7).

[0295] In another plasmid pool, oligonucleotides were cloned into the 3'UTR of an EGFP reporter containing a recombinase site for integrase, without a promoter within the plasmid itself. Pa01 integrase is known to have the highest integration rate among various homologs (Non-patent literature 0006 (Durrant et al., 2022)). Furthermore, a stable cell line was generated from a HEK293T cell line containing Pa01 integrase and BFP behind the EF1a promoter and a single recombination site inserted (Figure 8). This allows for the insertion of a single copy of the EGFP reporter into stabilized cells containing a single recombination site behind the EF1a promoter, and GFP fluorescence of the inserted reporter represents total gene expression within the cell. The GFP distribution in the inserted cell library was similar to that of a single cell line, and most elements were confirmed to have no effect on overall gene expression (Fig. 9). Based on this, only a small portion of the entire library was selected and enriched. High-expressing cells were classified into 0-0.5% (T) and 0.5-1.5% (mT), while low-expressing cells were classified into 0-1% (B) and 1-2% (mB) (Fig. 9). Elements from each cell group were amplified and sequenced.

[0296] A total of 270 and 118 positive expression elements were evaluated for each library (FDR<0.01, Figs. 4, 10), and the positive controls (K5 and K4) were confirmed to be enriched according to the classification method compared to the non-functional controls (K5m and K4m). In contrast, the negative controls were confirmed to be enriched in the B fraction (Figs. 4, 7, 10).

[0297]

[0298] Example 3. Verification of positive elements

[0299] The distributions of RNA / DNA ratios and HP / Free mRNA ratios of elements enriched in the T fraction suggest correlations between various methods for quantifying gene expression (Fig. 11). Furthermore, most elements located at the 3' end of the genome exhibited activity that promoted gene expression (Fig. 14).

[0300] For initial verification, a total of 47 segments were selected, with elements derived from VL.1 denoted as 'Q' and elements derived from VL.2 denoted as 'V'.

[0301] The selected elements are:

[0302] Q1 (SEQ ID NO: 1), Q2 (SEQ ID NO: 3), Q4 (SEQ ID NO: 5), Q5 (SEQ ID NO: 7), Q6 (SEQ ID NO: 9), Q7 (SEQ ID NO: 11), Q8 (SEQ ID NO: 13), Q11 (SEQ ID NO: 14), Q12 (SEQ ID NO: 16), Q17 (SEQ ID NO: 18), Q20 (SEQ ID NO: 69), Q21 (SEQ ID NO: 20), Q22 (SEQ ID NO: 21), Q23 (SEQ ID NO: 23), Q24 (SEQ ID NO: 25), Q26 (SEQ ID NO: 26), Q27 (SEQ ID NO: 27), V6 (SEQ ID NO: 28), V8 (SEQ ID NO: 30), Q3 (SEQ ID NO: 32), Q10 (SEQ ID NO: 33), Q13 (SEQ ID NO: 34), Q14 (SEQ ID NO: 35), Q18 (SEQ ID NO: 37), Q19 (SEQ ID NO: 38), Q28 (SEQ ID NO: 39), Q29 (SEQ ID NO: 40), Q30 (SEQ ID NO: 41), Q31 (SEQ ID NO: 42), Q32 (SEQ ID NO: 44), Q33 (SEQ ID NO: 45), V3 (SEQ ID NO: 46), V7 (SEQ ID NO: 47), V9 (SEQ ID NO: 48), Q9 (SEQ ID NO: 49), Q15 (SEQ ID NO: 50), Q16 (SEQ ID NO: 51), Q25 (SEQ ID NO: 52), Q34 (SEQ ID NO: 54), Q35 (SEQ ID NO: 55), Q36 (SEQ ID NO: 56), Q37 (SEQ ID NO: 57), Q38 (SEQ ID NO: 58), V1 (SEQ ID NO: 59), V2 (SEQ ID NO: 60), V4 (SEQ ID NO: 61), V5 (SEQ ID NO: 62).

[0303] The selected elements are those that were positive in the HCT116 and / or HEK293T data (Fig. 12, Fig. 17, Fig. 18).

[0304] As a result of verification using a luciferase reporter cloned into the 3'UTR of the corresponding element, it was confirmed that luciferase expression was significantly increased (Fig. 13).

[0305] Additionally, RNA expression levels were measured via qRT-PCR for the selected Q and V reporters, and RNA stability analysis using in vitro transcribed mRNA was performed, and it was confirmed that some elements regulate total gene expression through RNA stabilization (Fig. 15, Fig. 16).

[0306]

[0307] Example 4. Analysis of TENT4 dependence of viral elements

[0308] As a first step in exploring the TENT4 dependence of elements, we confirmed the TENT4 dependence of elements that had already been shown to be positive. To this end, we selected five key candidates (Q4, Q17, Q20, V6, and V8) and treated cells with the TENT4 inhibitor RG7834 or its R-isomer, RO0321. All five elements showed a decrease in luciferase expression following RG7834 treatment, highlighting the widespread and potent impact of TENT4-mediated mixed tailing (RNA activation) (Fig. 19).

[0309] To explore the TENT4 dependence of elements, we performed another enrichment process on the initially selected cells (the T fraction). Cells were then sorted into three fractions (L, M, and R) for each RO0321- or RG7834-treated cell library, and each fraction was sequenced (Fig. 20). To quantify the distribution of each element, cells containing the spike element were added to each fraction for normalization purposes after fractionation. As a result, most elements were found to form peaks in the M fraction under both RO0321 and RG7834 conditions. However, the TENT4-dependent controls (K5 and K4) exhibited a left-shifted distribution in the RG7834 sample (Fig. 21).

[0310] Predicted protein levels were calculated by weighting the ratios in each fraction and normalized to the calculated value of the control group. In conclusion, 53 and 32 tiles showing TENT4 dependence were identified in libraries 1 and 2, respectively (Figs. 22 and 23). Some of these tiles overlapped with each other, corresponding to 16 and 7 clusters, respectively, for a total of 23 candidate regions. Notably, 53 of the 85 candidates contained a CNGG motif, which is known to be recognized by ZCCHC14, a cofactor of TENT4.

[0311] ZCCHC14, one of the cofactors of TENT4, has a SAM domain, which is responsible for the CNGG(N) of mRNA. 0-3It is known to bind to stem-loops (Non-patent literature 0007 (Aviv et al. 2003)); (Non-patent literature 0008 (Green et al. 2003)); (Non-patent literature 0009 (She et al. 2017)). In addition, viral RNAs such as 1E and PRE are affected by ZCCHC14 and contain CNGG stem-loops (Non-patent literature 0010 (Aviv et al. 2006)); (Non-patent literature 0011 (Hyrina et al. 2019)).

[0312] Accordingly, elements including the CNGG stem-loop were verified to have both ZCCHC14 and TENT4 dependencies (Fig. 24, Fig. 25).

[0313] Meanwhile, candidate elements that did not contain a CNGG stem-loop also existed. When these elements were cloned as reporters, luciferase activity was reduced in HCT116 cells treated with RG7834 (Figs. 26, 27). To confirm their dependence on the TENT4 cofactor, the same analysis was performed using HCT116 ZCCHC14 knockout (KO) cells and HeLa ZCCHC2 KO cells. As a result, the RG7834 dependence of these elements remained unchanged, indicating that they were not dependent on ZCCHC2 or ZCCHC14 (Figs. 26, 28, 29).

[0314] Notably, the viral RNA contained CNGG stem-loop elements located near the TENT4-dependent regulatory region. Specifically, the K4 and Q7 regions contained adjacent CNGG motifs. To explore the functional role of these elements, we cloned a construct containing both regions and generated mutations at these sites to assess their contribution to RNA function (Fig. 30).

[0315] As a result of the TENT4 dependency analysis of the elements according to the present invention, it was confirmed that Q8, Q12, Q24, Q31, Q35, V4, V6, V8, CG1, CG2, CG3, CG4, CG5, and CG6 contain a CNGG stem-loop sequence and interact with TENT4 via ZCCHC14 to exhibit RNA stability-increasing activity, and Q4, Q6, Q7, Q17, and Q22 do not contain a CNGG stem-loop sequence but interact with TENT4 to exhibit RNA stability-increasing activity. Meanwhile, it was confirmed that the other Q and V elements exhibit RNA stability-increasing activity in a TENT4-independent manner (Fig. 31).

[0316]

[0317] Example 5. Mutation screening for positive elements

[0318] Mutation screening was performed to identify key nucleotides and core regions essential for the function of the positive element. This secondary screening was performed on the K4, Q4, Q17, Q20, V6, and V8 mutants (Fig. 32). For each mutant, single nucleotide substitutions and single or two consecutive nucleotide deletions were introduced. In addition, compensatory mutations that alter the base sequence while preserving the predicted double-helical structure were also introduced. A total of approximately 7,000 mutants were synthesized and cloned into the Pa01 MPRA vector.

[0319] After cloning and introduction into HEK293T cells, the cells were sorted into four fractions and each fraction was sequenced to calculate the total expression level (Fig. 33). The total expression level was calculated as the weighted sum of the abundance within each fraction (Fig. 34). Secondary screening results for K4, Q4, Q17, V6, and V8 confirmed that specific nucleotide substitutions, deletions, and paired mutations significantly affected the regulatory activity of the corresponding elements (Figs. 40, 49, 56, 37, 38, 42, 43, 47, 48, 52-55, 58-61). The mutations that affected the expression of these elements were located within the truncated elements, which represent the core regions of each element (indicated by red boxes (shaded boxes), Figs. 37, 38, 42, 43, 47, 48, 52 to 55, 58 to 61).

[0320] Truncated reporters were used to identify the minimal effective sequence (MES) for the V6, V8, Q4, Q17, and Q20 elements. The MES was defined by overlapping positive elements within the library. Luciferase activity was reduced in forms with the left 20 nt or right 20 nt truncated, thereby deriving the minimum range for each element (Figs. 35, 41, 45, 50, and 57).

[0321] Specifically, for V6, the core sequence (V6min, SEQ ID NO: 29) with minimal activity was identified as the 41st to 137th base sequence of V6 (SEQ ID NO: 28) (2766th to 2862nd base sequence of Tibetan frog hepatitis B virus) (Fig. 35, Fig. 37, Fig. 38, Fig. 40).

[0322] For V8, the core sequence (V8min, SEQ ID NO: 31) with minimal activity was identified as the 41st to 97th base sequence of V8 (SEQ ID NO: 30) (base sequence 2180th to 2236th of Heron hepatitis B virus) (Fig. 41, Fig. 42, Fig. 43). In addition, for the V8 element, it was confirmed that the CNGG motif was in the internal loop based on mutagenesis.

[0323] Below is the derived sequence and secondary structure of V8min. The CNGG motif is underlined.

[0324] Sequence: [UUAACACAUGGCGCAAUAUCCCAUAUCACCGGCGGGAGCGCAGUGUUUACCUUUUCA].

[0325] structure: [..(((((...((((......((........)).....)))).)))))..........].

[0326] The above structure is expressed in secondary structure notation (dot-bracket notation), where "." indicates an unpaired base (loop, bulge, or single-stranded region), "(" indicates left-hand pairing of the bases forming the stem, and ")" indicates right-hand pairing of the bases forming the stem.

[0327] For Q4, the core sequence (Q4min, SEQ ID NO: 6) with minimal activity was identified as the 81st to 157th base sequence of Q4 (SEQ ID NO: 5) (8172nd to 8248th base sequence of gallivirus A1) (Fig. 45, Fig. 47, Fig. 48, Fig. 49).

[0328] For Q17, the core sequence (Q17min, SEQ ID NO: 19) with minimal activity was identified as the 96th to 102nd base sequence of Q17 (SEQ ID NO: 18) (9602nd to 9697th base sequence of Melegrivirus A) (Fig. 50, Fig. 52 to Fig. 55, Fig. 56).

[0329] In the case of Q20, activity was confirmed in the 7th to 157th base sequence of Q20 (SEQ ID NO: 69) (7307th to 7457th base sequence of Eel picornavirus 1) (Q20min, SEQ ID NO: 70) (Fig. 57, Fig. 58 to Fig. 61).

[0330] In addition, for Q20, the core sequence (Q20min_79nt, SEQ ID NO: 71) with minimal activity was identified as the 72nd to 150th base sequence of Q20 (SEQ ID NO: 69) (base sequence 7372nd to 7450th of Eel picornavirus 1). Specifically, 150,000 HCT116 cells were transfected with a dual-luciferase reporter plasmid (100 ng) in which each element (Ctrl, Q20, Q20min_79nt) was inserted into the 3'UTR of Firefly, and the effect of each element on gene expression was measured at 48 hours. As a result, it was confirmed that the effect of increasing gene expression was maintained in Q20min_79nt, and the effect was further enhanced compared to Q20 (Fig. 62).

[0331] In addition, core sequences showing minimal activity were identified in elements Q1, Q2, Q5, Q6, Q7, Q11, Q12, Q22, Q23, Q14, Q31, and Q25. These elements were selected as representatives of adjacent elements whose activity was detected during the tiling process, and the corresponding core sequences correspond to the overlapping sections with adjacent elements. In addition, the same analysis as above was performed on this overlapping section, and it was confirmed that it had activity. Specifically, after transfecting 150,000 HCT116 cells with a dual-luciferase reporter plasmid (100 ng) in which each element was inserted into the 3'UTR of Firefly, the effect of each element on gene expression was measured at 48 hours. As a result, it was confirmed that the effect of increasing gene expression was maintained even in the core sequence (min) of each element (Fig. 63, Fig. 64).

[0332] Specifically, for Q1, the core sequence (Q1min, SEQ ID NO: 2) with minimum activity was identified as the overlapping region of Q1 (SEQ ID NO: 1), Q1' (SEQ ID NO: 75), and Q1'' (SEQ ID NO: 76), and was identified as the 21st to 197th base sequence of Q1 (10741st to 10917th base sequence of West Nile virus).

[0333] For Q2, the core sequence (Q2min, SEQ ID NO: 4) with minimum activity is the overlapping region of Q2 (SEQ ID NO: 3), Q2' (SEQ ID NO: 77), Q2'' (SEQ ID NO: 78), and Q2''' (SEQ ID NO: 79), and was identified as the 99th to 177th base sequence of Q2 (10539th to 10617th base sequence of dengue virus type I).

[0334] For Q5, the core sequence (Q5min, SEQ ID NO: 8) with minimal activity was identified as the overlapping region of Q5 (SEQ ID NO: 7), Q5' (SEQ ID NO: 80), and Q5'' (SEQ ID NO: 81), and was identified as the 41st to 197th base sequence of Q5 (7721st to 7877th base sequence of Canine picornavirus).

[0335] For Q6, the core sequence (Q6min, SEQ ID NO: 10) with minimum activity is the overlapping region of Q6 (SEQ ID NO: 9), Q6' (SEQ ID NO: 82), Q6'' (SEQ ID NO: 83), Q6''' (SEQ ID NO: 84), Q6'''' (SEQ ID NO: 85), and Q6''''' (SEQ ID NO: 86), and was confirmed to be the 21st to 97th base sequence of Q6 (7707th to 7777th base sequence of Passerivirus A1).

[0336] In the case of Q7, the core sequence (Q7min, SEQ ID NO: 12) with minimum activity is the overlapping region of Q7 (SEQ ID NO: 11) and Q7' (SEQ ID NO: 87), and among the overlapping regions, the 7th to 85th base sequence of Q7 (8967th to 9045th base sequence of Sicinivirus A) was identified.

[0337] In the case of Q11, the core sequence (Q11min, SEQ ID NO: 15) with minimum activity is the overlapping region of Q11 (SEQ ID NO: 14), Q11' (SEQ ID NO: 88), Q11'' (SEQ ID NO: 89), Q11''' (SEQ ID NO: 90), and Q11'''' (SEQ ID NO: 91), and was identified as the 32nd to 157th base sequence of Q11 (27412th to 27537th base sequence of Infectious bronchitis virus).

[0338] For Q12, the core sequence (Q12min, SEQ ID NO: 17) with minimum activity was identified as the overlapping region of Q12 (SEQ ID NO: 16), Q12' (SEQ ID NO: 92), and Q12'' (SEQ ID NO: 93), and as the 1st to 97th base sequence of Q12 (7201st to 7297th base sequence of Turkey calicivirus).

[0339] For Q22, the core sequence (Q22min, SEQ ID NO: 22) with minimum activity is the overlapping region of Q22 (SEQ ID NO: 21), Q22' (SEQ ID NO: 94), Q22'' (SEQ ID NO: 95), Q22''' (SEQ ID NO: 96), and Q22'''' (SEQ ID NO: 97), and was identified as the 73rd to 127th base sequence of Q22 (8387th to 8441st base sequence of cardiovirus C1).

[0340] For Q23, the core sequence (Q23min, SEQ ID NO: 24) with minimal activity was identified as the overlapping region of Q23 (SEQ ID NO: 23), Q23' (SEQ ID NO: 98), and Q23'' (SEQ ID NO: 99), and was identified as the 16th to 177th base sequence of Q23 (7346th to 7597th base sequence of Oscivirus A1).

[0341] For Q14, the core sequence (Q14min, SEQ ID NO: 36) with minimal activity was identified as the overlapping region of Q14 (SEQ ID NO: 35) and Q14' (SEQ ID NO: 100), and as the 1st to 177th base sequence of Q14 (641st to 817th base sequence of cardiovirus C1).

[0342] In the case of Q31, the core sequence (Q31min, SEQ ID NO: 43) with minimal activity was identified as the overlapping region of Q31 (SEQ ID NO: 42) and Q31' (SEQ ID NO: 101), and as the 1st to 157th base sequence of Q31 (base sequence 6652 to 6808 of gallivirus A1).

[0343] In the case of Q25, the core sequence (Q25min, SEQ ID NO: 53) with minimum activity is the overlapping region of Q25 (SEQ ID NO: 52) and Q25' (SEQ ID NO: 102), and was identified as the 81st to 197th base sequence of Q25 (7181st to 7297th base sequence of Rabbit hemorrhagic disease virus).

[0344] Consistent with the TENT4-dependent screening results, the TENT4-dependent V6, V8, Q4, and Q17 elements in HCT116 cells were found to have extended poly(A) tail lengths (Fig. 36, Fig. 46, Fig. 51).

[0345] The V6 and V8 elements are located in a similar position to the PRE, but are not homologous, and like the PRE, contain a CNGG motif. The V8 element does not rely on SAMD4, which is known to possess a SAM domain and recruit deadenylase (Figure 44) (Non-patent literature 0012 (Wang et al., 2021)).

[0346] For Q4 and Q17, the first stem-loop region for Q4 and the region from the single-stranded region to the second stem-loop for Q17 were important for element function, respectively (Figs. 45 to 49, 50 to 56).

[0347] For Q20 and K4, it was confirmed that most of the elements located along the stem as well as the loop were important for element function (Figs. 57 to 61).

[0348]

[0349] Example 6. Gene expression enhancement effect of synthetic mRNA by positive elements

[0350] To verify the practical applicability of these positive elements, the elements were inserted into synthetic mRNA and their effects on gene expression were evaluated. In vitro transcription (IVT) constructs containing Q4, Q17, Q20, and multiple Q elements significantly increased luciferase activity in HCT116 and HeLa cells compared to the control group (Figures 65-67).

[0351]

[0352] Through the above examples, this study expanded the application of the MPRA technique to diverse vertebrate virus families, thereby discovering novel transcriptional regulatory RNA elements. Furthermore, by utilizing vertebrate-derived viral oligonucleotides, it was suggested that regulatory functional elements that may function commonly across different viral strains could be identified.

[0353] The results presented in this invention can serve as a valuable resource for future functional studies and contribute to a deeper understanding of viral RNA regulation. The approach presented herein can also be applied to other viral groups, potentially leading to the discovery of additional regulatory elements and expanding our knowledge of virus-host interactions.

[0354] Future studies will explore the functional impact of the elements identified here on viral replication and pathogenicity, potentially suggesting novel therapeutic strategies targeting viral RNA elements.

[0355]

[0356] The sequence of the elements described in the present invention is as shown in Table 1 below. In Table 1 below, the sequence directions 'P' and 'N' represent the positive strand direction and the negative strand direction, respectively.

[0357]

[0358] 종류유래Accession numberSTARTEND서열방향길이서열 번호Q1West Nile virusNC_009942.11072110917P1971Q1minWest Nile virusNC_009942.11074110917P1772Q2dengue virus type INC_001477.11044110637P1973Q2mindengue virus type INC_001477.11053910617P794Q4gallivirus A1NC_018400.180928288P1975Q4mingallivirus A1NC_018400.181728248P776Q5Canine picornavirusNC_016964.176817877P1977Q5minCanine picornavirusNC_016964.177217877P1578Q6Passerivirus A1NC_014411.176817877P1979Q6minPasserivirus A1NC_014411.177017777P7710Q7Sicinivirus ANC_023861.189619157P19711Q7minSicinivirus ANC_023861.189679045P7912Q8Chicken astrovirusNC_003790.136813877P19713Q11Infectious bronchitis virusNC_001451.12738127577P19714Q11minInfectious bronchitis virusNC_001451.12741227537P12615Q12Turkey calicivirusNC_043516.172017397P197QTurkey16min calicivirusNC_043516.172017297P9717Q17Melegrivirus ANC_023858.195019697P19718Q17_minMelegrivirus ANC_023858.196029697P9619Q21anativirus A1NC_006553.180218217P19720Q22cardiovirus C1NC_038305.183158511P19721Q22mincardiovirus C1NC_038305.183878441P5522Q23Oscivirus A1NC_014412.174217617P19723Q23minOscivirus A1NC_014412.174367597P16224Q24Sicinivirus ANC_023861.188419037P19725Q26Melegrivirus ANC_023858.193819577P19726Q27Canine picornavirusNC_016964.160216217P19727V6Tibetan frog hepatitis B virusNC_030446.127262922P19728V6minTibetan frog hepatitis B virusNC_030446.127662862VP97298BHeron hepatitis virusNC_001486.121402336P19730V8minHeron hepatitis B virusNC_001486.121802236P5731Q3Equine rhinitis B virus 1NC_003983.186308826P19732Q10BtMr-AlphaCoV / SAX2011NC_028811.11720117397P19733Q13Nebraska virusNC_004064.172447440P19734Q14 C1NC_038305.1641837P19735Q14mincardiovirus C1NC_038305.1641817P17736Q18Porcine sapelovirus 1NC_003987.138814077P19737Q19Lloviu cuevavirusNC_016144.1989110087N19738Q28Paslahepevirus balayaniNC_001434.169807176P19739Q29Infectious bronchitis virusNC_001451.12062120817P19740ChinoCmon salmon bafinivirusNC_026812.179618157P19741Q31gallivirus A1NC_018400.166526848P19742Q31mingallivirus A1NC_018400.166526808P15743Q32Canine picornavirusNC_016964.131013297P19744Q33tremovirus A1NC_003990.167816977P19745V3Human gammaherpesvirus 4NC_007605.15015850354P19746V7Rous sarcoma virusNC_001407.187018897P19747V9Human gammaherpesvirus 4NC_007605.15745557651N19748Q9Marburg marburgvirusNC_001608.344214617P19749Q15Human respirovirus 3NC_001796.264416637P19750Q16Sunguru virusNC_025401.116411837P19751Q25Rabbit hemorrhagic disease virusNC_001543.171017297P19752Q25minRabbit hemorrhagic disease virusNC_001543.171817297P11753Q34Hepatovirus ANC_001489.1561757P19754Q35Teschovirus ANC_003985.168217017P19755Q36Eel virus European XNC_022581.132013397P19756Q37Sunguru virusNC_025401.140414237P19757Q38Tupaia virusNC_007020.160616257P19758V1Human betaherpesvirus 6BNC_000898.17152171717N19759V2Human gammaherpesvirus 4NC_007605.1146128146324P19760V4Human papillomavirus 5NC_001531.140734269P19761V5Snakehead retrovirusNC_001724.131613357P19762CG1Saffold virusNC_009448.278017997P19763CG2hunnivirus A1NC_018668.171817377P19764CG3Human betaherpesvirus 5NC_006273.2159353159549N19765CG4Atlantic salmon calicivirusNC_024031.126012797P19766CG5Astrovirus MLB1NC_011400.123612557P19767CG6Human betaherpesvirus 5NC_006273.2158673158869N19768Q20Eel picornavirus 1NC_022332.173017497P19769Q20minEel picornavirus 1NC_022332.173077457P15170Q20min_79ntEel picornavirus 1NC_022332.173727450P7971.

[0359]

[0360] From the above description, those skilled in the art will understand that the present invention can be implemented in other specific forms without altering its technical concept or essential characteristics. In this regard, it should be understood that the experimental examples and embodiments described above are illustrative in all respects and not restrictive. The scope of the present invention should be interpreted as encompassing all changes or modifications derived from the meaning and scope of the following claims and their equivalent concepts, rather than the detailed description above.

[0361]

[0362] Attach electronic files

Claims

1. A regulatory element comprising any one of the base sequences of SEQ ID NOs: 71, 6, 29, 31, 2, 4, 8, 10, 12, 15, 17, 22, 24, 36, 53, 13, 20, 25 to 27, 32 to 34, 37 to 41, 44 to 51, 54 to 62, 65 to 68, or a base sequence having at least 80% identity therewith.

2. In claim 1, (i) a fragment of the Eel picornavirus 1 gene (NCBI Reference Sequence: NC_022332.1), wherein the fragment comprises the base sequence of SEQ ID NO: 71; or a base sequence having at least 80% identity therewith; (ii) a fragment of the Gallivirus A1 gene (NCBI Reference Sequence: NC_018400.1), wherein the fragment comprises the base sequence of SEQ ID NO: 6; or a base sequence having at least 80% identity therewith; (iii) a fragment of a Tibetan frog hepatitis B virus gene (NCBI Reference Sequence: NC_030446.1), wherein the fragment comprises the base sequence of SEQ ID NO: 29; or a base sequence having at least 80% identity therewith; (iv) A fragment of a Heron hepatitis B virus gene (NCBI Reference Sequence: NC_001486.1), wherein the fragment comprises the base sequence of SEQ ID NO: 31; or a base sequence having at least 80% identity therewith; (v) a fragment of a West Nile virus gene (NCBI Reference Sequence: NC_009942.1), said fragment comprising the base sequence of SEQ ID NO: 2; or a base sequence having at least 80% identity therewith; (vi) a fragment of a Dengue virus type I gene (NCBI Reference Sequence: NC_001477.1), wherein the fragment comprises the base sequence of SEQ ID NO: 4; or a base sequence having at least 80% identity therewith; (vii) A fragment of a Canine picornavirus gene (NCBI Reference Sequence: NC_016964.1), wherein the fragment comprises the base sequence of SEQ ID NO: 8; or a base sequence having at least 80% identity therewith; (viii) A fragment of the Passerivirus A1 gene (NCBI Reference Sequence: NC_014411.1), wherein the fragment comprises the base sequence of SEQ ID NO: 10; or a base sequence having at least 80% identity therewith; (ix) A fragment of the Sicinivirus A gene (NCBI Reference Sequence: NC_023861.1), wherein the fragment comprises the base sequence of SEQ ID NO: 12; or a base sequence having at least 80% identity therewith; (x) A fragment of an infectious bronchitis virus gene (NCBI Reference Sequence: NC_001451.1), wherein the fragment comprises the base sequence of SEQ ID NO: 15; or a base sequence having at least 80% identity therewith; (xi) A fragment of a Turkey calicivirus gene (NCBI Reference Sequence: NC_043516.1), wherein the fragment comprises the base sequence of SEQ ID NO: 17; or a base sequence having at least 80% identity therewith; (xii) A fragment of the Cardiovirus C1 gene (NCBI Reference Sequence: NC_038305.1), wherein the fragment comprises the base sequence of SEQ ID NO: 22; or a base sequence having at least 80% identity therewith; (xiii) A fragment of the Oscivirus A1 gene (NCBI Reference Sequence: NC_014412.1), wherein the fragment comprises the base sequence of SEQ ID NO: 24; or a base sequence having at least 80% identity therewith; (xiv) a fragment of the Cardiovirus C1 gene (NCBI Reference Sequence: NC_038305.1), wherein the fragment comprises the base sequence of SEQ ID NO: 36; or a base sequence having at least 80% identity thereto; or (xv) A fragment of a Rabbit hemorrhagic disease virus gene (NCBI Reference Sequence: NC_001543.1), wherein the fragment comprises the base sequence of SEQ ID NO: 53; or a regulatory element comprising a base sequence having at least 80% identity thereto.

3. A regulatory element according to claim 1, wherein the base sequence comprises at least one base sequence selected from the group consisting of SEQ ID NOs: 1 to 17, 20 to 41, 44 to 62, and 65 to 71.

4. In claim 1, the control element is a control element that induces an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both.

5. In claim 1, the regulatory element is a regulatory element for increasing RNA stability or mRNA translation.

6. A construct comprising a target gene; and a regulatory element according to any one of claims 1 to 5.

7. A construct according to claim 6, wherein the target gene is at least one selected from the group consisting of a reporter, a protein, a physiologically active peptide, an antigen, or an antibody or a fragment thereof; or an antisense oligonucleotide, mRNA, dsRNA, shRNA, miRNA, siRNA, gRNA, saRNA, lncRNA, taRNA, ribozyme, ncRNA, exosomal RNA, and an aptamer.

8. In claim 6, the construct is an mRNA construct.

9. A vector comprising the construct of claim 6.

10. A recombinant host cell comprising the construct of claim 6 or a vector comprising the construct.

11. A composition comprising the construct of claim 6; a vector comprising the construct; or a recombinant host cell comprising the construct or vector.

12. A composition according to claim 11, wherein the composition is for preventing or treating a disease; or for producing an mRNA construct or a protein encoded by a target gene.

13. A method for increasing RNA stability or mRNA translation of a target gene, comprising the step of inserting or linking a regulatory element of any one of claims 1 to 5 into a UTR of the target gene.

Citation Information

Patent Citations

  • Releasing device for cover plate and electronic device including the same

    KR1020250066993A