Novel regulatory elements for enhancing RNA stability or mRNA translation, ZCCHC2 that interacts therewith, and uses thereof
A method for screening regulatory elements using virus sequence data enhances RNA stability and mRNA translation, leveraging ZCCHC2 interactions to increase target protein expression and address the challenge of functional annotation in virus sequence data.
Patent Information
- Application Number
- JP2024576351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-29
- Filing Date
- 2023-06-29
- Publication Date
- 2025-07-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The rapidly increasing number of virus sequences detected through metagenomics studies poses challenges for functional annotation, as most sequences lack clinical or industrial relevance, necessitating effective strategies for interpreting virus sequence data.
Development of a method to screen for regulatory elements that enhance RNA stability and mRNA translation using virus sequence data, utilizing ZCCHC2 to interact with these elements, and creating constructs, vectors, and recombinant host cells to increase target protein expression.
The method identifies novel regulatory elements that enhance RNA stability and mRNA translation, increasing target protein expression and providing tools for disease prevention and treatment.
Smart Images

Figure 2025520771000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to novel regulatory elements for enhancing RNA stability or mRNA translation, ZCCHC2 that interacts therewith, and uses thereof.
Background Art
[0002] Viruses have evolved various mechanisms to hijack the cellular gene expression machinery, and research in this field has greatly contributed to the development of RNA biology and biotechnology. For example, the 7-methylguanosine cap, internal ribosome entry sites, and RNA triple helices were first discovered in reovirus, poliovirus, and Kaposi's sarcoma-associated herpesvirus. Human immunodeficiency virus (HIV) is known to utilize TAR (transactivation response region) and RRE (rev-response element) to recruit cellular factors for viral transcription and RNA nuclear export, respectively (Vaishnav, et. al., New Biol., 1991, 3, 142-150; Dingwall, et. al., EMBO J., 1990, 9, 4145-4153). Hepatitis B virus (HBV) depends on PRE (posttranscriptional regulatory element) to elicit host nucleotidyltransferases that stabilize the viral transcriptome (Kim, et. al., Nat. Struct. Mol. Biol., 2020, 27, 581-588; Huang, et. al., Mol. Cell. Biol., 1993, 13, 7476-7486).
[0003] However, such discoveries have been made through analyses with low throughput for pathogenic viruses that represent only a tiny fraction of the entire virus. To date, 6,828 viruses have been named, and the NCBI Genome Database has 14,775 complete virus genome sequences (O’Leary, et. al., Nucleic Acids Res., 2016, 44, D733-D745). In recent years, metagenomics studies based on deep sequencing have detected hundreds of thousands of additional virus sequences from environmental and animal samples (Neri, et. al., Cell, 2022, 185, 4023-4037). Although the number of sequences available in this way is huge, sequences without clinical or industrial relevance remain mostly undeveloped. Thus, the rapidly increasing virus sequence data presents important challenges for functional annotation, and more effective strategies for interpreting virus sequence data are needed.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The inventors have developed a method for screening regulatory elements for enhancing RNA stability or mRNA translation using virus sequence data, discovered novel regulatory elements according to the method, and developed the use of ZCCHC2 that interacts with them.
Means for Solving the Problems
[0005] One object of the present invention is to provide a regulatory element for enhancing RNA stability and / or mRNA translation.
[0006] Another object of the present invention is to provide a construct, vector, or recombinant host cell containing the regulatory element in the gene of a target protein and its 3’untranslated region.
[0007] Another object of the present invention is to provide a composition comprising the construct, vector, or recombinant host cell.
[0008] Another object of the present invention is to provide a composition comprising ZCCHC2 that interacts with the regulatory element or a gene encoding the same.
[0009] Yet another object of the present invention is to provide a method for producing a target protein, comprising culturing the recombinant host cell and separating the target protein.
[0010] A further object of the present invention is to provide a method for producing an mRNA construct, comprising transcribing the construct in vitro using the construct or vector as a template and recovering the transcribed mRNA construct.
[0011] Still another object of the present invention is to provide the use of the construct, vector, recombinant host cell, or composition for enhancing RNA stability and / or mRNA translation.
[0012] Yet another object of the present invention is to provide the use of the construct, vector, recombinant host cell, or composition for preventing or treating diseases.
[0013] Still another object of the present invention is to provide the use of the construct, vector, recombinant host cell, or composition for producing an mRNA construct or a target protein.
Advantages of the Invention
[0014] The novel regulatory element for enhancing RNA stability or mRNA translation of the present invention and ZCCHC2 that interacts therewith can increase the expression of a target protein, and thus can be applied to various fields according to the use of the target protein.
Brief Description of the Drawings
[0015]
Figure 1A
Figure 1B
Figure 1C
Figure 1D
Figure 1E
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 3D
Figure 3E
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0016] Each description and embodiment disclosed in the present invention may be applicable to each other description and embodiment. That is, all combinations of various elements disclosed in the present invention belong to the scope of the present invention. Also, the scope of the present invention is not limited by the following specific description. Also, those having ordinary knowledge in the art can recognize or confirm many equivalents to the specific embodiments of the present invention described herein using only ordinary experiments. Also, such equivalents are intended to be included in the present invention.
[0017] Another aspect of the present invention is a regulatory element for enhancing RNA stability and / or mRNA translation. The regulatory element can enhance RNA stability and / or mRNA translation to increase protein expression. Specifically, the regulatory element of the present invention can include (i) the nucleotide sequence of a segment of the Aichi virus 1 gene (NCBI Reference Sequence: NC_001918.1) or the nucleotide sequence of its RNA; or (ii) a nucleotide sequence having at least 90% identity thereto; or (iii) a nucleotide sequence within the 3' untranslated region of a kobuvirus having at least 50% homology with the nucleotide sequence of (i).
[0018] The segment in the present invention can include, but is not limited to, more than 110 and less than or equal to 250 nucleotides continuously in the 5' direction from nucleotide 8251 within the Aichi virus 1 gene.
[0019] As one embodiment, the segment can include, but is not limited to, 120 to 240 (for example, 120, 130, 185, 240) nucleotides continuously in the 5' direction from nucleotide 8251 within the Aichi virus 1 gene.
[0020] The nucleotide sequence of (i) can include, but is not limited to, the nucleotide sequence of SEQ ID NO: 20, 94 or 95 or the nucleotide sequence of its RNA.
[0021] The nucleotide sequence of (ii) can include, but is not limited to, a nucleotide sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with the nucleotide sequence of (i). Specifically, the nucleotide sequence of (ii) can include, but is not limited to, a nucleotide sequence having one or more nucleotide substitutions, deletions, or both, in nucleotides 1 to 14 of the nucleotide sequence of SEQ ID NO: 20, or a nucleotide sequence of its RNA.
[0022] The nucleotide sequence of (iii) is a nucleotide sequence within the 3' untranslated region of Kobuvirus, and can include a nucleotide sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% homology with the nucleotide sequence of (i). Specifically, the nucleotide sequence of (iii) can include, but is not limited to, a nucleotide sequence having at least two hairpin structures within the 3' untranslated region of Kobuvirus. Also, the nucleotide sequence of (iii) can include, but is not limited to, the nucleotide sequence of any one of SEQ ID NOs: 98 to 140 or the nucleotide sequence of its RNA, or a nucleotide sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity therewith.
[0023] The nucleotide sequence of the virus used in the present invention is obtained from a known database (such as NCBI, etc.).
[0024] On the other hand, in the present invention, even if it is stated as "a regulatory element containing the nucleotide sequence of a specific SEQ ID NO." or "a regulatory element having the nucleotide sequence of a specific SEQ ID NO.", it is obvious that a regulatory element having a nucleotide sequence in which several sequences are deleted, modified, substituted or added can also be used in the present invention as long as it has the same or equivalent function as the regulatory element composed of the nucleotide sequence of the said SEQ ID NO.
[0025] For example, if it has the same or equivalent function as the regulatory element, it is obvious that a regulatory element in which a meaningless sequence is added to the inside or end of the sequence of the regulatory element of the sequence number, or a regulatory element in which a part of the sequence inside or end of the sequence of the regulatory element of the sequence number is deleted also belongs to the scope of the present invention.
[0026] Homology and identity mean the degree of relatedness to two given base sequences and can be expressed as a percentage. The terms homology and identity can often be used interchangeably.
[0027] Whether any two arrays have homology, similarity or identity can be determined using known computer algorithms such as the "FASTA" program with default parameters as in Pearson et al. (1988) [Proc. Natl. Acad. Sci. USA 85]: 2444. Or, it can be determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as executed by the needleman program of the EMBOSS package (EMBOSS: The European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16: 276-277) (version 5.0.0 or later versions) (including the GCG program package (Devereux, J., et al, Nucleic Acids Research 12: 387 (1984)), BLASTP, BLASTN, FASTA (Atschul, [S.] [F.,] [ET AL, J MOLEC BIOL 215]: 403 (1990); Guide to Huge Computers, Martin J. Bishop, [ED.,] Academic Press, San Diego, 1994, and [CARILLO ETA / .] (1988) SIAM J Applied Math 48: 1073). For example, the homology, similarity or identity of sequences can be determined using BLAST and ClustalW in the database of the National Center for Biotechnology Information.
[0028] The regulatory element of the present invention can interact with ZCCHC2 that interacts with TENT4 to induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail by mixed tailing, or both.
[0029] Another aspect of the present invention is a construct comprising the regulatory element of the present invention in the gene of the target protein and its 3'-untranslated region. Specifically, the construct can be a DNA construct or an RNA construct.
[0030] The target protein in the present invention is not particularly limited as long as the RNA stability and / or the translation of mRNA can be enhanced by the regulatory element of the present invention, and can be selected from a reporter, a bioactive peptide, an antigen, or an antibody or a fragment thereof.
[0031] The reporter in the present invention can be, but is not limited to, luciferase, a fluorescent protein, β-galactosidase, chloramphenicol acetyltransferase, or aequorin.
[0032] The bioactive polypeptide in the present invention is, but is not limited to, a hormone, a cytokine, a cytokine-binding protein, an enzyme, a growth factor, or insulin.
[0033] In the present invention, the antigen can be, but is not limited to, a vaccine antigen, a cancer-related antigen, or an allergy antigen.
[0034] In one embodiment, the construct of the present invention can further include at least one barcode sequence, a forward adapter sequence, a reverse adapter sequence, a poly(A) tail sequence, or a combination thereof, but is not limited thereto.
[0035] In one embodiment, the construct of the present invention may further include a promoter sequence. In this case, the target protein may or may not be operably linked to the promoter sequence, but is not limited thereto.
[0036] As one embodiment, the construct of the present invention can include, but is not limited to, 5' and 3' inverted repeats of a virus selected from the group consisting of adeno-associated virus, adenovirus, alphavirus, retrovirus (e.g., gammaretrovirus and lentivirus), parvovirus, herpesvirus, and SV40.
[0037] As one embodiment, the mRNA construct of the present invention can further include, but is not limited to, a 5' untranslated region, a 3' untranslated region, a poly(A) tail sequence, or a combination thereof.
[0038] Another aspect of the present invention is a vector containing the construct or a pool of the vectors.
[0039] The term "vector" as used in the present invention means a gene product containing a nucleotide sequence encoding a target protein operably linked to appropriate regulatory sequences so as to be able to express the target protein in a suitable host. The regulatory sequences can include, but are not limited to, a promoter capable of initiating transcription, any operator sequence for regulating such transcription, a sequence encoding an appropriate mRNA ribosome binding site, and sequences for regulating the termination of transcription and translation. After being introduced into an appropriate host cell, the vector can replicate or function independently of the host genome or can be integrated into the genome itself.
[0040] The vector used in the present invention is not particularly limited as long as it can be expressed in a host cell, and any vector known in the art can be used to introduce it into a host cell. Examples of commonly used vectors include plasmids, cosmids, viruses, and bacteriophages in their natural or recombinant states.
[0041] Furthermore, the term "operatively linked" means that the promoter sequence that initiates and mediates the transcription of the target protein gene is functionally linked to the gene sequence.
[0042] Another aspect of the present invention is a recombinant host cell containing the construct or vector.
[0043] The host cells in the present invention include all cells capable of expressing the target protein, including cells that have been genetically modified naturally or artificially. The host cells also include eukaryotic cells and prokaryotic cells, specifically, they may be eukaryotic cells or cells derived from mammals (e.g., humans), but are not limited thereto.
[0044] In the present invention, the method of introducing a construct or vector into a cell includes any method of introducing nucleic acid into a cell (e.g., transfection or transformation), and appropriate standard techniques known in the art can be selected and carried out according to the cells. For example, electroporation, precipitation of calcium phosphate (CaPO4), precipitation of calcium chloride (CaCl2), microinjection, polyethylene glycol (PEG) method, DEAE-dextran method, cationic liposome method, and lithium acetate-DMSO method, etc. can be mentioned, but are not limited thereto.
[0045] Another aspect of the present invention is a composition containing the construct, vector, or recombinant host cell. The construct, vector, recombinant host cell of the present invention, or a composition containing them can express the target protein in vitro, in vivo, or ex vivo.
[0046] As one embodiment, when the composition is administered to an individual, the target protein can be provided to the individual by the construct, vector, or recombinant host cell. Therefore, depending on the use of the provided target protein, it can exhibit a preventive or therapeutic effect on diseases (e.g., infectious diseases). Thus, the composition may be a pharmaceutical composition, but is not limited thereto.
[0047] Also, as one embodiment, the RNA construct or target protein of the present invention can be produced in vitro or ex vivo using the construct, vector, or recombinant host cell. Therefore, the composition may be a composition for producing the mRNA construct or target protein of the present invention, but is not limited thereto.
[0048] For example, when the target protein is a vaccine antigen, the construct, vector, recombinant host cell, or the composition itself can be used as a vaccine, or a vaccine antigen can be prepared using these.
[0049] As one embodiment, the construct or vector of the present invention further comprises a gene encoding ZCCHC2, a gene encoding TENT4, or a combination thereof, or the recombinant host cell or composition of the present invention further comprises ZCCHC2 or a gene encoding the same, TENT4 or a gene encoding the same, or a combination thereof, and through the interaction between ZCCHC2 and the regulatory element and the interaction between ZCCHC2 and TENT4, it can induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both, thereby enhancing RNA stability or mRNA translation, but is not limited thereto.
[0050] Another aspect of the present invention is a composition comprising ZCCHC2 or a gene encoding the same that interacts with the regulatory element. Specifically, the composition may be a composition for enhancing RNA stability or mRNA translation.
[0051] Through interaction with the regulatory element and TENT4, the ZCCHC2 can induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail by mixed tailing, or both, thereby enhancing RNA stability or mRNA translation. Therefore, the composition can increase the expression of the target protein of the present invention in vitro, in vivo, or ex vivo.
[0052] In one embodiment, for the expression of the target protein, the composition further comprises the construct, vector, and / or recombinant host cell of the present invention, or ZCCHC2 or the gene encoding it may be included in the construct, vector, and / or recombinant host cell of the present invention.
[0053] In one embodiment, the composition can further comprise TENT4 or the gene encoding it.
[0054] In one embodiment, depending on the use of the target protein whose in vivo expression is increased by the composition, the composition can exhibit a disease prevention or treatment effect. Therefore, the composition may be, but is not limited to, a pharmaceutical composition.
[0055] Also, in one embodiment, the RNA construct or target protein of the present invention can be produced in vitro or ex vivo using the composition. Therefore, the composition may be, but is not limited to, a composition for producing the RNA construct or target protein of the present invention.
[0056] For example, when the target protein is a vaccine antigen, the composition can increase the expression of the vaccine antigen in vivo, so the composition can be used as a vaccine composition, or the vaccine antigen can be prepared in vitro or ex vivo using the composition.
[0057] Another aspect of the present invention is a method for producing a target protein, comprising culturing the recombinant host cell and recovering the target protein.
[0058] The method for producing a target protein using the recombinant host cell in the present invention can be carried out using methods well known in the art. Specifically, the culturing can be continuously carried out in a batch process, a fed batch or a repeated fed batch process (but not limited thereto). The medium used for culturing can be appropriately selected by those skilled in the art according to the host cell. Specifically, the recombinant host cell of the present invention can be cultured in a normal medium containing an appropriate carbon source, nitrogen source, phosphorus source, inorganic compound, amino acid and / or vitamin, etc., under aerobic or anaerobic conditions while adjusting the temperature, pH, etc.
[0059] The method for producing the target protein can further include a further step after the culturing step. The further step can be appropriately selected according to the use of the target protein.
[0060] Specifically, the method for producing the target protein can include a step of recovering the target protein from one or more substances selected from the recombinant host cell, the dried product of the recombinant host cell, the extract of the recombinant host cell, the culture of the recombinant host cell, the supernatant of the culture, and the disrupted product of the recombinant host cell after the culturing step.
[0061] The method can further include a step of lysing the recombinant host cell before or simultaneously with the recovery step. The lysis of the recombinant host cell can be carried out by methods commonly used in the technical field to which the present invention pertains, such as a lysis buffer, a sonicator, heat treatment, and a French press. In addition, the lysis step can include, but is not limited to, enzymatic reactions such as cell wall / membrane degrading enzymes, nuclease, nucleic acid transferase, and / or protease.
[0062] In the present invention, the dried product of the recombinant host cell can be produced by drying the cells that have accumulated the target substance, but is not limited thereto.
[0063] In the present invention, the extract of the recombinant host cell can mean the material remaining after separating the cell wall / cell membrane in the cells. More specifically, it can mean the remaining components excluding the cell wall / cell membrane in the components obtained by lysing the cells. The cell extract contains the target protein, and as components other than the target protein, at least one component of cell proteins, carbohydrates, nucleic acids, and fibers may be included, but is not limited thereto.
[0064] In the recovery step in the present invention, the target protein can be recovered using suitable methods known in the art (for example, centrifugation, filtration, anion exchange chromatography, crystallization, HPLC, etc.).
[0065] In the present invention, the recovery step can include a purification step. The purification step may be one that separates and purely purifies only the target protein from the cells. Through the purification step, a purely purified target protein can be produced.
[0066] Another aspect of the present invention is a method for preparing an mRNA construct, including the step of transcribing the construct or vector in vitro and the step of recovering the transcribed mRNA construct.
[0067] Suitable methods known in the art can be used for the transcription method and the recovery method.
[0068] As an embodiment, the method can further include, but is not limited to, a step of treating with DNaseI after transcription to remove the DNA of the construct or vector used as a template, and / or a washing step.
[0069] Another aspect of the present invention is for use in enhancing the RNA stability and / or translation of the mRNA of the construct, vector, recombinant host cell, or composition.
[0070] Another aspect of the present invention is for use in the prevention or treatment of diseases of the construct, vector, recombinant host cell, or composition.
[0071] Another aspect of the present invention is for use in the preparation of a target protein of the construct, vector, recombinant host cell, or composition.
[0072] Hereinafter, the present invention will be described in more detail with reference to experimental examples and examples. However, these are merely for illustrative purposes of the present invention, and the scope of the present invention should not be construed as being limited by these experimental examples and examples.
Example
[0073] <Experimental Example>
[0074] 1. Culturing of cell lines
[0075] All cell lines used in the present invention have undergone mycoplasma negative tests. HeLa cells (deposited by C. H. Chun of Seoul National University, authenticated by ATCC (STR profiling)), Lenti-X 293T cells (manufactured by Clontech, 632180) and 293AV cells (manufactured by Cell biolabs, AAV-100) were cultured in DMEM medium containing 10% FBS (manufactured by Welgene, S001-01). HCT116 cells (manufactured by ATCC, CCL-247) were cultured in McCoy's 5A (manufactured by Welgene, LM 005-01) containing 10% FBS.
[0076] 2. Design of oligos for viromic screen
[0077] The genomic sequences of viruses capable of infecting humans were searched in the NCBI Virus Genome Browser (searched on January 10, 2020, 804 sequences, 504 viruses). For additional information on each virus, searches were conducted in the GenBank files of NCBI Nucleotide. Based on sequence similarity and virus classification, 143 representative virus species were selected, and Woodchuck hepatitis virus was added as a control. For tiling of RNA viruses, the entire genome of sequences in the positive-sense orientation was used for tiling. In the case of DNA viruses, the sequences of the 3' untranslated region of the coding transcriptome and all sequences of non-coding RNA were used for oligo design. When there was no annotation of the UTR, the UTR was predicted based on the annotation of the polyadenylation signal (PAS). When there was no annotation of the PAS, the PAS was predicted using Dragon PolyA Spotter ver. 1.2 within a range of 800 bp from the stop codon. When the PAS could not be predicted, a 390 bp region existing downstream of the stop codon for tiling was used. After determining the genomic regions for tiling, oligos with a sliding window of 130 nt and a shift size of 65 nt were designed. Subsequently, when a window containing the SacI or Notl restriction sites used for replication was included, the window was terminated at the restriction site to produce shorter segments. The next segment started from the restriction site to prevent cleavage of the segment by SacI or NotI during plasmid preparation. Therefore, the screening may miss some virus elements containing the restriction site sequence. Also, the design may miss some elements longer than 65 nt. For example, the leakage probability of an element with a size of 100 nt is about 50%.
[0078] Three barcodes of 7bp random sequences with at least 3 hamming distance were added to each oligo sequence. As a control, the 1E segment and its stem-loop mutants were added to the library. Additionally, the human hepatitis B virus PRE and its corresponding stem-loop mutants were included as controls. Positive and negative controls were each tiled. A total of 30,367 segments and 91,101 oligos were designed.
[0079] For the secondary screening, five classes of K5 mutations were designed. 1) In the case of single nucleotide substitutions, each site was converted to three base types with different bases across K5. 2) In the case of single nucleotide deletions, the bases at each site were removed. 3) In the case of two nucleotide deletions, two consecutive nucleotides were deleted at all sites. 4) To examine the significance of base pairs, the secondary structure was predicted with six other RNA secondary structure prediction softwares, and 38 predicted base pairs were mutated in a way that collected and conserved base pairs (AT / TA / GC / CG / GU / UG / del). 5) Two bases randomly selected in the predicted loop were mutated to generate other combinations. Also, homologs of K5 containing 88 homologous elements of other picornaviruses (including 45 of the genus Kobuvirus) were screened. When the homology was ambiguous, up to 130nt of 3’ was used for oligo design. Altogether, the library for the secondary screening generated a total of 3,864 oligos, including 1,288 elements for each of the three barcodes.
[0080] 3. Generation of plasmid pool
[0081] An oligo of 170 nt (including a 16 nt forward adapter sequence, a 17 nt reverse adapter sequence, and a 7 nt barcode sequence) was synthesized by Synbio Technologies. NotI and SacI restriction sites were added by 6 cycles of PCR using Q5 High-Fidelity 2X Master Mix (manufactured by NEB, M0492) and SacI-univ-F and NotI-univ-R primers. The amplification products were purified using a 6% Native PAGE gel and SYBRgold (manufactured by Invitrogen, S11494) staining. The purified amplification products and the pmirGLO-3XmiR-1 vector were digested with SacI-HF (manufactured by NEB, R3156S) and NotI-HF (manufactured by NEB, R3189S), and cloned into the 3' untranslated region of the firefly luciferase gene by DNA ligase (manufactured by NEB, M0202M) with T4. The ligation products were purified using the Zymo Oligo Clean & Concentrator kit (manufactured by Zymo Research, #D4061) and transformed into Lucigen Endura ElectroCompetent cell (manufactured by Lucigen, LU60242-2). The transformed bacteria were regenerated at 37 °C for 1 hour and then cultured with shaking at 30 °C for 14 hours. The number of colonies was confirmed to be approximately 1E7. The primer sequences used are shown in Table 1 below.
[0082]
Table 1A
Table 1B
Table 1C
Table 1D
Table 1E
Table 1F
Table 1G
Table 1H
Table 1I
Table 1J
[0083] 4. Construction of Library
[0084] For RNA stability screening, 4E5 HCT116 cells were seeded one day before transfection. 1.5 μg of plasmid pool was transfected using Lipofectamine 3000 (Invitrogen, L3000001) and p3000. RNA and DNA were extracted 8 hours after transfection using the Allprp RNA / DNA Mini Kit (Qiagen, 80004), and the remaining plasmid DNA was removed by treating the RNA with recombinant DNase I (RNase-free) (TAKARA, 2270A). RNA was reverse-transcribed using SSIV reverse transcriptase (Invitrogen, 18090010). The extracted DNA, cDNA obtained from the RNA, and the original plasmid pool were amplified by 14 cycles of PCR, and the mixed primers MPRAlib_N / NN / NNN_F and MPRAlib_N / NN / NNN_R were used (Table 1). Six cycles of the second PCR were performed using Illumina index primers. The PCR amplicon was sequenced by next-generation sequencing using the Illumina Novaseq 6000 platform.
[0085] To screen the nuclear / cytoplasmic fraction, the cytoplasm was obtained using cell lysis buffer (0.15 μg / μl digitonin [Merck, D141], 150 mM NaCl, 50 mM HEPES [pH 7.0 - 7.6], 20 U / ml RNase inhibitor [Ambion, AM2696], 1X protease inhibitor [Calbiochem, 535140], 1X phosphatase inhibitor [Merck, P0044]). The library production steps were performed as in the RNA stability screening.
[0086] For screening of the polysome fraction, Gradient Master TM(manufactured by Biocomp, B108-2) was used to prepare a 10 - 50% sucrose gradient. HCT116 cells with a scale three times that of RNA stability screening were treated with 100 μg / mL cycloheximide at 37°C for 1 minute, and then lysed on ice for 10 minutes with 150 μl of PEB (20 mM Tris-Cl, pH 7.5, 100 mM KCl, 5 mM MgCl2, 0.5% NP-40 [manufactured by Merck, 74385]), 100 U / ml RNase inhibitor, 1X protease inhibitor, and 1X phosphatase inhibitor, followed by centrifugation. The supernatant was layered onto the sucrose gradient and centrifuged at 36,000 rpm for 2 hours at 4°C using an SW41Ti rotor and a Beckman Coulter Ultracentrifuge Optima XE. Samples were obtained as 0.25 ml fractions using a BioLogic LP system coupled to a Model 2110 fraction collector (manufactured by Bio-Rad, 7318303) and a Model EM-1 Econo UV detector (manufactured by Bio-Rad). 0.75 ml of TRIzol TM LS Reagent (manufactured by Life Technologies) was added to each fraction. As shown in the graph of Figure 11, free mRNA, monosome, light polysome (LP; 2 - 3 ribosomes), medium polysome (MP; 4 - 8 ribosomes), and heavy polysome (HP; 9 or more ribosomes) were separated according to the absorbance trend at 254 nm using the polysome fractionation method and extracted using a Direct-Zol RNA Miniprep kit (manufactured by Zymo Research, R2052).
[0087] Next, the library manufacturing stage was carried out in a similar manner to the RNA stability screening. The sequencing data is available in the Zenodo database under the following identifier DOIs [DOI: 10.5281 / zenodo.6777910 (stability), 10.5281 / zenodo.6717932 (polysome), 10.5281 / zenodo.6696870 (secondary screening), 10.5281 / zenodo.7773943 (nuclear / cytoplasmic fraction).
[0088] 5. Data analysis
[0089] For all samples, reads were aligned to the oligos using bowtie 2.2.6 with the parameter -local. The aligned reads were filtered with unique alignments that exactly matched the barcode. Statistical tests were performed in MPRAnalyze using the mpralm function. Technical performance was evaluated using the Spearman correlation coefficient of the scipy module and the histogram plot. Normalized counts were used for visualization. In the case of polysome analysis, after transforming the variance stabilization using DESeq2, the average of the five fractions was subtracted, and then the relative distance of each fraction was calculated. The relative distance of each fraction was used to perform hierarchical clustering in the scipy module. For the quantification of another translation, the Mean ribosome load (MRL) was calculated using the following formula.
[0090] 1×p(Monosome)+2.5×p(Light polysome)+6×p(Medidum polysome)+11×p(Heavy polysome)
[0091] p(X): Ratio of sequencing reads of X (each fraction)
[0092] For the cutoff of mRNA stability, Log2(RNA / DNA) < -1 and p-value < 0.001 were used for negative regulatory elements, and Log2(RNA / DNA) > 0.5 and p-value < 0.05 were used for positive regulatory elements. For the cutoff of translation activation elements, Log2(heavy polysome / free mRNA) > 0.2 and / or MRL > 4.5 were used, and for the cutoff of translation downward regulatory elements, Log2(heavy polysome / free mRNA) < -0.2 and / or MRL < 3.5 were used.
[0093] For the second screening substitution data, the base-identity score of substitutions and deletions was calculated by the following formula.
[0094] A / mean(Stability x , Stability y , Stability z )(for substitution, x, y, z : substituted nucleotides) A / Stability for deletion (for deletion) A: Stability of wild-type K5
[0095] The base-pairing score for substitution data was calculated by the following formula.
[0096] mean(Stability of substitutions maintaining base pair)-mean(Stability of substitutions disrupting base pair)
[0097] The pair-deletion score for deletion data was calculated by the following formula.
[0098] A / Stability for pairwise deletion
[0099] To construct the picornavirus tree, the viral sequences retrieved from NCBI were aligned by ClusterOmega and visualized using FigTree v1.4.4. The conservation score was calculated as the number of nucleotides identical to the K5 element after aligning the multiple sequences across the top 33 species. For visualization of the RNA structure, the structure was predicted using RNAfold and visualized using forna.
[0100] 6. Construction of plasmids
[0101] The elements selected for the validation experiments were amplified by PCR in a plasmid library pool and cloned into the 3’untranslated region of the firefly luciferase gene in the pmirGLO-3XmiR-1 vector. In the case of the luciferase construct, the K5 element (8122 - 8251: NC_001918.1) was amplified from the plasmid pool library, and an additional 55 bp and 110 bp were added by PCR amplification to generate the eK5 element (8067 - 8251: NC_001918.1) and the entire untranslated region (8012 - 8251: NC_001918.1). The 120-K5 element (8132 - 8251: NC_001918.1), 110-K5 element (8142 - 8251: NC_001918.1), and the K5m element (8122 - 8251, 8185ΔG: NC_001918.1) were amplified from the pmirGLO-3XmiR-1 K5 plasmid, and the eK5m element (8067 - 8251: NC_001918.1) was amplified from the pmirGLO-3XmiR-1 eK5 plasmid. The K4 element (7931 - 8060: NC_009448.2) was amplified from the plasmid pool library, and an additional 50 bp was included by PCR amplification to generate the eK4 element (7881 - 8060: NC_009448.2). The 1E element (414 - 463: RNA2.7) was amplified from the pmirGLO-3XmiR-1 1E vector.
[0102] For AAV production, the pAAV-CAG-GFP (Addgene, Plasmid #37825) plasmid was used as a template. The K5 element (8122-8251: NC_001918.1), K5m element (8122-8251, 8185ΔG: NC_001918.1), eK5 element (8067-8251: NC_001918.1), and eK5m element (8067-8251, 8185ΔG: NC_001918.1) were amplified from the pmirGLO-3XmiR-1 eK5 and eK5m plasmids and the WPRE sequence in the pAAV-CAG-GFP plasmid was replaced by Gibson assembly. For the control plasmid, the WPRE sequence in the 3’untranslated region of the GFP gene in pAAV-CAG-GFP was removed by PCR-based amplification.
[0103] For the construction of the d2EGFP plasmid, the firefly luciferase gene from the pmirGLO-3XmiR-1 vector was replaced with the GBA 5’untranslated region, d2EGFP CDS, and GBA 3’untranslated region to generate a control plasmid. The UTRs from the luciferase construct were amplified and inserted into the d2EGFP control vector.
[0104] For tethering and construct building, pmirGLO-3xBoxB was generated from the pmirGLO-3xmir1-5xBoxB vector, and ZCCHC2 amplified from HCT116 cDNA in the case of the pGK-ZCCHC2 construct was subcloned into the pGK vector. Tethering constructs containing the ZCCHC2ΔC (1-375 a.a.) and ZCCHC2ΔN (201 aa-1,178 a.a.) structures were generated by subcloning ZCCHC2 with pGK-TEV-HA-λN. To prepare mutant versions of the ZCCHC2 zinc finger, the first and second cysteines of the zinc finger (CX2CX3GHX4C) were replaced with serine by mutagenesis PCR. In the case of the TNRC6B C-term structure, the C-term region (716-1,028 a.a.) of the TNRC6B gene was amplified with HCT116 cDNA and subcloned into the pGK and pGK-TEV-HA-λN vectors using Gibson assembly.
[0105] For the RaPID experiment, the EGFP CDS, 3xBoxB sequence, and eK5 sequence were amplified from the d2EGFP, pmirGLO-3xBoxB, and pmirGLO-3xmir-1-eK5 plasmids, respectively, and subcloned into the pCK vector using Gibson assembly.
[0106] A list of the plasmids obtained by the above method is shown in Table 1.
[0107] 7. Luciferase assay and transfection
[0108] The luciferase assay was performed as follows. For the luciferase reporter assay with Lipofectamine 3000, 2E5 of HeLa or HCT116 cells on a 24-well plate were transfected with 100 ng of pmirGLO-3XmiR-1 plasmid on day 0 and obtained on day 2. For the knockdown experiment, 100 ng of pmirGLO-3XmiR-1 K5 plasmid and 40 nM siRNA (Dharmacon siRNA smartpool) for each target gene were co-transfected with Lipofectamine 3000. For the ZCCHC2 structure experiment, 50 ng of pmirGLO-3XmiR-1 plasmid and 60 ng of pGK-null or pGK-ZCCHC2 or pGK-ZCCHC2 zinc finger mutant structure were co-transfected. For the de-ringing experiment, 50 ng of pmirGLO-3xBoxB plasmid and 60 ng of pGK-ZCCHC2 wild-type / mutant construct were co-transfected like with or without the flag of λN-HA-TEV. For the luciferase assay, cells were lysed and analyzed with the Dual-luciferase reporter assay system (Promega) according to the manufacturer's instructions.
[0109] 8. RT-qPCR
[0110] RNA was extracted with the RNeasy Mini Kit (Qiagen, 74106), treated with DNase (Qiagen, 79254), and reverse-transcribed with Primescript RTmix (Takara, RR036A). The mRNA level was measured with SYBR Green assays (Life Technologies, 4367659) and the StepOnePlus Real-Time PCR System (Applied Biosystems) or QuantStudio 3 (Applied Biosystems). The list of RT-qPCR primers is shown in Table 1.
[0111] 9. AAV Generation and Purification
[0112] AAV generation and purification were performed as follows. The 293AAV cell line (manufactured by Cell Biolabs, #AAV - 100) was cultured in DMEM containing 10% FBS, 0.1 mM MEM Non - essential Amino Acids (NEAA), and 2 mM L - glutamine. To produce AAV carrying the GFP protein, the 293 AAV cell line was dispensed into 150 - mm Petri dishes overnight, and when it grew to about 70%, it was co - transfected with the pAAV - CAG - GFP plasmid variant (manufactured by Addgene, 37825), pAdDelta6F6 (manufactured by Addgene, 112867), and pAAVDJ (manufactured by Cell Biolabs, VPK - 420 - DJ) plasmids using Lipofectamine 3000 and p3000. After transfection for 72 hours, the cells were harvested and resuspended in 2.5 ml of serum - free DMEM. Subsequently, cell lysis was performed by carrying out 4 cycles of freezing / thawing (freezing in alcohol / dry ice for 30 minutes and thawing in a 37°C water bath for 15 minutes for each cycle). The AAV supernatant was obtained by centrifuging at 10,000 Хg for 10 minutes at 4°C. ViraBind TM The AAV was purified using the ViraBind TM AAV Purification Kit (manufactured by Cell Biolabs), and the viral titer was measured using the QuickTiter
[0113] 10. Production of In Vitro Transcribed RNA
[0114] For in vitro transcribed RNA, the DNA template was prepared by performing PCR using a forward primer (T7 promoter + gene_specific_F) and a reverse primer (T120 + gene_specific_R, two nucleotides of 2'-O-methylated deoxyuridine at the 5' end). 250 ng of the DNA template was mMESSAGE mMACHINE TM transcribed in vitro using the T7 Transcription Kit (Invitrogen, AM1344) and components (7.5 mM each of ATP / CTP / UTP [New England Biolabs, N0450S], 1.5 mM GTP, and 6 mM CleanCap® (Reagent AG (3’ OMe) [TriLink Biotechnologies, N-7413-10]). The DNA template was removed using recombinant DNase I (RNase-free) and washed using the RNeasy MiniElute Cleanup Kit (Qiagen, 74204). The primers used for the production of the template for in vitro transcription are shown in Table 1.
[0115] 11. Preparation and analysis of mRNA-transfected samples
[0116] 2E5 HeLa cells on a 12-well plate were transfected with in vitro transcribed RNA using Lipofectamine MessengerMax. For samples transfected with luciferase mRNA, the cells were lysed and analyzed using the Dual-luciferase reporter assay system according to the manufacturer's instructions. For d2EGFP samples, the cells were lysed on ice for 10 minutes in RIPA lysis and extraction buffer (Thermo, 89901) containing 1X protease inhibitor and 1X phosphatase inhibitor, and then centrifuged. After boiling the samples in 5X SDS buffer, they were loaded onto a Novex SDS-PAGE gel (10-20%) using a ladder (Thermo, 26616). After transferring the gel to a PVDF membrane (Millipore) activated with methanol, the membrane was blocked with PBS-T containing 5% skim milk, probed with the primary antibody, and washed three times with PBS-T. The primary antibodies used were anti-EGFP (1:3,000, CAB4211, Invitrogen) and anti-alpha-TUBULIN (1:300, Abcam, ab52866) as the primary antibodies. The anti-mouse or anti-rabbit HRP-conjugated secondary antibody (Jackson ImmunoResearch Laboratories) was incubated for 1 hour and washed three times with PBS-T. Chemiluminescence was detected using West Pico or Femto Luminol reagent (Thermo), and the signal was detected using the ChemiDoc XRS+ System (Bio-Rad). For d2EGFP samples, the GFP signal was detected by flow cytometry (BD Accuri C6 Plus).
[0117] 12. Hire-PAT Analysis
[0118] Hire-PAT analysis and signal processing of capillary electrophoresis data were performed as described in the literature (Kim et al., Nat. Struct. Mol. Biol., 2020, 27, 581-588). The poly(A) site of the firefly luciferase gene was utilized as confirmed by Sanger sequencing in the said literature, and the forward PCR primer for the poly(A) site is shown in Table 1.
[0119] 13. Gene-specific TAIL-seq
[0120] For measuring the poly(A) tail length distribution during RG7834 treatment, HeLa cells were transfected with the pmirGLO-3XmiR-1 plasmid containing the K5 element in the 3’untranslated region of firefly luciferase treated with RO0321 (manufactured by Glixx Laboratories Inc, GLXC-11004) or RG7834 (manufactured by Glixx Laboratories Inc, GLXC-221188), and obtained within 2 days. To compare the poly(A) tail length distribution between parental cells and ZCCHC2 knockout cells, parental cells and ZCCHC2 knockout cells were prepared in the same manner as the RG7834-treated samples. For performing gene-specific TAIL-seq, total RNA depleted of rRNA (Truseq Strnd Total RNA LP Gold, manufactured by Illumina, 20020599) was ligated to a 3’adapter and partially fragmented with RNase T1 (manufactured by Ambion). After purification with a Urea-PAGE gel (300 - 1500 nt), the RNA was reverse transcribed and amplified by PCR. For PCR amplification of firefly luciferase, GS-TAIL-seq-FireflyLuc-F was used as the forward primer. The library was sequenced on an Illumina platform (Miseq) in a paired-end run (51X251 cycles) with the PhiX control library v.2 (manufactured by Illumina) containing a spike-in mixture. The TAIL-seq sequencing data was stored in the Zenodo database with the identifier of DOI: 10.5281 / zenodo.6786179.
[0121] TAIL-seq was analyzed using Tailseeker v.3.1.5. For each transcriptome, genes were confirmed by mapping read 1 to the constitutive sequence of firefly luciferase and the human transcriptome using bowtie2.2.6. Next, the poly(A) tail length and 3’ end modification of the relevant transcript were extracted with read 2. The ratio of mixed tailing was calculated from the transcriptome with a poly(A) tail length longer than 50 nt.
[0122] 14. Preparation of TENT4, ZCCHC2, and ZCCHC14 knockout (KO) cells
[0123] TENT4 dKO cells were prepared in the same manner as described in the literature by Kim et al. Also, ZCCHC2 and ZCCHC14 knockout cell lines were prepared according to the description in the literature by Kim et al. 4E5 HeLa cells on a 6-well plate and 1E5 HCT116 cells on a 24-well plate were each transfected with 300 ng of the pSpCas9(BB)-2A-GFP-px458 plasmid (Addgene, #48138) containing sgRNA for ZCCHC2 (ACCTCAGGACGGACTTACCG (SEQ ID NO: 96), PAM sequence: TGG) and sgRNA for ZCCHC14 (CAAGTGGGCAGCGCGCGCCGCC (SEQ ID NO: 97), PAM sequence: CGG) using Metafectene (Biontex, T020). After single-cell screening, Sanger sequencing and Western blot analysis were performed to identify the knocked-out strains. The parental cells and the modified genomic sequences are shown in Table 1, and the inserted sequences are shown in red.
[0124] 15. Analysis by RNA proximity-dependent labeling method
[0125] The RaPID (RNA-protein interaction detection) assay was performed as follows. Specifically, a lentiviral transport construct generated from Lenti-X 293T (manufactured by Clontech, 632180) and the BASU RaPID plasmid (manufactured by Addgene, #107250) was transduced to generate a BASU-expressing HeLa stable cell line. 1E7 cells from a 150 pi plate were transfected with 40 μg of the synthesized RNA using Lipofectamine mMAX (manufactured by Life Technologies, LMRNA015). After 16 hours, the cells were treated with 200 μM biotin (manufactured by Sigma, B4639) for 1 hour. The treated cells were lysed on ice for 10 minutes with RIPA lysis and extraction buffer (manufactured by Thermo, 89901) containing 1X protease inhibitor and 1X phosphatase inhibitor, and then centrifuged. The lysate was cultured with Pierce streptavidin beads (manufactured by Thermo, 88816) with rotation at 4°C overnight. The beads were washed three times with wash buffer 1 (1% SDS containing 1 mM DTT, protease and phosphatase inhibitor cocktail), once with wash buffer (1 μM EDTA containing 0.1% Na-DOC, 1% Triton X-100, 0.5 M NaCl, 50 mM HEPES [pH 7.5], 1 mM DTT, protease and phosphatase inhibitor cocktail), and then once with wash buffer 3 (1 μM EDTA containing 0.5% Na-DOC, 150 mM NaCl, 0.5% NP-40, 10 mM Tris-HCl, 1 mM DTT, protease and phosphatase inhibitor cocktail).
[0126] For Western blot, after eluting the protein with Elution buffer (1.5x Laemelli sample buffer, 0.02 mM DTT, 4 mM biotin), Western blot analysis was performed using primary antibodies against ZCCHC2 (1:250, Atlas Antibodies, HPA040943), TENT4A (1:500, Atlas Antibodies, HPA045487), anti-alpha-TUBULIN (1:300, Abcam, ab52866) and anti-HA (1:2000, Invitrogen, 715500). For LC-MS / MS, the samples were washed 6 times with digestion buffer (50 mM Tris [pH 8.0]) at 37 °C for 1 minute. After washing, the protein-binding beads were incubated with 180 μL of digestion buffer containing 2 μL of 1 M DTT at 37 °C for 1 hour, then 16 μL of 0.5 M IAA was added and incubated at 37 °C for 1 hour. Then, 2 μL of 0.1 g / L trypsin was added and incubated at 37 °C overnight. The remaining digest was removed using HiPPR (Thermo, 88305) and washed with ZipTip C18 resin (Millipore, ZTC18S960) before LC-MS / MS analysis.
[0127] LC-MS / MS analysis was performed using an Orbitrap Eclipse Tribrid (Thermo) combined with a nanoAcquity system (Waters). The capillary assay (inner diameter 75 μm × 100 cm) and trap (inner diameter 150 m × 3 cm) columns were packed with 3 um Jupiter C18 particle (Phenomenex). The LC flow was set at 300 nl / min with a 60-minute linear gradient ranging from 95% solvent A (0.1% formic acid (Merck)) to 35% solvent B (100% acetonitrile, 0.1% formic acid). Full MS scans (m / z 300 - 1,800) were acquired at 120k resolution (m / z 200). High-energy collision dissociation (HCD) fragmentation occurred at 30% normalized collision energy (NCE) including the 1.4th precursor isolation window. MS2 scans were acquired at 30k resolution.
[0128] The MS / MS raw data was analyzed using MSFragger1 (v3.7), IonQuant2 (v1.8.10), and Philosopher3 (v4.8.1) incorporated in FragPipe (v18.0). For the identification and quantification of unlabeled proteins, the built-in FragPipe workflow (LFQ-MBR) was used with trypsin as the specified enzyme. The target-derived database (including contaminants) was generated in Fragpipe with the Swiss-Prot human database4 (October 2022). The combined_protein.tsv file was used for additional analysis. For the enrichment cut-off, a Log2FC greater than 1 was used by at least two repeated experiments.
[0129] 16. Co-immunoprecipitation (co-IP) and Western blotting
[0130] For the co-IP experiment, parental cells and TENT4 dKO cells in 150 pi plates were lysed on ice for 20 minutes using Buffer A (100 mM KCl, 0.1 mM EDTA, 20 mM HEPES [pH 7.5], 0.4% NP-40, 10% glycerol) containing 1 mM DL-dithiothreitol (DTT), 1X protease inhibitor, and RNase A (Thermo, EN0531), and then centrifuged. 12.5 μg of antibodies (anti-TENT4A and anti-TENT4B from NMG) conjugated with Protein A and G Sepharose beads (1:1 mixture, total 20 μl) were used for immunoprecipitation with 1 mg of the above lysate. After culturing at 4°C for 2 hours, the beads were washed, boiled in 20 μl of 2X SDS buffer, and loaded onto a 4 - 12% SDS-PAGE gel (Novex) together with a ladder (Thermo, 26616 and 26619). For the domain co-IP experiment, full-length ZCCHC2, truncated ZCCHC2 construct, and flag-tagged negative construct were transfected into ZCCHC2KO cells, and the cells were lysed within 2 days. After adding 10 μl of ANTI-FLAG® M2 Affinity Gel (Merck, A2220-10ML) to 1 mg of the above lysate, the mixture was cultured at 4°C for 2 hours to perform immunoprecipitation. For the input sample, 50 μg of cell lysate was used. After transferring the gel to a methanol-activated PVDF membrane (Millipore), the membrane was blocked with PBS-T containing 5% skim milk, probed with the primary antibody, and then washed 3 times with PBS-T. Anti-ZCCHC2 (1:250, Atlas, HPA040943), anti-ZCCHC14 (1:1,000, Bethyl Laboratories, A303-096A), anti-TENT4A (1:500, Atlas Antibodies, HPA045487), anti-TENT4B (1:500, lab-made), anti-GAPDH (1:1,000, Santa Cruz, sc-32233), and anti-FLAG (1:1000, abcam, ab1162) were used as the primary antibodies.Cultured for 1 hour using anti-mouse or anti-rabbit HRP-conjugated secondary antibody (manufactured by Jackson ImmunoResearch Laboratories), and washed three times with PBS-T. Chemiluminescence was performed using West Pico or Femto Luminol reagent (manufactured by Thermo, 34580 and 34095), and the signal was detected using ChemiDoc XRS+ System (manufactured by Bio-Rad).
[0131] 17. Reanalysis of RNA pull-down - LC-MS / MS data
[0132] The MS / MS data was processed with MaxQuant v.1.5.3.30 using the human Swiss-Prot database v.12 / 5 / 2018 basic settings, and a protein-level FDR cutoff of 0.8% was applied.
[0133] From the MaxQuant output file, the MaxLFQ intensity values were extracted from the proteingroups.txt file. After adding an approximate value of 10,000 to the MaxLFQ intensity values, Limma was run, and significant genes were filtered with Log2FC < 0.8 and FDR > 0.1.
[0134] 18. Domain conservation analysis
[0135] Aligned ZCCHC2 (Q9C0B9), ZCCHC14 (A0A590UJW6) and GLS-1 (Q8I4M5) using the UniProt Align tool, and calculated the conservation scores of the three proteins.
[0136] 19. RNA immunoprecipitation
[0137] For immunoprecipitation of ZCCHC2, lentiviral transporters from constructs generated from Lenti-X 293T (manufactured by Clontech, 632180) cells were transduced to generate HeLa stable cell lines expressing EGFP with a K5 element in the 3' untranslated region. Also, the cells were lysed by treating them with lysis buffer (20 mM HEPES (pH 7.6) [manufactured by Ambion, AM9851 and AM9856], 0.4% NP-40, 100 mM KCl, 0.1 mM EDTA, 10% glycerol, 1 mM DTT, 1X protease inhibitor [manufactured by Calbiochem, 535140]) on ice for 30 minutes and then centrifuged to obtain cell lysates. As a negative control, 10 μg of normal rabbit IgG (manufactured by Cell Signaling, 2729S) was used, and 10 μg of ZCCHC2 antibody (manufactured by atlas, HPA040943) was used for ZCCHC2 immunoprecipitation. After the antibody was attached to protein A magnetic beads (manufactured by Life Technologies, 10002D), 1 mg of cell lysate was cultured with the antibody-conjugated beads for 2 hours and then washed with wash buffer (the same lysis buffer as above containing 0.2% NP-40). After adding 5 ng of firefly luciferase mRNA used for normalization to each sample by spike-in, the RNA was purified with TRIzol reagent (manufactured by Life Technologies) and used for RT-qPCR. The RT-qPCR primers are shown in Table 1.
[0138] 20. Subcellular fractionation
[0139] Subcellular fractionation was performed as follows. Specifically, to obtain the cytoplasmic fraction, cells were lysed with 200 μl of cell lysis buffer (0.2 μg / μl digitonin [Merck, D141], 150 mM NaCl, 50 mM HEPES [pH 7.0 - 7.6], 0.1 mM EDTA, 1 mM DTT, 20 U / ml RNase inhibitor, 1X protease inhibitor, 1X phosphatase inhibitor). For the cell membrane and nuclear fractions, they were fractionated using an intracellular protein fractionation kit (Thermo Scientific, 78840) according to the manufacturer's instructions. As the primary antibodies, anti-GM130 (1:500, BD Bioscience, 610822) and anti-histone (1:2000, cell signaling, 4499) were used.
[0140] The reagents and materials used in the experimental examples of the present invention are shown in Table 2 below.
[0141]
Table 2A
Table 2B
Table 2C
Table 2D
Table 2E
[0142] [Examples]
[0143] 1. Viromic screen to identify regulatory RNA elements
[0144] To construct a library of viral RNA elements, due to the technical limitations of oligo synthesis, a two-step approach was taken that involved an initial screening with human viruses and an expansion of the secondary screen to include other related species. To identify viruses that can infect humans, the NCBI database, which currently annotates 502 human viruses belonging to 114 genera and 40 families, was used.
[0145] As shown in Figure 1A and Table 3, after manual inspection, 143 species representing 96 genera and 37 families were selected, excluding species with close sequence similarity and those that were ambiguously classified or lacked clear evidence of human infection. The catalog of the present invention includes all seven groups of the Baltimore classification system. For RNA viruses, the full genome sequence was used. In general, for DNA with a larger genome, the untranslated region (UTR) and non-coding genes were included.
[0146]
Table 3A
Table 3B
Table 3C
Table 3D
Table 3E
Table 3F
Table 3G
Table 3H
[0147] As shown in FIG. 1B, the size of the sliding window was 130 nt, and the viral genome was tiled at a step size of 65 n to generate a total of 30,367 segments for designing screening oligos. Each segment was manufactured to have three different barcodes for reliable detection. As positive controls, four segments containing the "1E" element from lncRNA2.7 of HCMV (human cytomegalovirus) (FIG. 8) and one segment containing the WPRE (woodchuck posttranscriptional regulatory element) of woodchuck hepatitis virus known to improve gene expression were included. As a non-functional control, the inventors used a mutation (1Em) containing an inactivating mutation in the loop of 1E. After synthesizing the designed oligos, they were amplified by PCR and inserted into the 3' untranslated region of the luciferase reporter plasmid. The constructed library contained a total of 91,101 reporter plasmids and included 30,367 segments in 143 human viruses and woodchuck hepatitis virus of dogs.
[0148] For functional evaluation, plasmid pools were transfected into a human colon cancer cell line (HCT-116) to quantify the effect of each element on gene expression (FIG. 1B). To monitor the effect on RNA abundance, all plasmids and mRNAs were extracted, amplified, and sequenced to calculate the ratio between the ratio of mRNA reads and the ratio of transfected DNA reads ("RNA / DNA"). To explore translational regulatory elements, cytoplasmic extracts were separated into five fractions (free mRNA, monosomes, light polysomes (LP), intermediate polysomes (MP), and heavy polysomes (HP)) using sucrose gradient centrifugation, and the extracts were used for RNA extraction and sequencing to predict translational efficiency for each UTR.
[0149] 2. Identification of Regulatory RNA Elements
[0150] The following experiments were conducted to confirm the effects on the abundance of 30,302 viral segments (30,190 segments in which all three barcodes were detected) on mRNA abundance. The experimental results were reproducible among the four replicate experiments and the barcodes. Specifically, the positive controls across 1E and WPRE had increased mRNA levels compared to the 1E mutation (Figure 1C). 245 upregulated segments and 628 downregulated segments were identified. The segments that increased the abundance of mRNA contained the stem-loop α of the hepatitis B virus, which is part of the PRE known to improve mRNA stability. The negative elements included RNA cleaved by ribonucleases such as the self-cleaving ribozyme of hepatitis D virus (HDV), and human cytomegalovirus (also known as human betaherpesvirus 5) and Epstein-Barr virus microRNA loci that may be cleaved by DROSHA, resulting in the degradation of reporter mRNA (Figure 1C).
[0151] Therefore, segments that stabilized (Log2(RNA / DNA) > 0.5, p-value < 0.05) or destabilized (Log2(RNA / DNA) < -1, p-value < 0.001) RNA were effectively identified in the above experiments (Tables 4 and 5). The 50 segments in Table 4 were confirmed to have Log2(RNA / DNA) values similar to or higher than the positive controls WPRE or HCMV1E, indicating excellent RNA abundance (Figure 1C).
[0152] Segments that Stabilize RNA
[0153]
Table 4A
Table 4B
Table 4C
[0154] Segments that destabilize RNA
[0155]
Table 5
[0156] Also, using polysome profiling sequencing data (Figure 1D), the translation effects of 30,155 segments (29,786 segments where all three barcodes were detected) were evaluated. Non-mutated WPRE and 1E were abundantly present in the heavy polysome fraction consistent with a positive effect on translation (Figure 1E). 535 upregulated segments and 66 downregulated segments were identified, and the translation efficiency was estimated using the lead ratio between the heavy polysome and free mRNA fractions (Log2(HP / free mRNA) > 0.2) (Table 6). The 30 segments in Table 6 below were confirmed to be able to enhance mRNA translation as they were abundantly present in the heavy polysome fraction similar to the positive control WPRE and HCMV 1E (Figure 1E).
[0157]
Table 6A
Table 6B
[0158] 3. Verification of regulatory elements
[0159] The very weak correlation between the presumed mRNA abundance and the translation efficiency suggested that most viral elements affect either the mRNA abundance or translation. Nevertheless, several segments were confirmed to affect both. To verify this, 16 previously unstudied candidates that improved both RNA abundance and translation were selected (Figure 2A, Table 7; satisfying Log2(HP / free mRNA) > 0.2 and MRL > 4.5). Using 3’ untranslated region reporters and individual luciferase assays, it was confirmed that 15 out of 16 candidates had increased luciferase expression with statistical significance (p < 0.05) (Figure 2B).
[0160]
Table 7
[0161] The K4 element in the 3’ UTR of Saffold virus (7931 - 8060 NC_009448.2) and the K5 element in the 3’ UTR of Aichi virus 1 (AiV-1) (8122 - 8251: NC_001918.1) were further investigated (Figure 2C). Both of the two viruses belong to the family Picornaviridae, which have a single-stranded positive RNA genome encoding a single polypeptide, and the virus is processed into multiple fragments due to proteolysis. Saffold virus and AiV-1 belong to the genera Cardiovirus and Kobuvirus, respectively, and are widely distributed in viruses that cause relatively mild symptoms including gastroenteritis, and it was confirmed that they have not yet been investigated.
[0162] To map the boundaries of the elements, extended or truncated segments of K4 and K5 were examined. An extended 180 nt segment of K4 (「eK4」, 7881 - 8060) that contains the entire 3’ untranslated region of Saffold virus showed a similar effect as the existing K4 segment, confirming that 130 nt at the 3’ end is sufficient to convey the activity of K4. However, the extended form of K5 (「eK5」, 8067 - 8251, 185 nt, SEQ ID NO: 94) further enhanced luciferase expression and outperformed other elements including the existing K5, K4, and extended K4 (eK4) (Figure 2D). Also, a 120 nt segment (8132 - 8251, SEQ ID NO: 95) shorter than K5 showed higher activity than K5. Notably, K5 was selected as one of the top 25 candidates in both mRNA abundance and translation screening, suggesting that K5 is a particularly potent element. Cleavage experiments on K5 showed that elements exceeding 110 nt at the 3’ end (8142 - 8251) can constitute at least the K5 element (Figure 2E). The K5-containing segment increased the mRNA level, and more importantly, the protein level was consistent with the screening data.
[0163] 4. Characterization of the K5 Element
[0164] To characterize K5 in detail, a secondary high-throughput screening assay against K5 variants and homologs was performed (Figures 3A and 3B). Single nucleotide substitutions, single nucleotide deletions, and two consecutive nucleotide deletions were performed at all sites of the 130 nt K5 element for mutagenesis (Figure 3C). Also, the sequence was changed, but compensatory mutations were introduced that preserved the predicted duplex structure. Furthermore, the loop was substituted with up to two bases randomly selected in different combinations. A total of 1,201 mutants were synthesized with three barcodes. After cloning and transfection, the mRNA level relative to the transfected DNA level was measured to evaluate the effect of the mutation on the mRNA abundance (Figure 3B).
[0165] As shown in Figure 3D, to quantify the contribution of specific nucleotide sequences, a "base identity score" was calculated using single base substitution data. Also, a "base pair score" was calculated based on compensatory mutation data indicating the requirements for base pairs in the stem region. As a result, some mutations, particularly mutations in the first 14 nucleotides, slightly increased the mRNA level (Figure 3C), which showed self-inhibitory activity, and such results were consistent with the cleavage experiment (Figure 2E). Also, in the case of other mutants, the mRNA level was increased to be similar to or higher than that of K5 (Figure 3C). In contrast, mutations to the first hairpin (including a pyrimidine-rich terminal loop) and the second hairpin (including a G bulge) substantially decreased the mRNA level, thereby confirming the importance of these hairpins in K5 activity (Figure 3D). The above results were consistent with the results of deletions and compensatory mutations.
[0166] To examine the phylogenetic distribution of K5, the 3’ untranslated region segments of 88 picornavirus species (K5 and 87 different picornavirus elements) were included in the secondary screening. Among these picornaviruses, 43 kobuvirus segments (Table 8; having at least 59% homology with K5) upregulated mRNA levels far more than the non-functional control K5m, which has a deletion in the G overhang of the second hairpin (Figures 3D and 3E, Table 8), and upregulated mRNA levels similar to or higher than K5. The above results confirmed that K5 is conserved in the genus Kobuvirus. Some kobuvirus segments lacking the conserved 3’ sequence were found to be less active upon analysis. Such deletions in the 3’ sequence may be due to incomplete database annotation.
[0167]
Table 8A
Table 8B
[0168] The 3' untranslated regions of most picornaviruses except for the Kobuvirus genus did not increase the mRNA abundance (Figure 3E). However, there were some exceptions, particularly the segment of Boone cardiovirus 1 (NC_038305.1) (SEQ ID NO: 187; RNA / DNA ratio = 1.2433) containing the K4 positive element (RNA / DNA ratio = 1.514), which is related to the Saffold virus. These two viruses belong to the Cardiovirus genus. Therefore, the K4 and its homologous elements of Cardiovirus can form a separate conserved regulatory element group. Specifically, among the following base sequences of K4, the underlined base sequence (positions 7952 to 7988 of NC_009448.2) has 78.38% identity with the corresponding sequence (the following underlined base sequence) within the segment of Boone cardiovirus 1, which is a homolog. Therefore, it can be seen that a homolog containing a base sequence that is a base sequence within the 3' untranslated region of Cardiovirus and has at least 70% identity with the base sequence at positions 7952 to 7988 within the Saffold virus gene can increase the mRNA abundance like K4.
[0169] K4
[0170]
Number
[0171] Underline: 7952 - 7988 of NC_009448.2
[0172] Boone cardiovirus 1
[0173] TTCGGTTGAGCCCCCACCCGGTACAACGCTTTACCTTAGAAGCCACTAAGGTGTACGCGGTCATCGGGGACCCCTCCTGGCCTTTGGTTTATTGGTGAATTACTAGTTCAGTTAGGTTTTGTTAGTTAGG(SEQ ID NO: 187)
[0174] 5. Enhancement of gene expression from vectors and synthetic mRNAs by K5
[0175] To test whether K5 can function in other molecular backgrounds, an adeno-associated virus (AAV)-based vector system, which is a single-stranded DNA virus belonging to the parvovirus family that enables gene transfer with low toxicity and efficacy for human gene therapy, was used. As shown in Figure 4, WPRE enhanced gene expression in AAV35, but its use in AAV was limited due to its large size (~600 nt) and the limited packaging capacity of AV (1.7 - 3 kb).
[0176] The minimal K5 (120 nt) or eK5 (185 nt) sequences and inactive mutants (K5m and eK5m) and WPRE were evaluated as controls. Such segments were inserted downstream of the EGFP coding sequence in the AAV vector, and the effect on gene expression was measured (Figure 4A). As shown in Figures 4B and 4C, GFP expression from the AAV vector increased under two different transduction conditions by both K5 and eK5. In particular, the effect of eK5 (~3-fold) was confirmed to be superior to that of WPRE (~2-fold). This confirmed that eK5 can greatly improve the AAV vector while saving packaging space.
[0177] Also, the experiment was repeated using a lentiviral vector. As a result, it was confirmed that GFP expression increased by eK5 in the same manner as the AAV vector even when using the lentiviral vector (Figure 10).
[0178] In vitro transcribed (IVT) mRNA represents another important platform for gene delivery, as exemplified by COVID-19 vaccines. To test the effect of K5 on IVTmRNA, as shown in Figure 4D, luciferase-encoding mRNA was synthesized with or without functional eK5. The mRNA contained a cap-1 homolog, a 3' untranslated region sequence derived from the pmirGLO vector, and a 120-base poly(A) tail. The mRNA was transfected into HeLa cells and cultured for up to 72 hours. As shown in Figure 4E, in the absence of functional eK5, luciferase levels decreased rapidly over time, confirming that the lifespan of the transduced mRNA was short. However, when eK5 was included, it was confirmed that the duration of expression increased rapidly.
[0179] Similar observations were confirmed from another set of IVTmRNA containing the GFP-encoding sequence (EGFP) and α-globin 3'UTR (GBA) widely used for mRNA stabilization. As shown in Figures 4D and 4F, inclusion of eK5 substantially increased protein production from the α-globin 3′UTR-containing mRNA, regardless of its position within the 3' untranslated region. From the above results, K5 was activated in all modalities tested, including plasmids, AAV vectors, and synthetic mRNA, thereby confirming its broad regulatory activity and therapeutic potential.
[0180] 6. Induction of Mixed Tailing of K5 by TENT4
[0181] Expression of the extended protein in the time-course experiment using synthetic mRNA transfection (Figure 4E) confirmed that K5 acts by increasing the stability of mRNA at least in the cytoplasm. Eukaryotic mRNA stability is mainly determined at the deadenylation stage. Therefore, to understand the mechanism of K5, the length of the poly(A) tail was monitored through high-resolution poly(A) tail analysis (Hire-PAT). Hire-PAT used RT-PCR with a reverse primer that binds to the conjugate part between the gene-specific forward primer and the poly(A) and G / I sequences following G / I tailing. As shown in Figure 5A, it was confirmed that K5 increases the length of the steady-state poly(A) tail of the reporter mRNA. This implies a mechanism related to poly(A) tail regulation by suppressing deadenylation, extending the poly(A) tail, or both.
[0182] To test the possibility that the change is related to the extension of the tail catalyzed by terminal nucleotidyl transferases (TENTs), TENTs were depleted and a luciferase assay was performed with the K5 reporter construct. As shown in Figure 5B, knockdown of the TENT4 dual-specificity antibody (diabody) paralogs (TENT4A and TENT4B) specifically decreased K5 reporter expression, while other TENTs (TENT1, TENT2, TENT3A / B (also called TUT4 and TUT7), and TENT5A / B / C / D) had no significant effect on K5 activity. To further verify the involvement of TENT4, the chemical inhibitors of the TENT4 enzyme, RG7834 and the inactive control R-isomer RO0321, were used. As shown in Figure 5C, the poly(A) tail of the K5 reporter mRNA was specifically shortened by RG7834, confirming that TENT4 is actually required for K5 function.
[0183] TENT4A (also known as PAPD7, TRF4-1, and TUT5) and TENT4B (also known as PAPD5, TRF4-2, and TUT3) extend poly(A) tails with intermittent mixing of non-adenosine residues (a process known as "mixed tailing"). The resulting mixed tails prefer adenosine residues by the major deadenylase complex CCR4NOT, thus effectively inhibiting deadenylation and stabilizing the transcriptome. To measure the frequency of mixed tails and investigate their direct involvement, we developed a modified version of TAIL-seq (which we named "gene-specific TAIL-seq (GS-TAIL-seq)"). Specifically, RNA was conjugated to biotin and ligated to a partially fragmented 3' adapter. 3'-end fragments were enriched using streptavidin beads, reverse-transcribed with a primer that binds to the adapter, and then amplified by PCR with a gene-specific forward primer. The sequencing data confirmed that the K5 reporter mRNA mainly had non-adenosine residues at the terminal and second-to-last positions, as expected for mixed tails. As shown in Figure 5D, the frequency of mixed tailing decreased after RG7834 treatment, confirming that K5 induces mixed tailing via TENT4. As shown in Figure 5F, the GS-TAIL-seq data further confirmed the Hire-PAT data shown in Figure 5C by showing that the poly(A) tail of the K5 reporter was shortened in RG7834-treated cells.
[0184] Also, as shown in Figures 5E - 5G, when HeLa and HCT116 cells were treated with RG7834, the luciferase activity and mRNA abundance of the K5 and eK5 reporters decreased. Inactivating mutations of K5 and eK5 containing a single G deletion (K5m and eK5m) were not significantly affected by RG7834, demonstrating specificity. These results together support the mechanism by which K5 acts via mixed tailing catalyzed by TENT4.
[0185] However, intriguingly, we confirmed that K5 was maintained in a fully activated state in the absence of ZCCHC14, an adaptor protein known to recruit TENT4 to viral RNA. As shown in Figure 5G, we confirmed that ZCCHC14 was dispensable for K5 activity in both reporter expression and tail elongation. This lack of ZCCHC14 dependence suggested that other factors that recognize K5 might exist.
[0186] 7. Identification of a host factor (ZCCHC2) for K5
[0187] To identify potential K5 adaptors, we performed the "RNA-protein interaction detection (RaPID)" method. As shown in Figure 5H, in vitro transcribed mRNA (IVTmRNA) containing eK5 and the BoxB element was transfected into cells stably expressing BASU, a λN peptide-fused biotin ligase. After 16 hours, the cells were treated with biotin for 1 hour to biotinylate BASU-associated proteins, followed by cell lysis, streptavidin capture, and mass spectrometry of biotinylated proteins. As shown in Figure 5H, we identified two cytoplasmic proteins, ZCCHC2 and DNAJC21, with nucleic acid binding GO terms, in proteins enriched in eK5-containing mRNA compared to control RNA lacking eK5 (Figure 5, Table 9).
[0188]
Table 9A
Table 9B
Table 9C
Table 9D
[0189] The TENT4 complex that can be obtained by in vitro RNA-pulldown experiments using the orthogonally oriented HCMV 1E stem-loop (SL-2.7) as bait was examined. As a result, in addition to TENT4A, TENT4B, ZCCHC14, SAMD4A, and K0355, which are known to interact with 1E, ZCCHC2 was also found (Figure 9). Although the intensity of ZCCHC2 is low and not necessary for 1E activity, ZCCHC2 was specifically enriched in the pulldown experiment, and it was predicted that ZCCHC2 could be an unknown component of the TENT4 complex. In particular, ZCCHC2 was the only protein that was generally enriched in both the RaPID and RNA pulldown experiments.
[0190] To verify the interaction between ZCCHC2 and eK5, Western blotting was performed following the RaPID experiment to detect ZCCHC2 associated with the eK5 bait (Figure 5I). Note that although TENT4A is slightly rich, it shows that the stability of the association of TENT4A with eK5 is inferior to that of ZCCHC2.
[0191] 8. Characterization of ZCCHC2
[0192] As shown in Figure 6A, ZCCHC2 is a 126 kDa protein with a long region, a PX domain, and a CCHC-type zinc finger (ZnF2) domain in a substantially irregular state, and its characteristics are not well confirmed. ZCCHC2 has a low association with ZCCHC14, and it was confirmed that it does not have a SAM domain that is known to interact with the CNGGN pentaloop in 1E and PRE. In addition, the C. legans gls-1 protein was predicted to be associated with ZCCHC2 even when lacking the PX or ZnF domain. gls-1 was previously confirmed to interact with GLD-4, a homolog of TENT4.
[0193] To test whether ZCCHC2 binds to TENT4, co-immunoprecipitation was performed. As shown in Figure 6B, ZCCHC2 co-immunoprecipitated with antibodies against TENT4A and TENT4B in parental HeLa cells, but not in TENT4A / B double knockout cells. This interaction was detected under RNaseA-treated conditions, indicating an RNA-independent interaction between TENT4 and ZCCHC2. As shown in Figure 6C, intracellular fractionation showed that ZCCHC2 was limited to the cytoplasm, suggesting that ZCCHC2 forms a cytoplasmic complex with TENT4. For reference, the TENT4 protein was distributed in both the nucleus and cytoplasm, with TENT4A mainly in the cytoplasm and TENT4B mainly in the nucleus. In addition, RT-qPR (RIP-QPCR) was performed following RNA immunoprecipitation using a HeLa cell line stably expressing EGFP at eK5 within the 3' untranslated region. As shown in Figure 6D, ZCCHC2 specifically interacted with the mRNA of EGFP containing eK5, reconfirming the RaPID and RNA pull-down results shown in Figures 5H and 5I. From the above results, it was confirmed that ZCCHC2 interacts with both K5 and TENT4.
[0194] Next, to investigate the function of ZCCHC2 in K5-mediated regulation, the ZCCHC2 gene was excised in HeLa cells using CRISPR-Cas9. Hire-PAT analysis was performed by the knockout (KO) to examine the length distribution of the poly(A) tail. As shown in Figure 6E, the poly(A) tail of the K5 reporter mRNA was decreased in ZCCHC2 KO cells compared to parental cells. In contrast, the K5 mutation did not become shorter in ZCCHC2 KO cells and was confirmed to have a short tail in parental cells. Similar results were confirmed in the eK5 construct, confirming that ZCCHC2 is important for the effect of tail elongation. Also, as shown in Figure 6F, gene-specific TAIL-seq experiments have shown that ZCCHC2 KO results in a decrease in mixed tailing, thereby confirming that ZCCHC2 is required for the mixed tailing of the K5 reporter mRNA.
[0195] Similarly, luciferase assays and RT-qPCR using the K5 reporter confirmed that K5 could no longer improve reporter expression in the absence of ZCCHC2. The above results were confirmed using the longer eK5 construct. As shown in Figure 6G, it was confirmed that, unlike the RG7834 parental cells, K5 reporter expression was not significantly affected in ZCCHC2 KO cells. From the above results, it was confirmed that ZCCHC2 is an important element for K5, and that this function of ZCCHC2 requires the activity of TENT4.
[0196] To verify the role of ZCCHC2, a rescue experiment was performed by transfecting a ZCCHC2-expressing plasmid into ZCCHC2 KO cells. As shown in Figure 6H, it was confirmed that the ectopic expression of ZCCHC2 increased luciferase expression in the K5 and eK5 constructs, but not in the mutants. Therefore, it was confirmed that ZCCHC2 is actually a core element that mediates the function of K5. Mutations introduced into the ZnF2 domain of ZCCHC2 did not rescue the KO cells, indicating that the RNA-binding motif plays an important role. Also, as shown in Figure 6A, a deletion mutation (ΔN) lacking the N-terminal 200 amino acids containing a highly homologous region (hereinafter also referred to as "HS") between ZCCHC2 and the related proteins ZCCHC14 and gls-1 was generated. As shown in Figure 6I, the ΔN mutation failed to rescue the defects of ZCCHC2 KO cells, indicating that the N-terminal of ZCCHC2 has an important function.
[0197] To further confirm the direct activity of ZCCHC2 on the target RNA, a tethering experiment was performed using a luciferase reporter containing the BoxB element instead of K5. As shown in Figure 6J, when the ZCCHC2 protein was tethered by the λN tag, the expression of the reporter was specifically upregulated. As a control, the expression decreased when the TNRC6B protein was attached. As shown in Figures 6I and 6K, it was confirmed that the ZCCHC2 ZnF mutant, which was inactive in the rescue experiment, was operative when tethered to the reporter RNA via the λN-BoxB system. From the above results, it was confirmed that ZnF2 acts only as an RNA-binding module and is not essential for the activation function.
[0198] Next, it was confirmed which part of ZCCHC2 is responsible for TENT4 recruitment. As shown in Figure 6A, two deletion mutants of FLAG-tagged ZCCHC2 were generated, one of which had a C-terminal deletion (ΔC, containing amino acids 1-375 at the N-terminus) and the other was prepared to have an N-terminal deletion (ΔN, containing amino acids 201-1,178). As shown in Figure 6L, the anti-FLAG antibody co-precipitated both TENT4A and TENT4B in cells expressing the full-length and ΔC ZCCHC2 proteins, confirming the interaction between TENT4 and ZCCHC2. From the above results, it was confirmed that the C-terminal part containing the PX and ZnF2 domains is essential for TENT4 binding. In particular, as shown in Figures 6I and 6L, ΔN was unable to interact with TENT4A or TENT4B, indicating that TENT4 can be recruited via the ZCCHC2 N-terminus. The N-terminal part contains the HS region, and it was confirmed that the HS region has a sequence similar to the GLD4-binding region of gls-1, a distant homolog of ZCCHC2 in C. elegans (Figure 6A). Therefore, it was confirmed that the HS region can constitute a previously undefined conserved domain that mediates protein-protein interactions.
[0199] From the above results, it was confirmed that ZCCHC2 interacts with TENT4 and K5 using its N-terminus and C-terminus, respectively. As shown in Figure 7, it was confirmed that the interaction can induce mixed tailing via the recruitment of TENT4 to K5. In addition, it was confirmed that a long poly(A) tail can promote translation by recruiting the well-established cytoplasmic poly(A)-binding protein (PABPC) to interact with eIF4G, a component of the eukaryotic translation initiation factor complex (eIF4F). Without being mutually exclusive, it was further confirmed that an unknown factor can be involved in the translation activation induced by K5 and ZCCHC2.
[0200] From the above description, those skilled in the art to which the present invention pertains will be able to understand that the present invention can be implemented as other specific examples without changing its technical idea and essential features. In this regard, it should be understood that the aforementioned experimental examples and examples are illustrative in all aspects and not limiting. The scope of the present invention should be construed as including the meaning and scope of the claims described below, as well as all changes or modified forms derived from the equivalent concepts thereof, rather than the above detailed description.
Claims
1. (i) the nucleotide sequence of a segment of the Aichi virus 1 gene (NCBI Reference Sequence: NC_001918.1) or the nucleotide sequence of its RNA; or (ii) a nucleotide sequence having at least 90% identity therewith; or (iii) a nucleotide sequence within the 3'-untranslated region of a kobuvirus having at least 50% homology with the nucleotide sequence of (i), which is a regulatory element for enhancing RNA stability or mRNA translation, wherein the segment comprises more than 110 and not more than 250 nucleotides continuously in the 5'-direction from nucleotide 8251 within the Aichi virus 1 gene, which is a regulatory element for enhancing RNA stability or mRNA translation.
2. The regulatory element according to claim 1, wherein the nucleotide sequence of (i) is the nucleotide sequence of SEQ ID NO: 20, 94, or 95 or the nucleotide sequence of its RNA.
3. The regulatory element according to claim 1, wherein the nucleotide sequence of (iii) is a nucleotide sequence having at least two hairpin structures within the 3'-untranslated region of a kobuvirus.
4. The regulatory element according to claim 1, wherein the nucleotide sequence of (iii) is any one of the nucleotide sequences of SEQ ID NOs: 98 to 140 or the nucleotide sequence of its RNA.
5. The regulatory element according to claim 1, wherein the nucleotide sequence of (ii) is a nucleotide sequence having substitution, deletion, or both of any one or more nucleotides from positions 1 to 14 of the nucleotide sequence of SEQ ID NO: 20 or the nucleotide sequence of its RNA.
6. The regulatory element according to claim 1, wherein the regulatory element enhances RNA stability and mRNA translation to increase protein expression.
7. The regulatory element according to claim 1, wherein the regulatory element interacts with ZCCHC2 that interacts with TENT4 to induce an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both.
8. A construct comprising a gene of a target protein and the regulatory element according to any one of claims 1 to 7 in its 3'-untranslated region.
9. The construct according to claim 8, wherein the target protein is selected from a reporter, a bioactive peptide, an antigen, or an antibody or a fragment thereof.
10. The construct according to claim 8, wherein the construct is an mRNA construct.
11. A vector comprising the construct according to claim 8.
12. A recombinant host cell comprising the construct according to claim 8 or a vector comprising the construct.
13. The recombinant host cell according to claim 12, wherein the host cell further comprises ZCCHC2 or a gene encoding the same, TENT4 or a gene encoding the same, or a combination thereof.
14. A composition comprising the construct according to claim 8, a vector comprising the construct, or a recombinant host cell comprising the construct or the vector.
15. The composition according to claim 14, which is for the prevention or treatment of a disease or for the production of an mRNA construct or a target protein.
16. The construct or vector further comprises a gene encoding ZCCHC2, a gene encoding TENT4, or a combination thereof, or The recombinant host cell or composition according to claim 14, further comprises ZCCHC2 or a gene encoding the same, TENT4 or a gene encoding the same, or a combination thereof.
17. A composition for enhancing RNA stability or mRNA translation, comprising ZCCHC2 or a gene encoding the same, which interacts with the regulatory element according to any one of claims 1 to 7.
18. The composition according to claim 17, which induces an increase in the length of the poly(A) tail, an increase in the stability of the poly(A) tail, or both, to enhance RNA stability or mRNA translation.
19. The composition according to claim 17, further comprising TENT4 or a gene encoding the same.
Citation Information
Patent Citations
Genetic elements that drive circular RNA translation and methods of use
JP2023532663A
Genetic elements driving circular RNA translation and methods of use
WO2021263124A2