Target-molecule-specific functional protein control system

The target molecule-specific functional protein regulatory system addresses the limitations of existing systems by using a fusion protein with antibody variable domains to regulate functional proteins in response to any target molecule within a cell, enabling efficient cell control and gene expression.

WO2025100390A1PCT designated stage expired Publication Date: 2025-05-15KYOTO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/039227
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-06
Filing Date
2024-11-05
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

Existing systems are inadequate for specifically responding to any molecule within a cell, as they require two antibodies that bind to different locations on a target molecule, making it difficult to identify non-interfering antibodies and limiting their applicability to molecules with small surface areas or those outside the cell.

Method used

A target molecule-specific functional protein regulatory system is developed, where a fusion protein is expressed intracellularly with the light chain and heavy chain variable domains of an antibody specific to a target molecule bound to the N-terminal fragment of a divided functional protein, allowing for regulation of functional proteins in response to any target molecule.

Benefits of technology

This system enables detection of any molecule within a cell without destroying the cell and allows for cell control through gene expression, including molecular detection, sensor development, cell separation, and gene therapy, with simpler operations compared to existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024039227_15052025_PF_FP_ABST
    Figure JP2024039227_15052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a target-molecule-specific functional protein control system comprising either: (a) a first nucleic acid molecule including a nucleic acid sequence which codes for a first fusion protein that includes a first domain of a functional protein and a first target binding domain including a light chain variable domain of an antibody which specifically recognizes a target molecule and (b) a second nucleic acid molecule including a nucleic acid sequence which codes for a second fusion protein that includes a second domain of the functional protein and a second target binding domain including a heavy chain variable domain of the antibody which specifically recognizes the target molecule; or (c) the first fusion protein and (d) the second fusion protein, wherein the framework region of one of the light chain variable domain and the heavy chain variable domain includes the sequence of the framework region of trastuzumab or a mutation sequence thereof, and the target molecule, the first target binding domain, and the second target binding domain form a ternary complex.
Need to check novelty before this filing date? Find Prior Art

Description

Target molecule-specific functional protein regulatory systems

[0001] The present invention relates to a target molecule-specific functional protein regulation system and method.

[0002] Controlling gene expression in response to specific intracellular molecules is a powerful strategy for monitoring cellular states and controlling cellular programs. In nature, metabolic systems use transcriptional and translational regulators to monitor cellular metabolite consumption and precisely regulate metabolic pathways.

[0003] Artificial molecular systems capable of controlling gene expression in response to intracellular molecules have been developed. A technique for controlling a split RNA polymerase by adding an external compound is known (see, for example, Patent Document 1). A technique for performing transcription or translation in response to intracellular molecules using two antibodies that bind to target molecules is known (see, for example, Non-Patent Documents 1 to 3). A technique for performing genome editing in response to intracellular miRNA (see, for example, Non-Patent Document 4) and a technique for performing genome editing in response to extracellular molecules (see, for example, Non-Patent Document 5) are also known.

[0004] WO 2017 / 212400 A2

[0005] Cao, J., Zhong, N., Wang, G. et al. Nanobody-based sandwich reporter system for living cell sensing influenza A virus infection. Sci Rep 9, 15899 (2019).Hideyuki Nakanishi, Hirohide Saito, and Keiji Itaka, Versatile Design of Intracellular Protein-Responsive Translational Regulation System for Synthetic mRNA. ACS Synthetic Biology 11 (3), 1077-1085 (2022)Cella, F., Wroblewska, L., Weiss, R. et al. Engineering protein-protein devices for multilayered regulation of mRNA translation using orthogonal proteases in mammalian cells. Nat Commun 9, 4392 (2018).Moe Hirosawa and others, Cell-type-specific genome editing with a microRNA-responsive CRISPR-Cas9 switch, Nucleic Acids Research, Volume 45, Issue 13, 27 July 2017, Page e118,Toni A. Baeumler et al., Engineering Synthetic Signaling Pathways with Programmable dCas9-Based Chimeric Receptors. Cell Report, 20, 2639-2653, 2017

[0006] However, all of the above systems were insufficient as systems capable of responding to any intracellular molecule. That is, Patent Document 1 failed to achieve molecule-dependent control of split RNA polymerase present within cells. Non-Patent Documents 1 to 3 successfully achieved transcription or translation in response to intracellular molecules, but all required two antibodies that bind to the target molecule. When using two antibodies that bind to a target molecule, the antibodies must be bound to two different sites on the surface of the target molecule and designed so that they do not interfere with each other. There is no established method for efficiently identifying two antibodies that do not interfere with each other, making their preparation difficult. Furthermore, it is difficult to respond to molecules with small molecular surface areas, such as small molecules. Non-Patent Document 4 is a system that responds to intracellular miRNAs and cannot respond to other molecules. Non-Patent Document 5 is a system that can only respond to extracellular molecules and cannot respond to intracellular molecules.

[0007] It is necessary to establish a universal platform that can specifically respond to any molecule within a living cell.

[0008] As a result of extensive research, the present inventors have come up with the idea of ​​controlling a functional protein in response to an arbitrary target molecule within a cell by expressing a fusion protein in which the light chain variable domain and heavy chain variable domain of an antibody that specifically recognizes an arbitrary target molecule are linked to the N-terminal fragment and C-terminal fragment, respectively, of a split functional protein, and have thus completed the present invention.

[0009] That is, the present invention includes the following: [1] A target molecule-specific functional protein regulatory system comprising the following (a) and (b): (a) a first nucleic acid molecule comprising a nucleic acid sequence encoding a first fusion protein comprising a first target-binding domain comprising a light-chain variable domain of an antibody that specifically recognizes the target molecule and a first domain of a functional protein, (b) a second nucleic acid molecule comprising a nucleic acid sequence encoding a second fusion protein comprising a second target-binding domain comprising a heavy-chain variable domain of an antibody that specifically recognizes the target molecule and a second domain of a functional protein, or the following (c) and (d): (c) the first fusion protein, (d) the second fusion protein, wherein the first domain of the functional protein is either the N-terminal domain or the C-terminal domain of the functional protein, and the second domain of the functional protein is the other of the N-terminal domain and the C-terminal domain of the functional protein, and the framework region of either the light-chain variable domain or the heavy-chain variable domain comprises the sequence of the framework region of trastuzumab or a mutated sequence thereof, A system in which the target molecule, the first target-binding domain, and the second target-binding domain form a ternary complex. [2] The system described in [1], in which the functional protein includes a detection protein, a protein that controls RNA transcription and translation, an extracellularly secreted protein, a genome-editing protein, an enzyme, and a therapeutic protein. [3] The system described in [1], in which (e) a third nucleic acid molecule controlled by the functional protein is further included. [4] The system described in [3], in which the functional protein is a nucleic acid polymerase, and the third nucleic acid molecule includes a promoter sequence specific to the nucleic acid polymerase and a nucleic acid sequence encoding a target protein or a target RNA. [5] The system described in [4], in which the target protein includes a detection protein, a protein that controls RNA transcription and translation, an extracellularly secreted protein, a genome-editing protein, an enzyme, and a therapeutic protein.[6] The system described in [3], wherein the functional protein is a nucleic acid polymerase, and the third nucleic acid molecule comprises a promoter sequence specific to the nucleic acid polymerase and a nucleic acid sequence encoding a guide RNA that specifically recognizes a molecule to be genome edited. [7] The system described in [1], wherein the first or second nucleic acid molecule is a DNA construct or a synthetic mRNA molecule. [8] A method for regulating a functional protein specifically to a target molecule in a cell, comprising a step of introducing the system described in [1] into a cell. [9] The method described in [8], further comprising (e) a step of introducing into a cell a third nucleic acid molecule or an additional protein that is regulated by the functional protein.

[0010] The present invention enables the detection of any molecule, such as a protein, peptide, RNA, or small molecule, within a cell or cell-free extract. Intracellular detection can be performed without destroying the cell. Furthermore, cell control through gene expression after molecular detection is possible (e.g., cell identification, isolation / removal, genome editing, etc.). Furthermore, these operations can be completed simply by introducing genes into cells, simplifying the process and handling compared to existing similar technologies. Further applications include the development of molecular detection methods and sensors within living cells, cell isolation and control technologies, methods for target-cell-specific gene expression and gene therapy technologies, and the design of genetic circuit components in living cells, leading to the development of cell manipulation and cell production technologies.

[0011] Figure 1A shows a schematic diagram of the design and strategy of TdRNAP (target-dependent RNA polymerase). Split T7 RNAP assembles into a functional RNAP upon interaction of the fused VH and VL domains with a molecular target. The activated RNAP can transcribe a gene of interest (GOI) under the control of a T7 promoter, resulting in a variety of outputs and applications. Figure 1B shows a schematic diagram of the plasmids used to test GCN4-dRNAP in 293FT cells. N- and C-terminal split RNAP fragments (T7N and T7C) were fused to the VL and VH domains of an anti-GCN4 antibody, respectively. Figure 1C shows fluorescent images of 293FT cells transfected with GCN4-dRNAP and induced with EGFP or EGFP-GCN4. The scale bar represents 100 μm. Figure 1D shows the dosage-dependent transcriptional activation of GCN4-dRNAP. 293FT cells were transfected with reporter plasmids encoding firefly luciferase and Renilla luciferase under the control of a T7 promoter and a constitutive promoter, respectively. Cell luminescence was then analyzed. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 1E shows the dependence of antibody binding affinity on the transcriptional activity of GCN4-dRNAP (left panel), the CDR-H3 sequence, and the reported dissociation constant (Kd) of each VH mutant (right panel). For each VH mutant, induction by EGFP-GCN4 was compared with induction by EGFP, and the fold change was calculated. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 2A shows antibody replacement of TdRNAP with anti-FLAG and anti-EGFP antibodies to yield FLAG-dRNAP and EGFP-dRNAP. Figure 2B shows the structural stabilization of the framework region of the VL domain by CDR loop grafting and its effect on transcriptional activity. The CDR loops of the VL domain of an anti-FLAG antibody were grafted onto the framework regions of a mutated humanized trastuzumab (CDR-grafted VL domain).Values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by unpaired two-tailed t-test. *P < 0.05, **P < 0.01, ns indicates no significant difference (P > 0.05). Figure 2C shows the dosage- and affinity-dependent transcriptional activation of FLAG-dRNAP. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 2D shows the target-specific and dosage-dependent transcriptional activation of EGFP-dRNAP. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 2E shows the target specificity of GCN4-, FLAG-, and EGFP-dRNAP, as determined by induction with the corresponding molecular targets. Values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by one-way ANOVA with Bonferroni correction. *P < 0.05, **P < 0.01, ns indicates no significant difference (P > 0.05). Figure 3A shows the design of RNA- and small molecule-dependent RNAPs using anti-HCV IRES RNA and anti-fluorescein antibodies. Figure 3B shows fluorescence images of 293FT cells transfected with HCV IRES RNA-dRNAPs and induced with monocistronic or bicistronic constructs containing the HCV IRES. The scale bar represents 100 μm. Figure 3C shows a comparison of tdTomato fluorescence intensity between cells induced with monocistronic and bicistronic constructs. Fluorescence intensity was quantified from the integrated density of tdTomato fluorescence. Values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by unpaired two-tailed t-test. *P < 0.05, **P < 0.01, ns indicates no significant difference (P > 0.05). Figure 3D shows fluorescence images of 293FT cells transfected with fluorescein-dRNAP and induced with 0, 50, or 100 μg / ml of fluorescein. The scale bar represents 100 μm. Figure 3E shows the dose-dependent transcriptional activation of fluorescein-dRNAP.293FT cells were transfected with reporter plasmids encoding firefly luciferase and Renilla luciferase under the control of T7 and constitutive promoters, respectively. Values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by one-way ANOVA with Dunnett's multiple comparison test against 0 μg / mL fluorescein. *P < 0.05. Figure 4A shows the design of a transcription amplification system using CGG RNAP as an additional output of GCN4-dRNAP. This system includes a T7 promoter-driven CGG RNAP plasmid and a dual-promoter reporter plasmid encoding a luciferase reporter gene under the control of both the T7 and CGG promoters. Activated GCN4-dRNAP induces expression of both CGG RNAP and the reporter gene. Reporter gene expression is further enhanced by the induced CGG RNAP. Figure 4B shows validation of the amplification system in 293FT cells. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 4C shows the design of an orthogonal gene circuit consisting of EGFP-dependent T7 RNAP and GCN4-dependent CGG RNAP. Each TdRNAP independently controls the expression of a luciferase reporter gene under its corresponding promoter. EGFP-dependent T7 RNAP and GCN4-dependent CGG RNAP control the expression of firefly and Renilla luciferase, respectively. Figure 4D shows simultaneous monitoring of target-dependent reporter expression in the same 293FT cells. RLU is relative light unit. Values ​​represent the mean ± SE of n = 3 biological replicates. Figure 5A shows the application of TdRNAP to target-dependent genome editing in human cells. When cells express GCN4-fused EGFP (EGFP-GCN4), GCN4-dRNAP induces the expression of a guide RNA (gRNA), which then targets and knocks out the EGFP gene by forming a complex with Cas9. FIG. 5B shows the establishment of 293FT cell lines into which EGFP or EGFP-GCN4 was introduced at the AAVS1 locus.GCN4-dEGFP knockout appears to be preferentially induced in cells expressing EGFP-GCN4 rather than EGFP. Figure 5C shows flow cytometry analysis of GCN4-dependent EGFP knockout in an established 293FT cell line. Figure 5D shows the EGFP-negative cell population measured by flow cytometry in an EGFP knockout experiment using a T7 promoter for GCN4-dependent gRNA expression and 200 ng of Cas9 expression plasmid. Values ​​represent the mean ± SE of n = 7 biological replicates. Statistical analysis was performed by unpaired two-tailed t-test. **P < 0.01, ns indicates no significant difference (P > 0.05). Figure 5E shows the EGFP-negative cell population measured by flow cytometry in an EGFP knockout experiment using a U6 promoter for constitutive gRNA expression and 200 ng of Cas9 expression plasmid. Values ​​represent the mean ± SE of n = 5 biological replicates. Statistical analysis was performed by unpaired two-tailed t-test. **P < 0.01, ns indicates no significant difference (P > 0.05). Figure 6 shows the fusion pattern and linker length of TdRNAP. (a) shows the transcriptional response of GCN4-dRNAP in 293FT cells. The fusion orientation of the VH and VL domains to each domain of split T7 RNAP is indicated below the graph. Fifty induction plasmids were used, and each variable domain was fused with a 3xGS linker. RNAP(-) indicates transfection with an empty plasmid instead of the split T7 RNAP fragments (T7N and T7C). (b) shows the transcriptional response of GCN4-dRNAP in 293FT cells. The linker length between the variable domain and each half of the split T7 RNAP was varied as indicated. For both (a) and (b), values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by one-way ANOVA with Bonferroni correction. ns indicates no significant difference (P > 0.05). Figure 7 shows the evaluation of FLAG-dRNAPs with original and CDR loop-grafted variable domains.(a) Fluorescence images of 293FT cells transfected with the anti-FLAG antibody FLAG-dRNAP, consisting of the original VH and VL domains. The transfected cells were induced with EGFP, EGFP-1xFLAG, or EGFP-3xFLAG. The scale bar represents 100 μm. (b) shows the investigation of the VH and VL domain frameworks by CDR loop grafting. The CDR loops of the anti-FLAG antibody VH and VL domains were grafted onto the framework of mutant humanized trastuzumab (grafted VH and VL). Values ​​represent the mean ± SE of n = 3 biological replicates. Statistical analysis was performed by unpaired two-tailed t-test. *P < 0.05, ns indicates no significant difference (P > 0.05). Figure 8 shows the analysis of HCV IRES RNA-dependent transcriptional activation of HCV IRES RNA-dRNAP. (a) tdTomato mean fluorescence intensity measured by flow cytometry. Transfected 293FT cells were induced with an empty plasmid (control), a monocistronic construct, or a bicistronic construct. (b) Comparison of luciferase expression levels from the monocistronic and bicistronic constructs. For (a) and (b), values ​​represent the mean ± SE of n = 3 biological replicates. For (b), statistical analysis was performed by unpaired two-tailed t-test. ns indicates no significant difference (P > 0.05). Figure 9A shows the design of Hsp70-dRNAP and its application to monitor changes in endogenous Hsp70 expression levels. Transfected 293FT cells were stimulated at 42°C for 1 hour to increase endogenous Hsp70 expression. Figure 9B shows validation of Hsp70 promoter activation after heat shock stimulation using a plasmid encoding EGFP under the Hsp70 promoter. The scale bar represents 100 μm. Figure 9C shows fluorescence images of stimulated and unstimulated 293FT cells transfected with Hsp70-dRNAP or split RNAP lacking the VH and VL domains. The scale bar represents 100 μm.Figure 9D shows a comparison of tdTomato fluorescence intensity between stimulated and unstimulated 293FT cells transfected with Hsp70-dRNAP or split RNAP lacking the VH and VL domains. Fluorescence intensity was quantified from the integrated density of tdTomato fluorescence. Values ​​represent the mean ± SE of n = 3 biological replicates. Grubbs' test was used to detect and exclude outliers (α = 0.05). Figure 10A shows the alignment of the first target-binding domain (VL domain) used in the examples. Figure 10B shows the alignment of the second target-binding domain (VH domain) used in the examples.

[0012] Hereinafter, embodiments of the present invention will be described, but the present invention is not limited to the embodiments described below.

[0013] According to one embodiment, the present invention relates to a target molecule-specific functional protein regulation system, which comprises the following (a) and (b): (a) a first nucleic acid molecule comprising a nucleic acid sequence encoding a first fusion protein comprising a first target binding domain comprising a light chain variable domain of an antibody that specifically recognizes a target molecule and a first domain of a functional protein, (b) a second nucleic acid molecule comprising a nucleic acid sequence encoding a second fusion protein comprising a second target binding domain comprising a heavy chain variable domain of an antibody that specifically recognizes the target molecule and a second domain of a functional protein, or the following (c) and (d): (c) the first fusion protein, (d) the second fusion protein, wherein the first domain of the functional protein is either the N-terminal domain or the C-terminal domain of the functional protein, and the second domain of the functional protein is the other of the N-terminal domain and the C-terminal domain of the functional protein, and the framework region of either the light chain variable domain or the heavy chain variable domain comprises the sequence of the framework region of trastuzumab or a mutated sequence thereof, The system is one in which the target molecule, the first target binding domain, and the second target binding domain form a ternary complex.

[0014] The system according to this embodiment may optionally further include the following (e): (e) a third nucleic acid molecule controlled by the functional protein.

[0015] The term "target molecule-specific functional protein control system" particularly relates to a system that controls the activity of a functional protein in specific response to a target molecule. Controlling the activity of a functional protein in specific response to a target molecule means that the first and second domains of the functional protein are controlled to be reconstituted only in the presence of the target molecule, and the functional protein retains its original activity. More specifically, this means that in the absence of the target molecule, the functional protein has no activity, and in the presence of the target molecule, the split functional protein regains its activity.

[0016] The term "target molecule" refers to a molecule present within a system, preferably within a cell, into which the nucleic acid molecules (a) and (b) have been introduced at the time of use of the system. In some embodiments, it can refer to a molecule present within the cytoplasm or nucleus of a target living cell. The target molecule is not particularly limited as long as it is a molecule that can be present transiently or steadily within a cell and can be recognized by an antibody. Specific examples of target molecules include proteins, peptides, synthetic or natural low molecular weight compounds, synthetic or natural high molecular weight compounds, compounds containing nucleic acids such as DNA, RNA, and RNP, or fragments thereof. Target molecules can be selected appropriately depending on the purpose of control. For example, when controlling the expression of a functional protein in response to intracellular conditions, the molecule may be a protein or peptide, as long as it is specific to the intracellular condition to be detected. Furthermore, when controlling by introducing a target molecule from the outside, a low molecular weight compound that easily permeates the cell membrane can be selected. Furthermore, metabolic products resulting from natural or artificial reactions within cells can also be used as target molecules. In this case, the metabolic products may be low molecular weight compounds, etc. In another embodiment, the system of the present invention can also function in systems other than cells, such as cell-free translation systems using extracts of microorganisms or cells, reaction systems in artificial cells using liposomes, etc., and diagnosis by in vitro molecular detection, and in this case the target molecules are not particularly limited.

[0017] In this embodiment, a functional protein refers to any protein that exerts its function in response to a target protein. Functional proteins may be, for example, detection proteins, proteins that control RNA transcription and translation, extracellularly secreted proteins, genome editing proteins, enzymes, therapeutic proteins, etc., but are not limited to these. These functional proteins may exert their function alone, or may be substances that exert their detection, enzymatic activity, or therapeutic function together with another substance, or may act on another substance to exert their function. The classifications of detection proteins, genome editing proteins, enzymes, and therapeutic proteins are for the sake of convenience, and substances may fall into two or more of these categories.

[0018] A detector protein refers to any protein that can be translated and display detectable information. A detector protein may be a protein that can be visualized and quantified by fluorescence, luminescence, or color development, or by the assistance of fluorescence, luminescence, or color development. Examples of fluorescent proteins include blue fluorescent proteins such as Sirius and EBFP; cyan fluorescent proteins such as mTurquoise, TagCFP, AmCyan, mTFP1, MidoriishiCyan, and CFP; green fluorescent proteins such as TurboGFP, AcGFP, TagGFP, Azami-Green (e.g., hmAG1), ZsGreen, EmGFP, EGFP, GFP2, and HyPer; yellow fluorescent proteins such as TagYFP, EYFP, Venus, YFP, PhiYFP, PhiYFP-m, TurboYFP, ZsYellow, and mBanana; and Kusabira Orange. Examples of fluorescent proteins include, but are not limited to, orange fluorescent proteins such as TurboRFP, DsRed-Express, DsRed2, TagRFP, DsRed-Monomer, AsRed2, and mStrawberry; red fluorescent proteins such as TurboFP602, mRFP1, JRed, KillerRed, mCherry, HcRed, KeimaRed (e.g., hdKeimaRed), mRasberry, and mPlum; and near-infrared fluorescent proteins such as aequorin. Examples of proteins that assist fluorescence, luminescence, or color development include, but are not limited to, enzymes that decompose fluorescent, luminescent, or color development precursors, such as luciferase, phosphatase, peroxidase, and β-lactamase. When using an RNA molecule containing a nucleic acid sequence in its translation region that encodes a protein that supports fluorescence, luminescence, or color development, it must be used in a manner that allows contact between the corresponding precursor and the protein produced by translation of the RNA molecule. For example, the precursor can be contacted with a cell into which the RNA molecule has been introduced, or the corresponding precursor can be introduced into a cell into which the RNA molecule has been introduced.

[0019] Genome editing proteins are enzymes that can act on target genes to alter their function. Altering gene function includes, for example, reducing, losing, gaining, or enhancing gene function. More specifically, "editing" refers to enzymes that can cleave genes, activate or suppress gene expression, base edit, label (genome labeling), and perform epigenetic editing such as DNA methylation / demethylation and histone acetylation. Examples of such enzymes include nucleases and inactivating nucleases. Examples of nucleases include, but are not limited to, Clustered Regularly Interspaced Short Palindromic Repeats-Associated Proteins 9 (Cas9) transcription activator-like effector nucleases (TALENs), homing endonucleases, and zinc finger nucleases. For any nuclease, analogs with similar functionality, i.e., target gene cleavage activity, or derivatives with equivalent activity can also be used. For example, analogs of the Cas9 protein include Cas9 family proteins, including, but not limited to, Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Francisella novicida Cas9 (FnCas9), Campylobacter jejuni (CjCas9), Neisseria meningitidis Cas9 (NmCas9 or NmeCas9), Geobacillus stearothermophilus Cas9 (GeoCas9), and Streptococcus thermophilus CRISPR1-Cas9 (St1Cas9). Fusion proteins in which another protein is fused to a Cas9 protein, analog, or derivative are also included in nucleases.Examples of inactivated nucleases include inactivated nucleases fused to transcriptional activators, such as inactivated Cas9 (dCas9) nucleases fused to transcriptional activator VP64, such as dCas9-VP64, dCas9-VPR, dCas9-SunTag, dCas9-VP16, dCas9-VP160, and dCas9-P300, but are not limited to these.

[0020] Examples of enzymes include, but are not limited to, nucleic acid polymerases, reverse transcriptases, transcription and translation regulatory proteins, ribosomes, DNA / RNA modifying enzymes, oxidoreductases, phosphorylation enzymes, and nucleic acid-binding proteins. Genome editing proteins are also a type of enzyme and can be referred to as genome editing enzymes. Examples of nucleic acid polymerases include T7 RNA polymerase, Pfu DNA polymerase, and M-MLV reverse transcriptase. Examples of RNAs transcribed by RNA polymerase include, but are not limited to, guideRNA, miRNA, siRNA, lincRNA, snoRNA, tRNA, and ribozymes. Examples of oxidoreductases include luciferase and dihydrofolate reductase. Examples of phosphorylation enzymes include thymidine kinase. Examples of nucleic acid-binding proteins include endonucleases and PolyA-binding proteins.

[0021] A therapeutic protein is a protein that can be used to treat, prevent, or diagnose diseases or conditions by affecting cellular function. Affecting cellular function includes increasing, decreasing, or maintaining a specific cellular function within a certain range. Examples of therapeutic proteins include, but are not limited to, cell proliferation proteins, cell death proteins, cell signaling factors, drug resistance genes, transcriptional regulators, translational regulators, differentiation regulators, reprogramming inducers, RNA-binding protein factors, chromatin regulators, membrane proteins, and fragments or complexes thereof. These proteins can also be said to be capable of displaying detectable information by affecting cellular function, and can therefore be considered both therapeutic and detection proteins. Other therapeutic proteins include, but are not limited to, enzymes, growth factors, antibodies, antigens, proteins constituting viruses or parts thereof, proteins that inhibit viral production, genome editing proteins, and fragments or complexes thereof. For example, a cell proliferation protein functions as a marker by promoting the proliferation of only cells that express it and identifying the proliferated cells. Cell death proteins cause cell death in the cells that express them, killing the cells themselves whether they contain a specific molecule (target substance) or not, and function as markers that indicate cell viability. Cell signaling factors function as markers by identifying specific biological signals emitted by cells that express them. Examples of cell death proteins include RNA-degrading enzymes such as barnase from Bacillus amyloliquefaciens, HokB, Fst, GhoT (membrane disruption), HipA (inhibition of nucleic acid elongation by phosphorylation), RelE, YafO, VapC, MazF, MqsR, PemK, HicA (endonuclease), FicT (adenylation), oc (phosphorylation), CcdB, ParE (gyrase inhibitor), Tact (inhibitor of translation), and cbtA (inhibitor of cytoskeletalExamples of such proteins include, but are not limited to, toxins such as ATPase inhibitors (ATPase inhibitors), and apoptosis-inducing proteins such as Bax and Bim. For example, translational regulatory factors function as markers by recognizing and binding to the tertiary structure of specific RNAs, thereby controlling the translation of other mRNAs into proteins. The translational regulatory factors include 5R1, 5R2 (Nat Struct Biol. 1998 Jul; 5(7):543-6), B2 (Nat Struct Mol Biol. 2005 Nov;12(11):952-7), Fox-1 (EMBO J. 2006 Jan 11;25(1):163-73), GLD-1 (J Mol Biol. 2005 Feb 11;346(1):91-104), Hfq (EMBO J. 2004 Jan 28;23(2):396-405), HuD (Nat Struct Biol. 2001 Feb;8(2):141-5), SRP19 (RNA. 2005 Jul;11(7):1043-50), and L1 (Nat Struct Biol. 2003 Feb;10(2):104-8.), L11 (Nat Struct Biol. 2000 Oct;7(10):834-7.), L18 (Biochem J. 2002 Mar 15;362(Pt 3):553-60), L20 (J Biol Chem. 2003 Sep 19;278(38):36522-30.), L23 (J Biomol NMR. 2003 Jun;26(2):131-7), L25 (EMBO J. 1999 Nov 15;18(22):6508-21.), L30 (Nat Struct Biol. 1999 Dec;6(12):1081-3.), LicT(EMBO J. 2002 Apr 15;21(8):1987-97.), MS2 coat (FEBS J. 2006 Apr;273(7):1463-75.), Nova-2 (Cell. 2000 Feb 4;100(3):323-32), Nucleocapsid (J Mol Biol. 2000 Aug 11;301(2):491-511.), Nucleolin (EMBO J. 2000 Dec 15;19(24):6870-81.), p19 (Cell. 2003 Dec26;115(7):799-811), L7Ae (RNA. 2005 Aug;11(8):1192-200.), PAZ (PiWi Argonaut and Zwille) (Nat Struct Biol. 2003 Dec;10(12):1026-32.), RnaseIII (Cell. 2006 Jan 27;124(2):355-66), RR1-38 (Nat Struct Biol. 1998 Jul;5(7):543-6.), S15 (EMBO J. 2003 Apr 15;22(8):1898-908.), S4 (J Biol Chem. 1979 Mar 25;254(6):1775-7.), S8 (J Mol Biol. 2001 Aug 10;311(2):311-24.), SacY (EMBO J. 1997 Aug 15;16(16):5019-29.), SmpB (J Biochem (Tokyo). 2005 Dec;1'38(6):729-39), snRNP U1A (Nat Struct Biol. 2000 Oct;7(10):834-7.), SRP54 (RNA. 2005 Jul;11(7):l043-50), Tat (Nucleic Acids Res. 1996 Oct 15;24(20):3974-81.), ThrRS (Nat Struct Biol. 2002 May;9(5):343-7.), TIS11d (Nat Struct Mol Biol. 2004 Mar;11(3):257-64.), Virp1 (Nucleic Acids Res. 2003 Oct 1;31(19):5534-43.), Vts1P (Nat Struct Mol Biol. 2006 Feb;13(2):177-8.), and λN (Cell. 1998 Apr 17;93(2):289-99.) are exemplified.

[0022] It should be noted that there seems to be a small formatting issue in the original text where '1' is likely a typo in '1'38' in the "SmpB (J Biochem (Tokyo). 2005 Dec;1'38(6):729-39)" part. I've translated it as it is for the purpose of following the rules.Therapeutic proteins also include proteins that directly affect cellular function. Examples include, but are not limited to, cell proliferation proteins, cell death proteins, cell signaling factors, drug resistance genes, transcriptional regulators, translational regulators, differentiation regulators, reprogramming inducers, RNA-binding protein factors, chromatin regulators, and membrane proteins. For example, cell proliferation proteins function as markers by promoting the proliferation of only cells that express them and identifying the proliferated cells. Cell death proteins cause cell death in the cells that express them, thereby killing the cells themselves and functioning as markers indicating cell viability. Examples of cell death proteins include the RNA-degrading enzyme Barnase and the apoptosis-promoting proteins Bax and Bim.

[0023] In the system according to this embodiment, (a) a first nucleic acid molecule and (b) a second nucleic acid molecule are transcribed or translated in a cell or a cell-free translation system to express (c) a first fusion protein and (d) a second fusion protein. The system according to this embodiment may include a combination of (a) a first nucleic acid molecule and (b) a second nucleic acid molecule, or a combination of (c) a first fusion protein and (d) a second fusion protein, or both. The first or second fusion protein can be designed to match the target molecule and functional protein defined above to suit the intended use of the system. That is, the functional protein can be designed so that it loses activity in the absence of a target molecule and regains its activity as an output upon input of the target molecule. For this purpose, the first fusion protein includes a first target-binding domain including the light chain variable domain (VL domain) of an antibody that specifically recognizes the target molecule, and a first domain of a functional protein. The second fusion protein comprises a second target-binding domain comprising the heavy chain variable domain (VH domain) of an antibody that specifically recognizes the target molecule, and a second domain of a functional protein.

[0024] One of the first and second domains of a functional protein may be the N-terminal domain when the functional protein is split into two, and the other may be the C-terminal domain. Thus, the first fusion protein may contain the C-terminal domain and the second fusion protein may contain the N-terminal domain, or the first fusion protein may contain the N-terminal domain and the second fusion protein may contain the C-terminal domain. The location at which the functional protein is split into the N-terminal and C-terminal domains is not particularly limited as long as the following conditions are met: the N-terminal or C-terminal domain before reconstitution does not possess the activity of the protein on its own, and the functional protein is reconstituted with the VL domain and VH domain of an antibody that specifically recognizes a target molecule. Here, "reconstitution" of a functional protein means that the split functional protein is in a state where it is likely to transition back to its original protein structure before splitting, and the function (activity) of the functional protein is maintained.

[0025] The division site of a functional protein may be, for example, a loop portion that does not constitute an α-helix or β-sheet. Cleavable sites of known functional proteins are generally widely known from literature, etc., and a person skilled in the art can determine the division site based on this information. Alternatively, a suitable division site can be confirmed by a person skilled in the art through prior experiments or simulations.

[0026] Next, the design of the first target-binding domain and the second target-binding domain will be described. The first target-binding domain comprises the VL domain of an antibody that specifically recognizes the target molecule, and the second target-binding domain comprises the VH domain of an antibody that specifically recognizes the target molecule. The VL domain contained in the first target-binding domain and the VH domain contained in the second target-binding domain can be designed based on information obtained from monoclonal antibodies against the target molecule. Monoclonal antibodies against any target molecule can be obtained based on literature information, and some are commercially available. Alternatively, the gene sequence of an antibody against a target molecule can be identified by antibody screening using phage display or by collecting a B cell fraction from the peripheral blood of a rabbit immunized with the target molecule. Alternatively, antibodies can be designed and produced using in silico CDR loop design using artificial intelligence or molecular docking simulation.

[0027] Once the sequences of the VH domain and VL domain of a monoclonal antibody against a target molecule are obtained, they can be directly incorporated into the first target-binding domain and the second target-binding domain as the sequences constituting the first target-binding domain and the second target-binding domain. Alternatively, the sequences of one or both of the VH domain and the VL domain of the monoclonal antibody against the target molecule can be modified and then incorporated into the first target-binding domain and the second target-binding domain. When modified, at least one of the VH domain and the VL domain is modified to contain an amino acid sequence corresponding to the framework region of trastuzumab, preferably humanized trastuzumab, or a mutated sequence thereof. When the VH domain and the VL domain of the monoclonal antibody against the target molecule originally have an amino acid sequence corresponding to the framework region of trastuzumab, preferably humanized trastuzumab, or a mutated sequence thereof, modification is not necessary.

[0028] The VL domain and VH domain of the monoclonal antibody each comprise, from N- to C-terminus, the first framework region (FR1), the first complementarity-determining region (CDR loop 1), the second framework region (FR2), the second complementarity-determining region (CDR loop 2), the third framework region (FR3), the third complementarity-determining region (CDR loop 3), and the fourth framework region (FR4). In one aspect of this embodiment, at least one of the VL domain and the VH domain comprises at least one, preferably two, more preferably three, and most preferably all four of the four framework regions, each of which comprises an amino acid sequence corresponding to the corresponding framework region of humanized trastuzumab or a variant thereof. Even more preferably, all of the four framework regions comprise an amino acid sequence corresponding to the corresponding framework region of humanized trastuzumab or a variant thereof.

[0029] The amino acid sequence of the VL domain of humanized trastuzumab is shown in SEQ ID NO: 41, and the amino acid sequence of the VH domain is shown in SEQ ID NO: 42. In the amino acid sequence of SEQ ID NO: 41, FR1 is the amino acid sequence occurring at positions 1 to 24, FR2 is the amino acid sequence occurring at positions 35 to 47, FR3 is the amino acid sequence occurring at positions 57 to 87, and FR4 is the amino acid sequence occurring at positions 98 to 108. In the amino acid sequence of SEQ ID NO: 42, FR1 is the amino acid sequence occurring at positions 1 to 26, FR2 is the amino acid sequence occurring at positions 36 to 47, FR3 is the amino acid sequence occurring at positions 60 to 96, and FR4 is the amino acid sequence occurring at positions 109 to 120. The amino acid numbers described herein are numbered with the N-terminus being numbered 1, assuming that the total amino acid length of the VL domain is 108 and the total amino acid length of the VH domain is 120.

[0030] The Kabat numbering system may be used to designate framework regions and CDR loops in antibodies. This is because the length of the CDR loops varies depending on the antibody, which can change the overall amino acid length of each domain. FR1 to FR4 of the VL domain shown in SEQ ID NO: 41 according to the Kabat numbering system are designated in the same way as the previous numbering system. According to the Kabat numbering system, FR1 of the VH domain shown in SEQ ID NO: 42 is the amino acid sequence located at positions 1 to 26, FR2 is the amino acid sequence located at positions 36 to 47, FR3 is the amino acid sequence located at positions 59 to 92, and FR4 is the amino acid sequence located at positions 102 to 113.

[0031] The amino acid sequence of a preferred mutant sequence of the VL domain of humanized trastuzumab (SEQ ID NO: 1) and the amino acid sequence of a preferred mutant sequence of the VH domain (SEQ ID NO: 2) are shown in Table 1 below. Hereinafter, the present invention will be described using the preferred mutant sequences shown in SEQ ID NO: 1 and SEQ ID NO: 2 as examples. These sequences are referred to herein as "mutated humanized trastuzumab sequences." In both the VL domain and the VH domain, the underlined portions represent the complementarity-determining regions, which, from the N-terminus, represent CDR loop 1, CDR loop 2, and CDR loop 3. The region on the N-terminus of CDR loop 1 is FR1; the region on the C-terminus of CDR loop 1 that is the N-terminus of CDR loop 2 is FR2; the region on the C-terminus of CDR loop 2 that is the N-terminus of CDR loop 3 is FR3; and the region on the C-terminus of CDR loop 3 is FR4. Furthermore, within the framework regions, amino acid residues that are preferably not substituted are indicated in bold, and amino acid residues that may be substituted are indicated by boxed lines.

[0032]

[0033] Thus, the first target-binding domain of this embodiment preferably comprises a VL domain of an antibody that specifically recognizes a target molecule, with FR1 of the VL domain comprising the sequence of FR1 of the VL domain of a mutated humanized trastuzumab (DIQMTQSPSS LSASVGDRVT ITCR (SEQ ID NO: 3)), and / or FR2 of the VL domain comprising the sequence of FR2 of the VL domain of a mutated humanized trastuzumab (WYQQKP GKAPKLL (SEQ ID NO: 4)), and / or FR3 of the VL domain comprising the sequence of FR3 of the VL domain of a mutated humanized trastuzumab (GVPS RFSGSGSGTD FTLTISSLQP EDFATYY (SEQ ID NO: 5)), and / or FR4 of the VL domain comprising the sequence of FR4 of the VL domain of a mutated humanized trastuzumab (FGQ GTKVEIKR (SEQ ID NO: 6)).

[0034] The second target-binding domain of this embodiment comprises a VH domain of an antibody that specifically recognizes a target molecule, wherein FR1 of the VH domain preferably comprises the FR1 sequence of the VH domain of a mutated humanized trastuzumab (EVQLLESGGG LVQPGGSLRL SCAASG (SEQ ID NO: 7)), and / or FR2 of the VH domain preferably comprises the FR2 sequence of the VH domain of a mutated humanized trastuzumab (WVRQA PGKGLEW (SEQ ID NO: 8)), and / or FR3 of the VH domain preferably comprises the FR3 sequence of the VH domain of a mutated humanized trastuzumab (Y ADSVKGRFTI SRDNSKNTLY LQMNSLRAED TAVYYC (SEQ ID NO: 9)), and / or FR4 of the VH domain preferably comprises the FR4 sequence of the VH domain of a mutated humanized trastuzumab (DYW GQGTLVTVSS (SEQ ID NO: 10)). However, another embodiment is also possible in which the sequence of FR3 of the VH domain of the mutant humanized trastuzumab is Y ADSVKGRFTI SADNSKNTLY LQMNSLRAED TAVYYC (SEQ ID NO: 11).

[0035] The sequences of the framework regions of the VL domain and VH domain of mutant humanized trastuzumab (SEQ ID NOS: 3 to 11) may be partially substituted. For example, about 50% or less, preferably about 20% or less, and more preferably about 10% or less of the sequences of each of the framework regions may be substituted.

[0036] In one embodiment, the amino acid sequence of the mutant humanized trastuzumab VL domain may be independently substituted at about 5% in FR1, about 47% in FR2, about 17% in FR3, and about 10% in FR4, and the amino acid sequence of the VH domain may be independently substituted at about 16% in FR1, about 9% in FR2, about 46% in FR3, and about 9% in FR4.

[0037] In another embodiment, in each framework region of the sequences shown in SEQ ID NOs: 1 and 2, the amino acid residues shown in boxed letters may be substituted, and the substituted amino acid residues may be any. It is preferred that the amino acid residues shown in bold are not substituted.

[0038] Meanwhile, for both the VL domain of the first target-binding domain and the VH domain of the second target-binding domain, the CDR loops can be determined by numbering the amino acid sequences of the VH domain and VL domain of an antibody against a target molecule according to the Kabat numbering system. Specifically, CDR loop 1 of the VL domain of the first target-binding domain may be the amino acid sequence located at positions 25 to 34 of the VL domain of the antibody against the target molecule, and / or CDR loop 2 may be the amino acid sequence located at positions 48 to 56 of the VL domain of the antibody against the target molecule, and / or CDR loop 3 may be the amino acid sequence located at positions 88 to 97 of the VL domain of the antibody against the target molecule. Thus, at least one, preferably two, and more preferably all of CDR loop 1, CDR loop 2, and CDR loop 3 of the VL domain of the first target-binding domain should contain a sequence derived from the VL domain of an antibody against a target molecule as designated by the Kabat numbering system. Furthermore, about 20% or less, preferably about 10% or less, and more preferably about 5% or less of these sequences may be substituted.

[0039] Similarly, CDR loop 1 of the VH domain of the second target-binding domain may be the amino acid sequence located at positions 27 to 35 of the VH domain of an antibody specific for the target molecule, and / or CDR loop 2 may be the amino acid sequence located at positions 48 to 58 of the VH domain of an antibody specific for the target molecule, and / or CDR loop 3 may be the amino acid sequence located at positions 93 to 101 of the VH domain of an antibody specific for the target molecule. Thus, at least one, preferably two, and more preferably all of CDR loop 1, CDR loop 2, and CDR loop 3 of the VH domain of the second target-binding domain may contain a sequence derived from the VH domain of an antibody specific for the target molecule as designated by the Kabat numbering described above. Furthermore, these sequences may be substituted by about 20% or less, preferably about 10% or less, and more preferably about 5% or less.

[0040] The length of the CDR loop in both the VL domain and the VH domain is not particularly limited, and may be, for example, about 3 to 30 amino acids in length, about 5 to 20 amino acids in length, or about 8 to 15 amino acids in length.

[0041] The first target-binding domain comprises a VL domain but preferably does not comprise a constant domain or a VH domain, and the second target-binding domain comprises a VH domain but preferably does not comprise a constant domain or a VL domain, and does not comprise an intermolecular disulfide bond that is typically present in the VH and VL domains of monoclonal antibodies.

[0042] In addition to the structural features described above, the first and second target-binding domains form a ternary complex with the target molecule. "Forming a ternary complex" is defined as the first and second target-binding domains each binding to the same target molecule in the presence of the target molecule, thereby maintaining a stable interaction between the first and second target-binding domains. Preferably, the ternary complex has an equilibrium dissociation constant (Kd) of 50 nM or less. The equilibrium dissociation constant (Kd) is more preferably 10 nM or less, and most preferably 0.6 nM or less. The equilibrium dissociation constant (Kd) of the ternary complex can be determined by surface plasmon resonance (SPR) or isothermal titration calorimetry (ITC).

[0043] The domains designed as described above can be combined to design a first fusion protein and a second fusion protein. The combinations of the first fusion protein and the second fusion protein that enable the system of this embodiment to function are shown in Table 2 below. In Table 2 below, the domains are listed from left to right, from the N-terminus to the C-terminus of the fusion protein. However, when the functional protein is T7 RNA polymerase, combination numbers 1 or 3 shown in Table 2 below are preferred.

[0044]

[0045] In this manner, a first domain (F1) of a functional protein and a first target-binding domain (VL) comprising an antibody VL domain can be selected or designed, and a first fusion protein can be designed and its structure determined. Similarly, a second target-binding domain (VH) comprising an antibody VH domain and a second domain (F2) of a functional protein can be selected or designed, and a second fusion protein can be designed and its structure determined. In the first and second fusion proteins, the domains can be linked directly or via a linker. The length of the linker is not particularly limited, and may be, for example, approximately 3 to 50 amino acids, approximately 3 to 30 amino acids, or approximately 3 to 15 amino acids. The amino acid sequence of the linker is not particularly limited, as long as it does not self-associate on its own, does not have strong polarity that could inhibit the formation of a ternary complex between the VH, VL, and target molecule, and is not self- or enzymatically cleaved. Furthermore, the first and second fusion proteins may contain additional amino acid sequences of about 3 to 200 amino acids in length at the N-terminus and C-terminus of each domain.

[0046] Once the structures of the (c) first fusion protein and the (d) second fusion protein have been determined, the structures of the (a) first nucleic acid molecule and the (b) second nucleic acid molecule that express them can be determined. The (a) first nucleic acid molecule and the (b) second nucleic acid molecule may each independently be a DNA construct or a synthetic RNA molecule, and both may be used as appropriate depending on the purpose and application of the invention.

[0047] For example, vectors well known and commonly used in the art can be used for the DNA construct, including viral vectors, artificial chromosome vectors, plasmid vectors, and expression systems using transposons (sometimes referred to as transposon vectors). Examples of viral vectors include retroviral vectors, lentiviral vectors, adenoviral vectors, adeno-associated viral vectors, and Sendai viral vectors. Examples of artificial chromosome vectors include human artificial chromosomes (HAC), yeast artificial chromosomes (YAC), and bacterial artificial chromosomes (BAC, PAC). Plasmid vectors can be any mammalian plasmid, and may be, for example, an episomal vector. Examples of transposon vectors include expression vectors using piggyBac transposons. It is also possible to design a DNA construct that expresses the first and second fusion proteins as separate proteins from a single expression vector.

[0048] The synthetic RNA molecule may be a combination of two mRNA molecules: one mRNA molecule comprising, in a 5' to 3' orientation, a 5'-UTR, an open reading frame comprising a nucleic acid sequence encoding a first fusion protein, and a 3'-UTR; and the other mRNA molecule comprising, in a 5' to 3' orientation, a 5'-UTR, an open reading frame comprising a nucleic acid sequence encoding a second fusion protein, and a 3'-UTR. Alternatively, the synthetic RNA molecule may be a single mRNA molecule comprising, in a 5' to 3' orientation, a 5'-UTR, an open reading frame comprising a nucleic acid sequence encoding the second fusion protein, a self-cleaving sequence encoding a self-cleaving peptide, and a nucleic acid sequence encoding the first fusion protein, and a 3'-UTR.

[0049] The 5' UTR of mRNA contains a cap structure or cap analog at the 5' end. The cap structure may be 7-methylguanosine 5' phosphate. The cap analog is a modified structure that, like the cap structure, is recognized by the translation initiation factor eIF4E. Examples include, but are not limited to, Ambion's Anti-Reverse Cap Analog (ARCA), New England Biolabs' m7G(5')ppp(5')G RNA Cap Structure Analog, and TriLink's CleanCap. The cap analog may also be another modified structure recognized by a translation initiation factor. The 3' side of the cap structure or cap analog, but the 5' side of the start codon, may contain an arbitrary nucleic acid sequence of, for example, about 0 to 500 bases, preferably about 20 to 200 bases. These arbitrary nucleic acid sequences are preferably nucleic acid sequences that do not form secondary structures and do not specifically interact with other nucleic acid molecules. It is also preferable that the 5' UTR does not contain an AUG initiation codon.

[0050] The open reading frame, which is the coding region of the mRNA, contains an initiation codon, a nucleic acid sequence encoding the first or second fusion protein, and a stop codon. The nucleic acid sequence encoding the first fusion protein contains, in order from the 3' side of the initiation codon AUG to the 5' side of the stop codon, a nucleic acid sequence encoding the N-terminal domain of the fusion protein, a nucleic acid sequence encoding an optional linker amino acid sequence, and a nucleic acid sequence encoding the C-terminal domain of the fusion protein. The nucleic acid sequences encoding each domain can be determined from the structure of the first fusion protein designed above.

[0051] The 3'UTR of the mRNA contains a PolyA tail. The PolyA tail may be a sequence of approximately 50 to 250 adenine bases. However, the approximately 50 to 250 adenine bases do not need to be consecutively linked, and other nucleic acid bases may be included between the adenine bases as long as the stability of the mRNA is maintained.

[0052] The sugar residue (ribose) of each nucleotide in any of the mRNAs described above may be modified for purposes such as reducing cytotoxicity. Examples of sites of modification in the sugar residue include substitution of the hydroxyl group or hydrogen atom at the 2', 3', and / or 4' positions of the sugar residue with other atoms. Examples of types of modification include fluorination, alkoxylation (e.g., methoxylation, ethoxylation), O-allylation, S-alkylation (e.g., S-methylation, S-ethylation), S-allylation, and amination (e.g., -NH2). Such modifications of sugar residues can be carried out by methods known per se (see, for example, Sproat et al., (1991) Nucle. Acid. Res. 19, 733-738; Cotton et al., (1991) Nucl. Acid. Res. 19, 2629-2635; Hobbs et al., (1973) Biochemistry 12, 5138-5145).

[0053] Furthermore, sugar residues in mRNA can be bridged at the 2' and 4' positions to form BNA (bridged nucleic acid) (LNA). Such sugar residue modifications can be carried out by known methods (see, for example, Tetrahedron Lett., 38, 8735-8738 (1997); Tetrahedron, 59, 5123-5128 (2003); Rahman SMA, Seki S., Obika S., Yoshikawa H., Miyashita K., Imanishi T., J. Am. Chem. Soc., 130, 4886-4896 (2008)). Furthermore, mRNA can also have nucleic acid bases (e.g., purines and pyrimidines) modified (e.g., chemically substituted). Examples of such modifications include modification of the pyrimidine at position 5, modification of the purine at positions 6 and / or 8, modification with an exocyclic amine, substitution with 4-thiouridine, and substitution with 5-bromo- or 5-iodo-uracil. Furthermore, for the purpose of reducing cytotoxicity, modified bases such as pseudouridine (Ψ), N1-methylpseudouridine (N1mΨ), and 5-methylcytidine (5mC) may be contained in place of normal uridine or cytidine. The positions of the modified bases, whether uridine or cytidine, can be all or part of the bases independently, and if part of the bases are modified, they can be located at random positions in any proportion.

[0054] To enhance resistance to nucleases and hydrolysis, the phosphate groups (e.g., terminal phosphate residues) contained in mRNA may be modified. For example, the phosphate group P(O)O may be substituted with P(O)S (thioate), P(S)S (dithioate), P(O)NR2 (amidate), P(O)R, R(O)OR', CO, CH2 (formacetal), or 3'-amine (-NH-CH2-CH2-) (wherein each R or R' is independently H or substituted or unsubstituted alkyl (e.g., methyl, ethyl)). Examples of linking groups include -O-, -N-, and -S-, and adjacent nucleotides can be linked via these linking groups.

[0055] Once the molecular structure and nucleic acid sequence of the DNA construct or mRNA have been determined as described above, those skilled in the art can synthesize the DNA construct or mRNA using any known genetic engineering method. For example, mRNA can be obtained as a synthetic mRNA molecule by in vitro transcription using a template DNA containing a promoter sequence. Furthermore, (c) the first fusion protein and (d) the second fusion protein can also be obtained from these DNA constructs or mRNA.

[0056] The (e) third nucleic acid molecule that may be optionally included in the system according to this embodiment is a nucleic acid molecule that is controlled by a functional protein to produce coding or non-coding RNA or a desired protein. The third nucleic acid molecule can be designed in relation to the functional proteins encoded by the first and second nucleic acid molecules. Similarly, an additional protein that may be optionally included in the system according to this embodiment may be a protein that can be controlled by the functional protein.

[0057] For example, if the functional protein is a nucleic acid polymerase, the third nucleic acid molecule preferably includes a promoter sequence specific to the nucleic acid polymerase and a nucleic acid sequence encoding a target protein or target RNA. The target protein and target RNA are not particularly limited and can be selected and designed to suit the purpose of the system of this embodiment. The target protein can be selected from the functional proteins exemplified above, including proteins that control gene expression, such as transcription factors and translation factors; proteins applicable to vaccines, such as viral proteins and pathogenic proteins; and proteins applicable to cell selection, such as drug-degrading proteins and cell death-inducing proteins. The target RNA may be, for example, a guide RNA (gRNA) that specifically recognizes a molecule targeted for genome editing, or a switch mRNA, shRNA, tRNA, or rRNA that specifically responds to miRNA or proteins to control translation. The promoter sequence specific to the nucleic acid polymerase is also not particularly limited, and any known combination can be used. For example, a combination of T7 polymerase and a T7 promoter, or a mutant T7 polymerase and its corresponding promoter can be used, but is not limited to these. In this case, the third nucleic acid molecule can produce a protein or RNA of interest in specific response to reconstitution of the nucleic acid polymerase.

[0058] In particular, when the functional protein is a nucleic acid polymerase and the third nucleic acid molecule is a gRNA that specifically recognizes a molecule to be subjected to genome editing, the fourth nucleic acid molecule can include a nucleic acid molecule containing a nucleic acid sequence encoding the genome editing protein. In this system, the first fusion protein and the second fusion protein associate with each other in response to the target molecule, reconstituting the nucleic acid polymerase. The nucleic acid polymerase then produces a gRNA, and the gRNA and the genome editing protein produced by the fourth nucleic acid molecule cooperate to perform genome editing of the target molecule.

[0059] For example, when the functional protein is a genome editing protein, the third nucleic acid molecule preferably comprises a nucleic acid sequence encoding a gRNA that specifically recognizes a molecule to be edited by genome editing. In this case, the third nucleic acid molecule can cleave the molecule to be edited or inactivate other genome editing proteins that are already functioning in response to the reconstitution of the genome editing protein.

[0060] For example, when the functional protein is an RNA-binding protein, the third nucleic acid molecule preferably comprises a nucleic acid sequence encoding a switch mRNA that comprises a nucleic acid sequence that specifically recognizes the RNA-binding protein. In this case, the third nucleic acid molecule may be a molecule that can bind to the switch mRNA in specific response to reconstitution of the RNA-binding protein and promote or suppress translation from the switch mRNA.

[0061] The third nucleic acid molecule may also be a DNA construct or mRNA, but may need to be a DNA construct depending on the purpose. Once the molecular structure and nucleic acid sequence of the third nucleic acid molecule have been determined, it can also be synthesized by those skilled in the art by any known method in genetic engineering, similar to the first and second nucleic acid molecules.

[0062] Next, the present invention will be described from the perspective of a method. The method according to this embodiment is a method for regulating a target molecule-specific functional protein, and includes a step of introducing the first and second nucleic acid molecules contained in the aforementioned system into a cell or a cell-free translation system. Optionally, the method may further include a step of introducing a third nucleic acid molecule regulated by the functional protein into the cell or cell-free translation system.

[0063] In the method according to this embodiment, the term "cell" is not particularly limited and may refer to any cell. The cell may be a single cell or a "cell population," which is a collection of two or more cells. There is no theoretical upper limit to the number of "cell populations," and the term refers to a population consisting of any number of cells. For example, the "cell population" may be a cell population that may contain multiple different types of cells, and may include target cells and non-target cells.

[0064] The term "cells" may refer to cells collected from unicellular or multicellular species, or may be artificially engineered cells (including cell lines). Examples of such cells include yeast, insect cells, and animal cells, with animal cells being preferred. Examples of animal cells include cells derived from mammals (e.g., mice, rats, hamsters, guinea pigs, dogs, monkeys, orangutans, chimpanzees, and humans). Mammalian cells include cell lines such as monkey COS-7 cells, monkey Vero cells, Chinese hamster ovary (CHO) cells, dhfr gene-deficient CHO cells, mouse L cells, mouse AtT-20 cells, mouse myeloma cells, rat GH3 cells, human embryonic kidney-derived cells (e.g., HEK293 cells), human hepatoma-derived cells (e.g., HepG2), and human FL cells. Primary cultured cells prepared from human and other mammalian tissues are also useful. Furthermore, zebrafish embryos and Xenopus oocytes can also be used.

[0065] There are no particular limitations on the degree of differentiation of the cells or the age of the animal from which the cells are collected, and they may be any of (A) stem cells, (B) progenitor cells, (C) terminally differentiated somatic cells, or (D) other cells. Examples of (A) stem cells include, but are not limited to, embryonic stem (ES) cells, cloned embryonic stem (ntES) cells obtained by nuclear transfer, spermatogonial stem cells (GS cells), embryonic germ cells (EG cells), and induced pluripotent stem (iPS) cells. Examples of (B) progenitor cells include tissue stem cells (somatic stem cells) such as neural stem cells, hematopoietic stem cells, mesenchymal stem cells, and dental pulp stem cells. (C) Examples of somatic cells include keratinizing epithelial cells (e.g., keratinizing epidermal cells), mucosal epithelial cells (e.g., epithelial cells of the surface of the tongue), exocrine gland epithelial cells (e.g., mammary gland cells), hormone-secreting cells (e.g., adrenal medullary cells), metabolic and storage cells (e.g., hepatocytes), luminal epithelial cells that form the interface (e.g., type I alveolar cells), luminal epithelial cells of the inner duct (e.g., vascular endothelial cells), ciliated cells with transport capacity (e.g., airway epithelial cells), cells that secrete extracellular matrix (e.g., fibroblasts), contractile cells (e.g., smooth muscle cells), cells of the blood and immune system (e.g., T lymphocytes), cells related to sensory perception (e.g., rod cells), neurons and glial cells of the central and peripheral nervous systems (e.g., astrocytes), pigment cells (e.g., retinal pigment epithelial cells), and their precursor cells (tissue precursor cells). (D) Other cells include, for example, cells that have undergone differentiation induction, including progenitor cells and somatic cells induced to differentiate from pluripotent stem cells. They may also be cells induced by so-called "direct conversion (also called direct reprogramming or trans-differentiation)," which directly differentiates somatic or progenitor cells into desired cells without going through an undifferentiated state.

[0066] The following describes the introduction of (a) a first nucleic acid molecule and (b) a second nucleic acid molecule, and / or (c) a first fusion protein and (d) a second fusion protein into a cell, as well as the introduction of an optional third nucleic acid molecule or additional protein molecule into a cell. In the method for controlling a functional protein according to this embodiment, the introduction of the first and second nucleic acid molecules into a cell can be performed in vitro or in vivo, and is not particularly limited. The first and second nucleic acid molecules can be introduced into a cell separately, or as a single vector or RNA molecule. When a third nucleic acid molecule is used, the first, second, and third nucleic acid molecules can be introduced into a cell separately, or as a single vector or RNA molecule.

[0067] In vitro introduction can be achieved by any commonly used in vitro method for introducing DNA constructs, RNA, or fusion proteins or additional proteins into cells. Examples include, but are not limited to, lipofection, polymer injection, electroporation, calcium phosphate coprecipitation, DEAE-dextran injection, microinjection, and gene gun injection. The first and second fusion proteins can be introduced into cells by transfection using protein introduction reagents such as cationic lipids, membrane-permeable peptides, polymer capsules, and polymer nanogels. In vivo introduction can be achieved by any commonly used method for introducing RNA into cells in vivo. For example, in mammals, the DNA constructs, RNA, or the first and second fusion proteins can be directly introduced into cells using intramuscular injection, subcutaneous injection, intravenous injection, intra-articular injection, or other methods.

[0068] The amounts of the first and second nucleic acid molecules, the first and second fusion proteins, and the optional third nucleic acid molecule or additional protein introduced into cells vary depending on the type of target cell and the structure of the DNA / RNA, and are not limited to any specific amount. The amount of introduction that achieves the desired target control can be determined through preliminary experiments, etc. Furthermore, the first and second nucleic acid molecules, the first and second fusion proteins, and the optional third nucleic acid molecule or additional protein can be introduced into a cell-free translation system by adding the first and second nucleic acid molecules, the first and second fusion proteins, and the optional third nucleic acid molecule or additional protein to any system capable of cell-free translation using any method.

[0069] The first and second nucleic acid molecules introduced into a cell produce a first fusion protein and a second fusion protein within the cell. Furthermore, the first and second fusion proteins introduced into the cell exist as the first and second fusion proteins within the cell. When a target molecule is present within the cell, the first target binding domain, the second target binding domain, and the target molecule form a specific ternary complex. This reconstitutes the N-terminal domain of the functional protein and the C-terminal domain of the functional protein. The reconstituted functional protein performs a predetermined function within the cell. For example, it can acquire polymerase activity to initiate RNA synthesis, emit fluorescence, and initiate genome editing with genome editing activity. On the other hand, when a target molecule is not present within the cell, the first target binding domain and the second target binding domain do not interact with each other, and the first fusion protein and the second fusion protein exist within the cell as separate proteins. In other words, the functional protein is not reconstituted. And, if the first fusion protein and the second fusion protein are not reconstituted, the functional protein will not act with or against the optional third nucleic acid molecule or additional protein.

[0070] The present invention will be described in more detail below with reference to examples, which are not intended to limit the scope of the present invention.

[0071] Experimental Methods: Plasmid Construction. Plasmids for each TdRNAP were constructed using PCR products by HiFi assembly (NEB) or by Kunkel mutagenesis using pre-constructed plasmids. Genes for the N- and C-terminal fragments of T7 RNAP were amplified from the plasmid pCAG-T7pol (Addgene plasmid #59926). The evolved T7 RNAP N-terminal fragment gene, d5-19, was generated by multiple site-directed mutagenesis. Genes for the VH and VL domains of anti-GCN4 antibody were amplified from the plasmid pHR-scFv-GCN4-sfGFP-GB1-NLS-dWPRE (Addgene plasmid #60906). Genes for anti-EGFP antibody and anti-FLAG antibody were amplified from synthetic DNA fragments (Thermo Fisher Scientific). Genes for anti-HCV IRES RNA, anti-fluorescein, and anti-Hsp70 antibodies were generated by Kunkel mutagenesis using the anti-FLAG antibody gene as template DNA. All plasmids were constructed by HiFi assembly using PCR products, and details are listed in Tables 3A, B, and C. Standard PCR and mutagenesis PCR were all performed using PrimeSTAR Max DNA Polymerase (Takara Bio). The sequence numbers of individual proteins and peptides are listed in Table 4, and detailed sequences are shown in the Sequence Listing. The sequence numbers of the base sequences of the main genes and regulatory elements used in this study are also listed in Table 4, and detailed sequences are shown in the Sequence Listing.

[0072]

[0073]

[0074] Cell Culture and Stimulation. 293FT cells (Thermo Fisher Scientific) were cultured in DMEM medium (Nacalai Tesque) supplemented with 10% FBS (Biosera), MEM Non-Essential Amino Acids Solution (Thermo Fisher Scientific), 1 mM sodium pyruvate (Sigma-Aldrich), and 1 mM L-glutamine (Thermo Fisher Scientific) in a humidified incubator at 37°C with 5% CO2. Cells less than passage 30 were used for all experiments. Different passage numbers were used for each biological replicate. For fluorescein induction experiments, fluorescein diacetate (Fujifilm Wako) dissolved in ethanol was added at a final concentration of 50 μg / mL or 100 μg / mL 16 hours after transfection. Cells were cultured in fluorescein-supplemented medium until further analysis, at which point they were washed with PBS. For heat shock experiments, 16 hours after transfection, cells were exposed to 42°C for 1 hour.

[0075] Plasmid transfection: The day before transfection, 293FT cells were plated in a 24-well plate at 1 × 10 cells per well. 5 The cells were transfected with the plasmid mixture using 2.0 μL of Lipofectamine 2000 (Thermo Fisher Scientific). The amounts of transfected plasmids are shown in Table 5A, B, and C.

[0076]

[0077] Establishment of stable cell lines using the CRISPR-Cas9 system. Stable cell lines expressing EGFP-GCN4 or EGFP were generated using the CRISPR-Cas9 system. Each gene was cloned into a donor plasmid with homology arms targeting the endogenous AAVS1 locus on the genome. 293FT cells were seeded at 1 × 10 per well in a 24-well plate the day before transfection. 5 Cells were seeded in 1000 x 1000 cells. The cells were transfected with 1 μg of SpCas9 plasmid (Addgene plasmid #41815), 1 μg of AAVS1-targeting sgRNA plasmid (Addgene plasmid #41818), and 1 μg of donor plasmid using 2.0 μL of Lipofectamine 2000 (ThermoFisher Scientific) according to the manufacturer's protocol. Two days after transfection, the cells were trypsinized and transferred to 6-well plates in medium containing 1 μg / mL puromycin (Invivogen). Fluorescent-positive cells were then sorted and enriched using a FACSymphony S6 cell sorter (BD Biosciences). The sorted cells were then cultured in medium containing 0.5 μg / mL puromycin.

[0078] Luciferase Assays. All luciferase assays were performed using the Dual-Glo Luciferase Assay System (Promega) according to the manufacturer's protocol. 293FT cells were harvested and lysed two days after transfection. Firefly and Renilla luciferase activities were measured using a GloMax Navigator Microplate Luminometer (Promega). Relative luciferase activity (RLU) was calculated by normalizing the Firefly luciferase activity of each sample to the Renilla luciferase activity. Each biological experiment included two technical replicates.

[0079] Fluorescence reporter assay. Two days after transfection, cells were imaged using a CellVoyager CQ1 (Yokogawa Electric Corporation). Images were analyzed using ImageJ software (National Institutes of Health) and its plugin (https: / / github.com / yfujita-skgcat / image_converter). For quantification, the integrated signal of tdTomato-positive cells in each image was used as the fluorescence intensity, and analyzed using Image J. To calculate the mean tdTomato fluorescence intensity, cells were analyzed using a CytoFLEX S Flow Cytometer (Beckman Coulter), and the acquired data was analyzed using FlowJo software (BD Biosciences).

[0080] Gene Knockout Using CRISPR-Cas9: For knockout experiments using CRISPR-Cas9, we used the stable cell line, EGFP-GCN4 cells, and EGFP cells. Cells were seeded in 24-well plates and transfected with the plasmids listed in Supplementary Table 4. Three days after transfection, cells were trypsinized and transferred to 6-well plates in medium containing 0.5 μg / mL puromycin. Seven days after transfection, cells were analyzed using a CytoFLEX S flow cytometer. Acquired data were analyzed to calculate the EGFP-negative population using FlowJo software and the "flowCore" package in R (https: / / bioconductor.org / packages / release / bioc / html / flowCore.html).

[0081] In silico protein structure prediction: The structure of the CDR loop-grafted variable domain of the anti-FLAG antibody was predicted using ColabFold, a Google Colab-based protein structure prediction tool using AlphaFold2 and MMseqs2, with default settings. The predicted structure was visualized using PyMOL software and aligned to the original variable domain of the anti-FLAG antibody (PDB ID: 7BG1).

[0082] [Results] We envisioned the use of antibody variable domains as fusion proteins for split RNAPs to regulate gene expression in response to various intracellular molecules. To realize this concept, we aimed to develop a system for assembling split RNAPs based on the binding between the variable domains of a single antibody and a target molecule, rather than two antibodies. This would eliminate the need to screen two antibodies for split RNAP assembly and allow us to utilize existing antibodies to control gene expression. To realize this strategy, we focused on the structure of the variable domain. Because the variable domain is composed of a heavy chain and a light chain, it can be divided into two smaller domains: the heavy-chain variable domain and the light-chain variable domain (VH and VL). We hypothesized that the binding of the VH and VL domains to the target molecule would enable split RNAP assembly (Figure 1A). We named this system "target-dependent RNAP (TdRNAP)."

[0083] To test our hypothesis, we first designed a TdRNAP using the VH and VL domains of an anti-GCN4 antibody. This antibody has nanomolar affinity for the leucine zipper peptide of the yeast transcription factor GCN4, and its variable domains have optimized framework regions that allow folding without relying on intramolecular disulfide bonds, maintaining a stable structure even in the reducing environment of cells. We fused the VH and VL domains to the C- and N-terminal fragments of split T7 RNAP, respectively, to generate a GCN4-dependent RNAP (GCN4-dRNAP) (Fig. 1B). To test the GCN4-dRNAP, we used enhanced green fluorescent protein (EGFP) fused to the GCN4 peptide (EGFP-GCN4) as a targeting molecule and examined whether EGFP-GCN4 could induce assembly of the split RNAP. Human embryonic kidney 293FT cells were cotransfected with a GCN4-dRNAP expression plasmid, a reporter plasmid encoding near-infrared fluorescent protein 670 (iRFP670) under the control of a T7 promoter, and an inducer plasmid encoding EGFP-GCN4 or EGFP. Expression of iRFP670 was then monitored using a fluorescence microscope. Coexpression with EGFP-GCN4 enhanced iRFP670 fluorescence. This result indicated that GCN4 binding to the VH and VL domains can induce split RNAP assembly (Fig. 1C). This GCN4-dependent reporter expression also occurred when the VH and VL domains were fused to the reversed fusion pattern. Reporter assays using luciferase as a reporter gene demonstrated that GCN4-dRNAP exhibited slightly higher activity when the VH domain was fused to the C-terminus and the VL domain was fused to the N-terminus of RNAP (Fig. 6(a)).

[0084] Based on these results, we decided to use these fusion patterns (T7N-VL and VH-T7C) in subsequent experiments. We also evaluated the effect of linker length on GCN4-dRNAP activity. Luciferase assays revealed that GCN4-dRNAP activity was largely independent of linker length (Fig. 6(b)). We noted that T7 RNAP undergoes a large conformational change during the transition from transcription initiation to elongation. This is because large or highly charged target molecules may interfere with this conformational change. To minimize this adverse effect, we decided to use longer linkers in subsequent experiments.

[0085] Next, we investigated whether GCN4-dRNAP activity increased in a concentration-dependent manner with the target molecule. To examine concentration dependence, we performed luciferase assays with varying amounts of EGFP-GCN4. The results showed that induction by EGFP-GCN4 increased the reporter signal concentration-dependently, whereas induction by EGFP without GCN4 did not increase the reporter signal even at high concentrations (Fig. 1D). Finally, we investigated whether the binding affinity between the antibody and the target molecule affects TdRNAP activity. To examine this affinity dependence, we used three VH mutants of an anti-GCN4 antibody with amino acid mutations in the CDR loop that reduced the binding affinity for the GCN4 peptide. Luciferase assays showed that reduced binding affinity reduced RNAP activity (Fig. 1E). In particular, the VH mutant "GFA," which had a binding affinity more than 300-fold lower than that of WT, significantly reduced RNAP activity. On the other hand, the VH mutant "GLW," which had an affinity 1.7-fold lower than that of WT, exhibited activity equivalent to that of WT, indicating a 0.6 These results suggest that the nM dissociation constant is strong enough to maximize TdRNAP activity in mammalian cells. These results indicate that the activity of TdRNAP depends on the concentration and binding affinity of the intracellular target molecule. Taken together, these results demonstrate that the interaction of the VH / VL domain with its target molecule can trigger the association of split RNAP and induce target-dependent gene transcription.

[0086] Expanding the target molecule repertoire by antibody substitution. Next, we investigated whether antibody substitution could enable TdRNAP to induce reporter expression in response to its corresponding target molecule. To investigate this, we first designed a FLAG-dependent RNAP (FLAG-dRNAP) using the VH and VL domains of an anti-FLAG antibody (Fig. 2A, top). The anti-FLAG antibody used, "clone M2," recognizes the FLAG octapeptide (DYKDDDDK) with nanomolar affinity (Kd = 6.5 nM). However, the intracellular stability of its variable domain was unknown. First, when induced with FLAG-tagged EGFP (EGFP-1xFLAG), FLAG-dRNAP was able to induce reporter expression. However, despite its high binding affinity, the RNAP activity was very low (Fig. 7(a)). Furthermore, the use of a three-linked FLAG tag (3xFLAG) to increase binding affinity did not significantly improve RNAP activity (Fig. 7(a)). These results suggest that this variable domain requires an intramolecular disulfide bond for proper folding, and that its structure is unstable in the reducing environment of the cytoplasm.

[0087] To stabilize the structure, we grafted the CDR loops of an anti-FLAG antibody onto the stable framework region of the anti-HER2 antibody, mutated humanized trastuzumab (structure not shown). Grafting the CDR loops from the VL domain successfully enhanced FLAG-dRNAP activity. This result indicated that the structure of the VL domain of the anti-FLAG antibody is unstable in cells (Figure 2B). Grafting the CDR loops from both the VH and VL domains further increased reporter expression, even in the absence of a FLAG tag, suggesting that the VH and VL domains of the grafted anti-FLAG antibody bind weakly to each other (Figure 7(b)). Therefore, in subsequent experiments, we used a FLAG-dRNAP with the CDR loop sequences of both the VH and VL domains derived from the FLAG antibody, the framework region of the VL domain derived from the mutated humanized trastuzumab, and the framework region of the VH domain derived from the original FLAG antibody. We next investigated whether this FLAG-dRNAP exhibited concentration- and affinity-dependent increases in RNAP activity. Upon induction with EGFP-1xFLAG, a concentration-dependent increase in reporter signal was observed (Fig. 2C). The reporter signal was further enhanced upon induction with EGFP-3xFLAG, indicating that RNAP activity was also affinity-dependent. These findings are consistent with the results obtained with GCN4-dRNAP (Fig. 1D and 1E). These results demonstrate that TdRNAP can change its target molecule by antibody substitution. Furthermore, even if the variable domain of an antibody is unstable in cells, it can be adapted to TdRNAP by stabilizing the variable domain and using stable framework regions.

[0088] Based on the findings from the FLAG-dRNAP experiment, we next designed a protein-dependent RNAP that targets larger proteins rather than peptides. We chose an anti-EGFP antibody for the next substitution because commercially available anti-EGFP antibodies were generated using the stable framework of trastuzumab (Fig. 2A, bottom). To examine the target specificity of this EGFP-dependent RNAP (EGFP-dRNAP), we used EGFP and Azami-Green as target molecules. Azami-Green is a green fluorescent protein with a similar β-barrel structure to EGFP, but its sequence identity is low (29%). As a result, induction with EGFP increased the reporter signal in a concentration-dependent manner, whereas induction with Azami-Green did not increase the reporter signal even at high concentrations (Fig. 2D). These results clearly demonstrate that EGFP-dRNAP specifically recognizes and activates EGFP.

[0089] Next, we examined whether each TdRNAP could specifically react with its corresponding target without cross-reactivity. To examine cross-reactivity, we used GCN4-, FLAG-, and EGFP-dRNAPs and co-expressed them with their respective target molecules. Luciferase assays demonstrated that each TdRNAP was activated only in response to its corresponding target (Figure 2E), demonstrating that the variable domains of the fusion antibodies specifically recognized their corresponding target molecules in mammalian cells. These results indicated that the VH and VL domains retained their target specificity in mammalian cells, indicating that the CDR loop structures of the antibody variable domains were properly folded and stabilized. This suggests that stable framework regions and mammalian molecular chaperones contribute to the stabilized CDR loop structure and target specificity. Taken together, these results demonstrate that antibody substitution can expand the intracellular molecular target of TdRNAP, and that the substituted antibodies retain high specificity for their corresponding targets.

[0090] Design of RNA- and Small-Molecule-Dependent RNAPs Based on our previous results, we hypothesized that TdRNAPs might also be capable of controlling reporter expression in response to other types of molecules, such as RNA and small molecules. Several artificial RNA-based technologies have already been developed for RNA detection. However, these technologies require complementary base pairing to target RNA, making it difficult to target complex RNAs whose target sequences are masked by the RNA's secondary or tertiary structure. In contrast, antibodies exhibit structural rather than sequence specificity in RNA recognition. This structural specificity could be advantageous for targeting viral RNAs because, while viruses have a high mutation rate, they retain structural regions essential for viral replication and genomic RNA packaging. Therefore, we designed an RNA-dependent RNAP (RNA-dRNAP) using an antiviral RNA antibody and investigated whether the RNA-dRNAP could detect the structured regions of viral RNAs in mammalian cells.

[0091] To verify this, we used an anti-HCV IRES RNA antibody that specifically recognizes the structured RNA region within the internal ribosome entry site (IRES) of the hepatitis C virus (HCV) genomic RNA. The CDR loops of the VH and VL domains of the anti-HCV IRES RNA antibody were grafted onto a framework derived from mutant humanized trastuzumab. The resulting VH and VL domains were fused to a split RNAP to obtain IRES RNA-dRNAP (Figure 3A, top). The VH domain (SEQ ID NO: 2) derived from mutant humanized trastuzumab was used, but with the Arg at position 13 in FR3 substituted with Ala. This was because structural prediction by AlphaFold2 suggested that the Arg at position 13 in FR3 disrupts the structural arrangement of the CDR loop. In the VH domain of humanized trastuzumab (SEQ ID NO: 42), the Ala at position 13 in FR3 is used. The IRES RNA-dRNAP expression plasmid was cotransfected with a tdTomato reporter plasmid and a bicistronic plasmid carrying an HCV IRES (encoding Firefly and Renilla luciferase). Interestingly, tdTomato fluorescence was strongly enhanced upon induction with an IRES-containing bicistronic mRNA (Fig. 3B, 3C, and Fig. 8(a)). In contrast, tdTomato fluorescence was low upon induction with a monocistronic mRNA lacking the HCV IRES. Simultaneous luciferase assays confirmed the expression level of Renilla luciferase from the HCV IRES on the bicistronic mRNA. This indicates that the HCV IRES is functionally folded in living cells. Furthermore, luciferase assays showed that the expression levels of Firefly and Renilla luciferase were comparable between bicistronic and monocistronic mRNAs, and the translated luciferases did not affect RNAP activity (Fig. 8(b)). These results indicate that IRES RNA-dRNAP activates transcription through detection of the HCV IRES region in living cells.

[0092] Next, we used anti-small molecule antibodies to examine whether tdRNAP could respond to small molecules in living cells. For this test, we employed anti-fluorescein antibodies. The CDR loops of the heavy and light chains of the anti-fluorescein antibody were grafted onto a framework derived from mutant humanized trastuzumab, and the resulting VH and VL domains were fused to a split RNAP to design a fluorescein-dependent RNAP (Fluorescein-dRNAP) (Figure 3A, bottom). Cells were transfected with the Fluorescein-dRNAP expression plasmid and a tdTomato reporter plasmid. Next, transfected cells were treated with fluorescein diacetate (converted to fluorescein in living cells) and subjected to a fluorescent reporter assay. Fluorescein treatment increased reporter expression in a concentration-dependent manner. In contrast, solvent treatment only moderately induced reporter expression (Figures 3D and 3E). These results indicate that fluorescein promotes the association of the VH and VL domains and activates split RNAP. At the same time, they also show that the VH and VL domains associate weakly even in the absence of fluorescein. The VH and VL domains of the anti-fluorescein antibody interact with each other through several residues in the CDR loops, forming a binding pocket for fluorescein loading. This interaction may have induced the spontaneous assembly of split RNAP in the absence of fluorescein.

[0093] In conclusion, anti-RNA antibodies and anti-small molecule antibodies can also be applied to TdRNAP, thus enabling it to serve as a versatile platform for controlling gene expression in mammalian cells in response to a wide range of biochemical molecules.

[0094] Construction of Multilayer Gene Circuits Using TdRNAP Our results demonstrate that TdRNAP functions as a biochemical information converter, converting various inputs into desired genetic outputs. This functionality makes TdRNAP a versatile tool for constructing synthetic genetic circuits in living cells. We therefore aimed to develop multilayer genetic circuits that can simultaneously control the expression of multiple genes in response to corresponding target molecules by applying TdRNAP. First, we designed a transcriptional amplification system incorporating a T7 RNAP mutant, CGG-R12-KIRV (CGG RNAP), as a reporter gene. CGG RNAP is an evolved T7 RNAP with an amino acid mutation in its C-terminal region that specifically recognizes a mutated T7 promoter, known as the CGG promoter. We hypothesized that by placing CGG RNAP under the control of a T7 promoter and adding the CGG promoter to a conventional T7 promoter-driven reporter plasmid, CGG RNAP could be used to amplify reporter gene expression (Figure 4A). This transcriptional amplification system may be useful for detecting intracellular molecules with low abundance. To this end, we inserted a CGG promoter upstream of the T7 promoter to create a dual-promoter reporter plasmid expressing luciferase under the control of both the T7 and CGG promoters.

[0095] To test this transcriptional amplification system, we investigated whether GCN4-dRNAP could enhance reporter signal in conditions with low EGFP-GCN4 abundance. A GCN4-dRNAP expression plasmid was co-transfected into 293FT cells with a dual-promoter reporter plasmid and an additional reporter plasmid encoding CGG RNAP. Upon induction with 10 ng of EGFP-GCN4 expression plasmid, the transcriptional amplification system enhanced reporter expression 1.8-fold compared to the non-amplified system (Fig. 4B). The amplified signal was 1.3-fold higher than the signal saturation value observed in the non-amplified system when the amount of EGFP-GCN4 expression plasmid was 6-fold higher (60 ng) (Fig. 1D and Fig. 4B). This transcriptional amplification system further enhanced the reporter signal depending on the amount of T7 promoter-driven CGG RNAP plasmid, but also increased the background level. These results demonstrate that the additional reporter layer of CGG RNAP acts as a signal amplifier, enabling reporter signal amplification even at low concentrations of target molecules.

[0096] Next, we designed an orthogonal gene circuit using two types of TdRNAPs that could simultaneously control two reporter genes independently in response to two different target molecules. To demonstrate this concept, we designed an orthogonal system consisting of EGFP-dRNAP and GCN4-dRNAP. To independently control the two reporter genes, EGFP-dRNAP and GCN4-dRNAP were constructed with split T7 RNAP and split CGG RNAP, respectively, resulting in EGFP-dependent T7 RNAP and GCN4-dependent CGG RNAP (Fig. 4C). We investigated whether these two TdRNAPs exhibit orthogonal gene expression patterns in a target-dependent manner using Firefly and Renilla luciferase genes under the control of the T7 and CGG promoters, respectively. Luciferase assays demonstrated that the orthogonal system precisely controlled the expression of the two reporter genes in response to the presence of the corresponding target molecules (Fig. 4D). This result indicates that the two TdRNAPs simultaneously and in parallel regulate each reporter gene. Taken together, these results demonstrate that combining antibodies with T7 RNAP variants can also expand the pattern of output signals, providing a versatile tool for constructing multilayer genetic circuits in living cells.

[0097] Target-dependent genome editing in human cells. Controlling genome editing between target and non-target cells is one of the major challenges in gene therapy. To achieve such cell-specific genome editing, several biomarkers, including cell surface proteins and microRNAs, have been used to control the delivery and expression of genome editing devices. However, many of these biomarkers are by-products whose expression levels increase with disease progression, and the number of biomarkers available for conventional approaches is still limited. Ideally, genome editing should be driven autonomously by detecting gene products derived from target genes with genetic mutations. In this strategy, TdRNAP has the advantage of selectively targeting intracellular molecules that regulate gene expression. Furthermore, TdRNAP can also synthesize functional RNAs, such as guide RNAs (gRNAs), for genome editing. We therefore hypothesized that TdRNAP could autonomously trigger genome editing by directly recognizing disease-related proteins, such as mutant or fusion proteins, expressed from target genes in the genome.

[0098] To prove this concept, we designed a gene knockout experiment using the CRISPR-Cas9 system. Here, DNA cleavage is controlled by TdRNAP, which preferentially induces DNA cleavage in cells expressing the fusion gene (Fig. 5A). Therefore, we aimed to induce EGFP knockout in a GCN4-dependent manner using GCN4-dRNAP and an EGFP-targeting gRNA under the control of a T7 promoter. To verify this, we established two 293FT cell lines transfected with either the EGFP-GCN4 or EGFP gene at the AAVS1 locus of the genome, using the EGFP-GCN4 and EGFP genes as the fusion and normal genes, respectively (Fig. 5B). We cotransfected the GCN4-dRNAP and gRNA expression plasmids with the Cas9 expression plasmid into the EGFP-GCN4 and EGFP cell lines. The EGFP-negative cell population was then measured by flow cytometry. Flow cytometry analysis showed that GCN4-dRNAP preferentially induced EGFP knockout in the EGFP-GCN4 cell line compared with the EGFP cell line, demonstrating that EGFP-GCN4 itself served as a driver for knocking out its own gene (Figure 5C, Figure 5D). Notably, increasing the amount of Cas9 expression did not result in an increase in the EGFP-negative cell population in the EGFP cell line (Figure 5C). These results demonstrate that GCN4-dRNAP tightly controls the EGFP knockout event through intracellular regulation of gRNA expression in response to EGFP-GCN4. For comparison, we performed the same knockout experiment using a constitutive gRNA expression plasmid with a U6 promoter under GCN4-independent gRNA expression conditions. Under these constitutive gRNA expression conditions, the EGFP knockout efficiency was comparable between the two cell lines, indicating that the insertion of the GCN4 peptide gene sequence did not affect the knockout efficiency (Figure 5E). These results demonstrate that TdRNAP autonomously induces gRNA expression by directly recognizing aberrant proteins expressed from target genes in the genome, enabling precise genome editing in a cell-specific manner.

[0099] We demonstrated that TdRNAP induces genome editing in response to gene products expressed from a single genomic locus, indicating that TdRNAP is sensitive to intracellular gene products expressed at more moderate and uniform levels than transient overexpression via plasmid transfection. Furthermore, we found that TdRNAP engineered with an anti-Hsp70 antibody responded to increased endogenous Hsp70 expression after heat shock stimulation (Figures 9A, 9B, 9C, and 9D). This suggests that TdRNAP may be able to control transcriptional activity depending on the expression level of endogenous proteins. This dose-dependent activity is useful for targeting cancer cells with abnormal gene expression without affecting healthy cells. The ability to control gene expression may open new possibilities for personalized medicine and precision therapy. TdRNAP represents a robust and versatile strategy for engineering genetic circuits and cellular functions, and we anticipate applications in both bioengineering and therapeutics.

[0100] In this study, we demonstrate TdRNAP as a novel and universal platform for controlling gene expression in response to a wide variety of intracellular molecules. In the examples, we demonstrate that a single antibody variable domain can induce split T7 RNAP assembly by binding to its corresponding target. Using various identified antibodies against proteins / peptides, RNA, and small molecules, we also demonstrate that TdRNAP can expand its molecular repertoire for inducing gene expression. To our knowledge, this is the first demonstration of a single platform that transduces these three different types of biochemical information into transcriptional and translational output. In particular, our IRES RNA-dRNAP is the first example of RNA structural information being transferred to RNA synthesis. Furthermore, we applied TdRNAP to construct a multilayered genetic circuit for signal amplification and orthogonal signal transduction. Finally, we demonstrate cell-specific genome editing that autonomously triggers gene knockout by TdRNAP detecting intracellular gene products derived from target genes in the human genome.

[0101] This method also has the advantage that many identified antibodies can be used to expand the molecular targets for gene regulation. With established antibody screening methods and the advancement of machine learning-based antibody design techniques, antibodies with new molecular specificities can be easily identified and applied to the system of the present invention. Furthermore, compared to conventional split protein approaches, using antibody VL and VH domains eliminates the need to select the split site, allowing for easy targeting of various intracellular molecules and enabling gene regulation tailored to the cellular state. This advantage significantly expands the scope of application of molecular-guided gene regulation systems.

Claims

1. A target molecule-specific functional protein regulatory system comprising: (a) a first nucleic acid molecule comprising a nucleic acid sequence encoding a first fusion protein comprising a first target binding domain comprising a light chain variable domain of an antibody that specifically recognizes a target molecule and a first domain of a functional protein; (b) a second nucleic acid molecule comprising a nucleic acid sequence encoding a second fusion protein comprising a second target binding domain comprising a heavy chain variable domain of an antibody that specifically recognizes the target molecule and a second domain of a functional protein; or (c) and (d) below: (c) the first fusion protein; (d) the second fusion protein, wherein the first domain of the functional protein is either the N-terminal domain or the C-terminal domain of the functional protein, and the second domain of the functional protein is the other of the N-terminal domain or the C-terminal domain of the functional protein; a framework region of either the light chain variable domain or the heavy chain variable domain comprises a sequence of a framework region of trastuzumab or a mutant sequence thereof; The target molecule, the first target binding domain, and the second target binding domain form a ternary complex.

2. The system of claim 1, wherein the functional protein comprises a detection protein, a protein that controls RNA transcription and translation, an extracellular secretion protein, a genome editing protein, an enzyme, or a therapeutic protein.

3. The system of claim 1, further comprising (e) a third nucleic acid molecule or an additional protein that is controlled by the functional protein.

4. The system of claim 3, wherein the functional protein is a nucleic acid polymerase, and the third nucleic acid molecule comprises a promoter sequence specific for the nucleic acid polymerase and a nucleic acid sequence encoding a target protein or target RNA.

5. The system according to claim 4, wherein the target protein comprises a detection protein, a protein that controls RNA transcription and translation, an extracellular secretory protein, a genome editing protein, an enzyme, or a therapeutic protein.

6. The system of claim 3, wherein the functional protein is a nucleic acid polymerase, and the third nucleic acid molecule comprises a promoter sequence specific to the nucleic acid polymerase and a nucleic acid sequence encoding a guide RNA that specifically recognizes a molecule to be edited in genome editing.

7. The system of claim 1, wherein the first or second nucleic acid molecule is a DNA construct or a synthetic mRNA molecule.

8. A method for controlling a functional protein in a target molecule-specific manner in a cell, comprising the step of introducing the system according to claim 1 into a cell.

9. The method of claim 8, further comprising the step of: (e) introducing into the cell a third nucleic acid molecule or an additional protein that is regulated by the functional protein.

Citation Information

Patent Citations

  • IL-2 Compositions and Methods of Use Thereof

    JP2022533254A

  • Proximity-dependent split RNA polymerases as a versatile biosensor platform

    WO2017212400A2