Rapamycin-resistant cells

By integrating a naked FRB domain into cells using a DNA endonuclease and guide RNA system, rapamycin resistance is conferred, overcoming proliferation suppression and enabling effective intracellular signaling and cell expansion.

JP7894900B2Inactive Publication Date: 2026-07-24SEATTLE CHILDRENS HOSPITAL (DBA SEATTLE CHILDRENS RES INST)
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SEATTLE CHILDRENS HOSPITAL (DBA SEATTLE CHILDRENS RES INST)
Filing Date
2024-04-24
Publication Date
2026-07-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Rapamycin suppresses host cell proliferation and viability, limiting its usefulness for therapeutic and research purposes due to adverse effects on intracellular signaling.

Method used

A system comprising DNA endonuclease, guide RNA, and donor template is used to integrate a naked FKBP-rapamycin-binding (FRB) domain polypeptide into the target genomic locus of cells, enabling rapamycin resistance and intracellular signaling without cytotoxicity.

Benefits of technology

Genetically modified cells express the FRB domain, enhancing proliferation in the presence of rapamycin and allowing selective expansion and activation of CISC-expressing cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007894900000036
    Figure 0007894900000036
  • Figure 0007894900000037
    Figure 0007894900000037
  • Figure 0007894900000038
    Figure 0007894900000038
Patent Text Reader

Abstract

To provide compositions including proteins for expression in host cells to render them resistant to rapamycin, and methods of using the proteins, cells, and compositions disclosed herein for modulating cell signaling and for selective cell expansion.SOLUTION: Provided herein is a system comprising: a deoxyribonucleic acid (DNA) endonuclease or a nucleic acid encoding the DNA endonuclease; a guide RNA (gRNA) comprising a spacer sequence complementary to a target sequence within a target genomic locus in a cell, or nucleic acid encoding the gRNA; and a donor template including a donor cassette comprising a nucleic acid sequence encoding a naked FKBP-rapamycin binding (FRB) domain polypeptide.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Application No. 62 / 663,562, “Rapamycin-Resistant Cells,” filed on 27 April 2018, which is expressly incorporated herein by reference in its entirety.

[0002] Sequence listing reference This application was filed together with an electronic sequence listing. This sequence listing was provided as a file of approximately 120kb created on April 25, 2019, with the filename SCRI186WOSEQLISTING. The information contained herein is incorporated herein by reference in its entirety.

[0003] The present invention provides a chemically induced signaling complex in which two components combine in the presence of rapamycin or related compounds. This chemically induced signaling complex is an active signaling complex used in conjunction with a naked FKBP-rapamycin-binding protein expressed intracellularly to confer rapamycin resistance to host cells. [Background technology]

[0004] Rapamycin, also known as sirolimus, is a macrolide natural product with a complex structure, isolated from Streptomyces hygroscopicus, a bacterium discovered in 1975 from soil samples on Easter Island (also known as Rapa Nui) (Huang, S. et al. (2003). Cancer Biol. Ther., 2(3):222-232; Abraham, RT et al. (1996). Ann. Rev. Immunol., 14:483-510; Pollock, R. et al. (2002). Curr. Opin. Biotechnol., 13(5):459-467; Bayle, JH et al. (2006). Chem. Biol., 13(1):99-107). Rapamycin mediates the heterodimerization of the FKBP12 protein (FK506-binding protein 12) and the FRB (FKBP12 rapamycin-binding domain) (Huang, S. et al. (2003). Cancer Biol. Ther., 2(3):222-232). Due to its excellent physiological properties, including good solubility, membrane permeability across the blood-brain barrier, and good oral bioavailability (Abraham, RT et al. (1996). Ann. Rev. Immunol., 14:483-510), rapamycin is used as a low-molecular-weight dimerizer in a wide range of applications, including mammalian cells and other organisms (see, for example, Pollock, R. et al. (2002). Curr. Opin. Biotechnol., 13(5):459-467).

[0005] CISC (chemically induced signaling complex) is a multi-component synthetic protein complex configured to be co-expressed as two chimeric proteins in host cells, as described in international patent application PCT / US2017 / 065746 (this document is incorporated herein by reference). The two chimeric protein components of CISC are one extracellular domain constituting half of the rapamycin-binding complex, which is fused to the intracellular signaling complex constituting the other half of the rapamycin-binding complex. By delivering the nucleic acid encoding CISC to a host cell, intracellular signaling can be regulated in the presence of rapamycin or rapamycin-related compounds.

[0006] On the other hand, while intracellular signaling can be induced by rapamycin-induced dimerization of CISCs, the presence of rapamycin suppresses host cell proliferation and viability, thus limiting its usefulness for therapeutic and research purposes. Therefore, there is a need for novel compositions and methods that enable the use of rapamycin-mediated intracellular signaling by CISCs and mitigate the adverse effects of rapamycin or rapamycin-related compounds on host cell proliferation and viability. [Overview of the project] [Means for solving the problem]

[0007] In this specification, we provide compositions and methods for conferring rapamycin resistance to cells.

[0008] In one aspect, (i) Deoxyribonucleic acid (DNA) endonuclease, or nucleic acid encoding said DNA endonuclease; (ii) a guide RNA (gRNA) containing a spacer sequence complementary to a target sequence within a target genomic locus of a cell, or a nucleic acid encoding said gRNA; and (iii) Donor template including a donor cassette containing a nucleic acid sequence encoding a naked FKBP-rapamycin-binding (FRB) domain polypeptide. A system that includes, The present invention describes a system characterized in that the DNA endonuclease, the gRNA, and the donor template are configured such that a complex formed by the association of the DNA endonuclease and the gRNA promotes the targeted incorporation of the donor cassette into the target genomic locus of the cell, thereby producing a genetically modified cell capable of expressing the naked FRB domain polypeptide. In some embodiments, the DNA endonuclease is a Cas9 endonuclease. In some embodiments, the nucleic acid encoding the DNA endonuclease is codon-optimized for expression in the recombinant cell; the nucleic acid encoding the gRNA is codon-optimized for expression in the recombinant cell; and / or one or more coding sequences in the donor cassette are codon-optimized for expression in the recombinant cell. In some embodiments, the donor template is configured such that the integration of the donor cassette into the target genomic locus is performed by homologous recombination repair (HDR). In some embodiments, the donor template is configured such that the integration of the donor cassette into the target genomic locus is performed by non-homologous end joining (NHEJ). In some embodiments, the DNA endonuclease or the nucleic acid encoding the DNA endonuclease is formulated by encapsulation in liposomes or lipid nanoparticles. In some embodiments, the liposomes or lipid nanoparticles further comprise the gRNA or the nucleic acid encoding the gRNA. In some embodiments, the system further comprises a ribonucleoprotein (RNP) complex formed by the association of the DNA endonuclease and the gRNA.

[0009] In some embodiments, a naked FKBP-rapamycin-binding (FRB) domain polypeptide is provided. In some embodiments, the naked FRB polypeptide comprises the amino acid sequence shown in SEQ ID NO: 1 or SEQ ID NO: 2.

[0010] In some embodiments, the donor cassette further comprises one or more nucleic acid sequences encoding polypeptide elements of a dimerization-activatable chemical-induced signaling complex (CISC), The polypeptide elements of the aforementioned CISC are (i) A first CISC component comprising a first extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof; Furthermore (ii) A second CISC component comprising a second extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof. Includes, The first and second CISC components are configured such that, when expressed in cells, they can dimerize in the presence of rapamycin or rapalog to form a CISC with signal transduction capabilities.

[0011] In some embodiments, either the first extracellular binding domain or the second extracellular binding domain, or a functional derivative thereof, comprises an FK506-binding protein (FKBP) domain or a functional derivative thereof, and / or the other extracellular binding domain or a functional derivative thereof comprises an FRB domain or a functional derivative thereof. In some embodiments, the transmembrane domain of the first CISC component comprises the transmembrane domain of the IL-2 receptor, and / or the transmembrane domain of the second CISC component comprises the transmembrane domain of the IL-2 receptor. In some embodiments, either the signaling domain of the first CISC component or the signaling domain of the second CISC component, or a functional derivative thereof, comprises an IL-2 receptor subunit γ (IL2Rγ) domain or a functional derivative thereof, and / or the signaling domain of the other CISC component or a functional derivative thereof comprises an IL-2 receptor subunit β (IL2Rβ) domain or a functional derivative thereof. In some embodiments, the IL2Rβ domain polypeptide is cleaved. In some embodiments, the nucleic acid encoding the IL2Rβ domain comprises the nucleotide sequence shown in SEQ ID NO: 4. In some embodiments, the IL2Rβ domain includes the amino acid sequence shown in SEQ ID NO: 5.

[0012] In some embodiments, the rapalog is selected from the group consisting of everolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP1903, AP23573, and their metabolites and derivatives.

[0013] In some embodiments, the nucleic acid encoding the naked FRB domain is located downstream of one or more nucleic acid sequences encoding polypeptide elements of the CISC. In some embodiments, the donor cassette further includes (i) between each of the one or more nucleic acid sequences encoding each polypeptide element of the CISC; and / or (ii) between the nucleic acid encoding the naked FRB domain and the adjacent nucleic acid sequence encoding the polypeptide element of the CISC.

[0014] In some embodiments, each self-cleaving polypeptide encoded in the donor cassette is independently selected from the group consisting of P2A, T2A, E2A, and F2A. In some embodiments, the donor cassette further comprises a promoter operably linked to one or more coding sequences contained in the donor cassette. In some embodiments, the promoter is an inductive promoter or a constitutive promoter. In some embodiments, the promoter is an MND promoter. In some embodiments, the donor cassette further comprises a nucleic acid encoding a detection marker. In some embodiments, the detection marker is a green fluorescent protein (GFP) polypeptide, an mCherry polypeptide, or a low-affinity nerve growth factor receptor (LNGFR). In some embodiments, the nucleic acid encoding the naked FRB domain polypeptide does not have a nucleic acid encoding an endoplasmic reticulum localization signal peptide. In some embodiments, the donor cassette comprises a nucleic acid sequence contained in the nucleotide sequence shown in Sequence ID No. 3.

[0015] In some embodiments, the donor template is a viral vector. In some embodiments, the viral vector is a lentiviral vector, an adenovirus vector, or an adeno-associated virus (AAV) vector.

[0016] In one embodiment, a method for editing the genome of a cell, (i) Deoxyribonucleic acid (DNA) endonuclease, or nucleic acid encoding said DNA endonuclease; (ii) a guide RNA (gRNA) containing a spacer sequence complementary to a target sequence within a target genomic locus of a cell, or a nucleic acid encoding said gRNA; and (iii) Donor template including a donor cassette containing a nucleic acid sequence encoding a naked FKBP-rapamycin-binding (FRB) domain polypeptide. The process includes providing the cells with The present invention provides a method characterized in that the DNA endonuclease, the gRNA, and the donor template are configured such that a complex formed by the association of the DNA endonuclease and the gRNA promotes the targeted incorporation of the donor cassette into the target genomic locus of the cell, thereby producing a genetically modified cell capable of expressing the naked FRB domain polypeptide.

[0017] In one embodiment, genetically modified cells prepared according to a method for editing the genome of cells described herein are provided, the method is (i) Deoxyribonucleic acid (DNA) endonuclease, or nucleic acid encoding said DNA endonuclease; (ii) a guide RNA (gRNA) containing a spacer sequence complementary to a target sequence within a target genomic locus of a cell, or a nucleic acid encoding said gRNA; and (iii) Donor template including a donor cassette containing a nucleic acid sequence encoding a naked FKBP-rapamycin-binding (FRB) domain polypeptide. The process includes providing the cells with The DNA endonuclease, the gRNA, and the donor template are configured such that a complex formed by the association of the DNA endonuclease and the gRNA promotes the targeted integration of the donor cassette into the target genomic locus of the cell, thereby producing a genetically modified cell capable of expressing the naked FRB domain polypeptide. In some embodiments, the donor cassette is configured to express a naked FRB domain polypeptide within the cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a hematopoietic stem cell. In some embodiments, the cell is a lymphocyte. In some embodiments, the cell is a progenitor T cell or a regulatory T cell (T reg) In some embodiments, the cells are CD34+ cells, CD8+ cells, or CD4+ cells. In some embodiments, the cells are CD8+ cytotoxic T lymphocytes selected from the group consisting of naive CD8+ T cells, central memory CD8+ T cells, effector memory CD8+ T cells, and bulk CD8+ T cells. In some embodiments, the cells are CD4+ helper T lymphocytes selected from the group consisting of naive CD4+ T cells, central memory CD4+ T cells, effector memory CD4+ T cells, and bulk CD4+ T cells. In some embodiments, the cells have enhanced proliferation in the presence of rapamycin or rapalog compared to control cells in which the donor cassette does not contain nucleic acids encoding naked FRB domain polypeptides. In some embodiments, the concentration of rapamycin or rapalog is 0.1 nM to 100 nM or about 0.1 nM to about 100 nM.

[0018] In another embodiment, a method for activating genetically modified cells disclosed herein, comprising the step of contacting the cells with rapamycin or a rapalog, is described. In some embodiments, the rapalog is selected from the group consisting of everolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP1903, AP23573, and their metabolites, derivatives, and / or combinations. In some embodiments, the concentration of rapamycin or a rapalog contacted with the cells is 0.1 nM to 100 nM or about 0.1 nM to about 100 nM. In some embodiments, cells expressing components of a dimerization-activatable chemical-induced signaling complex (CISC) are selectively expanded and proliferated after contact with rapamycin or a rapalog.

[0019] In yet another embodiment, a method for selectively expanding and proliferating a population of genetically modified cells described herein that are included in a mixed cell population, The step includes contacting a mixed cell population with rapamycin or rapag, This invention describes a method characterized in which the activation and proliferation of genetically modified cells expressing components of a dimerization-activatable chemical-induced signaling complex (CISC) and a naked FRB domain are enhanced in vitro or in vivo compared to other cells in the mixed cell population. In some embodiments, the rapalog is selected from the group consisting of everolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP1903, AP23573, and their metabolites, derivatives, and / or combinations. In some embodiments, the concentration of rapamycin or rapalog brought into contact with the mixed cell population is 0.1 nM to 100 nM or about 0.1 nM to about 100 nM.

[0020] Each of the embodiments and models described herein can be used in combination unless expressly or explicitly excluded in those embodiments or models.

[0021] Throughout this specification, various patents, patent applications, and other types of publications (such as articles and contents of electronic databases) are cited. All patents, patent applications, and publications cited herein are incorporated herein by reference in their entirety for any purpose. [Brief explanation of the drawing]

[0022] [Figure 1]This is a schematic diagram of a lentiviral construct expressing an intracellular naked "decoy" FRB (FKBP12 rapamycin-binding domain). Constructs expressing an additional naked FRB* domain along with CISC are named "decoy-CISC" or "DISC". The asterisks in the diagram indicate that this sequence has a point mutation that enables interaction with the rapalog AP21967 and rapamycin (T2098L mutation in the amino acid sequence of mTOR (Bayle, JH et al. (2006). Chem. Biol., 13(1):99-107)).

[0023] [Figure 2] This is a conceptual diagram illustrating a hypothetical mechanism of action of naked FRB in T cells.

[0024] [Figure 3] The graph shows the growth and expansion of T cells expressing the CISC construct and T cells expressing the DISC construct, cultured in the presence of rapamycin or rapalog AP21967. The DISC construct promotes T cell growth and expansion better than cells transduced with the CISC construct alone at all rapamycin doses tested.

[0025] [Figure 4] This is a conceptual diagram showing the key amino acid changes introduced into various forms of the CISC construct that were tested. Four pairs of FRB-IL2Rβ / FKBP-IL2Rγ receptor proteins (V4-V7) were further constructed. These V4-V7 contained one or more amino acid sequences of the PAAL spacer and / or the GGS linker or GGSP linker. The amino acid sequences of the additional spacers and linkers were placed at the boundary between the extracellular FRB / FKBP domain and the IL2Rβ / IL2Rγ domain, or within the IL2Rβ domain.

[0026] [Figure 5]This graph shows the number of lentivirus-transduced T cells that have expanded and proliferated after being cultured for several weeks with the culture medium additives listed in the legend on the right.

[0027] [Figure 6] These two graphs show the time-dependent enrichment of lentivirus-transduced T cells (read as a percentage of mCherry+ cells) when cultured with rapamycin (upper panel) or rapalog AP21967 (lower panel).

[0028] [Figure 7] A conceptual diagram of a lentiviral construct used to directly compare the expansion and proliferation of rapamycin-induced T cells after transducing a DISC construct or a μDISC construct is shown (the initial CISC construct is included for reference).

[0029] [Figure 8] This graph shows the number of lentivirus-transduced T cells (mCherry+) over several days when cultured with the concentrations of rapamycin shown in the graph.

[0030] [Figure 9] This graph shows the enrichment of lentivirus-transduced T cells (read as a percentage of mCherry+ cells) after culturing for 9 days with various doses of rapamycin.

[0031] [Figure 10] This graph shows the growth rate of lentivirus-transduced (mCherry+) cells after 21 days of culture with various doses of rapamycin. The growth rate was calculated by dividing the number of mCherry+ cells on day 21 by the number of mCherry+ cells on day 0. The x-axis, from left to right, shows cells treated with 1 nM rapamycin (first three bars), cells treated with 5 nM rapamycin (bars 4-6), and cells treated with 10 nM rapamycin (bars 7-9).

[0032] [Figure 11] This is a schematic diagram of the experimental protocol used in a gene editing experiment to insert a DISC construct into the endogenous FOXP3 gene. It shows the AAV6 donor template (upper panel) and the gene editing operation using the AAV6-delivered donor template and CRISPR / CAS9 RNP (lower panel) to introduce ectopic promoter-induced DISC expression upstream of endogenous FOXP3. gRNA = guide RNA; 5'HA and 3'HA = human FOXP3 homologous arms.

[0033] [Figure 12] This is a schematic diagram showing the experimental design used in a gene editing experiment to insert a DISC construct into the endogenous FOXP3 gene of regulatory T cells. The MND promoter induces the expression of upstream DISC / μDISC of HA-tagged endogenous FOXP3. The 5' and 3' homologous arms at both ends are homologous regions of human FOXP3.

[0034] [Figure 13] Figure 13A is a schematic diagram showing the structure of the AAV construct used in a two-phase expansion culture protocol designed to improve the expansion and proliferation of CD4+ T cells successfully edited using DISC(edTreg). In this construct, the MND promoter induces cell surface expression of microDISC and intracellular expression of HA-tagged FOXP3. Figure 13B summarizes the results of the two-phase expansion culture experiment described herein, with the cell conditions for Phase 1 expansion culture shown on the X axis. The conditions for Phase 2 expansion culture are also shown.

[0035]

[0036] [Figure 14]Figures 14A and 14B are graphs summarizing the number of μDISC edTreg cells relative to the total number of cells in culture 10 and 15 days after editing. The culture conditions in Phase 0 expansion culture are shown on the X axis. The total number of cells in culture is shown in white, and the number of cells that were successfully edited (HA+ / μDISC+ and FOXP3+) is shown in black. Based on these findings, the second condition from the left was selected as the optimal condition for producing μDISC edTreg cells. Under this condition, 5 ng / ml of IL-2 is used during Phase 0 expansion culture.

[0037] [Figure 15] Figures 15A and 15B summarize the results of flow cytometry experiments conducted to compare the expansion and proliferation of μDISC GFP edTreg cells 7 days after transplantation into NSG mice, between those treated with rapamycin and those treated with a solvent. The mean ± sd values ​​for GFP+ cells (%) in 75 μL peripheral blood samples (Figure 15A) and the mean ± sd values ​​for GFP+ cell counts (Figure 15B) are shown in graphs. IR = radiation exposure; Rapa = intraperitoneal injection of 0.1 mg / kg rapamycin every two days throughout the experiment. P-values ​​were obtained using Student's t-test. In each figure, the two bars on the left represent samples from irradiated mice, and the two bars on the right represent samples from unirradiated mice.

[0038] [Figure 16]Figures 16A and 16B summarize the results of experiments conducted to compare the expansion and proliferation of μDISC GFP edTreg cells 14 days after transplantation into NSG mice, between those treated with rapamycin and those treated with a solvent. Figure 16A shows flow cytometry plots for viability analysis (Live / Dead) against human CD45, or flow cytometry plots for forward scattering (FCS) of cells against GFP, for representative mice from each cohort. The graph on the right side of Figure 16A summarizes the results for each cohort. Figure 16B shows the mean ± sd number of GFP+ cells in a 75 μL peripheral blood sample. IR = radiation exposure; Rapa = intraperitoneal injection of 0.1 mg / kg of rapamycin every two days throughout the experiment. P-values ​​were obtained using Student's t-test.

[0039] [Figure 17] This graph summarizes the percentage (%) and number of μDISC GFP-edTreg cells observed over time in the peripheral blood of in vivo recipient mice treated with or not treated with rapamycin. The plots summarize flow cytometry data, with each symbol representing an individual mouse, and the mean ± standard deviation shown for each bar. The upper panel plots GFP (%) in CD45+CD4+ gating, and the lower panel plots the number of GFP+ cells in peripheral blood samples.

[0040] [Figure 18] The Kaplan-Meier survival curves for an in vivo immunosuppressive model are shown. Teff = group with only T effectors. In these experiments, effector T cells were administered to all groups. [Modes for carrying out the invention]

[0041] This specification describes compositions and methods intended for use in conjunction with chemically induced signaling complexes (CISCs) polypeptides, which induce CISC-mediated rapamycin-mediated intracellular signaling in host cells without causing the cytotoxicity associated with the action of rapamycin, for example, by conferring resistance to growth inhibition by rapamycin or rapamycin-related compounds (e.g., Rapalog) to genetically engineered host cells. In some embodiments, the compositions and methods disclosed herein are intended for use in conjunction with CISC polypeptides, which are synthetic protein complexes consisting of multiple components, and which are configured to be co-expressed as two chimeric proteins in host cells. The CISC signaling system described in International Application No. PCT / US2017 / 065746 (this document is incorporated herein by reference) can be genetically engineered to induce intracellular signaling in response to rapamycin, a macrolide compound used in a variety of therapeutic and research applications. The two chimeric protein components of CISC are: one is the extracellular domain (either the FK506-binding protein (FKPB) domain or the FKBP-rapamycin-binding (FRB) domain) that constitutes half of the rapamycin-binding complex; this is fused to the intracellular signaling complex that constitutes the other half of the rapamycin-binding complex. When rapamycin binds to the extracellular domain of CISC, a CISC heterodimer is formed, and intracellular signaling is induced via the intracellular signaling complex portion of the chimeric polypeptide.

[0042] Although CISC-expressing cells are useful, it has been observed that exposure to rapamycin results in lower cell proliferation compared to when the rapamycin log AP21967 is used. Mammalian target of rapamycin (mTOR), also known as FK506-binding protein 12-rapamycin-related protein 1 (FRAP1), is a kinase encoded by the MTOR gene in humans. mTOR is a member of the phosphatidylinositol 3-kinase-related kinase family of protein kinases. mTOR is a growth regulator that stimulates cell proliferation by phosphorylating substrates that govern anabolic processes such as lipid synthesis and mRNA translation, and by delaying catabolic processes such as autophagy.

[0043] The applicant has devised a construct and method to overcome the problem of suppressed cell proliferation. As shown in Figure 2, the FKBP domain-containing protein is naturally expressed in the cytoplasm of cells, and mTOR contains an FRB domain. While not wishing to be bound by any particular theory, it is thought that the binding of the rapamycin-FKBP complex to the FRB domain of mTOR blocks or reduces intracellular signaling mediated by mTOR, resulting in decreased mRNA translation and cell proliferation. Based on this, although not wishing to be bound by any particular theory, the inventors of this application hypothesized that by expressing a naked (not bound to mTOR) FRB protein domain in cells and binding this naked FRB protein domain to the intracellular rapamycin-FKBP complex, the negative effect of rapamycin on the proliferation of CISC-expressing cells can be mitigated. Previous reports have indicated that expressing the naked FRB domain of mTOR in cells is toxic to mammalian cells, suggesting that it may not be suitable as a means of mitigating the interaction between the rapamycin-FKBP complex and the mTOR FRB domain in human cells (see Vilella-Bach, M. et al. (1999). J. Biol. Chem., 274(7):4266-4272). Furthermore, the T2098L mutation in the FRB domain (sequence numbering follows the amino acid sequence of mTOR) destabilizes the FRB domain compared to the wild-type protein, leading to faster intracellular degradation. However, this mutant protein is stabilized by binding to rapamycin (Stankunas, K. et al. (2003). Mol. Cell., 12(6):1615-1624).

[0044] In contrast to these previous reports, the inventors of this application have surprisingly discovered, as detailed herein, that the growth inhibitory effect of rapamycin can be effectively mitigated by intracellularly expressing a naked “decoy” FRB(FRB*) domain in CISC-expressing host cells. The construct expressing an additional naked FRB* domain along with CISC has been named “decoy-CISC” or “DISC”. Thus, by conferring rapamycin resistance with intracellularly expressed naked FRB domains, the usefulness of rapamycin-responsive CISC-expressing cells can be improved in both therapeutic applications and studies of intracellular signaling pathways.

[0045] Definition of Terms Unless otherwise stated, the technical and scientific terms used herein have the meanings generally understood by those skilled in the art in which this disclosure pertains. All patents, applications, published applications, and other publications referenced herein are expressly incorporated herein by reference unless otherwise stated. If there are multiple definitions for any term used herein, the definitions given in this section shall prevail unless otherwise stated.

[0046] In this specification, the singular forms "a," "an," and "the" include plural forms unless otherwise explicitly stated.

[0047] In this specification, the term “about” has the general and ordinary meaning as used herein, and may be used, for example, when describing a measured value, to mean a variation of ±20%, ±10%, ±5%, ±1%, or ±0.1% from a given value.

[0048] In this specification, "protein sequence" refers to the polypeptide sequence of amino acids, which is the primary structure of a protein. Furthermore, "upstream" refers to the 5' position on a polynucleotide and the position toward the N-terminus on a polypeptide. Similarly, "downstream" refers to the 3' position on a nucleotide and the position toward the C-terminus on a polypeptide. Therefore, "N-terminus" refers to the position of an element on a polynucleotide toward the N-terminus on a polypeptide, or a specific position.

[0049] "Nucleic acid" or "nucleic acid molecule" refers to polynucleotides, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), oligonucleotides, fragments obtained by polymerase chain reaction (PCR), and fragments obtained by ligation, cleavage, endonuclease activity, and exonuclease activity. Nucleic acid molecules may consist of monomers made of natural nucleotide monomers (DNA, RNA, etc.), or analogs of natural nucleotides (e.g., enantiomers of natural nucleotides), or combinations thereof. "Nucleic acid molecule" also includes so-called "peptide nucleic acids," which include natural nucleic acid bases or modified nucleic acid bases attached to a polyamide backbone. Nucleic acids may be single-stranded or double-stranded. In some embodiments, nucleic acid sequences encoding fusion proteins are provided. In some embodiments, the nucleic acid is RNA or DNA.

[0050] In this specification, "encodes" refers to the property that a specific nucleotide sequence in a polynucleotide, such as a gene, cDNA, or mRNA, functions as a template for synthesizing another macromolecule, such as a given amino acid sequence. Therefore, if mRNA corresponding to a particular gene is transcribed and translated to produce a protein in a cell or other biological system, that gene codes for this protein.

[0051] A "nucleotide sequence encoding a polypeptide" includes any degenerate nucleotide sequence that encodes the same amino acid sequence. In some embodiments, nucleic acids encoding fusion proteins are provided.

[0052] A “vector,” “expression vector,” or “construct” is a nucleic acid used to introduce heterologous nucleic acids into cells and, due to the presence of various regulatory factors, can express heterologous nucleic acids in cells. Examples of vectors include, but are not limited to, plasmids, minicircles, yeast, and viral genomes. In some embodiments, the vector is a plasmid, minicircle, yeast, or viral genome. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a lentivirus. In some embodiments, the vector is an adeno-associated virus (AAV) vector (e.g., AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, etc., but are not limited to these). In some embodiments, the vector is a vector for protein expression in bacterial systems such as E. coli.

[0053] In this specification, the terms “expression” or “protein expression” refer to the translation of a transcribed RNA molecule into a protein molecule. Protein expression may be characterized by its temporal, spatial, developmental, or morphological characteristics, and by quantitative or qualitative indicators. In some embodiments, a single protein or a group of proteins are expressed in a configuration (e.g., arrangement) that dimerizes in the presence of a ligand.

[0054] In this specification, a “fusion protein” or “chimeric protein” is a protein created by ligating two or more genes that originally encoded separate proteins or parts of separate proteins. A fusion protein can also be created from specific protein domains derived from two or more separate proteins. Translation of such a fusion gene yields a single polypeptide or a group of polypeptides possessing functional properties derived from each of the original proteins. Recombinant fusion proteins can be artificially created using recombinant DNA technology used in biological research or therapy. Methods for creating such fusion proteins are known to those skilled in the art. Some fusion proteins combine entire peptides and therefore may contain all domains of the original proteins, and in particular, all functional domains. However, other fusion proteins, especially those not found in nature, combine only parts of their coding sequences and therefore do not retain the original function of the parent genes from which such proteins originate.

[0055] In this specification, the term “regulatory element” refers to a DNA molecule that has gene regulatory activity, such as a DNA molecule that can influence the transcription and / or translation of a operably linked, transcribed DNA molecule. Regulators such as promoters (e.g., MND promoters, e.g., MND promoters containing the nucleic acid sequence of SEQ ID NO: 33), leaders, introns, and transcription termination regions are DNA molecules that have gene regulatory activity and play an essential role in the overall expression of genes in living cells. Therefore, isolated regulatory elements (such as promoters) that function in plants are useful for modifying the phenotype of plants by genetic engineering methods.

[0056] In this specification, the term "operably ligated" means that a first molecule is bound to a second molecule, and these molecules are arranged such that the first molecule influences the function of the second molecule. These two molecules may be part of a single, contiguous molecule or they may be adjacent. For example, if a promoter regulates the transcription of a transcribable DNA molecule of interest within a cell, the promoter is operably ligated to this transcribable DNA molecule.

[0057] In this specification, “dimeric chemical-induced signaling complex,” “dimerized CISC,” or “dimer” refers to two components that constitute a CISC, which may or may not bind to each other to form a fusion protein complex. “Dimerization” refers to the process by which two separate entities bind to each other to form a single entity, for example, in response to the binding of a ligand (e.g., rapamycin). In some embodiments, dimerization is stimulated by a ligand or drug. In some embodiments, “dimerization” refers to homodimerization, i.e., the binding of two identical entities to dimerize, for example, the binding of two identical CISC components to dimerize. In some embodiments, “dimerization” refers to heterodimerization, i.e., the binding of two different entities to dimerize, for example, the binding of two different separate CISC components to dimerize. In some embodiments, the dimerization of CISC components forms a cellular signaling pathway. In some embodiments, dimerization of CISC components enables selective expansion and proliferation of cells or cell populations. Further CISC systems may include CISC-gibberellin-based CISC dimerization systems or CISC-TMP-based CISC dimerization systems. Other chemically derivable dimerization (CID) systems and their components may also be used.

[0058] In this specification, “chemical-induced signaling complex” or “CISC” refers to a recombinant complex that initiates a signal within a cell and, as a direct result, undergoes dimerization upon ligand induction. A CISC may be a homodimer (a dimer of two identical components) or a heterodimer (a dimer of two different components). Therefore, in this specification, the term “homodimer” refers to a dimer consisting of two protein components described herein that have the same amino acid sequence. The term “heterodimer” refers to a dimer consisting of two protein components described herein that do not have the same amino acid sequence.

[0059] CISCs may be synthetic complexes, as described in further detail herein. In this specification, “synthetic” means a complex, protein, dimer, or composition described herein that is not natural and is not found in nature. In some embodiments, “IL2R-CISC” refers to a signaling complex comprising components of the interleukin-2 receptor. In some embodiments, “IL2 / 15-CISC” refers to a signaling complex comprising receptor signaling subunits shared by interleukin-2 (IL2) and interleukin-15 (IL15). In some embodiments, “IL7-CISC” refers to a signaling complex comprising components of the interleukin-7 receptor. Thus, CISCs may be named according to the constituent parts that make up a particular CISC. Those skilled in the art will recognize that the constituent parts of a chemically induced signaling complex may consist of natural or synthetic components useful for incorporation into a CISC. Therefore, these examples provided herein are not limiting to the invention.

[0060] In this specification, “cytokine receptor” refers to a receptor molecule that recognizes and binds to a cytokine. In some embodiments, cytokine receptors include modified cytokine receptor molecules (e.g., “cytokine receptor variants”) which involve substitutions, deletions, and / or additions to the amino acid sequence and / or nucleic acid sequence of the cytokine receptor. Thus, the term “cytokine receptor” is intended to include not only wild-type cytokine receptors but also recombinant cytokine receptors, synthetic cytokine receptors, and cytokine receptor variants. In some embodiments, the cytokine receptor is a fusion protein comprising an extracellular binding domain, a hinge domain, a transmembrane domain, and a signaling domain. In some embodiments, the components of the receptor (i.e., each domain of the receptor) are either native or synthetic. In some embodiments, each domain is of human origin.

[0061] In this specification, "FKBP" refers to the FK506-binding protein domain. FKBP refers to a family of proteins possessing prolyl isomerase activity, which are functionally related to cyclophyllin but not similar in terms of amino acid sequence. FKBP has been identified in many eukaryotes, from yeast to humans, and functions as a protein folding chaperone for proteins containing proline residues. FKBP belongs to the immunophilin family along with cyclophyllin. FKBP includes, for example, FKBP12, as well as proteins encoded by the AIP gene, AIPL1 gene, FKBP1A gene, FKBP1B gene, FKBP2 gene, FKBP3 gene, FKBP5 gene, FKBP6 gene, FKBP7 gene, FKBP8 gene, FKBP9 gene, FKBP9L gene, FKBP10 gene, FKBP11 gene, FKBP14 gene, FKBP15 gene, FKBP52 gene, and / or LOC541473 gene; including their homologs and functional protein fragments.

[0062] In this specification, “FRB” refers to the FKBP rapamycin-binding domain. The FRB domain is a polypeptide region (protein “domain”) configured to form a ternary complex with the FKBP protein and rapamycin or its rapalog. FRB domains are present in a variety of natural proteins, including mTOR proteins from humans and other species (also referred to herein as FRAP, RAPT1, or RAFT); yeast proteins containing Tor1 and / or Tor2; and FRAP homologs of the Candida genus. Both FKBP and FRB are major components of mammalian target of rapamycin (mTOR) signaling.

[0063] A "naked FKBP rapamycin-binding domain polypeptide" or "naked FRB domain polypeptide" (also called "FKBP rapamycin-binding domain polypeptide" or "FRB domain polypeptide") refers to a polypeptide consisting of the amino acid sequence of an FRB domain, or a protein in which approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the protein's amino acid sequence consists of the FRB domain sequence. Typically, such proteins can bind to and interact with the rapalog AP21967 or rapamycin. The FRB domain can be expressed as a 12kDa soluble protein (Chen, J. et al. (1995). Proc. Natl. Acad. Sci. USA, 92(11):4947-4951). The FRB domain forms a 4-helix bundle, which is a structural motif commonly found in globular proteins. The dimensions of the entire FRB domain are 30 Å × 45 Å × 30 Å, and the four helices are connected by short lower loops similar to the folding of cytochrome b562 (Choi, J. et al. (1996). Science, 273(5272):239-242). In some embodiments, the naked FRB domain contains the amino acid sequence of SEQ ID NO: 1 (MEMWHEGLEEASRLYFGERNVKGMFEVLEPLHAMMERGPQTLKETSFNQAYGRDLMEAQEWCRKYMKSGNVKDLTQAWDLYYHVFRRISK) or the amino acid sequence of SEQ ID NO: 2 (MEMWHEGLEEASRLYFGERNVKGMFEVLEPLHAMMERGPQTLKETSFNQAYGRDLMEAQEWCRKYMKSGNVKDLLQAWDLYYHVFRRISK).

[0064] As used herein, the “extracellular binding domain” refers to one of the domains constituting the complex, configured to bind to a specific atom or molecule, and located outside the cell. In some embodiments, the extracellular binding domain of the CISC is the FKBP domain or a functional derivative thereof. In some embodiments, the extracellular binding domain is the FRB domain or a functional derivative thereof. In some embodiments, the extracellular binding domain is configured to stimulate the dimerization of two CISC components by binding to a ligand or drug. In some embodiments, the extracellular binding domain is configured to bind to a cytokine receptor modulator.

[0065] In this specification, “cytokine receptor modulator” refers to an agent that modulates the phosphorylation of downstream targets of cytokine receptors, the activation of signaling pathways associated with cytokine receptors, and / or the expression of specific proteins such as cytokines. Such agents may directly or indirectly modulate the phosphorylation of downstream targets of cytokine receptors, the activation of signaling pathways associated with cytokine receptors, and / or the expression of specific proteins such as cytokines. Examples of cytokine receptor modulators include, but are not limited to, cytokines; cytokine fragments; fusion proteins; and / or antibodies or their binding sites that immune-specifically bind to cytokine receptors or fragments thereof. Furthermore, examples of cytokine receptor modulators include, but are not limited to, peptides, polypeptides (e.g., soluble cytokine receptors), fusion proteins, and / or antibodies or their binding sites that immune-specifically bind to cytokines or fragments thereof.

[0066] In this specification, “activate” means the enhancement of at least one biological activity of the protein of interest. Similarly, “activation” means the state of the protein of interest in which the activity is enhanced. “Activatable” means that the protein of interest is activatable in the presence of a signal, drug, ligand, compound, or stimulus. In some embodiments, the dimers described herein are activated in the presence of a signal, drug, ligand, compound, or stimulus to become a signal-signaling dimer. In this specification, “signal-signaling” means the ability or configuration of a dimer to initiate or sustain a downstream signaling pathway.

[0067] In this specification, “hinge domain” refers to a domain that may ligate an extracellular binding domain to a transmembrane domain, thereby conferring flexibility to the extracellular binding domain. In some embodiments, the hinge domain positions the extracellular domain closer to the cell membrane, minimizing the possibility of recognition by antibodies or their binding fragments. In some embodiments, the extracellular binding domain is located at the N-terminus of the hinge domain. In some embodiments, the hinge domain may be naturally occurring or synthetic.

[0068] In this specification, “transmembrane domain” or “TM domain” refers to a domain that is stable within a membrane, such as a cell membrane. The terms “transmembrane span,” “membrane endogenous protein,” and “membrane endogenous domain” are also used herein. In some embodiments, the hinge domain and extracellular domain are located at the N-terminus of the transmembrane domain. In some embodiments, the transmembrane domain is a native or synthetic domain. In some embodiments, the transmembrane domain is the IL-2 transmembrane domain.

[0069] In this specification, “signaling domain” refers to a domain of a fusion protein or a component of a CISC involved in a signaling cascade within a cell (such as a mammalian cell). “Signaling domain” refers to a signaling portion that provides a signal to a cell (such as a T cell) that mediates a cellular response (such as a T cell response), in addition to a primary signal provided by the CD3ζ chain of the TCR / CD3 complex, such as activation, proliferation, differentiation, and / or cytokine secretion. In some embodiments, the signaling domain is located at the N-terminal end of a transmembrane domain, hinge domain, and extracellular domain. In some embodiments, the signaling domain is a synthetic or native domain. In some embodiments, the signaling domain is a linked intracellular signaling domain. In some embodiments, the signaling domain is a cytokine signaling domain. In some embodiments, the signaling domain is an antigen signaling domain. In some embodiments, the signaling domain is an interleukin-2 receptor γ subunit (IL2Rγ or IL2Rg) domain. In some embodiments, the signaling domain is an interleukin-2 receptor β-subunit (IL2Rβ or IL2Rb) domain, or a cleaved IL2Rβ domain (e.g., a cleaved IL2Rβ domain containing the amino acid sequence of SEQ ID NO: 5). In some embodiments, binding of a drug or ligand to the extracellular binding domain causes dimerization of the CISC components, resulting in activation of the signaling pathway and signaling via the signaling domain. In this specification, "signaling" refers to the activation of the signaling pathway by binding of a ligand or drug to the extracellular domain. Binding of the ligand or drug to the extracellular domain causes dimerization of the CISC components and activation of the signal.

[0070] In this specification, “IL2Rb” or “IL2Rβ” refers to the interleukin-2 receptor β subunit. Similarly, “IL2Rg” or “IL2Rγ” refers to the interleukin-2 receptor γ subunit, and “IL2Ra” or “IL2Rα” refers to the interleukin-2 receptor α subunit. The IL-2 receptor has three forms: α-chain, β-chain, and γ-chain, which are also subunits of other cytokine receptors. IL2Rβ and IL2Rγ are members of the type I cytokine receptor family. In this specification, “IL2R” refers to the interleukin-2 receptor, which is involved in T cell-mediated immune responses. IL2R is involved in receptor-dependent endocytosis and the transmission of pro-mitotic signals from interleukin-2. Similarly, "IL-2 / 15R" refers to the receptor signaling subunit shared by IL-2 and IL-15, which may include an α subunit (IL2 / 15Ra or IL2 / 15Rα), a β subunit (IL2 / 15Rb or IL2 / 15Rβ), or a γ subunit (IL2 / 15Rg or IL2 / 15Rγ).

[0071] In some embodiments, the chemically induced signaling complex is a heterodimerization-activated signaling complex comprising two components. In some embodiments, the first component comprises an extracellular binding domain, an optional hinge domain, a transmembrane domain, and one or more linked intracellular signaling domains, which are one of the heterodimerization pairs. In some embodiments, the second component comprises an extracellular binding domain, an optional hinge domain, a transmembrane domain, and one or more linked intracellular signaling domains, which are the other of the heterodimerization pairs. Thus, in some embodiments, two recombination events occur. In some embodiments, these two CISC components are expressed in cells such as mammalian cells. In some embodiments, cells such as mammalian cells, or a population of cells such as a mammalian cell population, are brought into contact with a ligand or factor that induces heterodimerization, thereby initiating signaling. In some embodiments, the homodimerization pair dimerizes, thereby expressing a single CISC component in cells such as mammalian cells, and the homodimerized CISC component initiates signaling.

[0072] In this specification, “ligand” or “drug” refers to a molecule having a desired biological effect. In some embodiments, an extracellular binding domain recognizes and binds to the ligand, forming a ternary complex comprising the ligand and two binding CISC components. Ligands include, but are not limited to, proteinaceous molecules such as peptides, polypeptides, proteins, post-translationally modified proteins, antibodies, and their binding moieties; small molecules (less than 1000 daltons), inorganic or organic compounds; and nucleic acid molecules such as, but are not limited to, double-stranded or single-stranded DNA, double-stranded or single-stranded RNA (e.g., antisense RNA, RNAi, etc.), aptamers, and triple-helical nucleic acid molecules. Ligands may originate from, be obtained from, or be derived from, or be obtained from, synthetic molecular libraries. In some embodiments, the ligand is a protein, an antibody or its functional derivative, a small molecule, or a drug. In some embodiments, the ligand is rapamycin or a rapamycin analog (rapalog). In some embodiments, the rapalog may be a variant of rapamycin obtained by modifying rapamycin with one or more modifications, such as demethylation, removal, or substitution of methoxy at positions C7, C42, and / or C29; removal, derivatization, or substitution of hydroxyl at positions C13, C43, and / or C28; reduction, removal, or derivatization of ketone at positions C14, C24, and / or C30; substitution of a 6-membered pipecolate ring with a 5-membered prolyl ring; or another substitution on the cyclohexyl ring or substitution of the cyclohexyl ring with a substituted cyclopentyl ring.Therefore, in some embodiments, the rapalog is everolimus, merilimus, novolimus, pimecrolimus, ridaflorimus, tacrolimus, temsirolimus, umilolimus, zotarolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP23573 or AP1903, or any metabolite, derivative and / or combination thereof. In some embodiments, the ligand is an IMID drug (e.g., thalidomide, pomalidomide, lenalidomide or related analogues).

[0073] In this specification, "simultaneous binding" refers to the simultaneous binding of two or more CISC components to a ligand, and in some cases, the binding of two or more CISC components to a ligand substantially simultaneously, forming a complex consisting of multiple components including CISC components and ligand components, resulting in signal activation. For simultaneous binding to occur, the CISC components must be configured to spatially bind to a single ligand, and both CISC components must be configured to bind to the same ligand (or different parts thereof).

[0074] In this specification, “selective expansion and proliferation” refers to the ability to expand and proliferate a desired cell, such as mammalian cells, or a desired cell population, such as a mammalian cell population. In some embodiments, selective expansion and proliferation refers to the development or expansion of a population of pure cells (such as mammalian cells) in which two genes have been recombined. One component of a dimerized CISC is involved in one gene recombination, and the other component is involved in the other gene recombination. Thus, each component of a heterodimerized CISC is associated with each gene recombination. By exposing cells to a ligand, it becomes possible to selectively expand and proliferate only cells (such as mammalian cells) that have both desired modifications. Thus, in some embodiments, the only cells (e.g., mammalian cells) that can respond to contact with the ligand are those that express both components of the heterodimerized CISC.

[0075] In this specification, “host cell” includes any type of cell (e.g., mammalian cell) that is sensitive to transformation, transfection, or transduction by a nucleic acid construct or vector. In some embodiments, the host cell, such as a mammalian cell, is a T cell or a regulatory T cell (T reg) In this specification, “T cells” or “T lymphocytes” may be obtained from any mammal, for example, from primates or other species, including monkeys, dogs and humans. In some embodiments, the T cells are of the same species as the recipient (of the same species but from a different donor). In some embodiments, the T cells are autologous (the donor and recipient are the same). In some embodiments, the T cells are syngeneic (the donor and recipient are different but identical twins). In some embodiments, the host cells, such as mammalian cells, are hematopoietic stem cells. In some embodiments, the host cells are CD34+ cells, CD8+ cells or CD4+ cells. In some embodiments, the host cells are CD8+ cytotoxic T lymphocytes selected from the group consisting of naive CD8+ T cells, central memory CD8+ T cells, effector memory CD8+ T cells and bulk CD8+ T cells. In some embodiments, the host cells are CD4+ helper T lymphocytes selected from the group consisting of naive CD4+ T cells, central memory CD4+ T cells, effector memory CD4+ T cells, and bulk CD4+ T cells. In this specification, “cell population” refers to a group of cells (such as mammalian cells) comprising two or more types of cells. In some embodiments, cells (such as mammalian cells) containing the protein sequence described herein or an expression vector encoding the protein sequence are produced.

[0076] In this specification, “cytotoxic T lymphocytes (CTLs)” refers to T lymphocytes that express CD8 on their cell surface (e.g., CD8+ T cells). In some embodiments, such cells are, for example, “memory” T cells (T) that have experienced an antigen. M A cell is provided. In some embodiments, a cell for secreting a fusion protein is provided. In some embodiments, the cell is a cytotoxic T lymphocyte. In this specification, a "central memory" T cell (or "T") is used. CMA central memory T cell (T) is a cytotoxic T lymphocyte (CTL) that has experienced an antigen and, compared to naive cells, expresses CD62L, CCR-7, and / or CD45RO on its surface, but does not express CD45RA or has reduced CD45RA expression. In some embodiments, cells are provided for secreting a fusion protein. In some embodiments, the cells are central memory T cells (T). CM In some embodiments, central memory cells may be positive for CD62L, CCR7, CD28, CD127, CD45RO, and / or CD95 expression, but have reduced CD54RA expression, compared to naive cells. In this specification, “effector memory” T cells (or “T EM A T cell is an antigen-experienced T cell that, compared to a central memory cell, does not express CD62L on its surface or has reduced CD62L expression, and compared to a naive cell, does not express CD45RA or has reduced CD45RA expression. In some embodiments, cells for secreting a fusion protein are provided. In some embodiments, the cells are effector memory T cells. In some embodiments, the effector memory cells may be negative for CD62L and / or CCR7 expression and positive or negative for CD28 and / or CD45RA expression compared to a naive cell or a central memory cell.

[0077] As described herein, “naive” T cells are T lymphocytes that have not experienced an antigen and, compared to central memory cells or effector memory cells, express CD62L and / or CD45RA and do not express CD45RO. In some embodiments, cells (such as mammalian cells) for secreting fusion proteins are provided. In some embodiments, the cells (such as mammalian cells) are naive T cells. In some embodiments, naive CD8+ T lymphocytes are characterized by the expression of naive T cell phenotypic markers, such as CD62L, CCR7, CD28, CD127, and / or CD45RA.

[0078] The “effector” T cells described herein are cytotoxic T lymphocytes that have experienced an antigen and, compared to central memory T cells or naive T cells, do not express CD62L, CCR7, and / or CD28, or have reduced expression of CD62L, CCR7, and / or CD28, and are granzyme B positive and / or perforin positive. In some embodiments, cells (such as mammalian cells) for secreting fusion proteins are provided. In some embodiments, the cells (such as mammalian cells) are effector T cells. In some embodiments, the cells (such as mammalian cells) do not express CD62L, CCR7, and / or CD28, or have reduced expression of CD62L, CCR7, and / or CD28, and are granzyme B positive and / or perforin positive, compared to central memory T cells or naive T cells.

[0079] In this specification, “transformed” or “transfected” means a cell (such as a mammalian cell), tissue, organ, or organism into which a foreign polynucleotide molecule, such as a construct, has been introduced. The introduced polynucleotide molecule may be incorporated into the genomic DNA of the recipient cell (such as a mammalian cell), tissue, organ, or organism, thereby allowing the introduced polynucleotide molecule to be passed on to subsequent offspring. Furthermore, “transgenic” cells (such as mammalian cells) or “transgenic” organisms, or “transfected” cells (such as mammalian cells) or “transfected” organisms include offspring of transgenic cells or organisms or transfected cells or organisms, including offspring produced by breeding programs that use such transgenic organisms as parents in mating, and which exhibit altered phenotypes resulting from the presence of the foreign polynucleotide molecule. “Transgenic” also means bacteria, fungi, or plants containing one or more heterologous nucleic acid molecules.

[0080] In this specification, “transduction” refers to the introduction of genes into cells (such as mammalian cells) using a virus (for example, a lentivirus or adeno-associated virus).

[0081] In this specification, “subject” or “individual” refers to an animal that is the subject of treatment, observation, or experimentation. “Animals” include ectothermic vertebrates, warm-blooded vertebrates, and invertebrates such as fish, crustaceans, and reptiles, in particular mammals. “Mammals” include, but are not limited to, mice, rats, rabbits, guinea pigs, dogs, cats, sheep, goats, cattle, horses, primates (such as monkeys, chimpanzees, and apes), in particular humans. In some embodiments, the subject is a human.

[0082] The “marker sequences” described herein encode cells containing the target protein (such as mammalian cells) or proteins used to select or track the target protein. Embodiments described herein provide fusion proteins which may contain marker sequences that can be selected in experiments such as flow cytometry.

[0083] The “amino acid sequence identity (%)” described herein for CISC sequences or other polypeptide sequences (e.g., naked FRB domain polypeptide sequences) is defined as the percentage of amino acid residues in each of the extracellular binding domain, hinge domain, transmembrane domain, and / or signaling domain in the candidate sequence that match the amino acid residues in each domain in the reference sequence. This amino acid sequence identity is calculated after aligning the candidate sequence and the reference sequence and inserting gaps as necessary to calculate the maximum sequence identity (%), and conservative substitutions are not considered as part of the sequence identity. Alignment for determining amino acid sequence identity (%) can be performed in various ways that fall within the scope of the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, ALIGN-2, and Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for measuring alignment, such as any algorithm required to maximize alignment over the full length of multiple sequences being compared. For example, calculating amino acid sequence identity (%) using the WU-BLAST-2 computer program (Altschul, SF et al. (1996). Methods in Enzymol., 266:460-480) involves using several search parameters, most of which are set to their default values. Parameters not set to default values ​​(e.g., adjustable parameters) are set to overlap span=1, overlap fraction=0.125, word threshold (T)=11, and scoring matrix=BLOSUM62. In some embodiments of CISC, the CISC includes an extracellular binding domain, a hinge domain, a transmembrane domain, and a signaling domain, each of which may include a native or synthetic domain, or a mutant or cleaved form of the native domain.In some embodiments, a variant or cleavage of a given domain includes an amino acid sequence having sequence identity (%) within a range defined by 100%, 95%, 90%, or 85% sequence identity, or any two of the aforementioned percentages, with respect to the sequence shown in the sequence provided herein.

[0084] In this specification, “CISC variant polypeptide sequence” or “CISC variant amino acid sequence” means a protein sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% amino acid sequence identity (or amino acid sequence identity (%) within the range defined by any two of these percentages) with the protein sequences provided herein, as defined below, or a specifically derived fragment thereof, such as a protein sequence of an extracellular binding domain, hinge domain, transmembrane domain, and / or signaling domain. Typically, a polypeptide or fragment of a CISC variant exhibits at least 80% or approximately 80% amino acid sequence identity with the amino acid sequence of CISC or a fragment derived therefrom, at least 81% or approximately 81% amino acid sequence identity, at least 82% or approximately 82% amino acid sequence identity, at least 83% or approximately 83% amino acid sequence identity, at least 84% or approximately 84% amino acid sequence identity, at least 85% or approximately 85% amino acid sequence identity, at least 86% or approximately 86% amino acid sequence identity, at least 87% or approximately 87% amino acid sequence identity, at least 88% or approximately 88% amino acid sequence identity, and at least 89%. Alternatively, it has at least approximately 89% amino acid sequence identity, at least 90% or at least approximately 90% amino acid sequence identity, at least 91% or at least approximately 91% amino acid sequence identity, at least 92% or at least approximately 92% amino acid sequence identity, at least 93% or at least approximately 93% amino acid sequence identity, at least 94% or at least approximately 94% amino acid sequence identity, at least 95% or at least approximately 95% amino acid sequence identity, at least 96% or at least approximately 96% amino acid sequence identity, at least 97% or at least approximately 97% amino acid sequence identity, at least 98% or at least approximately 98% amino acid sequence identity, or at least 99% or at least approximately 99% amino acid sequence identity.The variant does not include the native protein sequence.

[0085] In this specification, “polypeptide sequence of a naked FRB domain variant” or “amino acid sequence of a naked FRB domain variant” means a protein sequence having at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, at least 98% or at least about 98%, or at least 99% or at least about 99% amino acid sequence identity (or amino acid sequence identity (%) within the range defined by any two of these percentages) of the protein sequence provided herein (e.g., SEQ ID NO: 1 or SEQ ID NO: 2), as defined below, or a specifically derived fragment thereof, such as a protein sequence of an extracellular binding domain, hinge domain, transmembrane domain and / or signaling domain.Typically, a polypeptide or fragment of a variant of the naked FRB domain exhibits at least 80% or approximately 80% amino acid sequence identity with the amino acid sequence of the naked FRB domain or a fragment derived therefrom, at least 81% or approximately 81%, at least 82% or approximately 82%, at least 83% or approximately 83%, at least 84% or approximately 84%, at least 85% or approximately 85%, at least 86% or approximately 86%, at least 87% or approximately 87%, and at least 88% or approximately 88% amino acid sequence identity. The variant has at least 89% or approximately 89% amino acid sequence identity, at least 90% or approximately 90% amino acid sequence identity, at least 91% or approximately 91% amino acid sequence identity, at least 92% or approximately 92% amino acid sequence identity, at least 93% or approximately 93% amino acid sequence identity, at least 94% or approximately 94% amino acid sequence identity, at least 95% or approximately 95% amino acid sequence identity, at least 96% or approximately 96% amino acid sequence identity, at least 97% or approximately 97% amino acid sequence identity, at least 98% or approximately 98% amino acid sequence identity, or at least 99% or approximately 99% amino acid sequence identity. The variant does not contain any native protein sequences.

[0086] In this specification, whether in the transitional clause or the body of a claim, the terms “comprise(s)” and “comprising” are to be interpreted as having an open-ended meaning. That is, these terms are to be interpreted as synonymous with the expressions “at least have” or “at least contain.” When the term “comprise(s)” is used in a context relating to a method, it means that the method includes at least the specified steps, and may include other steps. When the term “comprise(s)” is used in a context relating to a compound, composition, or apparatus, it means that the compound, composition, or apparatus includes at least the specified features or components, and may include other features or components.

[0087] CISC The one or more protein sequences may have a first sequence and a second sequence. In some embodiments, the first sequence encodes a first CISC component which may include a first extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof. In some embodiments, the second sequence encodes a second CISC component which may include a second extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a portion thereof. In some embodiments, the first and second CISC components may be configured to dimerize in the presence of a ligand when expressed. In some embodiments, the first and second CISC components may be configured to dimerize simultaneously in the presence of a ligand when expressed.

[0088] In some embodiments, the present invention provides protein sequences of two heterodimerizing CISC components, or sequences encoding two heterodimerizing CISC components. In some embodiments, the first CISC component is an IL2Rγ-CISC complex. In some embodiments, the IL2Rγ-CISC complex includes the amino acid sequence shown in SEQ ID NO: 9. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 9 is also included in the embodiments. In some embodiments, the IL2Rγ-CISC complex includes the amino acid sequence shown in SEQ ID NO: 10. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 10 is also included in the embodiments. In some embodiments, the IL2Rγ-CISC complex includes the amino acid sequence shown in SEQ ID NO: 11. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 11 is also included in the embodiments. In some embodiments, the IL2Rγ-CISC complex includes the amino acid sequence shown in SEQ ID NO: 12. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 12 is also included in the embodiments.

[0089] In some embodiments, the protein sequence of the first CISC component includes a protein sequence encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain. Nucleic acid sequences encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain are also included in the embodiments. In some embodiments, the protein sequence of the first CISC component, including the first extracellular binding domain, a hinge domain, a transmembrane domain, and / or a signaling domain, includes an amino acid sequence having sequence identity of 100%, 99%, 98%, 95%, 90%, 85%, or 80% of the sequence shown in SEQ ID NOs: 9, 10, 11, or 12, or sequence identity including all numerical values ​​and ranges included in these percentages.

[0090] In some embodiments, the second CISC component is a complex containing IL2Rβ. In some embodiments, the IL2Rβ-CISC complex contains the amino acid sequence shown in SEQ ID NO: 13. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 13 is also included in the embodiments. In some embodiments, the IL2Rβ-CISC complex contains the amino acid sequence shown in SEQ ID NO: 14. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 14 is also included in the embodiments. In some embodiments, the IL2Rβ-CISC complex contains the amino acid sequence shown in SEQ ID NO: 15. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 15 is also included in the embodiments. In some embodiments, the IL2Rβ-CISC complex contains the amino acid sequence shown in SEQ ID NO: 24. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 24 is also included in the embodiments. In some embodiments, the second CISC component is a complex containing IL7Rα. In some embodiments, the IL7Rα-CISC complex contains the amino acid sequence shown in SEQ ID NO: 16, 17, or 26. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 16, 17, or 26 is also included in the embodiments.

[0091] In another embodiment, IL2Rβ-CISC comprises a cleaved intracellular IL2Rβ domain or a nucleic acid sequence encoding a cleaved intracellular IL2Rβ domain. When the cleaved IL2Rβ domain forms heterodimerization with the IL2Rγ-CISC protein sequence, it retains the ability to activate downstream IL2 signaling. In some embodiments, the cleaved IL2Rβ comprises the amino acid sequence shown in SEQ ID NO: 5 (PAALGKDTIPWLGHLLVGLSGAFGFIILVYLLINCRNTGPWLKKVLKCNTPDPSKFFSQLSSEHGGDVQKWLSSPFPSSSFSPGGLAPEISPLEVLERDKVTQLLLQQDKVPEPASLSLNTDAYLSLQELQ; SEQ ID NO: 5). In some embodiments, the cleaved IL2Rβ domain of SEQ ID NO: 5 has one, two, three, four, five, six, seven, eight, nine, or ten amino acids deleted from its N-terminus. In another embodiment, an FRB-CISC having a cleavable intracellular IL2Rβ domain contains the amino acid sequence shown in SEQ ID NO: 7 (MALPVTALLLPLALLLHAARPILWHEMWHEGLEEASRLYFGERNVKGMFEVLEPLHAMMERGPQTLKETSFNQAYGRDLMEAQEWCRKYMKSGNVKDLLQAWDLYYHVFRRISKPAALGKDTIPWLGHLLVGLSGAFGFIILVYLLINCRNTGPWLKKVLKCNTPDPSKFFSQLSSEHGGDVQKWLSSP*WORFPSSSFSPGGLAPEISPLEVLERDKVTQLLLQQDKVPEPASLSLNTDAYLSLQELQ; SEQ ID NO: 7).

[0092] In some embodiments, the protein sequence of the second CISC component includes a protein sequence encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain. Nucleic acid sequences encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain of the second CISC component are also included in the embodiments. In some embodiments, the protein sequence of the second CISC component, including a second extracellular binding domain, a hinge domain, a transmembrane domain, and / or a signaling domain (including, in some embodiments, a cleaved IL2Rβ signaling domain), includes an amino acid sequence having sequence identity with the sequence shown in SEQ ID NOs. 5, 8, 13, 14, 15, 16, or 17, with sequence identity of 100%, 99%, 98%, 95%, 90%, 85%, or 80%, or sequence identity including all numerical values ​​and ranges included in these percentages.

[0093] In some embodiments, the protein sequence may include a linker. In some embodiments, the linker may contain one, two, three, four, five, six, seven, eight, nine, or ten amino acids (such as glycine), or a large number of amino acids (such as glycine), or a number of amino acids within a range defined by any two of these values. In some embodiments, the glycine spacer contains at least three glycine molecules. In some embodiments, the glycine spacer includes the sequence shown in SEQ ID NO: 18 (GGGS; SEQ ID NO: 18), the sequence shown in SEQ ID NO: 19 (GGGSGGG; SEQ ID NO: 19), the sequence shown in SEQ ID NO: 20 (GGG; SEQ ID NO: 20), the sequence shown in SEQ ID NO: 21 (GGS; SEQ ID NO: 21), the sequence shown in SEQ ID NO: 22 (GGSP; SEQ ID NO: 22), or the sequence shown in SEQ ID NO: 31 (PAAL; SEQ ID NO: 31). Nucleic acid sequences encoding SEQ ID NOs: 18-22 and 31 are also included in embodiments. In some embodiments, the transmembrane domain is located at the N-terminus of the signaling domain, the hinge domain is located at the N-terminus of the transmembrane domain, the linker is located at the N-terminus of the hinge domain, and the extracellular binding domain is located at the N-terminus of the linker.

[0094] In some embodiments, the present invention provides protein sequences of two homodimerizing CISC components, or sequences encoding two homodimerizing CISC components. In some embodiments, the first CISC component is an IL2Rγ-CISC complex. In some embodiments, the IL2Rγ-CISC complex includes the amino acid sequence shown in SEQ ID NO: 23. The embodiments also include nucleic acid sequences encoding the protein sequence of SEQ ID NO: 23.

[0095] In some embodiments, the protein sequence of the first CISC component includes a protein sequence encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain. Nucleic acid sequences encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain are also included in the embodiments. In some embodiments, the protein sequence of the first CISC component, including the first extracellular binding domain, a hinge domain, a transmembrane domain, and / or a signaling domain, includes an amino acid sequence having sequence identity with the sequence shown in SEQ ID NO: 23, with sequence identity of 100%, 99%, 98%, 95%, 90%, 85%, or 80%, or with sequence identity including all numerical values ​​and ranges included in these percentages.

[0096] In some embodiments, the second CISC component is a complex containing IL2Rβ or a complex containing IL2Rα. In some embodiments, the IL2Rβ-CISC complex contains the amino acid sequence shown in SEQ ID NO: 24. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 24 is also included in the embodiments.

[0097] In some embodiments, the IL2Rα-CISC complex includes the amino acid sequence shown in SEQ ID NO: 25. The embodiments also include the nucleic acid sequence encoding the protein sequence of SEQ ID NO: 25.

[0098] In some embodiments, the protein sequence of the second CISC component includes a protein sequence encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain. Nucleic acid sequences encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain of the second CISC component are also included in the embodiments. In some embodiments, the protein sequence of the second CISC component, including the second extracellular binding domain, a hinge domain, a transmembrane domain, and / or a signaling domain, includes an amino acid sequence having sequence identity with the sequence shown in SEQ ID NO: 24 or SEQ ID NO: 25, with sequence identity of 100%, 99%, 98%, 95%, 90%, 85%, or 80%, or sequence identity including all numerical values ​​and ranges included in these percentages.

[0099] In some embodiments, the protein sequence may include a linker. In some embodiments, the linker may contain one, two, three, four, five, six, seven, eight, nine, or ten amino acids (such as glycine), or a large number of amino acids (such as glycine), or a number of amino acids within a range defined by any two of these values. In some embodiments, the glycine spacer contains at least three glycine molecules. In some embodiments, the glycine spacer includes the sequence shown in SEQ ID NO: 18 (GGGS; SEQ ID NO: 18), the sequence shown in SEQ ID NO: 19 (GGGSGGG; SEQ ID NO: 19), the sequence shown in SEQ ID NO: 20 (GGG; SEQ ID NO: 20), the sequence shown in SEQ ID NO: 21 (GGS; SEQ ID NO: 21), the sequence shown in SEQ ID NO: 22 (GGSP; SEQ ID NO: 22), or the sequence shown in SEQ ID NO: 31 (PAAL; SEQ ID NO: 31). Nucleic acid sequences encoding SEQ ID NOs: 18-22 and 31 are also included in embodiments. In some embodiments, the transmembrane domain is located at the N-terminus of the signaling domain, the hinge domain is located at the N-terminus of the transmembrane domain, the linker is located at the N-terminus of the hinge domain, and the extracellular binding domain is located at the N-terminus of the linker.

[0100] In some embodiments, the sequences encoding the two homodimerizing CISC components incorporate an FKBP F36V domain that homodimerizes with the ligand AP1903.

[0101] In some embodiments, a protein sequence of a single homodimerizable CISC component or a sequence encoding a single homodimerizable CISC component is provided. In some embodiments, the single CISC component is an IL7Rα-CISC complex. In some embodiments, the IL7Rα-CISC complex includes the amino acid sequence shown in SEQ ID NO: 26. The embodiments also include a nucleic acid sequence encoding the protein sequence of SEQ ID NO: 26.

[0102] In some embodiments, the single CISC component is an MPL-CISC complex. In some embodiments, the MPL-CISC complex includes the amino acid sequence shown in SEQ ID NO: 27. The nucleic acid sequence encoding the protein sequence of SEQ ID NO: 27 is also included in the embodiments.

[0103] In some embodiments, the protein sequence of the single CISC component includes a protein sequence encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain. Nucleic acid sequences encoding an extracellular binding domain, a hinge domain, a transmembrane domain, or a signaling domain are also included in the embodiments. In some embodiments, the protein sequence of the first CISC component, comprising a first extracellular binding domain, a hinge domain, a transmembrane domain, and / or a signaling domain, includes an amino acid sequence having sequence identity within a range defined by 100%, 99%, 98%, 95%, 90%, 85%, or 80% sequence identity, or any two of these percentages, with respect to the sequence shown in SEQ ID NO: 26 or SEQ ID NO: 27.

[0104] In some embodiments, the sequence encoding a single homodimerizable CISC component incorporates an FKBP F36V domain that homodimerizes with the ligand AP1903.

[0105] vector Various vector combinations can be constructed to efficiently perform transduction and transgene expression. In some embodiments, the vector is a viral vector. In other embodiments, the vector may be a combination of a viral vector and a plasmid vector. Other viral vectors include formy virus, adenovirus vector, adeno-associated virus (AAV) vector, retrovirus vector, and / or lentivirus vector. In some embodiments, the vector is a lentivirus vector. In some embodiments, the vector is a formy virus vector, adenovirus vector, retrovirus vector, or lentivirus vector. In some embodiments, the vector is a vector for protein expression in bacterial systems such as E. coli. In other embodiments, the first vector may encode a first CISC component comprising a first extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof, and the second vector may encode a second CISC component comprising a second extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a portion thereof.

[0106] A vector may have one or more promoters for inducing the expression of a DNA sequence in the vector (for example, a DNA sequence encoding a naked FRB domain or a component of CISC). A “promoter” is a DNA region that initiates the transcription of a particular gene. The promoter may be located near the gene transcription start site or upstream within the same DNA strand (5' region of the sense strand). The promoter may be a conditionally inducible promoter or a constitutive promoter. The promoter may be specific to protein expression in bacterial cells, specific to protein expression in mammalian cells, or specific to protein expression in insect cells. In some embodiments, if a nucleic acid encoding a fusion protein is provided, the nucleic acid further comprises a promoter sequence. In some embodiments, the promoter is specific to protein expression in bacterial, mammalian, or insect cells. In some embodiments, the promoter is a conditional promoter or a constitutive promoter. In some embodiments, the promoter is an MND promoter (a synthetic promoter comprising the U3 region of a modified MoMuLV LTR and an enhancer of myeloproliferative sarcoma virus). In some embodiments, the MND promoter comprises the nucleic acid sequence of SEQ ID NO: 33.

[0107] As used herein, “conditional” or “inducible” refers to nucleic acid constructs such as promoters that express a gene in the presence of an inducing factor but substantially do not express the gene in the absence of the inducing factor.

[0108] In this specification, "constitutive" refers to a nucleic acid construct that expresses a continuously produced polypeptide, as it contains a constitutive promoter.

[0109] In some embodiments, the inducible promoter exhibits low basal-level activity. In some embodiments, when a lentiviral vector is used, the basal-level activity in cells where expression is not induced is a percentage within the range defined by 20%, 15%, 10%, 5%, 4%, 3%, 2%, 1%, or less (but not 0%) of the activity when gene expression is induced in the cell, or any two of these values. Basal-level activity can be determined by measuring the expression level of the transgene (e.g., a marker gene) in the absence of an inducer (e.g., a drug) using flow cytometry. In some embodiments described herein, expression is measured using a marker protein (e.g., mCherry).

[0110] In some embodiments, when the inducible promoter is expressed, it can induce higher activity compared to when expression is not induced or at the basal level. In some embodiments, the activity level when expression is induced is 2, 4, 6, 8, 9, 10 or more times higher than when expression is not induced, or within a range defined by any two of these values. In some embodiments, the transgene under the control of the inducible promoter is off for a period defined by less than 10 days, less than 8 days, less than 6 days, less than 4 days, less than 2 days, or less than 1 day, or any two of these periods, in the absence of a transactivator, but is not off for 0 days.

[0111] In some embodiments, the inducible promoter is designed and / or modified to have low activity at the basal level, induce high levels of expression, and / or be switchable on and off for a short period of time.

[0112] In some embodiments, the expression vector includes a nucleic acid encoding a protein sequence shown in one or more of SEQ ID NOs: 1, 2, 5, 7, 9, 10, 11, 12, 13, 14, 15, 16, and 17. In some embodiments, the expression vector includes a nucleic acid sequence shown in SEQ ID NO: 28. SEQ ID NO: 28 encodes a protein sequence shown in SEQ ID NOs: 12 and 16.

[0113] In some embodiments, the expression vector is a variant of SEQ ID NO: 28, shown in SEQ ID NO: 29. SEQ ID NO: 29 encodes the protein sequences shown in SEQ ID NO: 10 and SEQ ID NO: 14.

[0114] In some embodiments, the expression vector is a variant of SEQ ID NO: 28, shown in SEQ ID NO: 30. SEQ ID NO: 30 encodes the protein sequences shown in SEQ ID NO: 11 and SEQ ID NO: 15.

[0115] In some embodiments, the expression vector comprises a nucleic acid having at least 80%, 85%, 90%, 95%, 98%, or 99% nucleic acid sequence identity (or nucleic acid sequence identity (%) within the range defined by any two of these percentages) with the nucleotide sequence provided herein, or a specifically induced fragment thereof. In some embodiments, the expression vector comprises a promoter. In some embodiments, the expression vector comprises the nucleic acid encoding a fusion protein. In some embodiments, the vector is RNA or DNA.

[0116] Naked FRB domain We provide naked FKBP-rapamycin-binding (FRB) domain polypeptides for intracellular expression. In some embodiments, the naked FRB domain polypeptide is co-expressed with one or more protein sequences of the first and second CISC components, which are described in more detail herein. The naked FRB domain polypeptide may contain an amino acid sequence having 100%, 99%, 98%, 95%, 90%, 85%, or 80% sequence identity with the sequence shown in SEQ ID NO: 1 or 2, or sequence identity including all numerical values ​​and ranges within these percentages.

[0117] Genome editing system This disclosure provides a genome editing system for cells that expresses the naked FRB domain polypeptide described herein and optionally expresses the CISC described herein by targeting and incorporating a nucleic acid that encodes a naked FRB domain polypeptide and optionally further encodes a CISC into the cell genome. Furthermore, this disclosure provides a system for use in the selective activation and / or expansion of genetically modified cell populations, for example, a system used to prepare a cell population in which genetically modified cells are selectively enriched.

[0118] In some embodiments, i) Deoxyribonucleic acid (DNA) endonuclease, or nucleic acid encoding said DNA endonuclease; ii) A guide RNA (gRNA) containing a spacer sequence complementary to a target sequence within a target genomic locus of a cell, or a nucleic acid encoding said gRNA; and iii) Donor template including a donor cassette containing a nucleic acid sequence encoding a naked FKBP-rapamycin-binding (FRB) domain polypeptide. We provide a system that includes this. The DNA endonuclease, the gRNA, and the donor template are configured such that a complex formed by the association of the DNA endonuclease and the gRNA promotes the targeted incorporation of the donor cassette into the target genomic locus of the cell, thereby producing a genetically modified cell capable of expressing the naked FRB domain polypeptide.

[0119] In some embodiments, according to any of the systems described herein, the gRNA includes a spacer sequence complementary to a sequence within the cell's FOXP3 locus, AAVS1 locus, or TCR(TRAC) locus. In some embodiments, the gRNA includes the spacer sequence shown in any of SEQ ID NOs 40-57, or a variant of the spacer sequence having three or fewer mismatches compared to any of SEQ ID NOs 40-57.

[0120] In this specification, “Cas endonuclease” or “Cas nuclease” includes, for example, RNA-induced DNA endonuclease enzymes associated with the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) adaptive immune system. In this specification, “Cas endonuclease” refers to both natural Cas endonuclease and recombinant Cas endonuclease. In some embodiments, according to any of the systems described herein, the Cas DNA endonuclease is Cas9 endonuclease. In some embodiments, the Cas9 endonuclease is Cas9 endonuclease derived from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 is Cas9 derived from Staphylococcus lugdunensis (SluCas9).

[0121] In some embodiments, according to any of the systems described herein, the system comprises a nucleic acid encoding a DNA endonuclease. In some embodiments, the nucleic acid encoding the DNA endonuclease has codons optimized for expression in a host cell. In some embodiments, the nucleic acid encoding the DNA endonuclease has codons optimized for expression in a human cell. In some embodiments, the nucleic acid encoding the DNA endonuclease is DNA (such as a DNA plasmid). In some embodiments, the nucleic acid encoding the DNA endonuclease is RNA (such as mRNA).

[0122] In some embodiments, according to any of the systems described herein, the donor template is configured to incorporate a donor cassette into a genomic locus targeted by gRNA by homologous recombination repair (HDR) included in the system. In some embodiments, homologous arms corresponding to the sequence of the targeted genomic locus are positioned on both sides of the donor cassette. In some embodiments, the length of the homologous arms is at least 0.2 kb or at least about 0.2 kb (for example, at least 0.3 kb, at least 0.4 kb, at least 0.5 kb, at least 0.6 kb, at least 0.7 kb, at least 0.8 kb, at least 0.9 kb, at least 1 kb or at least more, or at least about 0.3 kb, at least about 0.4 kb, at least about 0.5 kb, at least about 0.6 kb, at least about 0.7 kb, at least about 0.8 kb, at least about 0.9 kb, at least about 1 kb or at least more). In some embodiments, the length of the homologous arm is at least 0.4kb or at least about 0.4kb, for example 0.45kb, 0.6kb, or 0.8kb. A typical homologous arm is further mentioned as a homologous arm contained in a donor template having the sequence of sequence number 32. Typical donor templates include those having the sequences of sequence numbers 3-4, 8, 28-30, 32, and 37-39. In some embodiments, the donor template is encoded in an adeno-associated virus (AAV) vector. In some embodiments, the AAV vector is an AAV2 vector, an AAV5 vector, or an AAV6 vector. In some embodiments, the AAV vector is an AAV6 vector.

[0123] In some embodiments, according to any of the systems described herein, the donor template is configured to incorporate the donor cassette into a genomic locus targeted by the gRNA included in the system via non-homologous end joining (NHEJ). In some embodiments, gRNA target sites are located adjacent to one or both sides of the donor cassette. In some embodiments, gRNA target sites are located adjacent to both sides of the donor cassette. In some embodiments, the gRNA target sites are target sites of the gRNA included in the system. In some embodiments, the gRNA target sites of the donor template are the reverse complementary strand of the cellular genomic gRNA target site targeted by the gRNA included in the system. In some embodiments, the donor template is encoded by an adeno-associated virus (AAV) vector. In some embodiments, the AAV vector is an AAV2 vector, an AAV5 vector, or an AAV6 vector. In some embodiments, the AAV vector is an AAV6 vector.

[0124] In some embodiments, according to any of the systems described herein, the DNA endonuclease or the nucleic acid encoding the DNA endonuclease is formulated by encapsulation in liposomes or lipid nanoparticles. In some embodiments, the liposomes or lipid nanoparticles further comprise gRNA. In some embodiments, the liposomes or lipid nanoparticles are lipid nanoparticles. In some embodiments, the system comprises lipid nanoparticles containing a DNA endonuclease and a nucleic acid encoding gRNA. In some embodiments, the nucleic acid encoding the DNA endonuclease is mRNA encoding the DNA endonuclease.

[0125] In some embodiments, according to any of the systems described herein, the DNA endonuclease associates with gRNA to form a ribonucleoprotein (RNP) complex.

[0126] Genome editing The present invention provides a method for genetically modifying a host cell or organism using gene editing, wherein the method involves expressing one of the naked FRB domain polypeptides disclosed herein and optionally expressing a chemically induced signaling complex polypeptide also disclosed herein. In some embodiments, the gene editing is performed using a CRISPR / Cas system.

[0127] CRISPR Endonuclease System CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) genomic loci can be found in the genomes of many prokaryotes (such as bacteria and archaea). In prokaryotes, CRISPR loci encode products that function as a type of immune system useful in defending prokaryotes from foreign invaders such as viruses and phages. The function of a CRISPR locus consists of three stages: incorporation of new sequences into the CRISPR locus, expression of CRISPR RNA (crRNA), and silencing of foreign invading nucleic acids. Five types of CRISPR systems (e.g., type I, type II, type III, type U, and type V) have been identified.

[0128] The CRISPR locus contains numerous short repeat sequences called "repetitive sequences." When expressed, these repeat sequences can form a secondary hairpin structure (e.g., a hairpin) and / or have a single-stranded sequence that does not have a fixed structure. Repetitive sequences usually occur as clusters and vary considerably between species. The repeat sequences are spaced apart by unique intervening sequences called "spacers," forming a locus with a repeat sequence-spacer-repetitive sequence structure. The spacers are either identical to or highly homologous to known invading sequences. The unit consisting of spacers and repeat sequences codes for crisprRNA (crRNA), which is processed to obtain the mature form of the unit consisting of spacers and repeat sequences. crRNA has a "seed" sequence or spacer sequence involved in targeting the target nucleic acid (in its natural form in prokaryotes, the spacer sequence targets the nucleic acid of the invader). The spacer sequence is located at the 5' or 3' end of the crRNA.

[0129] The CRISPR locus further contains polynucleotide sequences encoding CRISPR-related (Cas) genes. Cas genes encode endonucleases involved in the biosynthetic and interference stages of crRNA function in prokaryotes. Some Cas genes have homologous secondary and / or tertiary structures.

[0130] Genome-targeted nucleic acids form complexes by interacting with site-directed polypeptides (e.g., nucleic acid-inducible nucleases such as Cas9). The genome-targeted nucleic acid (e.g., gRNA) leads the site-directed polypeptide to the target nucleic acid.

[0131] As described above, in some embodiments, site-directed polypeptides and genome-targeted nucleic acids can be administered to cells or subjects separately. On the other hand, in some other embodiments, site-directed polypeptides can be pre-complexed with one or more guide RNAs, or site-directed polypeptides can be pre-complexed with tracrRNA and one or more crRNAs. The pre-complexed material can then be administered to cells or subjects. Such pre-complexed materials are known as ribonucleoprotein particles (RNPs).

[0132] Type II CRISPR System In the biosynthesis of crRNA in the naturally occurring type II CRISPR system, trans-activated CRISPR RNA (tracrRNA) is required. The tracrRNA is modified by endogenous RNase III and then hybridizes to the crRNA repeat sequence in the pre-crRNA array. Endogenous RNase III is recruited to cleave the pre-crRNA. The cleaved crRNA is trimmed by exoribonuclease (e.g., 5' end trimming) to produce the mature form of crRNA. The tracrRNA remains hybridized to the crRNA, and the tracrRNA and crRNA associate with a site-directed polypeptide (e.g., Cas9). In the crRNA-tracrRNA-Cas9 complex, the complex is led by the crRNA to a target nucleic acid that the crRNA can hybridize to. Hybridization of the crRNA to the target nucleic acid activates Cas9, which then cleaves the target nucleic acid. In the type II CRISPR system, the target nucleic acid is called a protospacer adjacent motif (PAM). In fact, PAM is essential for facilitating the binding of site-directed polypeptides (e.g., Cas9) to target nucleic acids. Type II systems (also known as Nmeni or CASS4) are further subdivided into type II-A (CASS4) and type II-B (CASS4a). Jinek, M. et al. (2012). Science, 337(6096):816-821 demonstrates the usefulness of the CRISPR / Cas9 system for RNA-programmable genome editing, and further, international patent application WO 2013 / 176772 describes numerous examples and applications of the CRISPR / Cas endonuclease system for site-directed gene editing.

[0133] V-type CRISPR system The type V CRISPR system has several key differences from the type II system. For example, Cpf1 is a single-chain RNA-inducible endonuclease, but unlike the type II system, it lacks tracrRNA. In fact, a CRISPR array associated with Cpf1 processes into mature crRNA without requiring further transactivated tracrRNA. The type V CRISPR array processes into short mature crRNAs of 42–44 nucleotides in length, each mature crRNA beginning with a 19-nucleotide direct repeat sequence followed by a 23–25 nucleotide spacer sequence. In contrast, mature crRNAs in the type II system begin with a 20–24 nucleotide spacer sequence followed by a 22-nucleotide or approximately 22-nucleotide direct repeat sequence. Furthermore, Cpf1 utilizes a T-rich protospacer flanking motif, allowing the target DNA following this short T-rich PAM to be efficiently cleaved by the Cpf1-crRNA complex. This is in contrast to the type II system, where a G-rich PAM following the target DNA is utilized. Therefore, the type V system cleaves the target at a position away from the PAM, while the type II system cleaves the target adjacent to the PAM. Furthermore, in contrast to the type II system, Cpf1 cleaves the DNA double strand at a shifted position, resulting in a 4-nucleotide or 5-nucleotide 5' end overhang. In contrast, double-strand breaks by the type II system result in blunt ends. Cpf1 is predicted to contain a RuvC-like endonuclease domain, similar to the type II system, but unlike the type II system, it lacks a second HNH endonuclease domain.

[0134] Site-directed polypeptide or DNA endonuclease Modification of target DNA by NHEJ and / or HDR can result in, for example, mutations, deletions, alterations, integrations, gene modifications, gene substitutions, gene tagging, transgene insertions, nucleotide deletions, gene disruption, translocations, and / or gene mutations. The integration of non-native nucleic acids into genomic DNA is an example of genome editing.

[0135] Site-directed polypeptides are nucleases used in genome editing to cleave DNA. Site-directed polypeptides can be administered to cells or subjects as one or more polypeptides, or as one or more mRNAs encoding such polypeptides.

[0136] In relation to the CRISPR / Cas system or CRISPR / Cpf1 system, the site-directed polypeptide binds to a guide RNA, thereby allowing the guide RNA to identify the site on the target DNA to which the site-directed polypeptide is directed. In embodiments of the CRISPR / Cas system or CRISPR / Cpf1 system described herein, the site-directed polypeptide is an endonuclease, such as a DNA endonuclease. In this specification, "Cas endonuclease" or "Cas nuclease" includes, for example, RNA-induced DNA endonuclease enzymes associated with the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) adaptive immune system. In this specification, "Cas endonuclease" refers to both natural Cas endonucleases and recombinant Cas endonucleases.

[0137] In some embodiments, the site-directed polypeptide has multiple nucleic acid cleavage domains (e.g., nuclease sites). Two or more nucleic acid cleavage domains can be linked via a linker. In some embodiments, the linker is flexible. The length of the linker may be 1 amino acid length, 2 amino acid length, 3 amino acid length, 4 amino acid length, 5 amino acid length, 6 amino acid length, 7 amino acid length, 8 amino acid length, 9 amino acid length, 10 amino acid length, 11 amino acid length, 12 amino acid length, 13 amino acid length, 14 amino acid length, 15 amino acid length, 16 amino acid length, 17 amino acid length, 18 amino acid length, 19 amino acid length, 20 amino acid length, 21 amino acid length, 22 amino acid length, 23 amino acid length, 24 amino acid length, 25 amino acid length, 30 amino acid length, 35 amino acid length, 40 amino acid length, or longer.

[0138] The naturally occurring wild-type Cas9 enzyme has two nuclease domains: an HNH nuclease domain and a RuvC domain. In this specification, "Cas9" refers to both natural Cas9 and recombinant Cas9. The Cas9 enzymes described herein have an HNH nuclease domain or an HNH-like nuclease domain, and / or a RuvC nuclease domain or a RuvC-like nuclease domain.

[0139] The HNH domain or HNH-like domain has McrA-like folding. The HNH domain or HNH-like domain has two antiparallel β-chains and one α-helix. The HNH domain or HNH-like domain has a metal-binding site (e.g., a divalent cation-binding site). The HNH domain or HNH-like domain can cleave a single strand of the target nucleic acid (e.g., the complementary strand of the target strand of crRNA).

[0140] RuvC domains or RuvC-like domains possess RNaseH folding or RNaseH-like folding. RuvC domains or RNaseH domains are involved in a diverse range of nucleic acid-based functions, including actions on both RNA and DNA. RNaseH domains have five β-chains surrounded by multiple α-helices. RuvC domains or RNaseH domains or RuvC-like domains or RNaseH-like domains possess metal-binding sites (e.g., divalent cation-binding sites). RuvC domains or RNaseH domains or RuvC-like domains or RNaseH-like domains can cleave a single strand of target nucleic acid (e.g., the non-complementary strand of a double-stranded target DNA).

[0141] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity with a typical wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes, SEQ ID NO: 8 described in US2014 / 0068797, or Cas9 described in Sapranauskas, R. et al. (2011). Nucl. Acids Res, 39(21): 9275-9282), and various other site-directed polypeptides).

[0142] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% amino acid sequence identity with the nuclease domain of a typical wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes (shown above)).

[0143] In some embodiments, the site-directed polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes) in a sequence of 10 amino acids. In some embodiments, the site-directed polypeptide has at most 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes) in a sequence of 10 amino acids. In some embodiments, the HNH nuclease domain of the site-directed polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes) in a sequence of 10 amino acids. In some embodiments, the RuvC nuclease domain of the site-directed polypeptide has at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% identity with the wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes) in a sequence of 10 amino acids.

[0144] In some embodiments, the site-directed polypeptide has a modified form of a typical wild-type site-directed polypeptide. The modified form of a typical wild-type site-directed polypeptide has mutations that reduce the nucleic acid cleavage activity of the site-directed polypeptide. In some embodiments, the modified form of a typical wild-type site-directed polypeptide has nucleic acid cleavage activity of less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity of a typical wild-type site-directed polypeptide (e.g., Cas9 from S. pyogenes mentioned above). In addition, the modified form of the site-directed polypeptide may substantially lack nucleic acid cleavage activity. When the site-directed polypeptide is a modified form that substantially lacks nucleic acid cleavage activity, such a modified form is referred to herein as "enzymatically inactive."

[0145] In some embodiments, the modification of the site-directed polypeptide has a mutation that can induce a single-strand break (SSB) on the target nucleic acid (for example, by cleaving only one of the sugar-phosphate backbones of the double-stranded target nucleic acid). In some embodiments, this mutation results in a nucleic acid cleavage activity of less than 90%, less than 80%, less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nucleic acid cleavage activity of one or more of the multiple nucleic acid cleavage domains of the wild-type site-directed polypeptide (e.g., Cas9 derived from S. pyogenes as described above). In some embodiments, the mutation reduces the ability of one or more of the multiple nucleic acid cleavage domains to cleave the non-complementary strand of the target nucleic acid while retaining the ability to cleave the complementary strand. In some embodiments, the mutation reduces the ability of one or more of the multiple nucleic acid cleavage domains to cleave the complementary strand of the target nucleic acid while retaining the ability to cleave the non-complementary strand. For example, mutations occur in which one or more nucleic acid cleavage domains (e.g., nuclease domains) are inactivated at residues of a typical wild-type S. pyogenes Cas9 polypeptide, such as Asp10, His840, Asn854, and Asn856. In some embodiments, the residues to which mutations are induced correspond to the Asp10, His840, Asn854, and Asn856 residues of a typical wild-type S. pyogenes Cas9 polypeptide (as determined, for example, by sequence and / or structural alignment). Examples of such mutations include, but are not limited to, D10A, H840A, N854A, or N856A. Those skilled in the art will understand that mutations other than alanine substitutions are appropriate.

[0146] In some embodiments, the D10A mutation, when combined with one or more of the H840A, N854A, and N856A mutations, generates a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the H840A mutation, when combined with one or more of the D10A, N854A, and N856A mutations, generates a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the N854A mutation, when combined with one or more of the H840A, D10A, and N856A mutations, generates a site-directed polypeptide that substantially lacks DNA cleavage activity. In some embodiments, the N856A mutation, when combined with one or more of the H840A, N854A, and D10A mutations, generates a site-directed polypeptide that substantially lacks DNA cleavage activity. Site-directed polypeptides having a substantially inactive single nuclease domain are called "nickases".

[0147] In some embodiments, variants of RNA-inducible endonucleases (e.g., Cas9) can be used to enhance the specificity of CRISPR-mediated genome editing. For example, wild-type Cas9 is typically led by a single-strand guide RNA designed to hybridize with a specific sequence of about 20 nucleotides in the target sequence (such as an endogenous genomic locus). However, since a few mismatches can be tolerated between the guide RNA and the target locus, the required homologous sequence length at the target site may be efficiently shortened to, for example, about 13 nt, thereby increasing the likelihood that the CRISPR / Cas9 complex will bind to another location in the target genome and cleave a double-strand nucleic acid. This is also known as an off-target cleavage. On the other hand, since each Cas9 nickase variant cleaves only one strand, a pair of nickases must bind in close proximity on opposite strands of the target nucleic acid to produce a double-strand break, thereby creating a pair of nicks and resulting in a double-strand break. This requires two separate guide RNAs (one for each nickase) to bind in close proximity on opposite strands of the target nucleic acid. Adhering to this requirement effectively doubles the minimum length of homologous sequence needed to produce a double-strand break, thus reducing the likelihood of the double-strand break occurring elsewhere in the genome, as the two guide RNA sites (if present) are less likely to be close enough to each other to produce a double-strand break. As reported in the art, nickases can also be used to facilitate HDR rather than NHEJ. By performing HDR using a specific donor sequence that effectively mediates the desired modification, the selected modification can be introduced into a target site in the genome. Descriptions of various CRISPR / Cas systems for use in gene editing are found, for example, in international patent application WO2013 / 176772 and Sander, JD et al. (2014). Nat. Biotechnol., 32(4):347-355, as well as the references cited in these publications.

[0148] In some embodiments, site-directed polypeptides (e.g., variants of site-directed polypeptides, variants of site-directed polypeptides, enzymatically inactive site-directed polypeptides, and / or site-directed polypeptides enzymatically inactive under specific conditions) target nucleic acids. In some embodiments, site-directed polypeptides (e.g., variants of endoribonucleases, variants of endoribonucleases, enzymatically inactive endoribonucleases, and / or enzymatically inactive endoribonucleases under specific conditions) target DNA. In some embodiments, site-directed polypeptides (e.g., variants of endoribonucleases, variants of endoribonucleases, enzymatically inactive endoribonucleases, and / or enzymatically inactive endoribonucleases under specific conditions) target RNA.

[0149] In some embodiments, the site-directed polypeptide has one or more non-natural sequences (for example, the site-directed polypeptide is a fusion protein).

[0150] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), a nucleic acid-binding domain, and two nucleic acid-cleaving domains (such as an HNH domain and a RuvC domain).

[0151] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), and two nucleic acid cleavage domains (such as an HNH domain and a RuvC domain).

[0152] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), and two nucleic acid cleavage domains, one or both of which have at least 50% amino acid identity with the nuclease domain of Cas9 derived from bacteria (e.g., S. pyogenes).

[0153] In some embodiments, the site-directed polypeptide comprises an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), two nucleic acid cleavage domains (e.g., an HNH domain and a RuvC domain), and a linker that anneals a non-natural sequence (e.g., a nuclear localization signal) or the site-directed polypeptide to the non-natural sequence.

[0154] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), and two nucleic acid cleavage domains (e.g., an HNH domain and a RuvC domain), with one or both of the nucleic acid cleavage domains having a mutation that reduces the cleavage activity of these nuclease domains by at least 50%.

[0155] In some embodiments, the site-directed polypeptide has an amino acid sequence having at least 15% amino acid identity with Cas9 derived from bacteria (e.g., S. pyogenes), and two nucleic acid cleavage domains (e.g., an HNH domain and a RuvC domain), with a mutation at the 10th aspartic acid position of one of the nuclease domains and / or a mutation at the 840th histidine position of one of the nuclease domains, the mutation reducing the cleavage activity of the nuclease domain by at least 50%.

[0156] In some embodiments, one or more site-directed polypeptides (e.g., DNA endonucleases) comprise two nickases that cooperate to produce one double-strand break at a specific locus in the genome, or one or more site-directed polypeptides comprise four nickases that cooperate to produce two double-strand breaks at a specific locus in the genome, or one site-directed polypeptide (e.g., DNA endonuclease) acts to produce one double-strand break at a specific locus in the genome.

[0157] In some embodiments, polynucleotides encoding site-directed polypeptides can be used for genome editing. In some such embodiments, the polynucleotide encoding the site-directed polypeptide is codon-optimized for expression in cells containing the target DNA of interest, according to methods standard in the art. For example, if the target nucleic acid of interest is present in human cells, it is conceivable that a polynucleotide encoding Cas9 would be optimized for human codons and used to construct the Cas9 polypeptide.

[0158] Cas gene / polypeptide and protospacer adjacent motif A typical CRISPR / Cas polypeptide is the Cas9 polypeptide, shown in Figure 1 of Fonfara, I. et al. (2014). Nucl. Acids Res., 42(4):2577-2590 (this document is incorporated herein by reference). The CRISPR / Cas gene naming system has undergone significant revisions since the discovery of the Cas gene. Figure 5 by Fonfara et al. (2014) shows PAM sequences for Cas9 polypeptides derived from various species.

[0159] nucleic acid Genome-targeted nucleic acids or guide RNA This disclosure provides genome-targeted nucleic acids that can direct the activity of a relevant polypeptide (e.g., a site-directed polypeptide or DNA endonuclease) to a specific target sequence (e.g., a gene of interest) within a target nucleic acid. In some embodiments, the genome-targeted nucleic acid is RNA. Hereinafter, the genome-targeted RNA is referred to as “guide RNA” or “gRNA”. The guide RNA has at least a spacer sequence and a CRISPR repeat sequence that hybridize to a target nucleic acid sequence of interest. In a type II system, the gRNA further has a second RNA called a tracrRNA sequence. In a type II guide RNA (gRNA), the CRISPR repeat sequence and the tracrRNA sequence hybridize to form a double helix. In a type V guide RNA (gRNA), the CRISPR repeat RNA (crRNA) forms a double helix. In either system, the double helix binds to a site-specific polypeptide to form a guide RNA-site-specific polypeptide complex. The genome-targeted nucleic acid confers target specificity to the complex formed by association with the site-specific polypeptide. Therefore, this genome-targeted nucleic acid confers directionality to the activity of site-specific polypeptides.

[0160] In some embodiments, the genome-targeting nucleic acid is a bimolecule guide RNA. In some embodiments, the genome-targeting nucleic acid is a single-molecule guide RNA. The bimolecule guide RNA has two RNA strands. The first strand has an optional spacer extension sequence, a spacer sequence, and a CRISPR minimal repeat sequence in the direction from the 5' end to the 3' end. The second strand has a tracrRNA minimal sequence (complementary to the CRISPR minimal repeat sequence), a 3' tracrRNA sequence, and an optional tracrRNA extension sequence. The single-molecule guide RNA (sgRNA) of the type II system has an optional spacer extension sequence, a spacer sequence, a CRISPR minimal repeat sequence, a single-molecule guide linker, a tracrRNA minimal sequence, a 3' tracrRNA sequence, and an optional tracrRNA extension sequence in the direction from the 5' end to the 3' end. The optionally provided tracrRNA extension sequence may have elements that confer further functionality (e.g., stability) to the guide RNA. The single-molecule guide linker links the CRISPR minimal repeat sequence and the tracrRNA minimal sequence to form a hairpin structure. An optional tracrRNA elongation sequence has one or more hairpins. The single-molecule guide RNA (sgRNA) of the V-type system has the CRISPR minimal repeat sequence and a spacer sequence in the direction from the 5' end to the 3' end.

[0161] As an example, guide RNAs used in CRISPR / Cas / Cpf1 systems, or other smaller RNAs, can be readily synthesized by chemical means known in the art and described later. Although chemical synthesis procedures are continuously being expanded, the length of the polynucleotides significantly exceeds approximately 100 nucleotides, making the purification of such RNAs by procedures such as high-performance liquid chromatography (HPLC without gels, such as PAGE) difficult. One approach used to produce long RNAs is to produce two or more molecules and then ligate them. Very long RNAs, such as RNA encoding Cas9 endonuclease or Cpf1 endonuclease, can be readily produced enzymatically. Various RNA modifications can be introduced during or after the chemical synthesis and / or enzymatic production of RNA, such as modifications that increase stability, modifications that reduce the likelihood or degree of innate immune response, and / or modifications that enhance other attributes, as reported in the art.

[0162] Spacer extension arrangement In some embodiments of genome-targeted nucleic acids, spacer extension sequences can modulate activity, confer stability, and / or provide sites for modifying genome-targeted nucleic acids. Spacer extension sequences can also modulate on-target or off-target activity or specificity. In some embodiments, spacer extension sequences are provided. The lengths of the spacer extension sequences are greater than 1 nucleotide, greater than 5 nucleotides, greater than 10 nucleotides, greater than 15 nucleotides, greater than 20 nucleotides, greater than 25 nucleotides, greater than 30 nucleotides, greater than 35 nucleotides, greater than 40 nucleotides, greater than 45 nucleotides, greater than 50 nucleotides, greater than 60 nucleotides, greater than 70 nucleotides, greater than 80 nucleotides, greater than 90 nucleotides, greater than 100 nucleotides, greater than 120 nucleotides, greater than 140 nucleotides, greater than 160 nucleotides, greater than 180 nucleotides, and greater than 200 nucleotides. The length may exceed the creotide length, exceed 220 nucleotides, exceed 240 nucleotides, exceed 260 nucleotides, exceed 280 nucleotides, exceed 300 nucleotides, exceed 320 nucleotides, exceed 340 nucleotides, exceed 360 nucleotides, exceed 380 nucleotides, exceed 400 nucleotides, exceed 1000 nucleotides, exceed 2000 nucleotides, exceed 3000 nucleotides, exceed 4000 nucleotides, exceed 5000 nucleotides, exceed 6000 nucleotides, or exceed 7000 nucleotides, or any other nucleotide length. The length of the spacer extension sequence may be 1 nucleotide or approximately 1 nucleotide, 5 nucleotides or approximately 5 nucleotides, 10 nucleotides or approximately 10 nucleotides, 15 nucleotides or approximately 15 nucleotides, 20 nucleotides or approximately 20 nucleotides,25 nucleotides or approximately 25 nucleotides, 30 nucleotides or approximately 30 nucleotides, 35 nucleotides or approximately 35 nucleotides, 40 nucleotides or approximately 40 nucleotides, 45 nucleotides or approximately 45 nucleotides, 50 nucleotides or approximately 50 nucleotides, 60 nucleotides or approximately 60 nucleotides, 70 nucleotides or approximately 70 nucleotides, 80 nucleotides or approximately 80 nucleotides, 90 nucleotides or approximately 90 nucleotides, 100 nucleotides or approximately 100 nucleotides, 120 nucleotides or approximately 120 nucleotides, 140 nucleotides or approximately 140 nucleotides, 160 nucleotides or approximately 160 nucleotides, 180 nucleotides or approximately 180 nucleotides, 200 nucleotides or approximately 200 nucleotides, 220 nucleotides or approximately 220 nucleotides, 240 nucleotides or approximately 24 The nucleotide length may be 0 nucleotides, 260 nucleotides or approximately 260 nucleotides, 280 nucleotides or approximately 280 nucleotides, 300 nucleotides or approximately 300 nucleotides, 320 nucleotides or approximately 320 nucleotides, 340 nucleotides or approximately 340 nucleotides, 360 nucleotides or approximately 360 nucleotides, 380 nucleotides or approximately 380 nucleotides, 400 nucleotides or approximately 400 nucleotides, 1000 nucleotides or approximately 1000 nucleotides, 2000 nucleotides or approximately 2000 nucleotides, 3000 nucleotides or approximately 3000 nucleotides, 4000 nucleotides or approximately 4000 nucleotides, 5000 nucleotides or approximately 5000 nucleotides, 6000 nucleotides or approximately 6000 nucleotides, or 7000 nucleotides or approximately 7000 nucleotides, or greater. The length of the spacer extension sequence is less than 1 nucleotide, less than 5 nucleotides, less than 10 nucleotides, less than 15 nucleotides, less than 20 nucleotides, less than 25 nucleotides, less than 30 nucleotides, less than 35 nucleotides, less than 40 nucleotides.The length may be less than 45 nucleotides, less than 50 nucleotides, less than 60 nucleotides, less than 70 nucleotides, less than 80 nucleotides, less than 90 nucleotides, less than 100 nucleotides, less than 120 nucleotides, less than 140 nucleotides, less than 160 nucleotides, less than 180 nucleotides, less than 200 nucleotides, less than 220 nucleotides, less than 240 nucleotides, less than 260 nucleotides, less than 280 nucleotides, less than 300 nucleotides, less than 320 nucleotides, less than 340 nucleotides, less than 360 nucleotides, less than 380 nucleotides, less than 400 nucleotides, less than 1000 nucleotides, less than 2000 nucleotides, less than 3000 nucleotides, less than 4000 nucleotides, less than 5000 nucleotides, less than 6000 nucleotides, less than 7000 nucleotides, or more. In some embodiments, the length of the spacer extension sequence is less than 10 nucleotides. In some embodiments, the length of the spacer extension sequence is 10 to 30 nucleotides. In some embodiments, the length of the spacer extension sequence is 30 to 70 nucleotides.

[0163] In some embodiments, the spacer extension sequence has another portion (e.g., a stability control sequence, an endoribonuclease binding sequence, a ribozyme, etc.). In some embodiments, the other portion targets a nucleic acid and reduces or increases the stability of the nucleic acid. In some embodiments, the other portion is a transcriptional terminator segment (e.g., a transcription termination sequence). In some embodiments, the other portion functions in eukaryotic cells. In some embodiments, the other portion functions in prokaryotic cells. In some embodiments, the other portion functions in both eukaryotic and prokaryotic cells. Examples of suitable parts include, but are not limited to, 5' end caps (e.g., 7-methylguanylate caps (m7G)); riboswitch sequences (e.g., those that stabilize under control and / or allow access by proteins or protein complexes under control); sequences that form dsRNA double strands (e.g., hairpins); sequences that target RNA to subcellular locations (e.g., nucleus, mitochondria, chloroplasts, etc.); modifications or sequences that enable tracking (e.g., direct binding to fluorescent molecules, binding to regions that facilitate fluorescence detection, sequences that enable fluorescence detection, etc.); and / or modifications or sequences that provide binding sites for proteins (e.g., DNA-acting proteins such as transcription activators, transcription repressors, DNA methyltransferases, DNA methyl-degrading enzymes, histone acetyltransferases, histone deacetylases, etc.).

[0164] Spacer array The spacer sequence hybridizes to the sequence within the target nucleic acid. The spacer in the genome-targeted nucleic acid interacts with the target nucleic acid in a sequence-specific manner through hybridization (e.g., base pairing). Therefore, the nucleotide sequence of the spacer varies depending on the sequence of the target nucleic acid.

[0165] In the CRISPR / Cas system described herein, the spacer sequence is designed to hybridize to a target nucleic acid located at the 5' end of the PAM in the Cas enzyme used in this system (e.g., the Cas9 enzyme). The spacer may be a perfect match or a mismatch with the target sequence. Each Cas enzyme has a specific PAM sequence that recognizes the target DNA. For example, Cas9 derived from Streptococcus pyogenes recognizes a PAM in the target nucleic acid having the sequence 5'-NRG-3' (where R represents A or G and N is any nucleotide adjacent to the 3' end of the target nucleic acid sequence targeted by the spacer sequence).

[0166] In some embodiments, the target nucleic acid sequence has 20 nucleotides. In some embodiments, the target nucleic acid has fewer than 20 nucleotides. In some embodiments, the target nucleic acid has more than 20 nucleotides. In some embodiments, the target nucleic acid has at least 5, at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, or at least more nucleotides. In some embodiments, the target nucleic acid has at most 5, at most 10, at most 15, at most 16, at most 17, at most 18, at most 19, at most 20, at most 21, at most 22, at most 23, at most 24, at most 25, at most 30, or at most more nucleotides. In some embodiments, the target nucleic acid sequence has 20 bases adjacent to the 5' end of the first nucleotide of the PAM. For example, 5'-NNNNNNNNNNNNNNNNNNNN NRGIn sequences having -3', the target nucleic acid has sequences corresponding to these Ns (where N is any nucleotide), and the underlined sequence NRG (where R is G or A) is a Cas9 PAM derived from Streptococcus pyogenes. In some embodiments, the PAM sequence used in the compositions and methods of this disclosure as a sequence recognized by Cas9 derived from Streptococcus pyogenes (Sp) is NGG.

[0167] In some embodiments, the length of the spacer sequence hybridizing to the target nucleic acid is at least 6 nucleotides (nt) or at least about 6 nucleotides (nt). The length of the spacer sequence may be at least 6 nt or at least about 6 nt, at least 10 nt or at least about 10 nt, at least 15 nt or at least about 15 nt, at least 18 nt or at least about 18 nt, at least 19 nt or at least about 19 nt, at least 20 nt or at least about 20 nt, at least 25 nt or at least about 25 nt, at least 30 nt or at least about 30 nt, at least 35 nt or at least about 35 nt, and or at least 40nt or at least approximately 40nt, 6nt-80nt or approximately 6nt-approximately 80nt, 6nt-50nt or approximately 6nt-approximately 50nt, 6nt-45nt or approximately 6nt-approximately 45nt, 6nt-40nt or approximately 6nt-approximately 40nt, 6nt-35nt or approximately 6nt-approximately 35nt, 6nt-30nt or approximately 6nt-approximately 30nt, 6nt-25nt or approximately 6nt-approximately 25nt, 6nt-20nt or approximately 6nt-approximately 20nt, 6nt-19nt or approximately 6nt-approximately 19nt, 10nt-50nt or approximately 10nt-50nt, 10nt-45nt or approximately 10nt-45nt, 10nt-40nt or approximately 10nt-40nt, 10nt-35nt or approximately 10nt-35nt, 10nt-30nt or approximately 10nt-30nt, 10nt-25nt or approximately 10nt-25nt, 10nt-20nt or approximately 10nt-20nt, 10nt-19nt or approximately 10nt-19nt, 19nt-25nt or approximately 19nt-25 nt, 19nt~30nt or approximately 19nt~approximately 30nt, 19nt~35nt or approximately 19nt~approximately 35nt, 19nt~40nt or approximately 19nt~approximately 40nt, 19nt~45nt or approximately 19nt~approximately 45nt, 19nt~50nt or approximately 19nt~approximately 50nt, 19nt~60nt or approximately 19nt~approximately 60nt, 20nt~25nt or approximately 20nt~approximately 25nt, 20nt~30nt or approximately 20nt~approximately 30nt, 20nt~35nt or approximately 20nt~approximately 35nt,The spacer sequence may be 20nt to 40nt or approximately 20nt to approximately 40nt, 20nt to 45nt or approximately 20nt to approximately 45nt, 20nt to 50nt or approximately 20nt to approximately 50nt, or 20nt to 60nt or approximately 20nt to approximately 60nt. In some embodiments, the spacer sequence is 20 nucleotides long. In some embodiments, the spacer is 19 nucleotides long. In some embodiments, the spacer is 18 nucleotides long. In some embodiments, the spacer is 17 nucleotides long. In some embodiments, the spacer is 16 nucleotides long. In some embodiments, the spacer is 15 nucleotides long.

[0168] In some embodiments, the complementarity (%) between the spacer sequence and the target nucleic acid is at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or at least 100%. In some embodiments, the complementarity (%) between the spacer sequence and the target nucleic acid is at most 30% or about 30%, at most 40% or about 40%, at most 50% or about 50%, at most 60% or about 60%, at most 65% or about 65%, at most 70% or about 70%, at most 75% or about 75%, at most 80% or about 80%, at most 85% or about 85%, at most 90% or about 90%, at most 95% or about 95%, at most 97% or about 97%, at most 98% or about 98%, at most 99% or about 99%, or at most about 100%. In some embodiments, the complementarity (%) between the spacer sequence and the target nucleic acid is 100% at six consecutive 5' terminal nucleotides of the target sequence in the complementary strand of the target nucleic acid. In some embodiments, the complementarity (%) between the spacer sequence and the target nucleic acid is at least 60% at 20 or approximately 20 consecutive nucleotides. In some embodiments, the length of the spacer sequence and the length of the target nucleic acid may differ by 1 to 6 nucleotides, and this difference can be considered as one or more bulges.

[0169] In some embodiments, spacer sequences are designed or selected using a computer program. The computer program may use variables, such as, for example, predicted melting temperature, secondary structure formation, predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, GC ratio (%), genomic frequency (e.g., changes at one or more locations caused by mismatches, insertions, or deletions in identical or similar sequences), methylation status, and the presence of SNPs.

[0170] CRISPR's minimum iteration array In some embodiments, the minimal repeat sequence of CRISPR is a sequence having at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 100% sequence identity with a reference CRISPR repeat sequence (e.g., crRNA from S. pyogenes).

[0171] In some embodiments, the minimal repeat sequence of CRISPR has nucleotides that can hybridize to the minimal sequence of tracrRNA in cells. The minimal repeat sequence of CRISPR and the minimal sequence of tracrRNA form a double helix, for example, a base-paired double-stranded structure. Together, the minimal repeat sequence of CRISPR and the minimal sequence of tracrRNA bind to a site-directed polypeptide. At least a portion of the minimal repeat sequence of CRISPR hybridizes to the minimal sequence of tracrRNA. In some embodiments, at least a portion of the minimal repeat sequence of CRISPR has at least 30% or about 30%, at least 40% or about 40%, at least 50% or about 50%, at least 60% or about 60%, at least 65% or about 65%, at least 70% or about 70%, at least 75% or about 75%, at least 80% or about 80%, at least 85% or about 85%, at least 90% or about 90%, at least 95% or about 95%, or at least 100% complementarity with the minimal sequence of tracrRNA. In some embodiments, at least a portion of the minimal repeat sequence of CRISPR has complementarity with the minimal sequence of tracrRNA of at most 30% or about 30%, at most 40% or about 40%, 50% or about 50%, at most 60% or about 60%, at most 65% or about 65%, at most 70% or about 70%, at most 75% or about 75%, at most 80% or about 80%, at most 85% or about 85%, 90% or about 90%, at most 95% or about 95%, or at most 100%.

[0172] The minimum repeat sequence length for CRISPR is 7 nucleotides to 100 nucleotides or approximately 7 nucleotides to approximately 100 nucleotides. For example, the minimum repeat sequence length for CRISPR is 7 nucleotides (nt) to 50 nt or approximately 7 nt to approximately 50 nt, 7 nt to 40 nt or approximately 7 nt to approximately 40 nt, 7 nt to 30 nt or approximately 7 nt to approximately 30 nt, 7 nt to 25 nt or approximately 7 nt to approximately 25 nt, 7 nt to 20 nt or approximately 7 nt to approximately 20 nt, 7 nt to 15 nt or approximately 7 nt to approximately 15 nt, 8 nt to 40 nt or approximately 8 nt to approximately 40 nt, 8 nt to 30 nt or approximately 8 nt to approximately 30 nt, 8 nt to 2 The minimum repeat length of CRISPR is approximately 5 nucleotides or 8 nucleotides to 25 nucleotides, 8 nucleotides to 20 nucleotides or 8 nucleotides to 20 nucleotides, 8 nucleotides to 15 nucleotides or 8 nucleotides to 15 nucleotides, 15 nucleotides to 100 nucleotides or 15 nucleotides to 100 nucleotides, 15 nucleotides to 80 nucleotides or 15 nucleotides to 80 nucleotides, 15 nucleotides to 50 nucleotides or 15 nucleotides to 50 nucleotides, 15 nucleotides to 40 nucleotides or 15 nucleotides to 40 nucleotides, 15 nucleotides to 30 nucleotides or 15 nucleotides to 30 nucleotides, or 15 nucleotides to 25 nucleotides or 15 nucleotides to 25 nucleotides. In some embodiments, the minimum repeat length of CRISPR is approximately 9 nucleotides. In some embodiments, the minimum repeat length of CRISPR is approximately 12 nucleotides.

[0173] In some embodiments, the CRISPR minimal repeat sequence has at least 60% or at least about 60% identity with a reference CRISPR minimal repeat sequence (e.g., wild-type crRNA from S. pyogenes) in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides. For example, the CRISPR minimal repeat sequence has at least 65% or at least about 65% identity with a reference CRISPR minimal repeat sequence in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides, at least 70% or at least about 70% identity, at least 75% or at least about 75% identity, at least 80% or at least about 80% identity, at least 85% or at least about 85% identity, at least 90% or at least about 90% identity, at least 95% or at least about 95% identity, at least 98% or at least about 98% identity, at least 99% or at least about 99% identity, or at least 100% identity.

[0174] Minimal sequence of tracrRNA In some embodiments, the minimum tracrRNA sequence is a sequence having at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 100% sequence identity with a reference tracrRNA sequence (e.g., wild-type tracrRNA from S. pyogenes).

[0175] In some embodiments, the minimal tracrRNA sequence has nucleotides that can hybridize to the minimal CRISPR repeat sequence in cells. The minimal tracrRNA sequence and the minimal CRISPR repeat sequence form a double helix, for example, a base-paired double-stranded structure. Together, the minimal tracrRNA sequence and the minimal CRISPR repeat sequence bind to a site-directed polypeptide. At least a portion of the minimal tracrRNA sequence can hybridize to the minimal CRISPR repeat sequence. In some embodiments, the minimal tracrRNA sequence has at least 30% or about 30%, at least 40% or about 40%, at least 50% or about 50%, at least 60% or about 60%, at least 65% or about 65%, at least 70% or about 70%, at least 75% or about 75%, at least 80% or about 80%, at least 85% or about 85%, at least 90% or about 90%, at least 95% or about 95%, or at least 100% complementarity with the minimal repeat sequence of CRISPR.

[0176] The minimum sequence length of tracrRNA is between 7 nucleotides and 100 nucleotides, or approximately 7 nucleotides and 100 nucleotides. For example, the minimum sequence length of tracrRNA is 7 nucleotides (nt) to 50 nt or approximately 7 nt to 50 nt, 7 nt to 40 nt or approximately 7 nt to 40 nt, 7 nt to 30 nt or approximately 7 nt to 30 nt, 7 nt to 25 nt or approximately 7 nt to 25 nt, 7 nt to 20 nt or approximately 7 nt to 20 nt, 7 nt to 15 nt or approximately 7 nt to 15 nt, 8 nt to 40 nt or approximately 8 nt to 40 nt, 8 nt to 30 nt or approximately 8 nt to 30 nt, and 8 nt to 25 nt. The minimum sequence length of tracrRNA may be approximately 8nt to 25nt, 8nt to 20nt, 8nt to 15nt, 15nt to 100nt, 15nt to 80nt, 15nt to 50nt, 15nt to 40nt, 15nt to 30nt, or 15nt to 30nt. In some embodiments, the minimum sequence length of tracrRNA is approximately 9 nucleotides. In some embodiments, the minimum sequence length of tracrRNA is approximately 12 nucleotides. In some embodiments, the minimum tracrRNA sequence consists of a 23-48 nt tracrRNA as described in Jinek, M. et al. (2012). Science, 337(6096):816-821.

[0177] In some embodiments, the minimal tracrRNA sequence has at least 60% or at least about 60% identity with a reference minimal tracrRNA sequence (e.g., wild-type tracrRNA from S. pyogenes) in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides. For example, the minimal tracrRNA sequence has at least 65% or at least about 65% identity with a reference minimal tracrRNA sequence in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides, at least 70% or at least about 70% identity, at least 75% or at least about 75% identity, at least 80% or at least about 80% identity, at least 85% or at least about 85% identity, at least 90% or at least about 90% identity, at least 95% or at least about 95% identity, at least 98% or at least about 98% identity, at least 99% or at least about 99% identity, or at least 100% identity.

[0178] In some embodiments, the double strand formed by the CRISPR RNA minimal sequence and the tracrRNA minimal sequence has a double helix structure. In some embodiments, the double strand formed by the CRISPR RNA minimal sequence and the tracrRNA minimal sequence has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or at least more nucleotides, or at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, or at least more nucleotides. In some embodiments, the double strand formed by the CRISPR RNA minimal sequence and the tracrRNA minimal sequence has at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, or at most more nucleotides, or at most about 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, or at most more nucleotides.

[0179] In some embodiments, the double helix has mismatches (for example, the two strands in the double helix do not have 100% complementarity). In some embodiments, the double helix has at least one, at least two, at least three, at least four, or at least five mismatches, or at least about one, at least about two, at least about three, at least about four, or at least about five mismatches. In some embodiments, the double helix has at most one, at most two, at most three, at most four, or at most five, or at most about one, at most about two, at most about three, at most about four, or at most about five mismatches. In some embodiments, the double helix has two or fewer mismatches.

[0180] 3' tracrRNA sequence In some embodiments, the 3' tracrRNA sequence has a sequence that has at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 100% sequence identity with a reference tracrRNA sequence (e.g., tracrRNA from S. pyogenes).

[0181] In some embodiments, the length of the 3' tracrRNA sequence is between 6 nucleotides and 100 nucleotides or approximately 6 nucleotides and approximately 100 nucleotides. For example, the length of the 3' tracrRNA sequence is between 6 nucleotides (nt) and 50 nt or approximately 6 nt and 50 nt, 6 nt and 40 nt or approximately 6 nt and 40 nt, 6 nt and 30 nt or approximately 6 nt and 30 nt, 6 nt and 25 nt or approximately 6 nt and 25 nt, 6 nt and 20 nt or approximately 6 nt and 20 nt, 6 nt and 15 nt or approximately 6 nt and 15 nt, 8 nt and 40 nt or approximately 8 nt and 40 nt, 8 nt and 30 nt or approximately 8 nt and 30 nt, 8 nt and 25 nt Alternatively, the length may be approximately 8nt to 25nt, 8nt to 20nt, 8nt to 20nt, 8nt to 15nt, 15nt to 100nt, 15nt to 80nt, 15nt to 80nt, 15nt to 50nt, 15nt to 40nt, 15nt to 40nt, 15nt to 30nt, 15nt to 30nt, or 15nt to 25nt. In some embodiments, the length of the 3' tracrRNA sequence is approximately 14 nucleotides.

[0182] In some embodiments, the 3' tracrRNA sequence has at least 60% or at least about 60% identity with a reference 3' tracrRNA sequence (e.g., a wild-type 3' tracrRNA sequence from S. pyogenes) in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides. For example, a 3' tracrRNA sequence has at least 60% or at least about 60% identity with a reference 3' tracrRNA sequence (e.g., a wild-type 3' tracrRNA sequence from S. pyogenes) in a sequence consisting of at least 6, 7, or 8 consecutive nucleotides, at least 65% or at least about 65% identity, at least 70% or at least about 70% identity, at least 75% or at least about 75% identity, at least 80% or at least about 80% identity, at least 85% or at least about 85% identity, at least 90% or at least about 90% identity, at least 95% or at least about 95% identity, at least 98% or at least about 98% identity, at least 99% or at least about 99% identity, or at least 100% identity.

[0183] In some embodiments, the 3' tracrRNA sequence has two or more double-stranded regions (e.g., hairpins, hybridized regions). In some embodiments, the 3' tracrRNA sequence has two double-stranded regions.

[0184] In some embodiments, the 3' tracrRNA sequence has a stem-loop structure. In some embodiments, the stem-loop structure of the 3' tracrRNA has at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least fifteen, at least twenty, or at least more nucleotides. In some embodiments, the stem-loop structure of the 3' tracrRNA has at most one, at most two, at most three, at most four, at most five, at most six, at most seven, at most eight, at most nine, at most ten, or at most more nucleotides. In some embodiments, the stem-loop structure has a functional moiety. For example, the stem-loop structure may have an aptamer, a ribozyme, a hairpin that interacts with a protein, a CRISPR array, an intron, or an exon. In some embodiments, the stem-loop structure has at least one, at least two, at least three, at least four, at least five or more functional parts, or at least about one, at least about two, at least about three, at least about four, at least about five or more functional parts. In some embodiments, the stem-loop structure has at most one, at most two, at most three, at most four, at most five or more functional parts, or at most about one, at most about two, at most about three, at most about four, at most about five or more functional parts.

[0185] In some embodiments, the hairpin of the 3' tracrRNA sequence has a P domain. In some embodiments, the P domain on the hairpin has a double-stranded region.

[0186] tracrRNA elongation sequence In some embodiments, the tracrRNA elongation sequence is provided when the guide RNA is a single-molecule guide or a bi-molecule guide. In some embodiments, the length of the tracrRNA elongation sequence is 1 nucleotide to 400 nucleotides or approximately 1 nucleotide to approximately 400 nucleotides. In some embodiments, the length of the tracrRNA elongation sequence is greater than 1 nucleotide, greater than 5 nucleotides, greater than 10 nucleotides, greater than 15 nucleotides, greater than 20 nucleotides, greater than 25 nucleotides, greater than 30 nucleotides, greater than 35 nucleotides, greater than 40 nucleotides, greater than 45 nucleotides, greater than 50 nucleotides, greater than 60 nucleotides, greater than 70 nucleotides, greater than 80 nucleotides, greater than 90 nucleotides, greater than 100 nucleotides, greater than 120 nucleotides, greater than 140 nucleotides, greater than 160 nucleotides, greater than 180 nucleotides, greater than 200 nucleotides, greater than 220 nucleotides, greater than 240 nucleotides, greater than 260 nucleotides, greater than 280 nucleotides, greater than 300 nucleotides, greater than 320 nucleotides, greater than 340 nucleotides, greater than 360 nucleotides, greater than 380 nucleotides, or greater than 400 nucleotides. In some embodiments, the length of the tracrRNA elongation sequence is 20 nucleotides to 5000 nucleotides or more, or approximately 20 nucleotides to approximately 5000 nucleotides or more. In some embodiments, the length of the tracrRNA elongation sequence exceeds 1000 nucleotides.In some embodiments, the length of the tracrRNA elongation sequence is less than 1 nucleotide, less than 5 nucleotides, less than 10 nucleotides, less than 15 nucleotides, less than 20 nucleotides, less than 25 nucleotides, less than 30 nucleotides, less than 35 nucleotides, less than 40 nucleotides, less than 45 nucleotides, less than 50 nucleotides, less than 60 nucleotides, less than 70 nucleotides, less than 80 nucleotides, less than 90 nucleotides, less than 100 nucleotides, less than 120 nucleotides, less than 140 nucleotides, less than 160 nucleotides, less than 180 nucleotides, less than 200 nucleotides, less than 220 nucleotides, less than 240 nucleotides, less than 260 nucleotides, less than 280 nucleotides, less than 300 nucleotides, less than 320 nucleotides, less than 340 nucleotides, less than 360 nucleotides, less than 380 nucleotides, less than 400 nucleotides, or less than a certain number of sequences. In some embodiments, the length of the tracrRNA elongation sequence may be less than 1000 nucleotides. In some embodiments, the length of the tracrRNA elongation sequence is less than 10 nucleotides. In some embodiments, the length of the tracrRNA elongation sequence is 10 to 30 nucleotides. In some embodiments, the length of the tracrRNA elongation sequence is 30 to 70 nucleotides.

[0187] In some embodiments, the tracrRNA elongation sequence has a functional moiety (e.g., a stability control sequence, a ribozyme, or an endoribonuclease binding sequence). In some embodiments, the functional moiety is a transcription terminator segment (e.g., a transcription termination sequence). In some embodiments, the total length of the functional moiety is 10 nucleotides (nt) to 100 nucleotides or about 10 nt to about 100 nt, 10 nt to 20 nt or about 10 nt to about 20 nt, 20 nt to 30 nt or about 20 nt to about 30 nt, 30 nt to 40 nt or about 30 nt to about 40 nt, 40 nt to 50 nt or about 40 nt to about 50 nt, 50 nt to 60 nt or about 50 nt to about 60 nt, 60 nt to 70 nt or about 60 nt to about The functional portion is 70nt, 70nt-80nt or approximately 70nt-80nt, 80nt-90nt or approximately 80nt-90nt, 90nt-100nt or approximately 90nt-100nt, 15nt-80nt or approximately 15nt-80nt, 15nt-50nt or approximately 15nt-50nt, 15nt-40nt or approximately 15nt-40nt, 15nt-30nt or approximately 15nt-30nt, or 15nt-25nt or approximately 15nt-25nt. In some embodiments, the functional portion functions in eukaryotic cells. In some embodiments, the functional portion functions in prokaryotic cells. In some embodiments, the functional portion functions in both eukaryotic and prokaryotic cells.

[0188] Examples of suitable functional portions of a tracrRNA elongation sequence include, but are not limited to, a 3'-terminal polyadenylated tail; riboswitch sequences (e.g., those that stabilize under control and / or allow access by proteins or protein complexes under control); sequences that form a dsRNA double strand (e.g., a hairpin); sequences that target RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.); modifications or sequences that enable tracking (e.g., direct binding to fluorescent molecules, binding to regions that facilitate fluorescence detection, sequences that enable fluorescence detection, etc.); and / or modifications or sequences that provide a binding site for proteins (e.g., DNA-acting proteins such as transcription activators, transcription repressors, DNA methyltransferases, DNA methyl-degrading enzymes, histone acetyltransferases, histone deacetylases, etc.). In some embodiments, the tracrRNA elongation sequence has a primer binding site or molecular index (e.g., a barcode sequence). In some embodiments, the tracrRNA elongation sequence has one or more affinity tags.

[0189] Bulge In some embodiments, a "bulge" exists in the double helix formed by the CRISPR RNA minimal sequence and the tracrRNA minimal sequence. The bulge is a nucleotide unpaired region within the double helix. In some embodiments, the bulge contributes to the binding of the double helix to a site-specific polypeptide. The bulge has an unpaired sequence 5'-XXXY-3' on one strand of the double helix (where X is any purine base and Y is a nucleotide that can form a fluctuating base pair with a nucleotide on the opposite strand) and an unpaired nucleotide region on the other strand. The number of unpaired nucleotides on each strand of the double helix may vary.

[0190] In one example, the bulge has an unpaired purine base (e.g., adenine) on the CRISPR minimal repeat strand that forms the bulge. In some embodiments, the bulge has an unpaired sequence 5'-AAGY-3' on the tracrRNA minimal sequence strand that forms the bulge (where Y is a nucleotide that can form a fluctuating base pair with a nucleotide on the CRISPR minimal repeat strand).

[0191] In some embodiments, the bulge on the CRISPR minimal repeat chain side forming the double helix has at least one, at least two, at least three, at least four, at least five, or at least more unpaired nucleotides. The bulge on the CRISPR minimal repeat chain side forming the double helix has at most one, at most two, at most three, at most four, at most five, or at most more unpaired nucleotides. In some embodiments, the bulge on the CRISPR minimal repeat chain side forming the double helix has one unpaired nucleotide.

[0192] In some embodiments, the bulge on the tracrRNA minimal sequence side forming the double helix has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or at least more unpaired nucleotides. In some embodiments, the bulge on the tracrRNA minimal sequence side forming the double helix has at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, or at most 10, or at most more unpaired nucleotides. In some embodiments, the bulge on the second strand forming the double helix (for example, on the tracrRNA minimal sequence side forming the double helix) has 4 unpaired nucleotides.

[0193] In some embodiments, the bulge has at least one fluctuating base pair. In some embodiments, the bulge has one or fewer fluctuating base pairs. In some embodiments, the bulge has at least one purine nucleotide. In some embodiments, the bulge has at least three purine nucleotides. In some embodiments, the bulge sequence has at least five purine nucleotides. In some embodiments, the bulge sequence has at least one guanine nucleotide. In some embodiments, the bulge sequence has at least one adenine nucleotide.

[0194] hairpin In various embodiments, one or more hairpins are present at the 3' end of the minimal tracrRNA sequence within the 3' tracrRNA sequence.

[0195] In some embodiments, the hairpin begins with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least fifteen, at least twenty or more nucleotides from the 3' end of the last paired nucleotide of the double helix consisting of the CRISPR minimal repeat sequence and the tracrRNA minimal sequence, or at least about one, at least about two, at least about three, at least about four, at least about five, at least about six, at least about seven, at least about eight, at least about nine, at least about ten, at least about fifteen, at least about twenty or more nucleotides. In some embodiments, the hairpin may begin with at least one, at most two, at most three, at most four, at most five, at most six, at most seven, at most eight, at most nine, at most ten or more nucleotides from the 3' end of the double-stranded combo consisting of the CRISPR minimal repeat sequence and the tracrRNA minimal sequence, or at most about one, at most about two, at most about three, at most about four, at most about five, at most about six, at most about seven, at most about eight, at most about nine, at most about ten or more nucleotides.

[0196] In some embodiments, the hairpin has at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least fifteen, at least twenty or more consecutive nucleotides, or at least about one, at least about two, at least about three, at least about four, at least about five, at least about six, at least about seven, at least about eight, at least about nine, at least about ten, at least about fifteen, at least about twenty or more consecutive nucleotides. In some embodiments, the hairpin has at most 1, at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, at most 9, at most 10, or at most 15 or more consecutive nucleotides, or at most about 1, at most about 2, at most about 3, at most about 4, at most about 5, at most about 6, at most about 7, at most about 8, at most about 9, at most about 10, at most about 15 or more consecutive nucleotides.

[0197] In some embodiments, the hairpin has a CC dinucleotide (for example, two consecutive cytosine nucleotides).

[0198] In some embodiments, the hairpin has a double-stranded nucleotide (for example, a hairpin nucleotide in which nucleotides hybridize with each other). For example, the hairpin has a CC dinucleotide that hybridizes to a GG dinucleotide in the hairpin double-stranded 3' tracrRNA sequence.

[0199] One or more hairpins can interact with the guide RNA interaction region of a site-specific polypeptide.

[0200] In some embodiments, there are two or more hairpins, and in some embodiments, there are three or more hairpins.

[0201] Linker sequences of single-molecule guides In some embodiments, the linker sequence length of a single-molecule guide nucleic acid is 3 to 100 nucleotides or approximately 3 to 100 nucleotides. For example, Jinek, M. et al. (2012). Science, 337(6096):816-821 used a simple "tetraloop" (-GAAA-) consisting of four nucleotides. The length of the linker is, for example, 3 nucleotides (nt) to 90 nt or approximately 3 nt to 90 nt, 3 nt to 80 nt or approximately 3 nt to 80 nt, 3 nt to 70 nt or approximately 3 nt to 70 nt, 3 nt to 60 nt or approximately 3 nt to 60 nt, 3 nt to 50 nt or approximately 3 nt to 50 nt, 3 nt to 40 nt or approximately 3 nt to 40 nt, 3 nt to 30 nt or approximately 3 nt to 30 nt, 3 nt to 20 nt or approximately 3 nt to 20 nt, or 3 nt to 10 nt or approximately 3 nt to 10 nt. For example, the linker length is 3nt~5nt or approximately 3nt~5nt, 5nt~10nt or approximately 5nt~10nt, 10nt~15nt or approximately 10nt~15nt, 15nt~20nt or approximately 15nt~20nt, 20nt~25nt or approximately 20nt~25nt, 25nt~30nt or approximately 25nt~30nt, 30nt~35nt or approximately 30nt~35nt, 3 The linker length may be 5nt to 40nt or approximately 35nt to 40nt, 40nt to 50nt or approximately 40nt to 50nt, 50nt to 60nt or approximately 50nt to 60nt, 60nt to 70nt or approximately 60nt to 70nt, 70nt to 80nt or approximately 70nt to 80nt, 80nt to 90nt or approximately 80nt to 90nt, or 90nt to 100nt or approximately 90nt to 100nt. In some embodiments, the linker length of the single-molecule guide nucleic acid is 4 to 40 nucleotides long.In some embodiments, the length of the linker is at least 100 nucleotides or at least about 100 nucleotides, at least 500 nucleotides or at least about 500 nucleotides, at least 1000 nucleotides or at least about 1000 nucleotides, at least 1500 nucleotides or at least about 1500 nucleotides, at least 2000 nucleotides or at least about 2000 nucleotides, at least 2500 nucleotides or at least about 2500 nucleotides, at least 3000 nucleotides or at least about 3000 nucleotides, at least 3500 nucleotides or less The length is at least approximately 3500 nucleotides, at least 4000 nucleotides or at least approximately 4000 nucleotides, at least 4500 nucleotides or at least approximately 4500 nucleotides, at least 5000 nucleotides or at least approximately 5000 nucleotides, at least 5500 nucleotides or at least approximately 5500 nucleotides, at least 6000 nucleotides or at least approximately 6000 nucleotides, at least 6500 nucleotides or at least approximately 6500 nucleotides, or at least 7000 nucleotides or at least approximately 7000 nucleotides, or at least that much longer.In some embodiments, the length of the linker is at most 100 nucleotides or at most about 100 nucleotides, at most 500 nucleotides or at most about 500 nucleotides, at most 1000 nucleotides or at most about 1000 nucleotides, at most 1500 nucleotides or at most about 1500 nucleotides, at most 2000 nucleotides or at most about 2000 nucleotides, at most 2500 nucleotides or at most about 2500 nucleotides, at most 3000 nucleotides or at most about 3000 nucleotides, at most 3500 nucleotides or The maximum length is approximately 3500 nucleotides, 4000 nucleotides or approximately 4000 nucleotides, 4500 nucleotides or approximately 4500 nucleotides, 5000 nucleotides or approximately 5000 nucleotides, 5500 nucleotides or approximately 5500 nucleotides, 6000 nucleotides or approximately 6000 nucleotides, 6500 nucleotides or approximately 6500 nucleotides, 7000 nucleotides or approximately 7000 nucleotides, or longer.

[0202] Linkers can have a variety of sequences, but in some embodiments, the linker does not have a sequence with a broad region homologous to other parts of the guide RNA, because the presence of such a homologous broad region can lead to intramolecular binding that may interfere with other functional regions of the guide RNA. In Jinek, et al. (2012). Science, 337(6096):816-821, a simple sequence of four nucleotides—GAAA—was used, but various other sequences, such as longer sequences, can be used as well.

[0203] In some embodiments, the linker sequence has functional portions. For example, the linker sequence may have one or more features such as an aptamer, a ribozyme, a hairpin that interacts with a protein, a protein binding site, a CRISPR array, an intron, or an exon. In some embodiments, the linker sequence has at least one, at least two, at least three, at least four, at least five, or at least more functional portions, or at least about one, at least about two, at least about three, at least about four, at least about five, or at least more functional portions. In some embodiments, the linker sequence has at most one, at most two, at most three, at most four, at most five, or at most more functional portions, or at most about one, at most about two, at most about three, at most about four, at most about five, or at most more functional portions.

[0204] Donor DNA or donor template Site-directed polypeptides, such as DNA endonucleases, can introduce double-strand or single-strand breaks into nucleic acids (e.g., genomic DNA). Double-strand breaks can stimulate intrinsic DNA repair pathways in cells (e.g., homology-dependent repair (HDR), non-homologous end joining, alternative non-homologous end joining (A-NHEJ), or microhomology-mediated end joining (MMEJ)). NHEJ can repair the cleaved target nucleic acid without requiring a homologous template. This can result in small deletions or insertions (indels) in the target nucleic acid at the cleavage site, potentially leading to disruption or alteration of gene expression. Homologous-dependent repair (HDR), also known as homologous recombination (HR), can occur when a homologous repair template or donor is available.

[0205] Homologous donor templates have sequences homologous to the sequences adjacent to the cleavage site of the target nucleic acid. Generally, sister chromatids are used by cells as repair templates. On the other hand, repair templates for genome editing are often provided as exogenous nucleic acids such as plasmids, double-stranded oligonucleotides, single-stranded oligonucleotides, double-stranded oligonucleotides, or viral nucleic acids. In exogenous donor templates, an additional nucleic acid sequence (such as an introduced gene) or modification (such as a change or deletion of one or more bases) is generally introduced between adjacent homologous regions so that the additional nucleic acid sequence or modified nucleic acid sequence is incorporated into the target gene locus. MMEJ yields genetic results similar to NHEJ in that small deletions and insertions can occur at the cleavage site. MMEJ achieves the desired end-join DNA repair result by utilizing homologous sequences consisting of several base pairs adjacent to the cleavage site. In some cases, the expected repair result can be predicted by analyzing the short homologous sequences (microhomology) predicted in the target region of the nuclease.

[0206] Therefore, in some cases, homologous recombination is used to insert an exogenous polynucleotide sequence into the cleavage site of the target nucleic acid. Hereinafter, the exogenous polynucleotide sequence is referred to as a donor polynucleotide (or donor, donor sequence, or donor template). In some embodiments, a donor polynucleotide, a portion of a donor polynucleotide, a copy of a donor polynucleotide, or a portion of a copy of a donor polynucleotide is inserted into the cleavage site of the target nucleic acid. In some embodiments, the donor polynucleotide is an exogenous polynucleotide sequence, for example, a sequence that is not naturally present at the cleavage site of the target nucleic acid.

[0207] When a sufficient concentration of exogenous DNA molecules is supplied to the nucleus of a cell where a double-strand break occurs, the exogenous DNA can be inserted into the double-strand break site during the NHEJ repair process, allowing it to be stably maintained within the genome and potentially permanently added to it. Such exogenous DNA molecules are referred to as donor templates in some embodiments. If the donor template contains the coding sequence for the gene or genomic region of interest, along with relevant regulatory sequences such as promoters, enhancers, polyA sequences, and / or splice acceptor sequences as needed, the gene of interest can be expressed from the integrated copy in the genome and may be permanently expressed for the duration of the cell's life. Furthermore, the integrated copy from the donor DNA template can be transmitted to daughter cells when the cell divides.

[0208] If a donor DNA template containing adjacent DNA sequences homologous to the DNA sequences on both sides of a double-strand break site (called homologous arms) is present in sufficient concentration, this donor DNA template can be incorporated via the HDR pathway. The homologous arms act as substrates for homologous recombination between the donor template and the sequences on both sides of the double-strand break site. This allows for error-free insertion of a donor template in which the sequences on both sides of the double-strand break site have not been altered from the sequences of the unmodified genome.

[0209] Donors provided for HDR editing are highly diverse, but generally, they contain the intended editing sequence and short or long homologous arms on either side, enabling annealing to genomic DNA. The homologous region adjacent to the introduced gene alteration site may be less than 30 bp, or it may be as large as a few kilobase cassette, which may also contain promoters or cDNA. Both single-stranded and double-stranded oligonucleotide donors can be used. The size of these oligonucleotides ranges from less than 100 nt to several kilobases or more, but longer ssDNA can also be generated and used. Double-stranded donors such as PCR amplicons, plasmids, and minicircles are commonly used. Generally, AAV vectors (however, the gene editing methods disclosed herein are not limited to AAV vectors) have been shown to be a very effective means of delivering donor templates, but the limit for packaging to individual donors is less than 5 kb. Active transcription of the donor has been shown to triple the HDR rate, indicating that conversion can be increased by including a promoter. Conversely, donor CpG methylation may reduce gene expression and HDR rates.

[0210] In some embodiments, donor DNA can be introduced alone or together with a nuclease by various methods, such as transfection, nanoparticles, microinjection, or viral transduction. In some embodiments, the availability of donor in HDR can be enhanced by using various methods for linking donor DNA and nuclease. Examples of such methods include linking the donor to a nuclease, linking it to a DNA-binding protein that binds to the vicinity of the donor and nuclease, or linking it to a protein involved in DNA end joining or DNA repair.

[0211] In addition to genome editing by NHEJ or HDR, site-directed gene insertion using both the NHEJ pathway and HR can be performed. Such a combined approach is applicable in certain situations that may involve intron / exon boundaries. NHEJ can be demonstrated to be effective for ligation in introns, while error-free HDR may be more suitable for coding regions.

[0212] Nucleic acids encoding site-directed polypeptides or DNA endonucleases In some embodiments, based on the above, in the genome editing method and the composition, a nucleic acid sequence (or oligonucleotide) encoding a site-directed polypeptide or a DNA endonuclease can be used. The nucleic acid sequence encoding the site-directed polypeptide can be DNA or RNA. When the nucleic acid sequence encoding the site-directed polypeptide is RNA, this RNA can be covalently bound to the gRNA sequence or can exist as a separate sequence. In some embodiments, the peptide sequence of the site-directed polypeptide or DNA endonuclease can be used instead of these nucleic acid sequences.

[0213] Gene editing vectors In another aspect, the present disclosure provides a nucleic acid having a nucleotide sequence encoding the genome-targeting nucleic acid of the present disclosure, the site-directed polypeptide of the present disclosure, and / or any nucleic acid or protein molecule necessary to implement an embodiment of the method of the present disclosure. In some embodiments, such a nucleic acid is a vector (e.g., a recombinant expression vector).

[0214] As expression vectors contemplated in the present invention, viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, simian virus 40 (SV40), herpes simplex virus, human immunodeficiency virus, retroviruses (e.g., murine leukemia virus; spleen necrosis virus; and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, mammary tumor virus), and other recombinant vectors are mentioned, but not limited thereto. Other vectors contemplated for use in eukaryotic target cells include, but are not limited to, the pXT1 vector, pSG5 vector, pSVK3 vector, pBPV vector, pMSG vector, and pSVLSV40 vector (Pharmacia). Further vectors contemplated for use in eukaryotic target cells include, but are not limited to, the pCTx-1 vector, pCTx-2 vector, and pCTx-3 vector. Other vectors can also be used as long as they are compatible with the host cell.

[0215] In some embodiments, the vector has one or more transcriptional control elements and / or translational control elements. Depending on the host / vector system utilized, various suitable transcriptional control elements and translational control elements such as constitutive promoters, inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc. can be used in the expression vector. In some embodiments, the vector is a self-inactivating vector that inactivates viral sequences or components of the CRISPR mechanism or other elements.

[0216] Examples of suitable eukaryotic promoters (e.g., promoters that function in eukaryotic cells) include, but are not limited to, the cytomegalovirus (CMV) promoter, the earliest thymidine kinase promoter of herpes simplex virus (HSV), early SV40, late SV40, retroviral long terminal repeats (LTRs), human elongation factor 1 promoter (EF1), a hybrid construct fused with a cytomegalovirus (CMV) enhancer to the chicken β-actin promoter (CAG), mouse stem cell virus promoter (MSCV), phosphoglycerate kinase 1 locus promoter (PGK), and mouse metallothionein I. In further embodiments, the promoter is an MND promoter (e.g., an MND promoter containing the nucleic acid sequence of SEQ ID NO: 33).

[0217] Various promoters, such as RNA polymerase III promoters like U6 and H1, can be useful for expressing small RNAs, including guide RNAs used with Cas endonucleases. Descriptions and parameters that facilitate the use of such promoters are well-known in the art, and new information and approaches are constantly being reported. See, for example, Ma, H. et al. (2014). Mol. Ther. - Nucleic Acids 3, e161, doi:10.1038 / mtna.2014.12.

[0218] The expression vector may further contain ribosome binding sites for translation initiation and transcription termination. It may also contain appropriate sequences for amplification of expression. Furthermore, because the expression vector is fused to a site-directed polypeptide, it may contain nucleotide sequences encoding non-natural tags (e.g., histidine tags, hemagglutinin tags, green fluorescent protein, etc.) that are expressed as part of the fusion protein.

[0219] In some embodiments, the promoter is an inductive promoter (e.g., a heat shock promoter, a tetracycline-regulating promoter, a steroid-regulating promoter, a metal-regulating promoter, an estrogen receptor-regulating promoter, etc.). In some embodiments, the promoter is a constitutive promoter (e.g., a CMV promoter, or a UBC promoter). In some embodiments, the promoter is a spatially restricted promoter and / or a temporally restricted promoter (e.g., a tissue-specific promoter, a cell-type-specific promoter, etc.). In some embodiments, if at least one gene expressed in a host cell is inserted into the genome and then expressed under the control of an endogenous promoter present in that genome, the vector does not contain a promoter for this gene.

[0220] In some embodiments, the vector comprises one or more nucleic acid sequences encoding one or more components of a CISC (for example, any of the CISC polypeptides disclosed herein, such as a CISC polypeptide containing a cleaved ILR2β intracellular signaling domain). The gene editing vectors disclosed herein may also comprise nucleic acid sequences encoding any of the naked FRB domain polypeptides disclosed herein. In some embodiments, the vector comprises the nucleic acid sequence of SEQ ID NO: 3. In further embodiments, the vector comprises the nucleic acid sequence of SEQ ID NO: 8.

[0221] A complex of genome-targeted nucleic acids and site-directed polypeptides Genome-targeted nucleic acids form complexes by interacting with site-directed polypeptides (e.g., nucleic acid-inducible nucleases such as Cas9). The genome-targeted nucleic acid (e.g., gRNA) leads the site-directed polypeptide to the target nucleic acid.

[0222] As described above, in some embodiments, site-directed polypeptides and genome-targeted nucleic acids can be administered to cells or subjects separately. On the other hand, in some other embodiments, site-directed polypeptides can be pre-complexed with one or more guide RNAs, or site-directed polypeptides can be pre-complexed with tracrRNA and one or more crRNAs. The pre-complexed material can be administered to cells or subjects. Such pre-complexed material is known as ribonucleoprotein particles (RNPs).

[0223] Method for generating cells expressing naked FRB domain polypeptides and / or CISC components. In some embodiments described herein, it may be desirable to induce rapamycin-regulated cytokine signaling and / or selectively expand cells expressing dimerized CISC components and intracellularly expressed naked FRB domain polypeptides by introducing protein sequences or expression vectors into host cells (e.g., mammalian cells (e.g., lymphocytes)). For example, dimerized CISCs, upon contact with a ligand, can enable cytokine signaling in cells into which the CISC components have been introduced, thereby transmitting signals within cells (such as mammalian cells), while intracellularly expressed naked FRB domain polypeptides can confer resistance to the adverse effects of rapamycin to cells. Furthermore, the selective expansion and proliferation of cells (such as mammalian cells) can be controlled so that only cells undergoing two specific recombination events described herein (e.g., recombination events via the CRISPR / Cas system) are selected. The preparation of such cells can be carried out according to known techniques readily understandable to those skilled in the art in light of this disclosure.

[0224] In some embodiments, a method is provided for producing cells (such as mammalian cells) having CISCs, wherein these cells express dimerized CISCs and intracellularly expressed naked FRB domain polypeptides. The method may include the step of delivering a protein sequence described in any one of the embodiments described herein or an expression vector described in the embodiments described herein to cells (such as mammalian cells). In some embodiments, the protein sequence comprises a first sequence and a second sequence. In another embodiment, the protein sequence comprises a first sequence, a second sequence, a third sequence and a fourth sequence. In some embodiments, the first sequence encodes a first CISC component comprising a first extracellular binding domain, a hinge domain, a linker of a predetermined length optimized for a specific length, a transmembrane domain and a signaling domain. In some embodiments, the second sequence encodes a second CISC component comprising a second extracellular binding domain, a hinge domain, a linker of a predetermined length optimized for a specific length, a transmembrane domain and a signaling domain. In some embodiments, the spacer is a length within the range defined by 1 amino acid length, 2 amino acid length, 3 amino acid length, 4 amino acid length, 5 amino acid length, 6 amino acid length, 7 amino acid length, 8 amino acid length, 9 amino acid length, 10 amino acid length, 11 amino acid length, 12 amino acid length, 13 amino acid length, 14 amino acid length, or 15 amino acid length, or any two of these lengths. In some embodiments, the signaling domain includes an interleukin-2 signaling domain, such as IL2Rβ (e.g., cleaved IL2Rβ) or IL2Rγ domain. In some embodiments, the extracellular binding domain is a binding domain that binds to rapamycin or rapalog, such as FKBP or FRB or functional derivatives thereof. In some embodiments, the cell is a CD8+ cell or a CD4+ cell. In some embodiments, the cell is a CD8+ cytotoxic T lymphocyte selected from the group consisting of naive CD8+ T cells, central memory CD8+ T cells, effector memory CD8+ T cells, and bulk CD8+ T cells.In some embodiments, the cells are CD4+ helper T lymphocytes selected from the group consisting of naive CD4+ T cells, central memory CD4+ T cells, effector memory CD4+ T cells, and bulk CD4+ T cells. In some embodiments, the cells are progenitor T cells. In some embodiments, the cells are stem cells. In some embodiments, the cells are hematopoietic stem cells. In some embodiments, the cells are B cells. In some embodiments, the cells are neural stem cells. In some embodiments, the cells are NK cells.

[0225] Genetically modified cells and genetically modified cell populations In one embodiment, this disclosure provides a method for producing genetically modified cells by editing the genome of a cell. In several embodiments, a population of genetically modified cells is provided. Thus, the genetically modified cells include cells having at least one genetic modification introduced by genome editing (e.g., genome editing using the CRISPR / Cas system). In some embodiments, the genetically modified cells are genetically modified lymphocytes, such as T cells, including human CD4+ T cells. Genetically modified cells incorporating nucleic acids encoding a naked FRB domain polypeptide and optionally encoding CISC are envisioned herein.

[0226] The compositions described herein provide genetically modified host cells (e.g., mammalian cells) comprising the protein sequence or expression vector described herein. Thus, cells (such as mammalian cells) for expressing dimerized CISCs and naked FRB domains expressed intracellularly are provided, the cells comprising the protein sequence according to any one of the embodiments described herein or the expression vector according to any one of the embodiments described herein. In some embodiments, the cells are bacterial cells or mammalian cells, such as lymphocytes. In some embodiments, the cells are Escherichia coli. In some embodiments, the cells are insect cells capable of expressing the protein. In some embodiments, the cells are lymphocytes.

[0227] In some embodiments, the host cells are progenitor T cells or regulatory T cells. In some embodiments, the cells are stem cells, such as hematopoietic stem cells. In some embodiments, the cells are NK cells. In some embodiments, the cells are CD34+ T lymphocytes, CD8+ T lymphocytes, and / or CD4+ T lymphocytes. In some embodiments, the cells are B cells. In some embodiments, the cells are neural stem cells.

[0228] In some embodiments, the host cells are CD8+ T-cytotoxic lymphocytes and may include naive CD8+ T cells, central memory CD8+ T cells, effector memory CD8+ T cells, or bulk CD8+ T cells. In some embodiments, the cells are CD4+ helper T lymphocytes and may include naive CD4+ T cells, central memory CD4+ T cells, effector memory CD4+ T cells, or bulk CD4+ T cells.

[0229] The lymphocytes (T lymphocytes) can be recovered by known techniques and enriched or removed by known techniques such as flow cytometry and / or immunomagnetic selection, or affinity binding with antibodies. After the enrichment and / or removal steps, the desired T lymphocytes can be expanded and proliferated in vitro by known techniques or variations thereof, which are readily understood by those skilled in the art. In some embodiments, the T cells are autologous T cells obtained from a subject.

[0230] For example, a desired T cell population or subpopulation can be expanded by adding a T lymphocyte population before proliferation to an in vitro culture medium, then adding feeder cells such as non-dividing peripheral blood mononuclear cells (PBMCs) to the medium (for example, adding feeder cells in such a ratio that the post-addition cell population contains at least 5, 10, 20, or 40 or more PBMC feeder cells per T lymphocyte in the initial population at the start of expansion culture), and incubating the medium (for example, for a time sufficient to sufficiently expand the T cell number). The non-dividing feeder cells may include PBMC feeder cells irradiated with gamma rays. In some embodiments, the PBMCs are irradiated with gamma rays at 3000-3600 rads to prevent cell division. In some embodiments, to prevent cell division of the PBMCs, the PBMCs are irradiated with gamma rays at 3000 rad, 3100 rad, 3200 rad, 3300 rad, 3400 rad, 3500 rad, or 3600 rad, or any other radiation value between two endpoints of these values. The order in which T cells or feeder cells are added to the culture medium may be changed as needed. Typically, the culture can be incubated under conditions suitable for T lymphocyte proliferation, such as a temperature. Temperatures for human T lymphocyte proliferation are typically, for example, at least 25°C or at least about 25°C, at least 30°C or at least about 30°C, or at least 37°C or at least about 37°C. In some embodiments, the temperature for human T lymphocyte proliferation is about 22°C, about 24°C, about 26°C, about 28°C, about 30°C, about 32°C, about 34°C, about 36°C, or about 37°C, or any other temperature between two endpoints of these values.

[0231] T lymphocytes can be isolated and sorted into naive T cell subpopulations, memory T cell subpopulations, and effector T cell subpopulations, respectively, before or after expansion and proliferation.

[0232] CD8+ cells can be obtained using standard methods. In some embodiments, CD8+ cells are further sorted into naive CD8+ cells, central memory CD8+ cells, and effector memory CD8+ cells by identifying the cell surface antigens associated with each of these cells. In some embodiments, memory T cells are present in both the CD62L+ subset and the CD62L- subset derived from CD8+ peripheral blood lymphocytes. PBMCs are sorted into CD62L-CD8+ fractions and CD62L+CD8+ fractions after staining with anti-CD8 and anti-CD62L antibodies. In some embodiments, central memory T CM The expression of phenotypic markers includes CD45RO, CD62L, CCR7, CD28, CD3, and / or CD127, and granzyme B is negative or shows low expression. In some embodiments, central memory T cells are CD45RO+, CD62L+, and / or CD8+ T cells. In some embodiments, effector T cells E These cells are negative for CD62L, CCR7, CD28, and / or CD127, and positive for granzyme B and / or perforin. In some embodiments, naive CD8+ T lymphocytes are characterized by the expression of naive T cell phenotypic markers, such as CD62L, CCR7, CD28, CD3, CD127, and / or CD45RA.

[0233] CD4+ helper T cells are sorted into naive cells, central memory cells, and effector cells by identifying cell populations having cell surface antigens. CD4+ lymphocytes can be obtained by standard methods. In some embodiments, naive CD4+ T lymphocytes are CD45RO−, CD45RA+, CD62L+ and / or CD4+ T cells. In some embodiments, central memory CD4+ cells are CD62L+ and / or CD45RO+. In some embodiments, effector CD4+ cells are CD62L− and / or CD45RO−.

[0234] Whether it be mammalian cells or a population of mammalian cells, these cells or populations are selected for expansion and proliferation based on whether or not they have undergone two different genetic recombination events. In some embodiments, the genetic recombination events occur via the CRISPR / Cas system. In some embodiments, the genetic recombination events occur via the CRISPR / Cas9 system. If a cell or a population of mammalian cells undergoes one or fewer genetic recombination events, dimerization does not occur upon addition of a ligand. However, if a cell or a population of mammalian cells undergoes two genetic recombination events, the addition of a ligand causes dimerization of the CISC components, followed by the generation of a signaling cascade. Therefore, cells or a population of mammalian cells may be selected based on their responsiveness to contact with a ligand. In some embodiments, the amount of ligand added is 0.01nM, 0.02nM, 0.03nM, 0.04nM, 0.05nM, 0.06nM, 0.07nM, 0.08nM, 0.09nM, 0.1nM, 0.2nM, 0.3nM, 0.4nM, 0.5nM, 0.6nM, 0.7nM, 0.8nM, 0.9nM, 1.0nM, 1.5nM, 2.0nM, 2.5nM, 3.0nM, 3.5nM, 4.0nM, 4.5nM, 5.0nM, 5.5nM, 6.0 The concentration may be within the range defined by nM, 6.5nM, 7.0nM, 7.5nM, 8.0nM, 8.5nM, 9.0nM, 9.5nM, 10nM, 11nM, 12nM, 13nM, 14nM, 15nM, 20nM, 25nM, 30nM, 35nM, 40nM, 45nM, 50nM, 55nM, 60nM, 65nM, 70nM, 75nM, 80nM, 85nM, 90nM, 95nM, or 100nM, or any two of these values.

[0235] In some embodiments, cells such as mammalian cells or cell populations such as mammalian cell populations may be determined to be positive for dimerized CISCs and positive for naked FRB domain polypeptides expressed intracellularly, based on markers expressed as a result of signaling pathways. Thus, whether a cell population is positive for dimerized CISCs may be determined by flow cytometry using staining with surface marker-specific antibodies and isotype-matched control antibodies. In some embodiments, the marker is a fluorescent or luminescent protein, such as GFP or mCherry. In some embodiments, the marker is low-affinity nerve growth factor receptor (LNGFR).

[0236] In some embodiments, the cells are not germ cells.

[0237] Methods for activating signals inside cells In some embodiments, a method is provided for activating a signal inside a cell, such as a mammalian cell. This method may include the step of providing a cell (such as a mammalian cell) as described herein, which contains the protein sequence or expression vector described herein. In some embodiments, the method further includes the step of expressing the protein sequence encoding the dimerized CISC described herein or the expression vector described herein. In some embodiments, the method includes the step of contacting the cell (such as a mammalian cell) with a ligand, thereby inducing dimerization of a first CISC component and a second CISC component, resulting in the transmission of a signal inside the cell. In some embodiments, the ligand is rapamycin or rapalog. In some embodiments, effective amounts of ligand to induce dimerization include 0.01nM, 0.02nM, 0.03nM, 0.04nM, 0.05nM, 0.06nM, 0.07nM, 0.08nM, 0.09nM, 0.1nM, 0.2nM, 0.3nM, 0.4nM, 0.5nM, 0.6nM, 0.7nM, 0.8nM, 0.9nM, 1.0nM, 1.5nM, 2.0nM, 2.5nM, 3.0nM, 3.5nM, 4.0nM, 4.5nM, 5.0nM, and 5.5nM. Ligands are provided in concentrations within the range defined by M, 6.0nM, 6.5nM, 7.0nM, 7.5nM, 8.0nM, 8.5nM, 9.0nM, 9.5nM, 10nM, 11nM, 12nM, 13nM, 14nM, 15nM, 20nM, 25nM, 30nM, 35nM, 40nM, 45nM, 50nM, 55nM, 60nM, 65nM, 70nM, 75nM, 80nM, 85nM, 90nM, 95nM, or 100nM, or any two of these values.

[0238] In some embodiments, rapamycin (including its analogs, derivatives, and pharmaceutically acceptable salts) may be used as a ligand or agent in the methods described herein for the chemical induction of signal transduction complexes. Rapamycin includes sirolimus (Rapamune®), (3S,6R,7E,9R,10R,12R,14S,15E,17E,19E,21S,23S,26R,27R,34aS)-9,10,12,13,14,21,22,23,24,25,26,27,32,33,34,34a-hexadehydro-9,27-dihydroxy-3-[(1R)-2-[(1S,3R,4R)-4-hydroxy-3-methoxycyclohexyl]-1-methylethyl]-10,21-dimethoxy-6,8,12,14,20,26-hexamethyl-23,27-epoxy-3H-pyrido[2,1-c][1,4]oxaazacyclotriacontine-1,5,11,28,29(4H,6H,31H)-pentone). Also included is everolimus, which includes its analogs, derivatives, and pharmaceutically acceptable salts. Everolimus includes RAD001, Zortress, Certican, Afinitor, Votubia, 42-O-(2-hydroxyethyl)rapamycin, (1R,9S,12S,15R,16E,18R,19R,21R,23S,24E,26E,28E,30S,32S,35R)-1,18-dihydroxy-12-[(2R)-1-[(1S,3R,4R)-4-(2-hydroxyethoxy)-3-methoxycyclohexyl]propan-2-yl]-19,30-dimethoxy-15,17,21,23,29,35-hexamethyl-11,36-dioxa-4-azatricyclo[30.3.1.0 4 , 9Examples include hexatriaconta-16,24,26,28-tetraene-2,3,10,14,20-pentone. Also, merilimus is mentioned, which includes its analogs, derivatives, and pharmaceutically acceptable salts. Examples of Merilimus include SAR943, 42-O-(tetrahydrofuran-3-yl)rapamycin (Merilimus-1); 42-O-(oxetan-3-yl)rapamycin (Merilimus-2); 42-O-(tetrahydropyran-3-yl)rapamycin (Merilimus-3); 42-O-(4-methyl,tetrahydrofuran-3-yl)rapamycin; 42-O-(2,5,5-trimethyl,tetrahydrofuran-3-yl)rapamycin; 42-O-(2,5-diethyl-2-methyl,tetrahydrofuran-3-yl)rapamycin; 42-O-(2H-pyran-3-yl,tetrahydro-6-methoxy-2-methyl)rapamycin; or 42-O-(2H-pyran-3-yl,tetrahydro-2,2-dimethyl-6-phenyl)rapamycin). Furthermore, nobolimus is mentioned, including its analogs, derivatives, and pharmaceutically acceptable salts. An example of nobolimus is 16-O-demethylrapamycin. Also mentioned is pimecrolimus, including its analogs, derivatives, and pharmaceutically acceptable salts. Examples of pimecrolimus include Elidel (registered trademark), (3S,4R,5S,8R,9E,12S,14S,15R,16S,18R,19R,26aS)-3-((E)-2-((1R,3R,4S)-4-chloro-3-methoxycyclohexyl)-1-methylvinyl)-8-ethyl-5,6,8,11,12,13,14,15,16,17,18,19,24,26,26a-hexadecahydro-5,19-epoxy-3H-pyrido(2,1-c)(1,4)oxazacyclotricosin-1,17,20,21(4H,23H)-tetron, and 33-epi-chloro-33-desoxiascomycin. Also mentioned is lidahororimus, which includes its analogs, derivatives, and pharmaceutically acceptable salts.As for ridaforolimus, AP23573, MK-8669, dehorolimus, (1R,9S,12S,15R,16E,18R,19R,21R,23S,24E,26E,28E,30S,32S,35R)-12-((1R)-2-((1S,3R,4R)-4-((dimethylphosphinoyl)oxy)-3-methoxycyclohexyl)-1-methylethyl)-1,18-dihydroxy-19,30-dimethoxy-15,17,21,23,29,35-hexamethyl-11,36-dioxa-4-azatricyclo(30.3.1.0. 4 , 9Examples include hexatriaconta-16,24,26,28-tetraen-2,3,10,14,20-pentone. Also, tacrolimus is mentioned, and tacrolimus includes its analogs, derivatives and pharmaceutically acceptable salts. Examples of tacrolimus include FK-506, Fujimycin, Prograf®, Advagraf®, Protopic, 3S-[3R*[E(1S*,3S*,4S*)],4S*,5R*,8S*,9E,12R*,14R*,15S*,16R*,18S*,19S*,26aR*5,6,8,11,12,13,14,15,16,17,18,19,24,25,26,26a-hex Examples include sadecahydro-5,19-dihydroxy-3-[2-(4-hydroxy-3-methoxycyclohexyl)-1-methylethenyl]-14,16-dimethoxy-4,10,12,18-tetramethyl-8-(2-propenyl)-15,19-epoxy-3H-pyrido[2,1-c][1,4]oxazazacyclotricosin-1,7,20,21(4H,23H)-tetron monohydrate. Also, temsirolimus is an example, including its analogs, derivatives, and pharmaceutically acceptable salts. Temsirolimus includes CCI-779, CCL-779, Torisel (registered trademark), and (1R,2R,4S)-4-{(2R)-2-[(3S,6R,7E,9R,10R,12R,14S,15E,17E,19E,21S,23S,26R,27R,34aS)-9,27-dihydroxy-10,21-dimethoxy-6,8,12,14,20,26-hexamethyl -1,5,11,28,29-pentaoxo-1,4,5,6,9,10,11,12,13,14,21,22,23,24,25,26,27,28,29,31,32,33,34,34a-tetracosahydro-3H-23,27-epoxypyrido[2,1-c][1,4]oxazacyclohentricontin-3-yl]propyl}-2-methoxycyclohexyl 3-hydroxy-2-(hydroxymethyl)-2-methylpropanoate is an example. Also, umilolimus is an example, including its analogs, derivatives and pharmaceutically acceptable salts.Examples of umilolimus include Biolimus, Biolimus A9, BA9, TRM-986, and 42-O-(2-ethoxyethyl)rapamycin. Zotarolimus is also included, encompassing its analogues, derivatives, and pharmaceutically acceptable salts. Examples of zotarolimus include ABT-578 and (42S)-42-deoxy-42-(1H-tetrazole-1-yl)-rapamycin. C20-metharylrapamycin is also included, encompassing its analogues, derivatives, and pharmaceutically acceptable salts. An example of C20-metharylrapamycin is C20-Marap. C16-(S)-3-methylindolerapamycin is also included, encompassing its analogues, derivatives, and pharmaceutically acceptable salts. Examples of C16-(S)-3-methylindolerapamycin include C16-iRap. Also, AP21967 is an example, and AP21967 includes its analogues, derivatives, and pharmaceutically acceptable salts. Examples of AP21967 include C-16-(S)-7-methylindolerapamycin. Also, sodium mycophenolate is an example, and sodium mycophenolate includes its analogues, derivatives, and pharmaceutically acceptable salts. Examples of sodium mycophenolate include CellCept®, Myfortic, and (4E)-6-(4-hydroxy-6-methoxy-7-methyl-3-oxo-1,3-dihydro-2-benzofuran-5-yl)-4-methylhexa-4-enoic acid. Also, benidipine hydrochloride is an example, and benidipine hydrochloride includes its analogues, derivatives, and pharmaceutically acceptable salts. An example of benidipine hydrochloride is Coniel. Also mentioned is AP1903, which includes its analogs, derivatives, and pharmaceutically acceptable salts.Examples of AP1903 include rimiducid, [(1R)-3-(3,4-dimethoxyphenyl)-1-[3-[2-[2-[[2-[3-[(1R)-3-(3,4-dimethoxyphenyl)-1-[(2S)-1-[(2S)-2-(3,4,5-trimethoxyphenyl)butanoyl]piperidine-2-carbonyl]oxypropyl]phenoxy]acetyl]amino]ethylamino]-2-oxoethoxy]phenyl]propyl] (2S)-1-[(2S)-2-(3,4,5-trimethoxyphenyl)butanoyl]piperidine-2-carboxylate. Furthermore, any combination of these can be listed.

[0239] In some embodiments, the ligand used in the method is rapamycin or rapalog, such as everolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP23573 or AP1903, or their metabolites, derivatives and / or any combination thereof. Further useful rapamycin variants include, for example, rapamycin variants modified by one or more of the following modifications to rapamycin: demethylation, removal, or substitution of methoxy at positions C7, C42, and / or C29; removal, derivatization, or substitution of hydroxyl at positions C13, C43, and / or C28; reduction, removal, or derivatization of ketone at positions C14, C24, and / or C30; substitution of a 6-membered pipecolate ring with a 5-membered prolyl ring; and / or another substitution on the cyclohexyl ring or substitution of the cyclohexyl ring with a substituted cyclopentyl ring. Further useful rapamycin variants include novolimus, pimecrolimus, ridaflorimus, tacrolimus, temsirolimus, umilolimus, or zotarolimus, or their derivatives, metabolites, and / or any combination thereof. In some embodiments, the ligand is an IMID drug (e.g., thalidomide, pomalidomide, lenalidomide, or related analogues).

[0240] In some embodiments, the detection of signals within cells (such as mammalian cells) can be achieved by methods that detect markers resulting from signaling pathways. Therefore, signals may be detected, for example, by measuring the amount of Akt or other signaling markers in cells (such as mammalian cells) using Western blotting, flow cytometry, or other protein detection or quantification methods. Examples of markers for detection include JAK, Akt, STAT, NF-κ, MAPK, PI3K, JNK, ERK, or Ras, or other cellular signaling markers that indicate cellular signaling events.

[0241] In some embodiments, signaling affects cytokine signaling. In some embodiments, signaling affects IL2R signaling. In some embodiments, signaling affects phosphorylation of downstream targets of cytokine receptors. In some embodiments, methods of activating the signaling induce proliferation of CISC-expressing cells (such as mammalian cells) and, consequently, suppress the proliferation of non-CISC-expressing cells.

[0242] For signal transduction to occur in cells, cytokine receptors must not only dimerize or heterodimerize, but also be in the correct configuration to undergo structural changes (Kim, MJet al. (2007). J. Biol. Chem., 282(19):14253-14261). Therefore, for proper signal transduction to occur, it is desirable for the signal transduction domain to dimerize in the correct three-dimensional configuration, because simply dimerizing or heterodimerizing the receptor is not enough to activate it. The chemically induced signal transduction complexes described herein are typically oriented in the correct direction for downstream signal transduction events to occur.

[0243] Selective expansion and proliferation method for cell populations In some embodiments, a method is provided for selectively expanding and proliferating a population of cells, such as mammalian cells. In some embodiments, the method may include a step of providing cells (such as mammalian cells) as described herein, the cells comprising a protein sequence or expression vector as described herein. In some embodiments, the method further includes a step of expressing a protein sequence encoding a naked FRB domain polypeptide and / or a dimerized CISC as described herein, or an expression vector as described herein. In some embodiments, the method includes a step of contacting cells (such as mammalian cells) with a ligand, thereby inducing dimerization of a first CISC component and a second CISC component, resulting in the transmission of a signal into the cell interior. In some embodiments, the ligand is rapamycin or rapalog (for example, any of the rapamycin or rapalog compounds disclosed herein). In some embodiments, the effective amount of ligand provided to induce dimerization is 0.01nM, 0.02nM, 0.03nM, 0.04nM, 0.05nM, 0.06nM, 0.07nM, 0.08nM, 0.09nM, 0.1nM, 0.2nM, 0.3nM, 0.4nM, 0.5nM, 0.6nM, 0.7nM, 0.8nM, 0.9nM, 1.0nM, 1.5nM, 2.0nM, 2.5nM, 3.0nM, 3.5nM, 4.0nM, 4.5nM, 5.0n The concentration is defined as M, 5.5nM, 6.0nM, 6.5nM, 7.0nM, 7.5nM, 8.0nM, 8.5nM, 9.0nM, 9.5nM, 10nM, 11nM, 12nM, 13nM, 14nM, 15nM, 20nM, 25nM, 30nM, 35nM, 40nM, 45nM, 50nM, 55nM, 60nM, 65nM, 70nM, 75nM, 80nM, 85nM, 90nM, 95nM, or 100nM, or any two of these values.

[0244] In some embodiments, the ligand used is rapamycin or rapalog, such as everolimus, CCI-779, C20-methallylrapamycin, C16-(S)-3-methylindolerapamycin, C16-iRap, AP21967, sodium mycophenolate, benidipine hydrochloride, AP23573 or AP1903, or their metabolites, derivatives and / or any combination thereof. Further useful rapamycin variants include, for example, rapamycin variants modified by one or more of the following modifications to rapamycin: demethylation, removal, or substitution of methoxy at positions C7, C42, and / or C29; removal, derivatization, or substitution of hydroxyl at positions C13, C43, and / or C28; reduction, removal, or derivatization of ketone at positions C14, C24, and / or C30; substitution of a 6-membered pipecolate ring with a 5-membered prolyl ring; and / or another substitution on the cyclohexyl ring or substitution of the cyclohexyl ring with a substituted cyclopentyl ring. Further useful rapamycin variants include novolimus, pimecrolimus, ridaflorimus, tacrolimus, temsirolimus, umilolimus, or zotarolimus, or their derivatives, metabolites, and / or any combination thereof.

[0245] In some embodiments, selective expansion and proliferation of a population of cells, such as mammalian cells, occurs only when two different recombination events (e.g., recombination via the CRISPR / Cas9 system) occur. One of these recombination events is a component of one dimeric chemo-inducible signaling complex, and the other is a component of the other dimeric chemo-inducible signaling complex. When both of these events occur in a cell population (such as a mammalian cell population), the components of the chemo-inducible signaling complex dimerize in the presence of ligands, resulting in an active chemo-inducible signaling complex that generates a signal within the cell.

[0246] Genome editing methods In some embodiments, methods for editing the genome of a cell are provided, more specifically, methods for editing the genome of a cell to i) express a naked FRB domain polypeptide in the cytoplasm of a cell, and optionally ii) express one or more polypeptide elements of a chemically induced signaling complex (CISC) that is activated by dimerization, wherein the signaling-capable CISC can generate stimulating signals in signaling pathways that promote cell survival and / or proliferation.

[0247] In one embodiment, a method for editing the genome of a cell, i) Deoxyribonucleic acid (DNA) endonuclease, or nucleic acid encoding said DNA endonuclease; ii) A guide RNA (gRNA) containing a spacer sequence complementary to a target sequence within a target genomic locus of a cell, or a nucleic acid encoding said gRNA; and iii) Donor template including a donor cassette containing a nucleic acid sequence encoding a naked FKBP-rapamycin-binding (FRB) domain polypeptide. The process includes providing the cells with The present invention provides a method characterized in that the DNA endonuclease, the gRNA, and the donor template are configured such that a complex formed by the association of the DNA endonuclease and the gRNA promotes the targeted incorporation of the donor cassette into the target genomic locus of the cell, thereby producing a genetically modified cell capable of expressing the naked FRB domain polypeptide.

[0248] In one embodiment, a method for editing the genome of a cell is provided, comprising the step of providing a cell with one or more donor templates comprising: a) a gRNA directional to a gene or genomic sequence of interest; b) an RNA-induced nuclease (RGEN) or nucleic acid encoding such RGEN according to any embodiment described herein; c) a nucleic acid encoding a naked FRB domain; and d) a nucleic acid encoding a first CISC component comprising i) a first extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof; and i) a second CISC component comprising a second extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof, wherein the first and second CISC components are configured (e.g., arranged) to dimerize in the presence of a ligand (e.g., rapamycin) when expressed by T cells to produce a CISC having signaling ability capable of generating downstream signals (e.g., survival signals or proliferation signals). In some embodiments, one of the CISC components comprises a cleaved IL2Rβ intracellular signaling domain.

[0249] In some embodiments, according to the method for editing the cellular genome described herein, one or more nucleic acids encode a first CISC component comprising a first extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof; and a second CISC component comprising a second extracellular binding domain or a functional derivative thereof, a hinge domain, a transmembrane domain, and a signaling domain or a functional derivative thereof, and one or more of the nucleic acids encoding the first and second CISC components are expressed in one or more vectors. The extracellular binding domain further comprises an endoplasmic reticulum signaling sequence that targets the protein to the extracellular space. In some embodiments, the vector comprises a nucleic acid sequence represented by one or more of SEQ ID NOs: 3, 8, 28, 29, and 30.

[0250] In some embodiments, according to the methods for editing the cell genome described herein, the RNA-inducible nuclease (RGEN) is Cas1 endonuclease, Cas1B endonuclease, Cas2 endonuclease, Cas3 endonuclease, Cas4 endonuclease, Cas5 endonuclease, Cas6 endonuclease, Cas7 endonuclease, Cas8 endonuclease, Cas9 endonuclease (also known as Csn1 and Csx12), Cas100 endonuclease, Csy1 endonuclease, Csy2 endonuclease, Csy3 endonuclease, Cse1 endonuclease, Cse2 endonuclease, Csc1 endonuclease, Csc2 endonuclease, Csa5 endonuclease, Csn2 endonuclease, Csm2 endonuclease, C sm3 endonuclease, Csm4 endonuclease, Csm5 endonuclease, Csm6 endonuclease, Cmr1 endonuclease, Cmr3 endonuclease, Cmr4 endonuclease, Cmr5 endonuclease, Cmr6 endonuclease, Csb1 endonuclease, Csb2 endonuclease, Csb3 endonuclease, Csx17 endonuclease, Csx14 endonuclease The RGEN is selected from the group consisting of rease, Csx10 endonuclease, Csx16 endonuclease, CsaX endonuclease, Csx3 endonuclease, Csx1 endonuclease, Csx15 endonuclease, Csf1 endonuclease, Csf2 endonuclease, Csf3 endonuclease, Csf4 endonuclease, and Cpf1 endonuclease, as well as functional derivatives thereof. In some embodiments, the RGEN is Cas 9. In some embodiments, the nucleic acid encoding the RGEN is a ribonucleic acid (RNA) sequence. In some embodiments, the RNA sequence encoding the RGEN is covalently linked to a first gRNA or a second gRNA. In some embodiments, the RGEN and the first gRNA and / or second gRNA are pre-complexed to form an RNP complex before being provided to cells.In some embodiments, the molar ratio of RGEN when pre-complexing the first gRNA and / or second gRNA is gRNA:RGEN = 1:1 to 20:1.

[0251] Targeted embedding In some embodiments, a sequence encoding a naked FRB domain polypeptide or a functional derivative thereof can be incorporated at a specific location in the host genome by the methods provided herein, and this method is referred to as “targeted incorporation.” In some embodiments, targeted incorporation using a sequence-specific nuclease, such as a site-specific polypeptide, such as a DNA endonuclease (e.g., a nucleic acid-inducible nuclease such as Cas9), can induce double-strand breaks in genomic DNA.

[0252] The CRISPR-Cas system, used in several embodiments, has the advantage of allowing for the identification of the optimal CRISPR-Cas design by rapidly screening a large number of genomic targets. The CRISPR-Cas system uses an RNA molecule called a single-strand guide RNA (sgRNA) that can target a specific sequence in DNA to the relevant Cas nuclease (e.g., Cas9 nuclease). This targeting occurs through the formation of a Watson-Crick base pair between the target sequence of the approximately 20 bp sgRNA and the sequence in the genome. Once the sgRNA binds to the target site, the Cas nuclease cleaves both strands of the genomic DNA to create a double-strand break. The only requirement for designing an sgRNA that targets a specific DNA sequence is that the target sequence must contain a protospacer adjacent motif (PAM) sequence at the 3' end of the sgRNA sequence that is complementary to the genomic sequence. In the case of Cas9 nuclease, the PAM sequence is NRG (where R is A or G and N is any base) or the more restrictive PAM sequence NGG. Therefore, by designing in silico such that a 20 bp sequence is adjacent to a PAM motif, sgRNA molecules targeting any region of the genome can be designed. PAM motifs in eukaryotic genomes occur, on average, at only 15 bp. However, because sgRNAs designed using in silico methods produce double-strand breaks with varying efficiencies in cells, it is not possible to predict the cleavage efficiency of a series of sgRNA molecules produced using in silico methods. Since sgRNAs can be rapidly synthesized in vitro, it is possible to rapidly screen any sequence that could be an sgRNA sequence in a specific genomic region, allowing for the identification of the sgRNA that yields the most efficient cleavage. Generally, when testing a series of sgRNAs within a specific genomic region in cells, the cleavage efficiency ranges from 0 to 90%. The algorithms used in silico and laboratory experiments can be used to measure the off-target effects of specific sgRNAs.In most eukaryotic genomes, a perfect match for the 20bp recognition sequence of an sgRNA occurs only once; however, numerous other sites in the genome produce one or more base pair mismatches with the sgRNA. These sites are cleaved at varying frequencies, and predicting their frequency based on the number or location of mismatches is difficult. Further off-target cleavage may occur at sites not identified by in silico analysis. Therefore, screening a large number of sgRNAs in relevant cell types and identifying those with the most favorable off-target profiles is crucial for selecting the optimal sgRNA for therapeutic use. A favorable off-target profile is determined not only by the actual number of off-target sites and the frequency of cleavage at these sites, but also by considering their location in the genome. For example, off-target sites located near or within functionally important genes (particularly oncogenes or tumor suppressor genes) are considered less favorable than sites within gene regions with no known function. Therefore, identifying the optimal sgRNA cannot be predicted solely by in silico analysis of the organism's genome sequence and requires experimentation. While in silico analysis is useful for narrowing down the guides to be tested, it cannot predict guides that produce high-frequency on-target cleavage or desired guides with low off-target effects. Experimental data have shown that the cleavage efficiency of sgRNAs with perfect matching to the target genomic region (e.g., intron 1 of fibrinogen α) varies widely, from no cleavage to over 90% cleavage, and is unpredictable by any known algorithm. The ability of a particular sgRNA to promote cleavage by Cas enzymes is related to the accessibility of a specific site in genomic DNA, which can be determined by the chromatin structure of that site. In differentiated quiescent cells, the majority of genomic DNA resides in highly condensed heterochromatin, but actively transcribed regions reside in more open chromatin states, which are known to be easily accessible to larger molecules such as proteins like Cas proteins.Certain regions of DNA, even within actively transcribed genes, are more easily accessible than other regions due to the presence or absence of binding of transcription factors or other regulatory proteins. Predicting specific sites within the genome, within a particular genomic locus, or within a specific region of a genomic locus is impossible and therefore must be determined experimentally in the relevant cell types. If a particular site is selected as an insertable site, several modifications can be made to the selected site, for example, by moving a few nucleotides upstream or downstream of the selected site, with or without experimentation.

[0253] In some embodiments, the gRNAs that can be used in the methods disclosed herein include spacer sequences complementary to sequences within the FOXP3, AAVS1, or TCR(TRAC) loci of a cell. In some embodiments, the gRNAs that can be used in the methods disclosed herein include one or more spacer sequences represented by any one of the nucleotide sequences of SEQ ID NOs. 40-57, or derivatives of spacer sequences having at least 85% or at least about 85% nucleotide sequence identity with any one of the nucleotide sequences of SEQ ID NOs. 40-57.

[0254] Nucleic acid modification In some embodiments, as described in detail herein and as known in the art, the polynucleotides introduced into cells have one or more modifications that can be used individually or in combination for, for example, enhancing activity, stability or specificity, altering delivery, reducing or otherwise enhancing the innate immune response of the host cell.

[0255] In certain embodiments, modified polynucleotides are used in a CRISPR / Cas9 / Cpf1 system, in which case the guide RNA (single-molecule guide or bi-molecule guide) and / or the DNA or RNA encoding the Cas endonuclease or Cpf1 endonuclease introduced into the cell can be modified as described and illustrated below. Such modified polynucleotides can be used in a CRISPR / Cas9 / Cpf1 system to edit one or more genomic loci.

[0256] When using the CRISPR / Cas9 / Cpf1 system for such an example (but not limited to this), modifying the guide RNA can improve the formation or stability of the CRISPR / Cas9 / Cpf1 genome editing complex containing the guide RNA, which may be formed from a single-molecule guide or a bi-molecule guide and a Cas endonuclease or Cpf1 endonuclease. Alternatively, modifying the guide RNA can improve the initiation, stability, or dynamics of the interaction between the genome editing complex and the target sequence in the genome, which can be used, for example, to enhance on-target activity. Alternatively, modifying the guide RNA can improve specificity, for example, by improving the relative genome editing rate at the on-target site compared to the effect at other sites (off-target).

[0257] Alternatively, or in addition to the above, the stability of guide RNA can be improved by modifying its resistance to degradation by ribonucleases (RNases) present in cells, thereby extending the half-life of the guide RNA in cells. Modifications that extend the half-life of guide RNA may be particularly useful in embodiments in which Cas endonuclease or Cpf1 endonuclease is introduced into cells and edited using RNA that needs to be translated to produce endonuclease, because the extended half-life of the guide RNA introduced simultaneously with the endonuclease-encoding RNA allows for an extension of the time that the guide RNA and the encoded Cas or Cpf1 endonuclease coexist in the cell.

[0258] Alternatively, or in addition to the foregoing, modifications can be utilized to reduce the likelihood or extent to which RNA introduced into cells induces an innate immune response. As will be discussed later herein and as is known in the art, such immune responses have been well evaluated with respect to RNA interference (RNAi), such as small interfering RNAs (siRNAs), and tend to be associated with shortening of the RNA half-life and / or induction of cytokines or other factors related to the immune response.

[0259] It is also possible to modify the RNA encoding the endonuclease introduced into the cell with one or more modifications, including, but not limited to, modifications that improve the stability of the RNA (such as modifications that increase degradation by RNAse present in the cell), modifications that enhance the translation of the product (e.g., endonuclease), and / or modifications that reduce the likelihood or degree to which the RNA introduced into the cell induces an innate immune response.

[0260] Various modifications, such as those mentioned above and other modifications, can be used in combination. For example, in the case of CRISPR / Cas9 / Cpf1, one or more modifications can be added to the guide RNA (including those exemplified above), and / or one or more modifications can be added to the RNA encoding the Cas endonuclease (including those exemplified above).

[0261] For example, guide RNAs used in CRISPR / Cas9 / Cpf1 systems, or other smaller RNAs, can be readily synthesized by chemical means, as is well known in the art and will be discussed later, thereby allowing for the easy incorporation of various modifications. While chemical synthesis procedures are continuously being expanded, the length of these RNAs, which is significantly longer than approximately 100 nucleotides, tends to make purification by procedures such as high-performance liquid chromatography (HPLC without gels, such as PAGE) difficult. One approach used to produce chemically modified long RNAs is to create two or more molecules and then ligate them. Very long RNAs, such as the RNA encoding the Cas9 endonuclease, can be readily synthesized enzymatically. Enzymatically synthesized RNAs usually have fewer types of modifications available, but modifications that enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes can still be used, as reported in the art and will be discussed later, and new types of modifications are constantly being developed.

[0262] As an example, various types of modifications, particularly those frequently used in the chemosynthesis of smaller RNAs, may be one or more nucleotides with modifications at the 2' position of a sugar, and in some embodiments, these may be 2'-O-alkyl-modified nucleotides, 2'-O-alkyl-O-alkyl-modified nucleotides, or 2'-fluoro-modified nucleotides. In some embodiments, RNA modifications include 2'-fluoro, 2'-amino, or 2'O-methyl modifications at the pyrimidine ribose, debasal residues, or inverted bases at the 3' end of the RNA. Such modifications are incorporated into oligonucleotides by conventional methods, and these oligonucleotides have been shown to have a higher Tm (e.g., higher target binding affinity) for specific targets than 2'-deoxyoligonucleotides.

[0263] It has been shown that incorporating various nucleotide and nucleoside modifications into oligonucleotides can increase their resistance to nuclease digestion compared to natural oligonucleotides. These modified oligonucleotides remain intact for longer periods than unmodified oligonucleotides. Specific examples of modified oligonucleotides include those with a modified skeleton, such as those with phosphorothioates, phosphotriesters, methylphosphonates, short-chain alkyl or cycloalkyl intersaccharide bonds, or short-chain heteroatoms or heterocyclic intersaccharide bonds. Some oligonucleotides have a phosphorothioate skeleton and heteroatom skeletons, particularly the CH2-NH-O-CH2 skeleton, CH skeleton, -N(CH3)-O-CH2 skeleton (known as the methylene (methylimino) skeleton, or MMI skeleton), CH2-ON(CH3)-CH2 skeleton, CH2-N(CH3)-N(CH3)-CH2 skeleton and ON(CH3)-CH2-CH2 skeleton (the natural phosphodiester skeleton is represented as OPO-CH); amide skeleton (see De Mesmaeker, A. et al. (1995). Acc. Chem. Res., 28:366-374); morpholino skeleton structure (see Summerton and Weller, U.S. Patent No. 5,034,506); peptide nucleic acid (PNA) skeleton (where the phosphodiester skeleton of the oligonucleotide is substituted by a polyamide skeleton, and the nucleotide is directly or indirectly bonded to the aza nitrogen atom of the polyamide skeleton; Nielsen, PEet See al. (1991).Science, 254(5037):1497-1500).Phosphorus-containing bonds include, but are not limited to, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkyl phosphotriesters, methylalkyl phosphonates and other alkyl phosphonates having a 3'-alkylene phosphonate and a chiral phosphonate, phosphinates, phosphoramidates having a 3'-aminophosphoramidate and an aminoalkylphosphoramidate, thionophosphoramidates, thionoalkyl phosphonates, thionoalkyl phosphotriesters, boranophosphates having a normal 3'-5' bond, their 2'-5' bond analogues, and those having reverse polarity where adjacent nucleoside unit pairs are linked by a 3'-5' bond and a 5'-3' bond, or by a 2'-5' bond and a 5'-2' bond. U.S. Patent No. 3,687,808; U.S. Patent No. 4,469,863; U.S. Patent No. 4,476,301; U.S. Patent No. 5,023,243; U.S. Patent No. 5,177,196; U.S. Patent No. 5,188,897; U.S. Patent No. 5,264,423; U.S. Patent No. 5,276,019; U.S. Patent No. 5,278,302; U.S. Patent No. 5,286,717; U.S. Patent No. 5,321,131; U.S. Patent No. 5,399,676; U.S. Patent No. 5,405,9 See U.S. Patent No. 39; U.S. Patent No. 5,453,496; U.S. Patent No. 5,455,233; U.S. Patent No. 5,466,677; U.S. Patent No. 5,476,925; U.S. Patent No. 5,519,126; U.S. Patent No. 5,536,821; U.S. Patent No. 5,541,306; U.S. Patent No. 5,550,111; U.S. Patent No. 5,563,253; U.S. Patent No. 5,571,799; U.S. Patent No. 5,587,361; and U.S. Patent No. 5,625,050.

[0264] Morphorino-based oligomeric compounds are described in Braasch, DA et al. (2002). Biochemistry, 41(14):4503-4510; Genesis, Volume 30, Issue 3, (2001); Heasman, J. (2002). Dev. Biol., 243(2):209-214; Nasevicius, A. et al. (2000). Nat. Genet., 26(2):216-220; Lacerra, G. et al. (2000). Proc. Natl. Acad. Sci. USA, 97(17):9591-9596; and U.S. Patent No. 5,034,506 issued on July 23, 1991.

[0265] Oligonucleotide mimetics containing cyclohexenyl nucleic acids are described in Wang, J. et al. (2000). J. Am. Chem. Soc., 122(36):8595-8602.

[0266] Modified oligonucleotide skeletons that do not contain phosphorus atoms have skeletons formed by short-chain alkylnucleoside bonds or short-chain cycloalkylnucleoside bonds, alkylnucleoside bonds or cycloalkylnucleoside bonds having heteroatoms, or one or more short-chain heteroatom nucleoside bonds or short-chain heterocyclic nucleoside bonds. Such skeletons further include skeletons having monophorino bonds (partially formed from the sugar moiety of the nucleoside); siloxane skeletons; sulfide skeletons, sulfoxide skeletons and sulfone skeletons; formacetyl skeletons and thioformacetyl skeletons; methyleneformacetyl skeletons and methylenethioformacetyl skeletons; alkene-containing skeletons; sulfamate skeletons and methyleneimino skeletons and methylenehydrazino skeletons; sulfonate skeletons and sulfonamide skeletons; amide skeletons; and other skeletons having a mixture of N, O, S and CH2 as components. U.S. Patent No. 5,034,506; U.S. Patent No. 5,166,315; U.S. Patent No. 5,185,444; U.S. Patent No. 5,214,134; U.S. Patent No. 5,216,141; U.S. Patent No. 5,235,033; U.S. Patent No. 5,264,562; U.S. Patent No. 5,264,564; U.S. Patent No. 5,405,938; U.S. Patent No. 5,434,257; U.S. Patent No. 5,466,677; U.S. Patent No. 5,470,967; U.S. Patent No. 5,489,677; U.S. Patent No. 5,541,307; U.S. Patent No. 5,56 See U.S. Patent No. 1,225; U.S. Patent No. 5,596,086; U.S. Patent No. 5,602,240; U.S. Patent No. 5,610,289; U.S. Patent No. 5,602,240; U.S. Patent No. 5,608,046; U.S. Patent No. 5,610,289; U.S. Patent No. 5,618,704; U.S. Patent No. 5,623,070; U.S. Patent No. 5,663,312; U.S. Patent No. 5,633,360; U.S. Patent No. 5,677,437; and U.S. Patent No. 5,677,439 (these documents are incorporated herein by reference).

[0267] It may contain one or more substituted sugar moieties, for example, OH, SH, SCH3, F, OCN, OCH3, OCH3O(CH2)nCH3, O(CH2)nNH2, or O(CH2)nCH3 (where n is 1 to 10 or 1 to about 10); C1 to C 10 The 2' position may contain one of the following substituents: lower alkyl, alkoxyalkoxy, substituted lower alkyl, alkaryl, or aralkyl; Cl; Br; CN; CF3; OCF3; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; SOCH3; SO2CH3; ONO2; NO2; N3; NH2; heterocycloalkyl; heterocycloalkaryl; aminoalkylamino; polyalkylamino; substituted silyl; RNA cleavage group; reporter group; intercalator; group that improves the pharmacokinetic properties of oligonucleotides; or one of the other substituents having similar properties to groups that improve the pharmacokinetic properties of oligonucleotides. In some embodiments, modifications include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl)) (Martin, P. et al. (1995). Helv. Chim. Acta, 78(2):486-504). Other modifications include 2'-methoxy (2'-O-CH3), 2'-propoxy (2'-OCH2CH2CH3), and 2'-fluoro (2'-F). Similar modifications may be added to other positions on the oligonucleotide, particularly at the 3' position of the sugar on the 3' terminal nucleotide and at the 5' position of the 5' terminal nucleotide. The oligonucleotide may also have sugar mimetic properties; for example, it may have cyclobutyl instead of a pentofuranosyl group.

[0268] In some embodiments, the sugar and nucleoside bonds (such as the backbone) of the nucleotide unit are both replaced with new groups. The base unit is retained for hybridization with a suitable nucleic acid target compound. Oligonucleotide mimetic compounds, which have been shown to have excellent hybridization properties as one such oligomeric compound, are called peptide nucleic acids (PNAs). In PNA compounds, the sugar backbone of the oligonucleotide is replaced with an amide-containing backbone, such as an aminoethylglycine backbone. The nucleic acid base is retained and is directly or indirectly bonded to the aza nitrogen atom of the amide portion of the backbone. Representative U.S. patent specifications containing teachings on the preparation of PNA compounds include, but are not limited to, U.S. Patents 5,539,082; 5,714,331; and 5,719,262. Further teachings on PNA compounds are found in Nielsen, PE et al. (1991). Science, 254(5037):1497-1500.

[0269] In some embodiments, the guide RNA may further or otherwise include modifications or substitutions with nucleic acid bases (often simply referred to in the art as “bases”). Examples of “unmodified” or “natural” nucleic acid bases as used herein include adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Modified nucleic acid bases include those rarely or transiently found in natural nucleic acids, such as hypoxanthine, 6-methyladenine, 5-methylpyrimidine, especially 5-methylcytosine (also known as 5-methyl-2'deoxycytosine and often referred to as 5-Me-C in this art), 5-hydroxymethylcytosine (HMC), glycosyl HMC and gentobiosyl HMC; and synthetic nucleic acid bases, such as 2-aminoadenine, 2-(methylamino)adenine, 2-(imidazolylalkyl)adenine, 2-(aminoalkylamino)adenine or other heterosubstituted alkyladenines, 2-thiouracil, 2-thiothymine, 5-bromouracil, 5-hydroxymethyluracil, 8-azaguanine, 7-deazaguanine, N 6 Examples include (6-aminohexyl)adenine and 2,6-diaminopurine. Kornberg, A. et al. (1980). DNA Replication (2 nd (ed., pp.75-77). San Francisco, CA: WH Freeman & Co.; Gebeyehu, G. et al. (1987). Nucl Acids Res., 15(11):4513-4534. This may include "universal" bases known in the art, such as inosine. Substitution with 5-Me-C has been shown to increase the stability of nucleic acid double helix by 0.6-1.2°C (Sanghvi, YS (1993). Antisense Research and Applications, (pp. 276-278). Crooke, S. T. and Lebleu, B., (Eds.), Boca Raton, FL: CRC Press), and is an embodiment of base substitution.

[0270] In some embodiments, modified nucleic acid bases include other synthetic nucleic acid bases and native nucleic acid bases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl derivatives of adenine and guanine and other alkyl derivatives, 2-propyl derivatives of adenine and guanine and other alkyl derivatives, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and 5-halocytosine, 5-propynyluracil and 5-propynylcytosine, 6-azouracil, 6-azocytosine and 6-azothimine, 5-uracil (pseudolacil), 4-thiouracil, 8-haloadenine and 8-haloguani Examples include 8-aminoadenine and 8-aminoguanine, 8-thioladenine and 8-thiolguanine, 8-thioalkyladenine and 8-thioalkylguanine, 8-hydroxyladenine and 8-hydroxylguanine, and other 8-substituted adenines and guanines, 5-halouracil and 5-halocytosine, in particular 5-bromouracil and 5-bromocytosine, 5-trifluoromethyluracil and 5-trifluoromethylcytosine, and other 5-substituted uracils and 5-substituted cytosines, 7-methylguanine and 7-methyladenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.

[0271] Furthermore, examples of nucleic acid bases include those disclosed in U.S. Patent No. 3,687,808; those disclosed in Kroschwitz, J. (1990). Concise Encyclopedia of Polymer Science And Engineering, (pp. 858-859) New York, NY: Wiley; those disclosed in Englisch, U. et al. (1991). Angewandte Chemie International Edition, 30(6):613-722; and those disclosed in Sanghvi, YS (1993). Chapter 15, Antisense Research and Applications, (pp. 289-302), Crooke, ST and Lebleu, B. (Eds), Boca Raton, FL: CRC Press. Certain of these nucleic acid bases are particularly useful for improving the binding affinity of the oligomeric compounds of this disclosure. Examples of such nucleic acid bases include 5-substituted pyrimidines, 6-azapyrimidines, N-2 substituted purines, N-6 substituted purines and O-6 substituted purines, nucleic acid bases containing 2-aminopropyladenine, nucleic acid bases containing 5-propynyluracil, and nucleic acid bases containing 5-propynylcytosine. Substitution with 5-methylcytosine has been shown to increase the stability of nucleic acid double helix by 0.6 to 1.2°C (Sanghvi, YS (1993). Antisense Research and Applications, (pp. 276-278). Crooke, S. T. and Lebleu, B., (Eds.), Boca Raton, FL: CRC Press), and is an embodiment of base substitution, and the stability of nucleic acid double helix is ​​further increased, especially when combined with 2'-O-methoxyethyl sugar modification.Modified nucleic acid bases are listed in U.S. Patent Nos. 3,687,808; 4,845,205; 5,130,302; 5,134,066; 5,175,273; 5,367,066; 5,432,272; 5,457,187; 5,459,255; 5,484,908; 5,502,177; U.S. Patent It is described in U.S. Patent No. 5,525,711; U.S. Patent No. 5,552,540; U.S. Patent No. 5,587,469; U.S. Patent No. 5,596,091; U.S. Patent No. 5,614,617; U.S. Patent No. 5,681,941; U.S. Patent No. 5,750,692; U.S. Patent No. 5,763,588; U.S. Patent No. 5,830,653; U.S. Patent No. 6,005,096; and U.S. Patent Application No. 2003 / 0158403.

[0272] In some embodiments, mRNA (or DNA) and / or guide RNA encoding an endonuclease are chemically linked to one or more moieties or conjugates that enhance the activity, intracellular distribution, or cellular uptake of the oligonucleotide. Such components include lipid portions such as cholesterol (Letsinger, RL et al. (1989). Proc. Natl. Acad. Sci. USA, 86(17):6553-6556); cholic acid (Manoharan, M. et al. (1994). Bioorg. Med. Chem. Let., 4(8):1053-1060); thioethers, such as hexyl-S tritylthiol (Manoharan, M. et al. (1992). Ann. NY Acad. Sci., 660(1):306-309; and Manoharan, M. et al. (1993). Bioorg. Med. Chem. Let., 3(12):2765-2770); thiocholesterol (Oberhauser, B. et al. (1992). Nucl. Acids Res., 20(3):533-538); fatty acid chains, e.g., dodecanediol residues, undecyl residues (Kabanov, AV et al. (1990). FEBS Lett., 259(2):327-330 and Svinarchuk, FP et al. (1993). Biochimie, 75(1-2):49-54)); phospholipids, e.g., di-hexadecyl-rac-glycerol, 1,2-di-O-hexadecyl-rac-glycero-3-H-triethylammonium phosphonate (Manoharan et al. (1995). Tetrahedron Lett., 36(21):3651-3654 and Shea, RG et al. (1990). Nucl. Acids Res., 18(13):3777-3783); polyamine or polyethylene glycol chains (Manoharan, M. et al. al. (1995). Nucleos. Nucleot. Nucl., 14(3-5): 969-973); adamantane acetate (Manoharan, M. et al. (1995). Tetrahedron Lett., 36(21):3651-3654); palmityl moiety (Mishra, RK et al. (1995). Biochim. Biophys. Acta, 1264(2):229-237); or octadecylamine moiety or hexylamino-carbonyl-t-oxycholesterol moiety (Crooke, ST et al. (1996). J. Pharmacol. Exp. Ther., 277(2):923-937) are listed, but are not limited to these. Furthermore, U.S. Patent Nos. 4,828,979; 4,948,882; 5,218,105; 5,525,465; 5,541,313; 5,545,730; 5,552,538; 5,578,717; 5,580,731; 5,580,731; 5,591,584; 5,109,124; 5,118,802; 5,138,045; 5,414, No. 077; US No. 5,486,603; US No. 5,512,439; US No. 5,578,718; US No. 5,608,046; US No. 4,587,044; US No. 4,605,735; US No. 4,667,025; US No. 4,762,779; US No. 4,789,737; US No. 4,824,941; US ​​No. 4,835,263; US No. 4,876,335; US No. 4,904,582; US No. 4,958,013; US No. 5,082, No. 830; US No. 5,112,963; US No. 5,214,136; US No. 5,082,830; US No. 5,112,963; US No. 5,214,136; US No. 5,245,022; US No. 5,254,469; US No. 5,258,506; US No. 5,262,536; US No. 5,272,250; US No. 5,292,873; US No. 5,317,098; US No. 5,371,241; US ​​No. 5,391,723; US No. 5,416 See also U.S. Patent Nos. 203, 5,451,463, 5,510,475, 5,512,667, 5,514,785, 5,565,552, 5,567,810, 5,574,142, 5,585,481, 5,587,371, 5,595,726, 5,597,696, 5,599,923, 5,599,928, and 5,688,941.

[0273] In some embodiments, sugars and other portions can be used to target protein and nucleotide-containing complexes (e.g., cationic polysomes and liposomes) to specific sites. For example, hepatocytes can be targeted for transfusion via the asialoglycoprotein receptor (ASGPR). See, for example, Hu, J. et al. (2014). Protein Pept. Lett., 21(10):1025-1030. Biomolecules and / or complexes used in the present invention can be targeted to specific target cells of interest using systems known in the art and other systems that are constantly being developed.

[0274] In some embodiments, these targeting moieties or conjugates may include conjugated groups covalently bonded to functional groups such as primary or secondary hydroxyl groups. Examples of conjugated groups in this disclosure include intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, oligomers, and groups that enhance the pharmacokinetic properties of oligomers. Typical conjugated groups include cholesterol, lipids, phospholipids, biotin, phenazine, folic acids, phenanthridines, anthraquinones, acridines, fluorescein, rhodamine, coumarins, and dyes. Examples of groups that enhance pharmacokinetic properties in this disclosure include groups that improve uptake, enhance resistance to degradation, and / or enhance sequence-specific hybridization with target nucleic acids. Examples of groups that enhance pharmacokinetic properties in this disclosure include groups that improve the uptake, distribution, metabolism, or efflux of the compounds of this disclosure. Representative conjugated groups are disclosed in International Patent Application PCT / US92 / 09196, filed on October 23, 1992, and U.S. Patent No. 6,287,860 (both of which are incorporated herein by reference). Conjugated moieties include, but are not limited to, cholesterol moieties and lipid moieties such as cholic acid; thioethers such as hexyl-5-tritylthiol and thiocholesterol; aliphatic chains such as dodecanediol and undecyl residues; phospholipids such as di-hexadecyl-rac-glycerol or 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonic acid triethylammonium; polyamine or polyethylene glycol chains; or adamantane acetate, palmityl moieties, or octadecylamine or hexylamino-carbonyl-oxycholesterol moieties.For example, U.S. Patent No. 4,828,979; U.S. Patent No. 4,948,882; U.S. Patent No. 5,218,105; U.S. Patent No. 5,525,465; U.S. Patent No. 5,541,313; U.S. Patent No. 5,545,730; U.S. Patent No. 5,552,538; U.S. Patent No. 5,578,717; U.S. Patent No. 5,580,731; U.S. Patent No. 5,580,731; U.S. Patent No. 5,591,584; U.S. Patent No. 5,109,124; U.S. Patent No. 5,118,802; U.S. Patent No. 5,138,045; U.S. Patent No. 5,414 ,077; US Patent No. 5,486,603; US Patent No. 5,512,439; US Patent No. 5,578,718; US Patent No. 5,608,046; US Patent No. 4,587,044; US Patent No. 4,605,735; US Patent No. 4,667,025; US Patent No. 4,762,779; US Patent No. 4,789,737; US Patent No. 4,824,941; US ​​Patent No. 4,835,263; US Patent No. 4,876,335; US Patent No. 4,904,582; US Patent No. 4,958,013; US Patent No. 5,082 ,830; US Patent No. 5,112,963; US Patent No. 5,214,136; US Patent No. 5,082,830; US Patent No. 5,112,963; US Patent No. 5,214,136; US Patent No. 5,245,022; US Patent No. 5,254,469; US Patent No. 5,258,506; US Patent No. 5,262,536; US Patent No. 5,272,250; US Patent No. 5,292,873; US Patent No. 5,317,098; US Patent No. 5,371,241; US ​​Patent No. 5,391,723; US Patent No. 5,41 See U.S. Patent No. 6,203; U.S. Patent No. 5,451,463; U.S. Patent No. 5,510,475; U.S. Patent No. 5,512,667; U.S. Patent No. 5,514,785; U.S. Patent No. 5,565,552; U.S. Patent No. 5,567,810; U.S. Patent No. 5,574,142; U.S. Patent No. 5,585,481; U.S. Patent No. 5,587,371; U.S. Patent No. 5,595,726; U.S. Patent No. 5,597,696; U.S. Patent No. 5,599,923; U.S. Patent No. 5,599,928 and U.S. Patent No. 5,688,941.

[0275] Long polynucleotides, which are not readily synthesized chemically and are generally produced by enzymatic synthesis, can also be modified by various means. Such modifications include, for example, the introduction of specific nucleotide analogs, the incorporation of specific sequences or other parts at the 5' or 3' end of the molecule, and other modifications. As an example, the mRNA encoding Cas9 is approximately 4kb long and can be synthesized by in vitro transcription. Modifications to mRNA may be applied, for example, to increase its translation or stability (e.g., by increasing resistance to degradation in cells), or to reduce the tendency to elicit an innate immune response, which is often observed in cells after the introduction of exogenous RNA, particularly long RNAs such as the exogenous RNA encoding Cas9.

[0276] Numerous such modifications have been reported in the art, including, for example, poly-A tails, 5' cap analogues (e.g., Anti-Reverse Cap Analog (ARCA) or m7G(5')ppp(5')G (mCAP)), modifications of the 5' untranslated region (UTR) or 3' untranslated region (UTR), and the use of modified bases (e.g., pseudo-UTP, 2-thio-UTP, 5-methylcytidine-5'-triphosphate (5-methyl-CTP) or N 6 Examples include methyl-ATP and phosphatase treatment to remove the 5' terminal phosphate. These and other modifications are well known in the art, and new RNA modifications are constantly being developed.

[0277] Numerous commercial suppliers offer modified RNA, including TriLink Biotech, AxoLabs, Bio-Synthesis, and Dharmacon. As described by TriLink, for example, 5-methyl-CTP can be used to confer desirable characteristics such as increased nuclease stability, increased translation, or reduced interaction between innate immune receptors and in vitro transcribed RNA. 5-methylcytidine-5'-triphosphate (5-methyl-CTP), N 6methyl-ATP, as well as pseudo-UTP and 2-thio-UTP, have been shown to reduce innate immune stimulation while enhancing translation in culture and in vivo, as demonstrated in the publications by Kormann et al. (2011) and Warren et al. (2010), cited below.

[0278] Chemically modified mRNA delivered in vivo has been shown to be usable to achieve improved therapeutic efficacy. See, for example, Kormann, MSD et al. (2011). Nat. Biotechnol., 29:154-157. Such modifications can be used, for example, to increase the stability of RNA molecules and / or decrease their immunogenicity. Pseudo U, N 6 Regarding the use of chemical modifications such as methyl-A, 2-thio-U, and 5-methyl-C, it has been found that substituting only one-quarter of uridine and cytidine residues with 2-thio-U and 5-methyl-C, respectively, can significantly reduce mRNA recognition via Toll-like receptors (TLRs) in mice. These modifications can be used to effectively improve in vivo mRNA stability and lifespan by reducing the activation of the innate immune system. See, for example, Kormann et al. (2011).

[0279] It has also been shown that repeated administration of synthetic messenger RNA incorporating modifications designed to circumvent the natural response to viruses can reprogram differentiated human cells into pluripotency. See, for example, Warren, L. et al. (2010). Cell Stem Cell, 7(5):618-630. Such modified mRNAs, primarily acting as reprogramming proteins, can be an efficient means of reprogramming various types of human cells. These cells are called induced pluripotent stem cells (iPSCs), and RNA enzymatically synthesized with the incorporation of 5-methyl-CTP, pseudo-UTP, or Anti Reverse Cap Analog (ARCA) has been found to be effective in evading the cell's antiviral response. See, for example, Warren et al. (2010).

[0280] Other polynucleotide modifications reported in the art include, for example, the use of poly(A) tails, the addition of 5' cap analogs (e.g., m7G(5')ppp(5')G(mCAP)), modification of the 5' untranslated region (UTR) or 3' untranslated region (UTR), and phosphatase treatment to remove the 5' terminal phosphate, with new approaches constantly being developed.

[0281] The various compositions and techniques applicable to the preparation of modified RNA used herein have been developed in relation to the modification of RNA interference (RNAi), such as small interfering RNA (siRNA). The effect of mRNA interference-mediated siRNA on gene silencing is usually transient and may require repeated administration, which poses particular problems in vivo. Furthermore, siRNA is double-stranded RNA (dsRNA), and mammalian cells have immune responses that have evolved to detect and neutralize dsRNA, which often arises as a byproduct of viral infections. Therefore, there are mammalian enzymes such as PKR (dsRNA-responsive kinase), as well as retinoic acid-inducible gene I (RIG-I), which can mediate the cellular response to dsRNA, and Toll-like receptors (e.g., TLR3, TLR7, and TLR8), which can induce cytokines in response to dsRNA. For example, see the reviews by Angart, P. et al. (2013). Pharmaceuticals, 6(4):440-468; Kanasty, RL et al. (2012). Mol. Ther., 20(3):513-524; Burnett, JC et al. (2011). Biotechnol. J., 6(9):1130-1146; and Judge, AD (2008). Hum. Gene Ther., 19(2):111-124, as well as the references cited in these publications.

[0282] As reported in the following literature, a wide variety of modifications have been developed and applied to achieve other benefits that may be useful in relation to enhancing RNA stability, suppressing innate immune responses, and / or introducing polynucleotides into human cells. For example, Whitehead, KA et al. (2011). Ann. Rev. Chem. Biomolec. Eng., 2:77-96; Gaglione, M. et al. (2010). Mini Rev. Med. Chem., 10(7):578-595; Chernolovskaya, EL et al. 12(2):158-167;Deleavey, GG et al. (2009). Curr. Protoc. Nucleic Acid Chem., 39(1):16.3.1-16.3.22;Behlke, MA (2008). Oligonucleotides, 18(4):305-319;Fucini, RV et al. (2012). Nucleic Acid Ther., 22(3):205-210;Bremsen, Please refer to the review article by JB et al. (2012). Front. Genet., 3:154.

[0283] As mentioned above, there are numerous commercial suppliers of modified RNA, and many of these modified RNAs are specialized in modifications designed to improve the efficacy of siRNA. Various approaches are offered based on various findings reported in the literature. For example, Dharmacon states that it is using sulfur-based non-crosslinked oxygen substitution (phosphorothioates (PS)) on a large scale to improve the nuclease resistance of siRNA, as reported by Kole, R. (2012). Nat. Rev. Drug Disc., 11(2):125-140. It has also been reported that modification at the 2' position of ribose increases double-strand stability (Tm) while improving the nuclease resistance of internucleotide phosphate bonds, and this modification has been shown to also provide protection from immune activation. Combining 2'-substitutions with small, well-tolerated substituents (2'-O-methyl, 2'-fluoro, 2'-hydro) with moderate PS backbone modification yields highly stable siRNAs in vivo, as reported by Soutschek, J. et al. (2004). Nature, 432:173-178. 2'-O-methyl modification has been reported to be effective in improving stability, as reported by Volkov, AA et al. (2009). Oligonucleotides, 19:191-202. Regarding the reduction of innate immune response induction, it has been reported that modifying specific sequences with 2'-O-methyl, 2'-fluoro, or 2'-hydro suppresses interaction with TLR7 / TLR8 while significantly retaining silencing activity. For example, see Judge, AD et al. (2006). Mol. Ther., 13:494-505; and Cekaite, L. et al. (2007). J. Mol. Biol., 365(1):90-108. 2-thiouracil, pseudouracil, 5-methylcytosine, 5-methyluracil, and N 6Other modifications, such as methyladenosine, have also been shown to minimize the immune effects mediated by TLR3, TLR7, and TLR8. See, for example, Kariko, K. et al. (2005). Immunity, 23(2):165-175.

[0284] Various conjugates known and commercially available in the art can be used to enhance the delivery and / or cellular uptake of polynucleotides (such as RNA) used herein. Such conjugates include, for example, cholesterol, tocopherols and folic acid, lipids, peptides, polymers, linkers, and aptamers. See, for example, Winkler, J. (2013). Ther. Deliv., 4(7):791-809, and the references cited in this document.

[0285] Delivery In some embodiments, nucleic acid molecules used in the methods provided herein, for example, nucleic acids encoding genome-targeted nucleic acids and / or site-directed polypeptides of this disclosure, are packaged inside or on the surface of a delivery carrier for delivery to cells. Possible delivery carriers include, but are not limited to, nanospheres, liposomes, quantum dots, nanoparticles, polyethylene glycol particles, hydrogels, and micelles. Various targeting moieties can be used to enhance the selective interaction between such carriers and desired cell types or locations, as reported in the Art.

[0286] The introduction of the complexes, polypeptides, or nucleic acids of this disclosure into cells can be carried out by viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, nucleofection, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, or nucleic acid delivery via nanoparticles.

[0287] In some embodiments, guide RNA polynucleotides (RNA or DNA) and / or endonuclease polynucleotides (RNA or DNA) can be delivered by a delivery carrier using a virus known in the art or a non-virus delivery carrier. Alternatively, one or more endonuclease polypeptides can be delivered by a delivery carrier using a virus known in the art or a non-virus delivery carrier, such as electroporation or lipid nanoparticles. In some embodiments, DNA endonucleases can be delivered as one or more polypeptides alone, or pre-complexed with one or more guide RNAs, or pre-complexed with tracrRNA and one or more crRNAs.

[0288] In embodiments, polynucleotides can be delivered by non-viral delivery carriers, including, but not limited to, nanoparticles, liposomes, ribonucleoproteins, positively charged peptides, small RNA conjugates, aptamer-RNA chimeras, and RNA fusion protein complexes. Several specific examples of non-viral delivery carriers are described in Peer, D. et al. (2011). Gene Ther., 18: 1127-1133 (this document focuses on non-viral delivery carriers for siRNA, which are also useful for the delivery of other polynucleotides).

[0289] In embodiments, polynucleotides such as guide RNA, sgRNA, and mRNA encoding endonucleases can be delivered to cells or targets by lipid nanoparticles (LNPs).

[0290] While non-viral nucleic acid delivery methods have been tested in both animal and human models, the most well-developed system is lipid nanoparticles. Lipid nanoparticles (LNPs) generally consist of an ionizable cationic lipid and three or more additional components, which are typically cholesterol, DOPE, and lipid-containing polyethylene glycol (PEG). The cationic lipid can bind to positively charged nucleic acids to form a dense complex that protects the nucleic acid from degradation. As they pass through a microfluidic system, each component self-assembles to form particles 50–150 nM in size, with the nucleic acid encapsulated in a core complexed with the cationic lipid, surrounded by a lipid bilayer-like structure. After endocytosis, the LNPs reside in endosomes. The encapsulated nucleic acid escapes the endosome due to the ionizable nature of the cationic lipid. This delivers the nucleic acid to the cytoplasm, where the mRNA is translated into the encoded protein. Therefore, in some embodiments, both components can be efficiently delivered to cells by encapsulating the mRNA and gRNA encoding Cas9 in LNPs and injecting them intravenously. After escaping from the endosome, Cas9 mRNA can be translated into Cas9 protein and form a complex with gRNA. In some embodiments, nuclear translocation of the Cas9 protein / gRNA complex is facilitated by including a nuclear localization signal in the Cas9 protein sequence. Alternatively, a small gRNA can pass through the nuclear pore complex and form a complex with the Cas9 protein in the nucleus. Once the gRNA / Cas9 complex enters the nucleus, it scans the genome to find homologous target sites and selectively creates double-strand breaks at the desired target sites in the genome. The half-life of RNA molecules in vivo is generally short, ranging from a few hours to a few days. Similarly, the half-life of proteins also tends to be short, ranging from a few hours to a few days. Therefore, in some embodiments, delivery of gRNA and Cas9 mRNA using LNPs can result in only transient expression and activity of the gRNA / Cas9 complex. This can offer the advantage of reducing the frequency of off-target cleavage and, therefore, minimizing the risk of genotoxicity in some embodiments.LNPs are generally less immunogenic than viral particles. While many humans already have immunity to AAV, they do not have pre-existing immunity to LNPs. Furthermore, the likelihood of an adaptive immune response to LNPs is low, making repeated administration of LNPs possible.

[0291] Several types of ionizable cationic lipids have been developed for use in LNPs. These include C12-200 (Love, KT et al. (2010). Proc. Nat. Acad. Sci. USA, 107(5):1864-1869), MC3, LN16, and MD1. In one type of LNP, the GalNac moiety is attached to the outside of the LNP and functions as a ligand for uptake via the asialoclycoprotein receptor. LNPs are formulated using one of these cationic lipids to deliver gRNA and Cas9 mRNA to cells.

[0292] In some embodiments, LNPs refer to particles having a diameter of less than 1000 nm, less than 500 nm, less than 250 nm, less than 200 nm, less than 150 nm, less than 100 nm, less than 75 nm, less than 50 nm, or less than 25 nm. Alternatively, the size of the nanoparticles may be in the range of 1-1000 nm, 1-500 nm, 1-250 nm, 25-200 nm, 25-100 nm, 35-75 nm, or 25-60 nm.

[0293] LNPs can be prepared from cationic lipids, anionic lipids, or neutral lipids. Neutral lipids such as the fusion phospholipid DOPE and the membrane component cholesterol can be included in LNPs as "helper lipids" to enhance transfection activity and nanoparticle stability. Limitations of cationic lipids include low efficacy due to low stability and rapid clearance, and the possibility of inducing inflammatory or anti-inflammatory reactions. LNPs can also contain hydrophobic lipids, hydrophilic lipids, or both hydrophobic and hydrophilic lipids.

[0294] Any lipid or combination known in the art can be used to prepare LNPs. Examples of lipids used to prepare LNPs include DOTMA, DOSPA, DOTAP, DMRIE, DC-cholesterol, DOTAP-cholesterol, GAP-DMORIE-DPyPE, and GL67A-DOPE-DMPE-polyethylene glycol (PEG). Examples of cationic lipids include 98N12-5, C12-200, DLin-KC2-DMA (KC2), DLin-MC3-DMA (MC3), XTC, MD1, and 7C1. Examples of neutral lipids include DPSC, DPPC, POPC, DOPE, and SM. Examples of PEG-modified lipids include PEG-DMG, PEG-CerC14, and PEG-CerC20.

[0295] In this embodiment, LNPs can be prepared by combining multiple types of lipids in any molar ratio. Furthermore, LNPs can also be prepared by combining single or multiple polynucleotides with single or multiple lipids in various molar ratios.

[0296] In embodiments, site-directed polypeptides and genome-targeted nucleic acids can be administered to cells or targets separately. Alternatively, site-directed polypeptides can be pre-complexed with one or more guide RNAs, or with tracrRNA and one or more crRNAs. The pre-complexed material can then be administered to cells or targets. Such pre-complexed materials are known as ribonucleoprotein particles (RNPs).

[0297] RNA can form specific interactions with other RNA or DNA. While this property is utilized in many biological processes, it also carries the risk of indiscriminate interactions in nucleic acid-rich cellular environments. One solution to this problem is to form ribonucleoprotein particles (RNPs) in which RNA is pre-complexed with endonucleases. Another advantage of RNPs is the protection of RNA from degradation.

[0298] In some embodiments, the endonuclease contained in the RNP may or may not be modified. Similarly, the gRNA, crRNA, tracrRNA, or sgRNA may or may not be modified. Various modifications are known in the art and can be used.

[0299] Endonucleases and sgRNAs can typically be combined in a 1:1 molar ratio. Alternatively, endonucleases, crRNAs, and tracrRNAs can typically be combined in a 1:1:1 molar ratio. However, RNPs can be prepared using various molar ratios.

[0300] Gene editing elements can be delivered to the cell nucleus via several mechanisms. These mechanisms can generally be classified into viral and nonviral delivery. A combination of viral and nonviral delivery can be used in various applications.

[0301] In some embodiments, delivery can be carried out using recombinant adeno-associated virus (AAV) vectors. Methods, techniques, and systems suitable for the production of rAAV particles for providing cells with the functionality of an AAV genome and helper virus, packaged to contain the polynucleotides, rep gene, and cap gene to be delivered. The production of rAAV typically requires the presence of the rAAV genome, the AAV rep gene and cap gene isolated from this rAAV genome (e.g., not contained internally), and the functionality of the helper virus in a single cell (referred to herein as a packaging cell). Common approaches for producing rAAV include transient transfection, double infection with insect baculovirus, and monoinfection with insect baculovirus. Further information on these topics can be found, for example, in Ayuso, E. et al. (2010). Curr. Gene Ther., 10(6):423-436; Mietzsch, M. et al. (2014). Hum. Gene Ther., 25(3):212-222; Mietzsch, M. et al. (2015). Hum. Gene Ther., 26(10):688-697; and Mietzsch, M. et al. (2017). Hum. Gene Ther. Method, 28(1):15-22.

[0302] Typically, the AAV rep and cap genes may be derived from a serotype of AAV, such as one from which recombinant viruses can be produced, or from a serotype of AAV different from the ITR in the rAAV genome. Examples of AAV serotypes include, but are not limited to, AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-7, AAV-8, AAV-9, AAV-10, AAV-11, AAV-12, AAV-13, and AAV rh.74. Furthermore, various natural and / or synthetic serotypes of viruses may be used to target specific types of tissues. The production of pseudotyped rAAVs is disclosed, for example, in international patent application WO01 / 83692. Table 1A shows the AAV serotypes and Genbank accession numbers of several selected AAVs. [Table 1]

[0303] In some embodiments, the method for producing packaging cells involves creating a cell line that stably expresses all the components necessary for the production of AAV particles. For example, a single plasmid (or multiple plasmids) containing an rAAV genome lacking AAV rep and cap genes, and AAV rep and cap genes separate from this rAAV genome, along with a selection marker such as a neomycin resistance gene, is incorporated into the cell genome. The AAV genome is introduced into the bacterial plasmid by techniques such as GC tailing (Samulski, RJ et al. (1982). Proc. Natl. Acad. Sci. USA, 79(6):2077-2081), addition of a synthetic linker containing restriction endonuclease cleavage sites (Laughlin, CA et al. (1983). Gene, 23(1):65-73), or direct blunt-end ligation (Senapathy, P. (1984). J. Biol. Chem., 259:4661-4666). Next, the packaging cell line is infected with a helper virus such as adenovirus. Advantages of this method include the ability to select cells, and the cells being suitable for mass production of rAAV. Another preferred method involves introducing the rAAV genome and / or rep and cap genes into the packaging cells using adenovirus or baculovirus instead of plasmids.

[0304] General principles for the production of rAAVs are reviewed, for example, in Carter, BJ (1992). Curr. Opin. Biotechnol., 3(5):533-539; and Muzyczka, M. (1992). Curr. Topics Microbial. Immunol., 158:97-129.Various approaches are described in Tratschin, JD et al. (1984). Mol. Cell. Biol., 4(10):2072-2081; Hermonat, PL et al. (1984). Proc. Natl. Acad. Sci. USA, 81(2):6466-6470; Tratschin, JD et al. (1985). Mo1. Cell. Biol., 5(11):3251-3260;McLaughlin, SK et al. (1988). J. Virol., 62(6):1963-1973;Lebkowski, JS et al. (1988). Mol. Cell. Biol., 8(10):3988-3996. Samulski, RJ et al. (1989). J. Virol. 63(9):3822-3828); US Patent No. 5,173,414; WO 95 / 13365 and its corresponding US Patent No. 5,658,776; WO 95 / 13392; WO 96 / 17947; PCT / US98 / 18600; WO 97 / 09441 (PCT / US96 / 14423); WO 97 / 08298 (PCT / US96 / 13872); WO 97 / 21825 (PCT / US96 / 20777); WO 97 / 06243 (PCT / FR96 / 01064); WO 99 / 11764; Perrin, P. et al. (1995) Vaccine, 13(13):1244-1250; Paul, RW et al. (1993). Hum. Gene Ther., 4(5):609-615; Clark, KR et al. (1996) Gene Ther., 3(12):1124-1132; as described in U.S. Patent Nos. 5,786,211, 5,871,982, and 6,258,595.

[0305] The serotype of the AAV vector can be matched to the target cell type. For example, typical cell types described below can be transduced with AAV serotypes described herein. For example, AAV3, AAV5, AAV8, and AAV9 are suitable serotypes of AAV vectors for liver tissue / hepatocytes, but are not limited to these. In some cases, a serotype of AAV with low immunogenicity in the population or individual being treated may be selected.

[0306] In addition to adeno-associated virus vectors, other viral vectors can also be used. Such viral vectors include, but are not limited to, lentiviruses, alphaviruses, enteroviruses, pestiviruses, baculoviruses, herpesviruses, Epstein-Barr virus, papovaviruses, poxviruses, vaccinia viruses, and herpes simplex viruses.

[0307] In some embodiments, Cas9 mRNA, sgRNA targeting one or two loci of the fibrinogen α gene, and donor DNA are formulated by encapsulating each separately in lipid nanoparticles, or all of them together in a single lipid nanoparticle, or all of them together in two or more lipid nanoparticles.

[0308] In some embodiments, Cas9 mRNA is formulated by encapsulation in lipid nanoparticles, while sgRNA and donor DNA are delivered by integration into an AAV vector.

[0309] Cas9 nuclease can be delivered as a DNA plasmid, mRNA, or protein. Guide RNA can be expressed from the same DNA or delivered as RNA. This RNA can be chemically modified to alter or increase its half-life and / or reduce the likelihood or extent of an immune response. Endonuclease proteins can also be complexed with gRNA before delivery. Viral vectors enable efficient delivery. Split Cas9 and smaller orthologs of Cas9 can be packaged in AAVs, as well as HDR donors. Various nonviral delivery methods exist for delivering each of these components, and nonviral and viral delivery methods can be used in combination. For example, nanoparticles can be used to deliver the protein and guide RNA, and AAVs can be used to deliver the donor DNA.

[0310] In some embodiments relating to the delivery of genome editing components for therapeutic procedures, at least two components, namely a sequence-specific nuclease (e.g., site-directed polypeptide, e.g., DNA endonuclease, e.g., nucleic acid-inducible nuclease) and a DNA donor template, are delivered into the nucleus of the cell to be transformed (e.g., a lymphocyte). In some embodiments, the donor DNA template is packaged in lymphocyte-directed adeno-associated virus (AAV). In some embodiments, the AAV is selected from serotypes AAV8, AAV9, AAVrh10, AAV5, AAV6, and AAV-DJ. In some embodiments, the AAV-packaged DNA donor template is first administered to the subject (e.g., a patient) by peripheral intravenous injection, followed by the administration of the sequence-specific nuclease (e.g., DNA endonuclease). The advantage of initially delivering the donor DNA template packaged in AAV is that the delivered donor DNA template is stably maintained in the nucleus of the transduced cell, thereby enabling the subsequent administration of a sequence-specific nuclease (e.g., DNA endonuclease), which creates double-strand breaks in the genome and facilitates the integration of the DNA donor by HDR or NHEJ. In some embodiments, it is desirable that the sequence-specific nuclease (e.g., DNA endonuclease) remain active in the target cell for only the time required to promote targeted integration of the transgene, at a level sufficient to exert the desired therapeutic effect. If the sequence-specific nuclease (e.g., DNA endonuclease) remains active in the cell for a long period, this increases the frequency of off-target double-strand breaks. Typically, the frequency of off-target breaks is a function of the off-target break efficiency multiplied by the duration of nuclease activity. Because mRNA and the proteins translated from it are short-lived within cells, delivering sequence-specific nucleases (e.g., DNA endonucleases) in mRNA form shortens the duration of nuclease activity, ranging from a few hours to several days.Therefore, it is expected that delivering sequence-specific nucleases (e.g., DNA endonucleases) to cells already containing the donor template will maximize the ratio of targeted integration to off-target integration. Furthermore, delivery of the donor DNA template into the cell nucleus via peripheral intravenous injection via AAV is time-consuming, typically taking about 1–14 days, because the virus must infect the cell, bypass endosomes, migrate to the nucleus, and be converted from single-stranded AAV genome to double-stranded DNA molecules by host cell components. Thus, in at least some embodiments, since nuclease elements are typically active for about 1–3 days, delivery of the donor DNA template into the nucleus should be completed before supplying the CRISPR-Cas element.

[0311] In some embodiments, the sequence-specific nuclease (e.g., DNA endonuclease) is CRISPR-Cas9, consisting of a Cas9 nuclease and sgRNA directed to a DNA sequence within intron 1 of the fibrinogen α gene. In some embodiments, the Cas9 nuclease is delivered as mRNA encoding the Cas9 protein, operably fused to one or more nuclear localization signals (NLS). In some embodiments, the sgRNA and Cas9 mRNA are delivered to cells packaged in lipid nanoparticles. In some embodiments, the lipid nanoparticles contain C12-200 lipids (Love, KT et al. (2010). Proc. Nat. Acad. Sci. USA, 107(5):1864-1869). In some embodiments, the ratio of sgRNA to Cas9 mRNA packaged in the LNP is 1:1 (mass ratio) to maximize DNA cleavage in vivo in mice. In another embodiment, the mass ratio of sgRNA to Cas9 mRNA packaged in the LNP can be varied, for example, 10:1, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, or 2:1, or the reverse of these ratios. In some embodiments, Cas9 mRNA and sgRNA are packaged in separate LNP formulations, and the Cas9 mRNA-containing LNP is delivered to the subject approximately 1 to 8 hours before the delivery of the sgRNA-containing LNP, thereby optimizing the time for Cas9 mRNA to be translated before sgRNA delivery.

[0312] In some embodiments, LNP preparations containing gRNA and Cas9 mRNA ("LNP-nuclease preparations") are administered to subjects (e.g., patients) who have already received a DNA donor template packaged in an AAV. In some embodiments, the LNP-nuclease preparations are administered to subjects within 1 to 28 days, 7 to 28 days, or 7 to 14 days after administration of the AAV-donor DNA template. The optimal timing for delivering the LNP-nuclease preparation after administration of the AAV-donor DNA template can be determined using techniques known in the art, such as studies conducted in animal models, including mice and monkeys.

[0313] In some embodiments, DNA donor templates are delivered to the target's (e.g., patient) cells using nonviral delivery methods. While some targets (approximately 30%) already possess neutralizing antibodies against the most commonly used serotypes of AAV, thus hindering effective gene delivery by AAV, the use of nonviral delivery methods allows for the treatment of a larger number of targets. Several nonviral delivery methods are known in this field. In particular, lipid nanoparticles (LNPs) are known to efficiently deliver encapsulated cargoes into the cytoplasm of cells by intravenous injection in animals and humans. These LNPs are actively taken up by cells via receptor-mediated endocytosis.

[0314] In some embodiments, to promote the nuclear localization of the donor template, a 366 bp region consisting of the origin of replication and initial promoter of Simian virus 40 (SV40), which can promote plasmid nuclear localization, can be added to the donor template. Other DNA sequences that bind to cellular proteins can also be used to improve DNA nuclear translocation.

[0315] In some embodiments, the expression level or activity of the sequence encoding the introduced target protein (e.g., FVIII) is measured in the blood of the subject (e.g., patient) after administration of the AAV-donor DNA template and after the first administration of an LNP-nuclease preparation (e.g., containing gRNA and Cas9 nuclease or mRNA encoding Cas9 nuclease). If the amount of the target protein is insufficient for disease cure, for example, if the amount of the target protein is found to be at least 5-50% of the normal value, and especially if the amount of the target protein is found to be 5-20% of the normal value, a second or third administration of the LNP-nuclease preparation may be performed to promote further targeted integration of the fibrinogen α gene into intron 1. Whether it is possible to administer the LNP-nuclease preparation multiple times to achieve the desired therapeutic dose of the target protein may be tested and optimized using techniques known in the art, such as tests using animal models, such as mouse models or monkey models.

[0316] In some embodiments, according to any of the methods herein, which include the steps of i) administering an AAV-donor DNA template containing a donor cassette and ii) administering an LNP-nuclease preparation to a subject, the initial administration of the LNP-nuclease preparation to the subject is performed within 1 to 28 days after administration of the AAV-donor DNA template to the subject. In some embodiments, the initial administration of the LNP-nuclease preparation to the subject is performed after sufficient time has elapsed for the donor DNA template to be delivered into the nucleus of the target cell. In some embodiments, the initial administration of the LNP-nuclease preparation to the subject is performed after sufficient time has elapsed for the single-stranded AAV genome to be converted into double-stranded DNA molecules in the nucleus of the target cell. In some embodiments, after the initial administration of the LNP-nuclease preparation, the LNP-nuclease preparation is administered one or more times (e.g., two, three, four, five or more times). In some embodiments, the LNP-nuclease preparation is administered to the subject one or more times until the targeted integration of the donor cassette reaches a target amount and / or the expression of the donor cassette reaches a target amount. In some embodiments, the method further includes the steps of measuring the amount of targeted integration of the donor cassette and / or the expression of the donor cassette after each administration of the LNP-nuclease preparation, and administering the LNP-nuclease preparation one or more times if the targeted integration of the donor cassette has not reached a target amount and / or the expression of the donor cassette has not reached a target amount. In some embodiments, at least one of the one or more further administrations of the LNP-nuclease preparation is the same as the initial dose. In some embodiments, at least one of the one or more further administrations of the LNP-nuclease preparation is less than the initial dose. In some embodiments, at least one of the one or more further administrations of the LNP-nuclease preparation is more than the initial dose.

[0317] Therapeutic approach In one embodiment, a gene therapy approach is provided for treating a subject by adoptive cell therapy using cells edited according to any of the methods described herein. For example, in some embodiments, the disorder or health condition is an autoimmune disease (e.g., IPEX syndrome) or a disorder resulting from organ transplantation (e.g., GVHD), and the edited cells have a Treg phenotype. In some embodiments, the gene therapy approach incorporates nucleic acids comprising a naked FRB domain polypeptide and a sequence encoding CISC into the genome of a relevant type of cell of the subject, thereby providing a stable treatment such as alleviating one or more symptoms of the disorder or health condition over a long period of time, and / or providing a cure or alleviation of the disorder or health condition. In some embodiments, the cell type targeted by the gene therapy approach into which the naked FRB domain polypeptide and the sequence encoding CISC are incorporated is lymphocytes, for example, CD4+ T cells.

[0318] In some embodiments, the ex vivo cell therapy is carried out using lymphocyte cells isolated from a subject, for example, autologous CD4+ T cells derived from umbilical cord blood. The chromosomal DNA of the cells is then edited using the systems, compositions, and methods described herein. Finally, the edited cells are transplanted into the subject.

[0319] One advantage of the ex vivo cell therapy approach is that a comprehensive analysis of the therapeutic agent can be performed before administration. Nuclease-based therapies always come with some degree of off-target effects. By performing gene modification ex vivo, the characteristics of the modified cell population can be comprehensively confirmed before transplantation. Aspects of this disclosure include sequencing the entire genome of the modified cells and, if off-target cuts are present, confirming that these cuts are located in genomic locations that minimize risk to the target. Furthermore, specific cell populations, such as clonal populations, can be isolated before transplantation.

[0320] Another embodiment of such a method is an in vivo treatment. In this method, the chromosomal DNA of cells in a subject is modified using the systems, compositions and methods described herein. In some embodiments, the cells are lymphocytes, such as CD4+ cells, such as T cells.

[0321] One advantage of in vivo gene therapy is the ease of manufacturing and administering the drug. The same therapeutic approach and method can be used to treat one or more subjects, for example, multiple subjects with the same or similar genotype or allele. In contrast, ex vivo cell therapy typically uses the subject's own cells, isolating them, genetically modifying them, and then returning them to the same subject.

[0322] Another embodiment relates to genome-edited recombinant cells obtained by any one of the methods described herein for use in the suppression or treatment of FOXP3-related diseases or conditions such as inflammatory diseases or autoimmune diseases. Further embodiments relate to the use of genome-edited recombinant cells obtained by any one of the methods described herein as pharmaceuticals.

[0323] Cell transplantation to the target In some embodiments, the ex vivo methods of the present disclosure include transplanting genome-edited cells into a subject requiring such a method. This transplantation step can be carried out using any transplantation method known in the art. For example, the genetically modified cells can be injected directly into the subject's blood or administered to the subject by other means.

[0324] In some embodiments, the methods disclosed herein include “administration,” but this term can be used interchangeably with “introduction” and “transplantation,” and the methods disclosed herein include administering genetically modified therapeutic cells to a subject by a method or route that allows at least a portion of the introduced cells to be localized to a desired site in order to obtain a desired single or multiple effect. The therapeutic cells or progeny cells differentiated therefrom can be administered via a preferred route that allows delivery to a desired site in the subject’s body in which at least a portion of the transplanted cells or at least a portion of their components can remain viable. The viability of the cells after administration to the subject may range from a short period of time (e.g., several hours to several days) to several years or even the lifetime of the subject (e.g., long-term engraftment).

[0325] When therapeutic cells described herein are provided for prophylactic purposes, such therapeutic cells can be administered to a subject before any symptoms of the disease or condition being treated develop. Therefore, in some embodiments, prophylactic administration of a recombinant stem cell population is used to prevent the development of symptoms of the disease or condition.

[0326] In some embodiments, when genetically modified stem cells are provided for therapeutic purposes, the genetically modified stem cells are provided at (or after) the onset of symptoms or signs of a disease or condition, for example, at the onset of the disease or condition.

[0327] The effective amount of therapeutic cells (e.g., genome-edited stem cells) used in the various embodiments described herein is at least 10 2 each, at least 5 × 10 2 pieces, at least 10 3 each, at least 5 × 10 3 pieces, at least 10 4 each, at least 5 × 10 4 pieces, at least 10 5 each, at least 2 × 10 5 each, at least 3 × 10 5 each, at least 4 × 105 each, at least 5 × 10 5 each, at least 6 × 10 5 each, at least 7 × 10 5 each, at least 8 × 10 5 each, at least 9 × 10 5 each, at least 1 × 10 6 each, at least 2 × 10 6 each, at least 3 × 10 6 each, at least 4 × 10 6 each, at least 5 × 10 6 each, at least 6 × 10 6 each, at least 7 × 10 6 each, at least 8 × 10 6 each, at least 9 × 10 6 The number of cells may be one or multiples thereof. The therapeutic cells may be derived from one or more donors or may be autologous. In some embodiments described herein, the therapeutic cells are cultured and grown before being administered to a subject requiring the administration of the therapeutic cells.

[0328] In embodiments, at least a portion of a therapeutic cell composition (e.g., a composition comprising multiple cells of any of the cells described herein) can be localized to a desired site by delivering the composition to a target by a specific method or route. The cell composition can be administered by any suitable route that provides effective treatment to the target, for example, by administration of the cell composition, at least a portion of the composition (e.g., at least 1 × 10⁶ cells) can be localized to a desired site. 4Cells can be delivered to a desired site within the target body over a specified period of time. Methods of administration include injection, intravenous infusion, intra-infusion, and oral ingestion. "Injection" includes, but is not limited to, intravenous injection, intramuscular injection, intra-arterial injection, intracerebroventricular injection, intra-articular injection, intraorbital injection, intracardiac injection, intradermal injection, intraperitoneal injection, transtracheal injection, subcutaneous injection, subepidermal injection, intra-articular injection, subcapsular injection, subarachnoid injection, intraspinal injection, intracerebrospinal injection, and intrasternal injection, as well as intravenous infusion. In some embodiments, the route of administration is intravenous. Cell delivery can be performed by injection or intravenous infusion.

[0329] In one embodiment, the cells are administered systemically; in other words, the therapeutic cell population enters the target circulatory system and undergoes metabolism and other similar processes, rather than being directly administered to a target site, target tissue, or target organ.

[0330] The effectiveness of a treatment using a composition for the treatment of a disease or condition can be determined by a skilled clinician. However, a treatment is considered effective if one or all of the signs, symptoms, or markers of the disease are improved or alleviated. Effectiveness can also be measured by determining that the deterioration of the individual has stopped (e.g., the progression of the disease has stopped or at least slowed) based on an assessment of the need for hospitalization or medical intervention. Methods for measuring these indicators are known to those skilled in the art and / or are described herein. Treatment includes any treatment of a disease in an individual or animal (including, but not limited to, humans or mammals) and includes (1) suppression of the disease, e.g., cessation or delay of the progression of symptoms, or (2) alleviation of the disease, e.g., induction of regression of symptoms, and (3) prevention of the onset of symptoms or reduction of the likelihood of the onset of symptoms.

[0331] composition In one embodiment, the present disclosure provides compositions for carrying out the methods disclosed herein. The compositions may include one or more of the following: genome-targeted nucleic acids (e.g., gRNA); site-directed polypeptides (e.g., DNA endonucleases) or nucleotide sequences encoding such site-directed polypeptides; and polynucleotides (e.g., donor templates) to be inserted to achieve the desired genetic recombination by the methods disclosed herein.

[0332] In some embodiments, the composition comprises a nucleotide sequence encoding a genome-targeted nucleic acid (e.g., gRNA).

[0333] In some embodiments, the composition comprises a site-directed polypeptide (e.g., a DNA endonuclease). In some embodiments, the composition comprises a nucleotide sequence encoding a site-directed polypeptide.

[0334] In some embodiments, the composition comprises a polynucleotide (e.g., a donor template) to be inserted into the genome.

[0335] In some embodiments, the composition comprises (i) a nucleotide sequence encoding a genome-targeting nucleic acid (e.g., gRNA), and (ii) a site-directed polypeptide (e.g., a DNA endonuclease), or a nucleotide sequence encoding the site-directed polypeptide.

[0336] In some embodiments, the composition comprises (i) a nucleotide sequence encoding a nucleic acid (e.g., gRNA) that targets the genome, and (ii) a polynucleotide (e.g., a donor template) to be inserted into the genome.

[0337] In some embodiments, the composition comprises (i) a site-directed polypeptide (e.g., a DNA endonuclease) or a nucleotide sequence encoding the site-directed polypeptide, and (ii) a polynucleotide (e.g., a donor template) to be inserted into the genome.

[0338] In some embodiments, the composition comprises (i) a nucleotide sequence encoding a nucleic acid (e.g., gRNA) that targets the genome, (ii) a site-directed polypeptide (e.g., a DNA endonuclease), or a nucleotide sequence encoding the site-directed polypeptide, and (iii) a polynucleotide (e.g., a donor template) to be inserted into the genome.

[0339] In some embodiments of the composition, the composition comprises a genome-targeted nucleic acid using a single-molecule guide. In some embodiments of the composition, the composition comprises a genome-targeted nucleic acid using a bimolecule guide. In some embodiments of the composition, the composition comprises two or more bimolecule guides or single-molecule guides. In some embodiments, the composition comprises a vector encoding a nucleic acid that targets a nucleic acid. In some embodiments, the genome-targeted nucleic acid is a DNA endonuclease, particularly Cas9.

[0340] In some embodiments, the composition may include one or more gRNAs that can be used for genome editing, particularly for the insertion of naked FRB domain polypeptides and sequences encoding CISCs into the cellular genome. The one or more gRNAs may target a genomic site on the endogenous FOXP3 gene, a genomic site within the endogenous FOXP3 gene, or a genomic site near the endogenous FOXP3 gene. Thus, in some embodiments, the one or more gRNAs may have a spacer sequence complementary to the genomic sequence on the FOXP3 gene, a spacer sequence complementary to the genomic sequence within the FOXP3 gene, or a spacer sequence complementary to the genomic sequence near the FOXP3 gene.

[0341] In some embodiments, the gRNA for the composition includes spacer sequences shown in SEQ ID NOs. 40-57, and a spacer sequence selected from any one of the variants having at least 50% or about 50%, at least 55% or about 55%, at least 60% or about 60%, at least 65% or about 65%, at least 70% or about 70%, at least 75% or about 75%, at least 80% or about 80%, at least 85% or about 85%, at least 90% or about 90%, or at least 95% or about 95% identity or homology with any one of SEQ ID NOs. In some embodiments, the gRNA variant for the kit includes a spacer sequence having at least 85% or about 85% homology with any one of SEQ ID NOs. 40-57.

[0342] In some embodiments, the gRNA for the composition has a spacer sequence complementary to the target site in the genome. In some embodiments, the spacer sequence is 15 to 20 base pairs long. In some embodiments, the complementarity between the spacer sequence and the genome sequence is at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 100%.

[0343] In some embodiments, the composition may have a donor template having a DNA endonuclease or a nucleic acid encoding the DNA endonuclease, and / or a naked FRB domain polypeptide and a nucleic acid sequence encoding CISC. In some embodiments, the nucleic acid sequences encoding naked FRB domain polypeptides and CISCs have at least 70% or at least about 70% sequence identity with the sequences shown in SEQ ID NOs. 3-4, 8, 28-30, 32 and 37-39, for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least more sequence identity. In some embodiments, the nucleic acid sequence encoding the naked FRB domain polypeptide and CISC has at least 70% or at least about 70% sequence identity with the sequence shown in SEQ ID NO: 32, for example, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least more sequence identity, or at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least more sequence identity. In some embodiments, the DNA endonuclease is Cas9. In some embodiments, the nucleic acid encoding the DNA endonuclease is DNA or RNA.

[0344] In some embodiments, one or more nucleic acids for the kit can be encoded in an adeno-associated virus (AAV) vector. Therefore, in some embodiments, gRNA can be encoded in an AAV vector. In some embodiments, nucleic acids encoding DNA endonucleases can be encoded in an AAV vector. In some embodiments, donor templates can be encoded in an AAV vector. In some embodiments, two or more nucleic acids can be encoded in a single AAV vector. Therefore, in some embodiments, nucleic acids encoding DNA endonucleases and gRNA sequences can be encoded in a single AAV vector.

[0345] In some embodiments, the composition may have liposomes or lipid nanoparticles. Therefore, in some embodiments, any compound of the composition (e.g., DNA endonuclease or the nucleic acid, gRNA, and donor template encoding the DNA endonuclease) can be formulated by encapsulating it in liposomes or lipid nanoparticles. In some embodiments, one or more such compounds are bound to the liposomes or lipid nanoparticles via covalent or non-covalent bonds. In some embodiments, any of the compounds can be encapsulated separately or together in liposomes or lipid nanoparticles. Therefore, in some embodiments, each of the DNA endonuclease or the nucleic acid, gRNA, and donor template encoding the DNA endonuclease is formulated by encapsulating them separately in liposomes or lipid nanoparticles. In some embodiments, the DNA endonuclease is formulated by encapsulating it together with the gRNA in liposomes or lipid nanoparticles. In some embodiments, the DNA endonuclease or the nucleic acid, gRNA, and donor template encoding the DNA endonuclease are formulated together by encapsulating them in liposomes or lipid nanoparticles.

[0346] In some embodiments, the composition further comprises one or more additional reagents, such additional reagents selected from buffers, buffers for introducing polypeptides or polynucleotides into cells, wash buffers, control reagents, control vectors, control RNA polynucleotides, reagents for producing polypeptides from DNA in vitro, sequencing adapters, etc. The buffer may be a stabilizing buffer, a reconstitution buffer, a dilution buffer, etc. In some embodiments, the composition may further comprise one or more components used to promote or enhance on-target binding, cleave DNA by endonucleases, or improve targeting specificity.

[0347] In some embodiments, any of the components of the composition are formulated with pharmaceutically acceptable additives such as carriers, solvents, stabilizers, adjuvants, and diluents, depending on the specific method of administration and dosage form. In embodiments, the guide RNA composition is typically formulated to achieve a physiologically compatible pH, which is in the range of 3 to 11 or about 3 to about 11, or 3 to 7 or about 3 to about 7, depending on the formulation and route of administration. In some embodiments, the pH is adjusted to a range of pH 5.0 to pH 8 or about 5.0 to pH about 8. In some embodiments, the composition comprises a therapeutically effective amount of at least one compound described herein and one or more pharmaceutically acceptable additives. Optionally, the composition may have a combination of compounds described herein, or may contain a second active ingredient useful for treating or preventing bacterial growth (e.g., antimicrobial agents or antimicrobial agents), or may contain a combination of reagents of the present disclosure. In some embodiments, the gRNA is formulated with one or more other nucleic acids, such as a nucleic acid encoding a DNA endonuclease and / or a donor template. Alternatively, the nucleic acid encoding a DNA endonuclease and the donor template may be formulated separately or in combination with another nucleic acid in the manner described above for the gRNA formulation.

[0348] Suitable additives include carrier molecules containing macromolecules with slow metabolic progression, such as proteins, polysaccharides, polylactic acid, polyglycolic acid, high molecular weight amino acids, amino acid copolymers, and inactive virus particles. Other typical additives include antioxidants (e.g., ascorbic acid), chelating agents (e.g., EDTA), carbohydrates (e.g., dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (e.g., oils, water, saline solution, glycerol, and ethanol), wetting agents or emulsifiers, or pH buffering agents.

[0349] In some embodiments, any of the compounds in the composition (e.g., a DNA endonuclease or the nucleic acid encoding the DNA endonuclease, gRNA, and donor template) can be delivered to cells by chemical transfection (e.g., lipofection) or by transfection such as electroporation. In some embodiments, the DNA endonuclease can be pre-complexed with the gRNA to form a ribonucleoprotein (RNP) complex before being delivered to cells. In some embodiments, the RNP complex is delivered to cells by transfection. In such embodiments, the donor template is delivered to cells by transfection.

[0350] In some embodiments, “composition” refers to a therapeutic composition comprising therapeutic cells used in ex vivo therapeutic treatments.

[0351] In some embodiments, the therapeutic composition comprises a cell composition and a physiologically acceptable carrier, and may also contain at least one further bioactive agent described herein, dissolved or dispersed in the composition, as an active ingredient. In some embodiments, the therapeutic composition is substantially non-immunogenic when administered to mammalian or human subjects for therapeutic purposes, except when immunogenicity is desired.

[0352] Typically, the recombinant therapeutic cells described herein are administered as a suspension containing a pharmaceutically acceptable carrier. Those skilled in the art will understand that the pharmaceutically acceptable carrier used in the cell composition does not contain amounts of buffers, compounds, cryopreservants, preservatives, or other agents that substantially inhibit the viability of the cells delivered to the target. The cell-containing formulation may, for example, contain an osmotic buffer to maintain the integrity of the cell membrane, and may optionally contain nutrients to maintain cell viability at administration or to enhance cell engraftment. Such formulations and suspensions are known to those skilled in the art and / or can be adapted for use with progenitor cells according to the description herein using routine experiments.

[0353] In some embodiments, the cell composition can be emulsified, or provided as a liposome composition, provided that the emulsification process does not adversely affect the viability of the cells. The cells and other active ingredients can be mixed with one or more pharmaceutically acceptable and compatible additives in amounts suitable for use in the therapeutic methods described herein.

[0354] Further agents included in the cell composition include pharmaceutically acceptable salts of the components in the cell composition. Examples of pharmaceutically acceptable salts include acid addition salts (formed with free amino groups of polypeptides) formed by inorganic acids such as hydrochloric acid and phosphoric acid, or organic acids such as acetic acid, tartaric acid, and mandelic acid. Salts formed with free carboxyl groups can be derived from inorganic bases such as sodium, potassium, ammonium, calcium, and iron hydroxide, and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, and procaine.

[0355] Physiologically acceptable carriers are well known in the art. Typical liquid carriers are sterile aqueous solutions containing no materials other than the active ingredient and water, or sterile aqueous solutions containing a buffer such as sodium phosphate or physiological saline with a physiological pH value, or both (such as phosphate-buffered saline). Furthermore, aqueous carriers may contain one or more buffer salts, salts such as sodium chloride or potassium chloride, dextrose, or polyethylene glycol or other solutes. Liquid compositions may also contain another liquid phase in addition to water, or may contain another liquid phase after removing water. Examples of such other liquid phases include glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of active compound used in a cell composition effective for treating a particular disorder or condition varies depending on the characteristics of the disorder or condition and can be determined by known clinical techniques.

[0356] Kits and Systems This specification provides kits and systems comprising cells, expression vectors and protein sequences provided and described herein, and instructions for manufacturing and using them. For example, a kit comprising one or more of the protein sequences, expression vectors, and / or cells described herein is provided. Furthermore, a system for selectively activating a signal into rapamycin-resistant cells is provided, comprising the cells described herein, wherein the cells comprise the expression vector described herein, and the expression vector comprises a nucleic acid encoding the protein sequence described herein.

[0357] Some embodiments provide kits comprising one or more components of the CRISPR / Cas system for genome editing described herein.

[0358] In some embodiments, the kit may have one or more additional therapeutic agents that can be administered simultaneously or sequentially with other compositions within the kit for a desired purpose, such as genome editing or cell therapy.

[0359] In some embodiments, the kit may further include instructions for carrying out the method using each component of the kit. Instructions for carrying out the method are typically recorded on a suitable recording medium. For example, the instructions may be printed on a substrate such as paper or plastic. The instructions may also be included in the kit as an accompanying document on a label of the container containing the kit or its components (e.g., a label affixed to the packaging or sub-packaging). Alternatively, the instructions may be included as an electronic storage data file on a suitable computer-readable storage medium (e.g., a CD-ROM, diskette, flash drive). In some cases, the actual instructions may not be included in the kit, and a means for obtaining the instructions from a remote source (e.g., via the Internet) may be provided. An example of this embodiment is a kit that includes a web address from which the instructions can be viewed and / or downloaded. Similar to the instructions, the means for obtaining the instructions may be recorded on a suitable substrate.

[0360] As used herein, plural and / or singular terms may be interpreted by those skilled in the art as plural terms being singular and / or singular terms being plural, in accordance with the descriptions and / or uses herein. Various singular and plural terms are intentionally used to clarify the present invention.

[0361] Those skilled in the art will understand that the terms used herein, in particular in the appended claims (for example, the body of the appended claims), are generally "open-ended" terms (for example, the term "including" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "at least having," and the term "include" should be interpreted as "including, but not limited to").

[0362] Any feature of the embodiments described in the first to eleventh embodiments is applicable to all embodiments and models described herein. Furthermore, any part or all of the features of the embodiments described in the first to eleventh embodiments can be independently combined with any other embodiments described herein, for example, by combining all or part of one, two, three or more embodiments. Furthermore, any feature of the embodiments described in the first to eleventh embodiments may be adopted as an optional feature in other embodiments or models. Although various features, aspects and functions have been described with respect to various typical embodiments and models, the various features, aspects and functions described in individual embodiments are not limited to their application in the specific embodiments in which they are described, and regardless of whether a specific embodiment is described or whether a specific feature is presented as part of an embodiment of the present invention, the various features, aspects and functions described in individual embodiments may be used individually or in various combinations for application to one or more other embodiments of the present invention. Therefore, the scope of the present invention is not limited by the typical embodiments described above.

[0363] The examples and embodiments described herein are for illustrative purposes only. Those skilled in the art will be shown that these examples and embodiments can be modified and altered in various ways, and such modifications and alterations are included within the scope of the spirit and scope of this application and the appended claims. All publications, patents and patent applications cited herein are incorporated herein by reference in their entirety for all purposes.

[0364] Some embodiments relating to the disclosures provided herein will be further illustrated by the following examples, but the present invention is not limited thereto. [Examples]

[0365] Unless otherwise stated, the present invention is carried out using conventional techniques in molecular biology, microbiology, cell biology, biochemistry, nucleic acid chemistry and immunology that are well known to those skilled in the art. Such techniques are described in Sambrook, J., & Russell, DW (2012).; Molecular Cloning: A Laboratory Manual (4 th ed.). Cold Spring Harbor, NY: Cold Spring Harbor Laboratory and Sambrook, J., & Russel, DW (2001).;Molecular Cloning: A Laboratory Manual (3 rd Cold Spring Harbor, NY: Cold Spring Harbor Laboratory (jointly referred to herein as "Sambrook"); Ausubel, FM (1987).Current Protocols in Molecular Biology New York, NY: Wiley (including supplements through 2014);Mullis, KB, Ferre, F. & Gibbs, R. (1994).PCR: The Polymerase Chain Reaction, Boston: Birkhauser Publisher; Harlow, E., & Lane, D. (1999).Antibodies: A Laboratory Manual (2 ndThis is adequately explained in literature such as: (ed.) New York, NY: Cold Spring Harbor Laboratory Press; Beaucage, SL et al. (2000). Current Protocols in Nucleic Acid Chemistry, New York, NY: Wiley, (including supplements through 2014); and Makrides, SC (2003). Gene Transfer and Expression in Mammalian Cells, Amsterdam, NL: Elsevier Sciences BV).

[0366] Example 1: Fabrication and Characterization of DISC Constructs This example describes the preparation and testing of a lentiviral construct encoding CISC for intracellular expression of a naked "decoy" FRB domain in the cytoplasm of a host cell.

[0367] A CISC-containing lentiviral construct for regulating IL-2 signaling was modified with an additional naked FRB* domain located at the 3' end of the CISC chimeric receptor protein (Figure 1). The construct expressing the naked FRB* domain along with CISC was named "decoy-CISC" or "DISC". While we do not wish to be constrained by any particular theory, it is thought that the naked FRB domain competes with the endogenous FRB domain of mTOR for binding to rapamycin, thereby enabling CI...

Claims

1. A method for preparing genetically modified mammalian cells in vitro or ex vivo, comprising the steps of introducing into mammalian cells (i) a viral vector comprising a nucleotide sequence encoding a naked FK506-binding protein-rapamycin-binding (FRB) domain polypeptide and (ii) one or more polynucleotides encoding components of a chemically induced signaling complex (CISC), The naked FRB domain polypeptide comprises an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in Sequence ID No. 1, The polynucleotides, individually or in combination of two or more, encode components of the CISC. A method characterized in that the CISC is a signal transduction complex (IL2R-CISC) containing components of the interleukin-2 receptor, wherein in the presence of rapamycin, the FK506-binding protein (FKBP) domain and the FRB domain heterodimerize to transmit the IL-2 signal into the mammalian cell.

2. The method according to claim 1, wherein the naked FRB domain polypeptide comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence shown in SEQ ID NO:

1.

3. The method according to claim 1, wherein the naked FRB domain polypeptide comprises the amino acid sequence shown in SEQ ID NO:

1.

4. The method according to claim 1, wherein the naked FRB domain polypeptide comprises the amino acid sequence shown in SEQ ID NO:

1.

5. The method according to claim 1, wherein the naked FRB domain polypeptide comprises the amino acid sequence shown in SEQ ID NO:

2.

6. The method according to claim 1, wherein the naked FRB domain polypeptide comprises the amino acid sequence shown in SEQ ID NO:

2.

7. The method according to any one of claims 1 to 6, characterized in that the viral vector further comprises a promoter, the promoter being operably linked to the nucleotide sequence encoding the naked FRB domain polypeptide.

8. The method according to claim 7, wherein the nucleotide sequence encoding the naked FRB domain polypeptide is inserted into the genome of the mammalian cell such that the promoter is operably linked to the first exon of the FOXP3 gene in the genome of the mammalian cell.

9. The method according to claim 7 or 8, wherein the promoter is a constitutive promoter.

10. The method according to any one of claims 7 to 9, wherein the promoter is an MND promoter.

11. The method according to any one of claims 1 to 10, wherein the viral vector is a lentiviral vector.

12. The method according to any one of claims 1 to 10, wherein the viral vector is an adeno-associated virus vector.

13. The method according to any one of claims 1 to 12, further comprising the step of contacting the mammalian cell with (i) a DNA endonuclease or a nucleic acid encoding the DNA endonuclease and (ii) a guide RNA (gRNA) having a spacer sequence complementary to a sequence in the genome of the mammalian cell or a nucleic acid encoding the gRNA.

14. The method according to claim 13, wherein the DNA endonuclease or the nucleic acid encoding the DNA endonuclease and the gRNA or the nucleic acid encoding the gRNA are contained in lipid nanoparticles.

15. The method according to any one of claims 1 to 14, further comprising the step of contacting the mammalian cells with rapamycin.

16. One or more of the polynucleotides encoding the components of the IL2R-CISC are present in the viral vector. The polynucleotides, individually or in combination of two or more, encode components of the IL2R-CISC. The first CISC component comprises a first extracellular binding domain, a first transmembrane domain, and an interleukin-2 (IL-2) receptor subunit γ (IL2Rγ) cytoplasmic signaling domain. The second CISC component comprises a second extracellular binding domain, a second transmembrane domain, and an IL-2 receptor subunit β (IL2Rβ) cytoplasmic signaling domain. Here, one of the first CISC component and the second CISC component is characterized in that one contains an FK506-binding protein (FKBP) domain as an extracellular binding domain, and the other contains an FRB domain as an extracellular binding domain. The method according to any one of claims 1 to 15, characterized in that the first CISC component and the second CISC component heterodimerize in the presence of rapamycin to form a CISC having IL-2 signaling ability.

17. The method according to claim 16, wherein the first transmembrane domain comprises an IL2Rγ transmembrane domain and the second transmembrane domain comprises an IL2Rβ transmembrane domain.

18. The method according to claim 16 or 17, wherein the first extracellular binding domain comprises the FK506-binding protein (FKBP) domain, and the second extracellular binding domain comprises the FRB domain.

19. The first CISC component includes a first hinge domain, and the second CISC component includes a second hinge domain. The first hinge domain is located between the first extracellular binding domain and the first transmembrane domain. The method according to any one of claims 16 to 18, characterized in that the second hinge domain is located between the second extracellular binding domain and the second transmembrane domain.

20. The method according to any one of claims 16 to 19, wherein the first CISC component comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence shown in SEQ ID NO: 12, and / or the second CISC component comprises an amino acid sequence having at least 95% sequence identity with the amino acid sequence shown in SEQ ID NO:

13.

21. A genetically modified mammalian cell prepared by the method of any one of claims 1 to 15, comprising the nucleotide sequence encoding the naked FRB domain polypeptide and one or more polynucleotides encoding the components of the IL2R-CISC, A genetically modified mammalian cell characterized in that the polynucleotide, either individually or in combination of two or more, encodes a component of the IL2R-CISC.

22. The genetically modified mammalian cell according to claim 21, which is a T lymphocyte or a precursor T cell.

23. The genetically modified mammalian cell according to claim 21 or 22, which is a CD4+ T cell or a CD8+ T cell.

24. A genetically modified mammalian cell according to any one of claims 21 to 23, which is a regulatory T cell.

25. A genetically modified mammalian cell according to any one of claims 21 to 24, which is a FOXP3+ regulatory T cell.

26. A pharmaceutical composition for the treatment of a disease or disorder, comprising genetically modified mammalian cells according to any one of claims 21 to 25 and a pharmaceutically acceptable additive.

27. The pharmaceutical composition according to claim 26, wherein the disease or disorder is an inflammatory disease, an autoimmune disease, a foxp3-related disease or foxp3-related condition, or a disorder resulting from organ transplantation.