Compositions and methods for generating cells with reduced immunogenicity
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2026-03-18
AI Technical Summary
Existing cell therapy products, such as CAR T cells, are time-consuming and costly to collect, modify and amplify, affect the success rate of treatment, and have immunogenic problems such as transplantation and anti-response.
Gene editing is performed through the CRISPR-Cas system, specifically reduces HLA-I, HLA-II and TCR expression, thereby reducing the immunogenicity of cells and introducing CAR genes to enhance the therapeutic ability of cells.
It has achieved the reduction of the immunogenicity of cell therapy, improved the success rate and safety of treatment, and reduced production costs, and can produce relatively economical all-human cell therapy products on demand.
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 322,634, filed March 22, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0002] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Background technology]
[0003] Current cell therapy products, such as CAR T cells, involve collecting cells from a prospective patient, then modifying and optionally expanding these cells before using them for one or more treatments. The entire process can be time-consuming, negatively impact the success and outcome of the treatment, and expensive. As a result, there is a strong need to develop on-demand, reasonably priced, allogeneic cell therapy products that exhibit reduced immunogenicity, e.g., reduced graft-versus-host and / or host-versus-graft responses. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] U.S. Patent No. 10,266,850 [Patent Document 2] U.S. Patent No. 8,906,616 [Patent Document 3] WO2021 / 067788 [Patent Document 4] U.S. Patent No. 9,790,490 [Patent Document 5] U.S. Patent No. 10,113,179 [Patent Document 6] WO2021 / 158918 [Patent Document 7] U.S. Patent No. 9,982,279 [Patent Document 8] U.S. Patent No. 9,896,696 [Patent Document 9] WO2021 / 108324 [Patent Document 10] U.S. Patent Application Publication No. 2014 / 0242664 [Patent Document 11] U.S. Patent No. 10,900,034 [Patent Document 12] U.S. Patent No. 10,767,175 [Patent Document 13] U.S. Patent Application Publication No. 2018 / 0119140 [Patent Document 14] U.S. Patent No. 8,697,359 [Patent Document 15] U.S. Patent No. 10,113,167 [Patent Document 16] U.S. Patent No. 10,570,418 [Patent Document 17] U.S. Patent No. 10,829,787 [Patent Document 18] U.S. Patent No. 11,118,194 [Patent Document 19] U.S. Patent No. 11,125,739 [Patent Document 20] U.S. Patent Application Publication No. 2015 / 0344912 [Patent Document 21] U.S. Patent Application Publication No. 2018 / 0119140 [Patent Document 22] U.S. Patent Application Publication No. 2018 / 0282763 [Patent Document 23] U.S. Patent No. 11,118,194 [Patent Document 24] U.S. Patent No. 11,125,739 [Patent Document 25] WO2016 / 164356 [Patent Document 26] U.S. Patent No. 9,890,396 [Patent Document 27] WO2017 / 053729 [Patent Document 28] U.S. Patent No. 9,982,278 [Patent Document 29] WO2015 / 148863 [Patent Document 30] U.S. Patent No. 7,446,190 [Patent Document 31] U.S. Patent No. 8,399,645 [Patent Document 32] U.S. Patent No. 8,906,682 [Patent Document 33] U.S. Patent No. 9,181,527 [Patent Document 34] U.S. Patent No. 9,272,002 [Patent Document 35] U.S. Patent No. 9,266,960 [Patent Document 36] U.S. Patent No. 10,253,086 [Patent Document 37] U.S. Patent No. 10,640,569 [Patent Document 38] U.S. Patent No. 10,808,035 [Patent Document 39] WO2013 / 142034 [Patent Document 40] WO2015 / 120180 [Patent Document 41] WO2015 / 188141 [Patent Document 42] WO2016 / 120220 [Patent Document 43] WO2017 / 040945 [Patent Document 44] WO2017 / 017184 [Patent Document 45] WO2013 / 126794 [Patent Document 46] WO2013 / 163628 [Patent Document 47] WO2015 / 048577 [Patent Document 48] WO2015 / 070083 [Patent Document 49] WO2015 / 089354 [Patent Document 50] WO2015 / 134812 [Patent Document 51] WO2015 / 138510 [Patent Document 52] WO2015 / 148670 [Patent Document 53] WO2015 / 148860 [Patent Document 54] WO2015 / 153780 [Patent Document 55] WO2015 / 153789 [Patent Document 56] WO2015 / 153791 [Patent Document 57] U.S. Patent No. 8,383,604 [Patent Document 58] U.S. Patent No. 8,859,597 [Patent Document 59] U.S. Patent No. 8,956,828 [Patent Document 60] U.S. Patent No. 9,255,130 [Patent Document 61] U.S. Patent No. 9,273,296 [Patent Document 62] U.S. Patent Application Publication No. 2009 / 0222937 [Patent Document 63] U.S. Patent Application Publication No. 2009 / 0271881 [Patent Document 64] U.S. Patent Application Publication No. 2010 / 0229252 [Patent Document 65] U.S. Patent Application Publication No. 2010 / 0311124 [Patent Document 66] U.S. Patent Application Publication No. 2011 / 0016540 [Patent Document 67] U.S. Patent Application Publication No. 2011 / 0023139 [Patent Document 68] U.S. Patent Application Publication No. 2011 / 0023144 [Patent Document 69] U.S. Patent Application Publication No. 2011 / 0023145 [Patent Document 70] U.S. Patent Application Publication No. 2011 / 0023146 [Patent Document 71] U.S. Patent Application Publication No. 2011 / 0023153 [Patent Document 72] U.S. Patent Application Publication No. 2011 / 0091441 [Patent Document 73] U.S. Patent Application Publication No. 2012 / 0159653 [Patent Document 74] U.S. Patent Application Publication No. 2013 / 0145487 [Non-patent literature]
[0005] [Non-Patent Document 1] Makarova et al. (2017) Cell, 168: 328 [Non-patent document 2] Wang et al. (2016) Annu. Rev. Biochem., 85: 227 [Non-patent document 3] Zetsche et al. (2015) Cell, 163: 759 [Non-patent document 4] Makarova et al. (2017) Cell, 168: 328 [Non-patent document 5] Shmakov et al. (2015) Mol. Cell, 60: 385 [Non-patent document 6] Zetsche et al. (2015) Cell, 163: 759 [Non-Patent Document 7] Yamano et al. (2016) Cell, 165: 949 pages [Non-patent document 8] Gaoら(2016) Cell Res., 26: 901 pages [Non-licensed Document 9] Kimら(2017) ACS Synth. Biol.6(7): pages 1273~82 [Non-licensed Document 10] Zhangら(2017) Cell Discov. 3:17018 pages [Non-licensed Document 11] Gaoら(2017) Nat. Biotechnol., 35: 789 pages [Non-licensed Document 12] Jayavaradhan(2019) Nat. Commun. 10(1): 2866 pages [Non-licensed Document 13] Janssen (2019) Mol. Ther. Nucleic Acids 16: 141-54 [Non-licensed Document 14] Kocak(2019) Nat. Biotech. 37: Pages 657~66 [Non-licensed Document 15] Gruberら(2008) Nucleic Acids Res.、36(Web Server issue): W70-W74 [Non-licensed Document 16] rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi [Non-licensed Document 17] Zuker and Stiegler (Nucleic Acids Res.9 (1981), pages 133~148) [Non-licensed Document 18] AR Gruber, 2008, Cell 106(1): pages 23~24 [Non-licensed Document 19] PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151~62 pages [Non-licensed Document 20] Parkら(2018) Nat. Commun.9: 3313 pages [Non-licensed Document 21] Wuら(2018) Cell. Mol. Life Sci., 75(19): 3593~3607 pages [Non-licensed Document 22] Gruber(2008) Nucleic Acids Res.、36: W70 [Non-licensed Document 23] Maruyamaら(2015) Nat Biotechnol. 33(5): pages 538~42 [Non-licensed Document 24] Chu (2015) Nat Biotechnol. 33(5): 543~48 pages [Non-licensed Document 25] Yuら(2015) Cell Stem Cell 16(2): pages 142~47 [Non-licensed Document 26] Pinder (2015) Nucleic Acids Res. 43(19): 9379~92 pages [Non-licensed Document 27] Yagizら(2019) Commun. Biol. 2: 198 pages [Non-licensed Document 28] Wattsら(2008) Drug Discov. Today 13: 842~55 pages [Non-licensed Document 29] Hendelら(2015) Nat. Biotechnol. 33: 985 pages [Non-licensed Document 30] Dangら(2015) Genome Biol. 16: 280 pages [Non-licensed Document 31] Kocazら(2019) Nature Biotech. 37: Pages 657~66 [Non-licensed Document 32] Liu (2019) Nucleic Acids Res. 47(8): 4169~4180 pages [Non-licensed Document 33] Schubertら(2018) J. Cytokine Biol. 3(1): 121 pages [Non-licensed Document 34] Tengら(2019) Genome Biol. 20(1): 15 pages [Non-licensed Document 35] Wattsら(2008) Drug Discov. Today 13(19~20): pages 842~55 [Non-licensed Document 36] Piccirilli (1990) Nature, 343: 33 pages [Non-licensed Document 37] Rappaport (1993) Biochemistry, 32: 3047 pages [Non-licensed Document 38] Pardridgeら(2010) Cold Spring Harb. Protoc.、doi:10.1101 / pdb.prot5407 [Non-licensed Document 39] Shalekら(2012) Nano Letters, 12: 6498 pages [Non-licensed Document 40] Van Brunt (1988) Biotechnology, 6: 1149 pages [Non-licensed Document 41] Anderson (1992) Science, 256: 808 pages [Non-licensed Document 42] Nabel & Feigner (1993) TIBTECH, 11: 211 pages [Non-licensed Document 43] Mitani & Caskey (1993) TIBTECH, 11: 162 pages [Non-licensed Document 44] Dillon (1993) TIBTECH, 11: 167 pages; Miller (1992) Nature, 357: 455 pages. [Non-licensed Document 45] Vigne, (1995) Restorative Neurology and Neuroscience, 8: 35 pages [Non-licensed Document 46] Kremer & Perricaudet (1995) British Medical Bulletin, 51: 31 pages [Non-licensed Document 47] Haddada (1995) Current Topics in Microbiology and Immunology, 199: 297 pages Yu (1994) Gene Therapy, 1: 13 pages [Non-licensed Document 48] Doerfler and Bohm (Eds.) (2012) The Molecular Repertoire of Adenoviruses II: Molecular Biology of Virus-Cell Interactions [Non-licensed Document 49] Goeddel, Gene Expression Technology: Methods in Enzymology, 185, Academic Press, San Diego, Calif. (1990) [Non-licensed Document 50] Takebe (1988) Mol. Cell. Biol., 8: 466 pages [Non-licensed Document 51] O'Hare (1981) Proc. Natl. Acad. Sci. USA., 78: 1527 pages [Non-licensed Document 52] kazusa.or.jp / codon / [Non-licensed Document 53] Nakamura (2000) Nucl. Acids Res., 28: 292 pages [Non-licensed Document 54] Changら(1987) Proc. Natl. Acad Sci USA, 84: 4959 pages [Non-licensed Document 55] Nehlsら(1996) Science, 272: 886 pages [Non-licensed Document 56] Savicら(2018) eLife 7:e33761 [Non-licensed Document 57] Lazzarotto (2018) Nat Protoc. 13(11): 2615~42 pages [Non-licensed Document 58] Wienertら(2019) Science 364(6437): pages 286~89 [Non-licensed Document 59] Kleinstiverら(2016) Nat. Biotech. 34: pp. 869~74 [Non-licensed Document 60] Kocak(2019) Nat. Biotech. 37: Pages 657~66 [Non-licensed Document 61] crispr.mit.edu (Hsuら(2013) Nat. Biotech. 31: 827~832 pages) [Non-licensed Document 62] Martin, Remington's Pharmaceutical Sciences, 15th ed., Mack Publ. Co., Easton, PA (1975) [Non-licensed Document 63] Remington's Pharmaceutical Sciences, 18th Edition (Mack Publishing Company, 1990) [Non-licensed Document 64] Anselmoら(2016) Bioeng. Transl. Med. 1: 10~29 pages [Non-licensed Document 65] Remington: The Science and Practice of Pharmacy, Mack Publishing Co., 20th Edition, 2000 [Non-licensed Document 66] Sustained and Controlled Release Drug Delivery Systems, JR Robinson (ed.), Marcel Dekker, Inc., New York, 1978 [Non-licensed Document 67] Gruppら(2015) Blood、126: 4983 pages [Non-licensed Document 68] Park (2015) J. Clin. Oncol., 33: 7010 pages [Non-licensed Document 69] Locke(2015) Blood、126: 3991 pages [Non-licensed Document 70] Hale (2017) Mol Ther Methods Clin Dev., 4: 192 pages [Non-licensed Document 71] MacLeod ら(2017) Mol Ther, 25: 949 pages [Non-licensed Document 72] Eyquemら(2017)Nature、543: 113 pages [Non-licensed Document 73] Liuら(2017) Cell Res, 27: 154 pages [Non-licensed Document 74] Renら(2017) Clin Cancer Res、23: 2255 pages [Non-licensed Document 75] Cooperら(2018) Leukemia、32: 1970 pages [Non-licensed Document 76] Renra (2017) Oncotarget, 8: 17002 pages [Non-licensed Document 77] Suら(2016) Oncoimmunology、6: e1249558 [Non-licensed Document 78] Zhangら(2017) Front Med, 11: 554 pages [Non-licensed Document 79] Zenatti (2011) Nat. Genet. 43(10):932~39 pages [Non-licensed Document 80] kumc.edu / gec / support, genome.gov / 10001200 [Non-licensed Document 81] ncbi.nlm.nih.gov / books / NBK22183 / [Brief explanation of the drawings]
[0006] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings. [Figure 1A] FIG. 1A shows a schematic diagram illustrating the structure of an exemplary single-guide VA-type CRISPR system. [Figure 1B] FIG. 1B shows a schematic diagram illustrating the structure of an exemplary dual-guide VA-type CRISPR system. [Figure 2A] FIG. 2A shows a series of diagrams illustrating the incorporation of protecting groups (e.g., protective nucleotide sequences or chemical modifications) into a VA-type CRISPR-Cas system. [Figure 2B] Figure 2B shows a series of diagrams illustrating the incorporation of a donor template recruitment sequence into a VA-type CRISPR-Cas system. [Figure 2C] Figure 2C shows a series of diagrams illustrating the incorporation of editing enhancers into a Type VA CRISPR-Cas system. While these additional elements are shown in the context of a dual-guide Type VA CRISPR system, it is understood that they may also be present in other CRISPR systems, including single-guide Type VA CRISPR systems, single-guide Type II CRISPR systems, or dual-guide Type II CRISPR systems. [Figure 3] Figure 3 shows the percentage of treated cell populations measured by flow cytometry where (A) TCR, HLA-I, and HLA-II were triple knocked out, or (B) TCR, HLA-I, and HLA-II were triple KO and CAR was inserted after treatment; FL = full length, ldsPLA074 = linear DNA used to insert the CAR. [Figure 4]Figure 4 shows the reduction of HLA-I, HLA-II, and / or TCR surface expression (y-axis) in cells transfected with RNPs containing nucleic acid-guided nucleases complexed with various gCD3D gNAs. [Figure 5] Figure 5 shows the reduction of HLA-I, HLA-II, and / or TCR surface expression (y-axis) in cells treated with various RNPs containing nucleic acid-guided nucleases complexed with CD247, CD3G, or TRAC gNA. [Figure 6A] FIG. 6A shows the reduction in TCR surface expression (y-axis) in cells transfected with RNP containing nucleic acid-guided nuclease complexed with TRBC gNA. [Figure 6B] FIG. 6B shows simultaneous TRBC KO and CAAR KI (CAAR expression, y-axis) in cells transfected with RNP containing TRBC gNA and a nucleic acid-guided nuclease complexed with a repair template. [Figure 7] Figure 7 shows the reduction of TRC surface expression in cells transfected with RNPs containing nucleic acid-guided nuclease complexed with CD3E gNA (7A, y-axis); and simultaneous CD3E KO and CAR KI (CAR expression, y-axis, 7B) in cells transfected with RNPs containing nucleic acid-guided nuclease complexed with TRBC gNA and repair template. DETAILED DESCRIPTION OF THE INVENTION
[0007] overview I. Cells with reduced immunogenicity A. Compositions Comprising Cells 1. Cells containing genome modifications 2. Cell populations containing genome modifications 3. Complex of guide nucleic acid and nucleic acid-guided nuclease for generating genome modifications B. Methods for reducing the immunogenicity of cells II. Engineered, non-naturally occurring dual-guide CRISPR-cas systems A. Cas proteins B. Guide Nucleic Acid C.gNA modification III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA A. Ribonucleoprotein (RNP) Delivery and "casRNA" Delivery B. CRISPR expression system C. Donor template D. Efficiency and Specificity E. Multiplicity F. Genomic Safe Harbor IV. Pharmaceutical Compositions V. Therapeutic Use A. Gene Therapy VI. Kit VII. Embodiments VIII. Working Examples IX. Equivalents
[0008] I. Cells with reduced immunogenicity The immune system recognizes specific antigen patterns on cell surfaces, e.g., human leukocyte antigen (HLA) proteins in humans. These patterns of protein antigens are genetically determined and vary among individuals, with an individual's immune system recognizing its own specific antigen pattern as "self" and different antigen patterns as "non-self" or "foreign." Typically, foreign cells, e.g., allogeneic cells (cells derived from genetically dissimilar individuals) and / or those displaying a different HLA pattern than expected, provoke one or more immune responses in the host. In the context of cell therapy applications, this immune response, referred to as "host versus graft" (HvG), can hinder and / or reduce the effectiveness of one or more therapeutic agents because the body recognizes the therapeutic agent as foreign and targets it for elimination.
[0009] Furthermore, engineered cells, e.g., modified cells, used in cell therapy can recognize the antigenic patterns of the host cells as foreign and elicit an immune response. This immune response, referred to herein as "graft versus host" (GvH), can result in the therapy having negative and / or detrimental effects on the recipient.
[0010] Provided herein are compositions, methods, and / or kits for generating cells that exhibit reduced immunogenicity. In certain embodiments, provided herein are cells comprising one or more modifications that result in reduced HvG, GvH, and / or both. In certain embodiments, the cells comprise eukaryotic cells. In certain embodiments, the cells comprise human cells. In certain embodiments, the cells comprise human immune cells, such as neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, lymphocytes, or combinations thereof, e.g., T cells. In preferred embodiments, the cells comprise T cells. In certain embodiments, the cells comprise engineered immune cells, e.g., chimeric antigen receptor (CAR)-T cells comprising one or more CAR polypeptides or portions thereof and / or dual CARs. In certain embodiments, the cells comprise human stem cells, such as human pluripotent or pluripotent stem cells, embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells, CD34+ cells, or combinations thereof. In preferred embodiments, the human stem cells comprise hematopoietic stem cells, CD34+ stem cells, and / or induced pluripotent stem cells (iPSCs). In certain embodiments, the cells comprise allogeneic cells. As used herein, the term "allogeneic" includes cells derived from the same species that are genetically dissimilar and therefore immunologically incompatible with the host.
[0011] In certain embodiments, provided herein are compositions, methods, and / or kits comprising dual CARs, e.g., CAR fusion proteins or two separate CARs. As used herein, the term "dual CAR" includes polypeptides comprising a first CAR or portion thereof and a second CAR or portion thereof, either separate or connected by one or more polypeptide linkers. In certain embodiments, the second CAR or portion thereof targets the same antigen as the first CAR or portion thereof. In certain embodiments, the second CAR or portion thereof targets a different antigen than the first CAR or portion thereof. Additionally, polypeptides comprising any number of CARs or portions thereof, either separate or connected by one or more polypeptide linkers, are disclosed herein. In certain embodiments, the cell comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 and / or no more than 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 CARs or portions thereof, for example, 1 to 15, preferably 1 to 10, more preferably 2 to 10, even more preferably 2 to 7, and even more preferably 2 to 5 CARs or portions thereof, separate or connected by one or more polypeptide linkers. The polypeptide linker may comprise any suitable linker, including natural or non-naturally occurring amino acids.
[0012] In certain embodiments, cells can be engineered to contain one or more genomic modifications. In certain embodiments, cells can be engineered to contain one or more genomic modifications that reduce the immunogenicity of the cells, e.g., the engineered cells elicit little or no immune response in vitro and / or in vivo. In certain embodiments, allogeneic cells with respect to the host (recipient, patient, or suitable option) can be engineered to contain one or more genomic modifications that reduce the immunogenicity of one or more allogeneic cells in the host. In certain embodiments, cells can be engineered to elicit an immune response of 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1% or less compared to unengineered equivalents. In certain embodiments, cells can be engineered not to elicit an immune response in the host. The immune response can be measured using any suitable technique, for example, flow cytometry or ELISA.
[0013] In certain embodiments, the cells comprise (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-1 protein, (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-2 protein or transcription factors that regulate the expression of one or more subunits of an HLA-2 protein, and / or (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of a TCR protein. In preferred embodiments, the cells comprise all three genomic modifications. In certain embodiments, the one or more genomic modifications completely inactivate one or more genes. In certain embodiments, the one or more genomic modifications at least partially or completely eliminate surface expression of an active (immunogenic) protein. In certain embodiments, the one or more genomic modifications completely eliminate surface expression of an active (immunogenic) protein. In certain embodiments, cells comprising one or more genomic modifications may further comprise one or more additional modifications, including, but not limited to, the introduction of one or more heterologous genes, e.g., transgenes. The one or more transgenes can be introduced into any suitable location in the genome. In certain embodiments, the one or more transgenes are introduced into a safe harbor site (SHS), e.g., a safe harbor such as those discussed in the Genomic Safe Harbor section below. In certain embodiments, the one or more transgenes are introduced into one or more sites comprising genomic modifications (1)-(3). For example, a CAR transgene can be introduced into one or more genes encoding a subunit of a TCR protein, e.g., the TRAC gene, and / or a B2M-HLA-E and / or B2M HLA-G fusion protein can be introduced into one or more genes encoding a subunit of an HLA-1 protein, e.g., the B2M gene.
[0014] In certain embodiments, provided herein are compositions comprising one or more populations of cells having the genetic modifications described herein. In certain embodiments, the composition comprises a single cell population, each cell comprising the same set of genomic modifications (1) through (3). In certain embodiments, provided herein are compositions comprising multiple cell populations, each cell population comprising a different set of genomic modifications. Generally, at least one cell population comprises cells comprising all of: (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of an HLA-1 protein; (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein; and (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of a TCR protein, in addition to one or more additional cell populations that do not comprise all three genetic modifications. In certain embodiments, the one or more additional cell populations comprise cells comprising (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-1 protein, (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-2 protein or transcription factors that regulate the expression of one or more subunits of an HLA-2 protein, and / or (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of a TCR protein, but not all of (1) through (3). In a preferred embodiment, the subunit of the HLA-1 protein comprises B2M. In a preferred embodiment, the transcription factor that regulates the expression of one or more subunits of an HLA-2 protein comprises CIITA. In certain embodiments, the subunit of the TCR protein is an alpha subunit or a beta subunit. In a preferred embodiment, the gene encoding a subunit of a TCR protein is the TRAC gene.In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In more preferred embodiments, at least one cell population comprises one or more additional cell populations, cells that contain one or more, but not all three, genomic modifications in addition to all of: (1) one or more genomic modifications that partially or completely inactivate the B2M gene; (2) one or more genomic modifications that partially or completely inactivate the CIITA gene; and (3) one or more genomic modifications that partially or completely inactivate a TRC subunit gene, e.g., the TRAC gene. In certain embodiments, the one or more genomic modifications at least partially or completely eliminate surface expression of the active (immunogenic) protein. In certain embodiments, the one or more genomic modifications completely eliminate surface expression of the active (immunogenic) protein. In certain embodiments, the one or more cells comprising one or more genomic modifications may further comprise one or more additional modifications, including, but not limited to, the introduction of one or more heterologous genes, e.g., transgenes. The one or more transgenes can be introduced into any suitable location in the genome. In certain embodiments, the one or more transgenes are introduced into a safe harbor site (SHS), e.g., a safe harbor, such as those discussed in the Genomic Safe Harbor section below. In certain embodiments, the one or more transgenes are introduced into one or more sites comprising genomic modifications (1)-(3). For example, a CAR transgene can be introduced into one or more genes encoding a subunit of a TCR protein, e.g., the TRAC gene, and / or a B2M-HLA-E and / or B2M HLA-G fusion protein can be introduced into one or more genes encoding a subunit of an HLA-1 protein, e.g., the B2M gene.In certain embodiments, the plurality of cell populations comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or 45 and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 cell populations, e.g., 1-50 cell populations.
[0015] Cells can be engineered using any suitable composition and method. In certain embodiments, cells can be engineered by delivering a composition comprising a site-specific nuclease and / or one or more polynucleotides encoding the site-specific nuclease to the cell. The site-specific nuclease may be any suitable nuclease, such as a homing endonuclease, TALEN, meganuclease, Argonaute, and / or CRISPR / Cas nuclease, i.e., a nucleic acid-guided nuclease. In a preferred embodiment, the site-specific nuclease comprises a nucleic acid-guided nuclease. The site-specific nuclease can hydrolyze the backbone, i.e., generate one or more breaks or strand breaks in the DNA duplex at or near the nuclease's recognition site, i.e., target site. One or more strand breaks in at least one strand of the DNA can be repaired by any suitable innate cellular repair mechanism, such as non-homologous recombination (NHEJ) and / or homology-directed repair (HDR). In certain embodiments, repair of one or more strand breaks in at least one strand of DNA by NHEJ results in one or more genomic modifications, such as insertions and / or deletions (INDELs). In certain embodiments, heterologous DNA, e.g., one or more portions of a donor template, can be introduced into a cell, and at least a portion of the heterologous DNA can be inserted by the cell at or near one or more strand breaks in the DNA by HDR.
[0016] In certain embodiments, the site-specific nuclease comprises a nucleic acid-guided nuclease, e.g., a CRISPR / Cas nuclease. In certain embodiments, the nucleic acid-guided nuclease comprises one or more engineered, non-naturally occurring components. In certain embodiments, the nucleic acid-guided nuclease comprises a Class 1 or Class 2 Cas nuclease, such as Type VA, Type VB, Type VC, Type VD, or Type VE. In certain embodiments, the nucleic acid-guided nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, MAD20, ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART11 * In a preferred embodiment, the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the amino acid sequence of a MAD, ART, or ABW nuclease, such as MAD2, MAD7, ART11, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, and / or ART35 nuclease. * In a more preferred embodiment, the nucleic acid-guided nuclease comprises MAD2, MAD7, ART2, ART11, or ART11 nuclease. *In certain embodiments, the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 37. In even more preferred embodiments, the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, 99%, or 100% identical to the amino acid sequence of SEQ ID NO: 37. In certain embodiments, the nucleic acid-guided nuclease comprises one or more nuclear localization signals (NLSs), e.g., one, four, or five nuclear localization signals, such as one to five NLSs at the carboxy terminus, one to five NLSs at the amino terminus, or a combination thereof. In certain embodiments, the nucleic acid-guided nucleases provided herein comprise one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the one or more NLSs comprise SEQ ID NOs: 40, 51, and 56. Additional nucleases and modifications thereof can be found in the Cas protein section below.
[0017] In certain embodiments, the nucleic acid-guided nuclease further comprises a guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid. In certain embodiments, the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence. In certain embodiments, the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence, and optionally a 5' sequence. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In certain embodiments, the guide nucleic acid comprises an engineered, non-naturally occurring guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the guide nucleic acid section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, when the guide nucleic acid is a dual guide nucleic acid, the stem of the targeter nucleic acid and the stem of the modulator nucleic acid hybridize. In certain embodiments, the dual guide nucleic acid can bind to and activate a nucleic acid-guided nuclease that, in a naturally occurring system, is activated by a single cRNA in the absence of tracrRNA.
[0018] In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below. In certain embodiments, the chemical modification comprises 2'-O-alkyl, 2'-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-deoxy-3'-phosphonoacetate, 2'-deoxy-3'-thiophosphonoacetate, or a combination thereof.
[0019] In certain embodiments, provided herein are guide nucleic acids that include a spacer sequence that is at least partially complementary to a site within: (1) one or more genes encoding a subunit of an HLA-1 protein; (2) one or more genes encoding a subunit of an HLA-2 protein or a transcription factor that regulates expression of one or more subunits of an HLA-2 protein; and / or (3) one or more genes encoding a subunit of a TCR protein.
[0020] In certain embodiments, one or more guide nucleic acids can be complexed with one or more nucleases, e.g., nucleic acid-guided nuclease complexes. In certain embodiments, provided herein are nucleic acid-guided nuclease complexes comprising a nucleic acid-guided nuclease and a compatible guide nucleic acid comprising a spacer sequence at least partially complementary to a site within one or more genes encoding a subunit of an HLA-1 protein, (2) one or more genes encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein, and / or (3) one or more genes encoding a subunit of a TCR protein. In certain embodiments, the one or more guide nucleic acids, the one or more nucleic acid-guided nucleases, and / or the one or more nucleic acid-guided nucleases may further comprise one or more additives that stabilize the nucleic acid-guided nuclease complex. Such cells and / or populations of cells with reduced immunogenicity can be used for a variety of purposes, one such purpose may be as CAR T cells.
[0021] A. Compositions Comprising Cells 1. Cells containing genome modifications In certain embodiments, compositions are provided herein that include cells containing one or more genomic modifications that reduce or eliminate immune responses against the cells in an allogeneic host. The one or more genomic modifications can, for example, partially or completely inactivate a gene encoding the antigen or a portion of the antigen, thereby altering the surface expression of one or more antigens, which affects the immunogenicity of the one or more modified cells. In certain embodiments, cells containing one or more genomic modifications are generated from initial cells, such as primary cells or stem cells, that do not contain genomic modifications that affect immunogenicity. In certain embodiments, the initial, unmodified cells are modified so that all desired genetic modifications are introduced into the cells. In other embodiments, a sequential process is used, for example, cells are modified to introduce some of the desired modifications, and then one or more of their progeny are further modified; this sequential approach can be two, three, four, or more steps. That is, cells containing one or more genomic modifications can be expanded and used as a starting point for the introduction of one or more additional genomic modifications, if desired. In certain embodiments, when the cells comprise stem cells, the stem cells can be differentiated before and / or after the introduction of one or more genomic modifications. Further methods are described below in Methods for Reducing the Immunogenicity of Cells. In certain embodiments, the composition comprising one or more cells comprising one or more genomic modifications further comprises a pharmaceutically acceptable excipient.
[0022] a. Cells containing modifications that result in partial or complete inactivation of genes encoding subunits of HLA-1 In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a subunit of an HLA-1 protein. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding a subunit of an HLA-1 protein. In certain embodiments, the first genomic modification completely inactivates the gene encoding a subunit of an HLA-1 protein. In certain embodiments, the first genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-1 protein. In certain embodiments, the first genomic modification completely eliminates surface expression of active (immunogenic) HLA-1 protein. In certain embodiments, the gene encoding a subunit of an HLA-1 protein comprises a B2M gene. In certain embodiments, the first genomic modification comprises a substitution, insertion, deletion, nonsense mutation, truncation, or a combination thereof. In certain embodiments, the first genomic modification comprises insertion of heterologous DNA, e.g., a transgene, e.g., a transgene comprising a polynucleotide encoding a B2M-fusion protein, such as a B2M-HLA fusion protein, e.g., a B2M-HLA-E fusion protein or a B2M-HLA-G fusion protein.
[0023] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0024] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof.
[0025] In certain embodiments, the cell further comprises a second genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In preferred embodiments, the first transgene is inserted into a gene encoding a subunit of the HLA-1 protein, e.g., the B2M gene. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or portion thereof, e.g., a fusion protein of a CAR or portion thereof. In certain embodiments, the dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124.In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0026] b. Cells containing modifications that result in partial or complete inactivation of genes encoding HLA-1 and HLA-2 subunits In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a subunit of the HLA-1 protein and a second genomic modification in a gene encoding a subunit of the HLA-2 protein or a transcription factor regulating the expression of one or more subunits of the HLA-2 protein. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding the subunit of the HLA-1 protein, and / or the second genomic modification partially or completely inactivates the gene encoding the subunit of the HLA-2 protein or a transcription factor regulating the expression of one or more subunits of the HLA-2 protein. In certain embodiments, the first genomic modification completely inactivates the gene encoding the subunit of the HLA-1 protein, and / or the second genomic modification partially or completely inactivates the gene encoding the transcription factor regulating the expression of the subunit of the HLA-2 protein or one or more subunits of the HLA-2 protein. In certain embodiments, the first and / or second genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-1 and / or HLA-2 proteins. In certain embodiments, the first and / or second genomic modification completely eliminates surface expression of active (immunogenic) HLA-1 and / or HLA-2 proteins. In certain embodiments, the gene encoding a subunit of the HLA-1 protein comprises a B2M gene. In certain embodiments, the gene encoding a transcription factor that regulates expression of one or more subunits of the HLA-2 protein comprises CIITA. In certain embodiments, the first and / or second genomic modification comprises a substitution, insertion, deletion, nonsense mutation, truncation, or a combination thereof. In certain embodiments, the first genomic modification comprises insertion of heterologous DNA, e.g., a transgene, e.g., a transgene comprising a polynucleotide encoding a B2M-fusion protein, such as a B2M-HLA fusion protein, e.g., a B2M-HLA-E fusion protein or a B2M-HLA-G fusion protein.
[0027] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0028] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof. In certain embodiments, the cell further comprises a third genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In preferred embodiments, the first transgene is inserted into a gene encoding a subunit of the HLA-1 protein, e.g., the B2M gene. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or portion thereof, e.g., a fusion protein of a CAR or portion thereof. In certain embodiments, the dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124.In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0029] c. Cells containing modifications that result in partial or complete inactivation of genes encoding HLA-1, HLA-2, and TCR subunits In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a subunit of an HLA-1 protein, a second genomic modification in a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates expression of one or more subunits of an HLA-2 protein, and a third genomic modification in a gene encoding a subunit of a TCR protein, as described above. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding a subunit of an HLA-1 protein, the second genomic modification partially or completely inactivates the gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates expression of one or more subunits of an HLA-2 protein, and / or the third genomic modification partially or completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification completely inactivates a gene encoding a subunit of an HLA-1 protein, the second genomic modification partially or completely inactivates a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein, and / or the third genomic modification partially or completely inactivates a gene encoding a subunit of a TCR protein. In certain embodiments, the first, second, and / or third genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-1, HLA-2, and / or TCR proteins. In certain embodiments, the first, second, and / or third genomic modification completely eliminates surface expression of active (immunogenic) HLA-, HLA-2, and / or TCR proteins. In certain embodiments, the gene encoding a subunit of an HLA-1 protein comprises the B2M gene. In certain embodiments, the gene encoding a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein comprises the CIITA gene. In certain embodiments, the subunits of the TCR protein comprise an alpha or beta subunit.In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In certain embodiments, the subunit of the TCR protein comprises an alpha subunit. In certain embodiments, the gene encoding the subunit of the TCR protein comprises a TRAC gene. In certain embodiments, the first, second, and / or third genomic modification comprises a substitution, insertion, deletion, nonsense mutation, truncation, or a combination thereof. In certain embodiments, the first genomic modification comprises insertion of heterologous DNA, e.g., a transgene, e.g., a transgene comprising a polynucleotide encoding a B2M-fusion protein, such as a B2M-HLA fusion protein, e.g., a B2M-HLA-E fusion protein or a B2M-HLA-G fusion protein. In certain embodiments, the third genomic modification comprises insertion of heterologous DNA, e.g., a transgene, e.g., a polynucleotide encoding a CAR protein or a dual CAR protein.
[0030] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0031] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof.
[0032] In certain embodiments, the cell further comprises a fourth genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In preferred embodiments, the first transgene is inserted into a gene encoding a subunit of the HLA-1 protein, e.g., the B2M gene. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In a preferred embodiment, the transgene comprising a polynucleotide encoding a CAR or a portion thereof is inserted into a gene encoding a subunit of a TCR protein, e.g., the TRAC gene. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, the dual CAR comprises a first CAR or a portion thereof and a second CAR or a portion thereof, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. In certain embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-104 or 116-124.
[0033] d. Cells containing modifications that result in partial or complete inactivation of genes encoding HLA-1 and TCR subunits In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a subunit of an HLA-1 protein and a second genomic modification in a gene encoding a subunit of a TCR protein, as described above. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding a subunit of an HLA-1 protein, and / or the second genomic modification partially or completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification completely inactivates the gene encoding a subunit of an HLA-1 protein, and / or the second genomic modification partially or completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first and / or second genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-1 and / or TCR proteins. In certain embodiments, the first and / or second genomic modification completely eliminates surface expression of active (immunogenic) HLA-1 and / or TCR proteins. In certain embodiments, the gene encoding a subunit of the HLA-1 protein comprises a B2M gene. In certain embodiments, the subunit of the TCR protein comprises an alpha or beta subunit. In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In certain embodiments, the subunit of the TCR protein comprises an alpha subunit. In certain embodiments, the gene encoding a subunit of the TCR protein comprises a TRAC gene. In certain embodiments, the first and / or second genomic modification comprises a substitution, insertion, deletion, nonsense mutation, truncation, or a combination thereof. In certain embodiments, the first genomic modification comprises the insertion of heterologous DNA, e.g., a transgene, e.g., a transgene comprising a polynucleotide encoding a CAR protein or a dual CAR protein.
[0034] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0035] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof. In certain embodiments, the cell further comprises a third genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In preferred embodiments, the first transgene is inserted into a gene encoding a subunit of the HLA-1 protein, e.g., the B2M gene. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In a preferred embodiment, the transgene comprising a polynucleotide encoding a CAR or a portion thereof is inserted into a gene encoding a subunit of a TCR protein, e.g., the TRAC gene. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, the dual CAR comprises a first CAR or a portion thereof and a second CAR or a portion thereof, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. In certain embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-104 or 116-124.
[0036] e. Cells containing modifications that result in partial or complete inactivation of genes encoding subunits of HLA-2. In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein. In certain embodiments, the first genomic modification partially or completely inactivates a gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein. In certain embodiments, the first genomic modification completely inactivates a gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein. In certain embodiments, the first genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-2 protein. In certain embodiments, the first genomic modification completely eliminates surface expression of active (immunogenic) HLA-2 protein. In certain embodiments, the gene encoding a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein comprises the CIITA gene. In certain embodiments, the first genomic modification comprises a substitution, an insertion, a deletion, a nonsense mutation, a truncation, or a combination thereof.
[0037] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0038] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof.
[0039] In certain embodiments, the cell further comprises a second genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or portion thereof, e.g., a fusion protein of a CAR or portion thereof. In certain embodiments, the dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124.In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0040] f. Cells containing modifications that result in partial or complete inactivation of genes encoding HLA-2 and TCR subunits In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein, and a second genomic modification in a gene encoding a subunit of a TCR protein, as described above. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein, and / or the second genomic modification partially or completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification completely inactivates the gene encoding a transcription factor that regulates the expression of a subunit of an HLA-2 protein or one or more subunits of an HLA-2 protein, and / or the second genomic modification completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first and / or second genomic modification reduces or eliminates surface expression of active (immunogenic) HLA-2 and / or TCR proteins. In certain embodiments, the first and / or second genomic modification completely eliminates surface expression of active (immunogenic) HLA-2 and / or TCR proteins. In certain embodiments, the gene encoding a transcription factor that regulates expression of one or more subunits of the HLA-2 protein comprises the CIITA gene. In certain embodiments, the subunit of the TCR protein comprises an alpha or beta subunit. In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In certain embodiments, the subunit of the TCR protein comprises an alpha subunit. In certain embodiments, the gene encoding the subunit of the TCR protein comprises the TRAC gene. In certain embodiments, the first genomic modification comprises a substitution, insertion, deletion, nonsense mutation, truncation, or a combination thereof.In certain embodiments, the second genomic modification comprises the insertion of heterologous DNA, e.g., a transgene, e.g., a polynucleotide encoding a CAR protein or dual CAR proteins.
[0041] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0042] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof.
[0043] In certain embodiments, the cell further comprises a second genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the first transgene is inserted into the TRAC gene. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, the dual CAR comprises a first CAR or a portion thereof and a second CAR or a portion thereof, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. In certain embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124.In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0044] g. Cells containing modifications that result in partial or complete inactivation of genes encoding TCR subunits In certain embodiments, provided herein are compositions comprising cells comprising a first genomic modification in a gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification partially or completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification completely inactivates the gene encoding a subunit of a TCR protein. In certain embodiments, the first genomic modification reduces or eliminates surface expression of an active (immunogenic) TCR protein. In certain embodiments, the first genomic modification completely eliminates surface expression of an active (immunogenic) TCR protein. In certain embodiments, the subunit of the TCR protein comprises an alpha or beta subunit. In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In certain embodiments, the subunit of the TCR protein comprises an alpha subunit. In certain embodiments, the gene encoding a subunit of a TCR protein comprises a TRAC gene. In certain embodiments, the first genomic modification comprises a substitution, an insertion, a deletion, a nonsense mutation, a truncation, or a combination thereof. In certain embodiments, the first genomic modification comprises the insertion of heterologous DNA, e.g., a transgene, e.g., a polynucleotide encoding a CAR protein or a dual CAR protein.
[0045] In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cell is an induced pluripotent stem cell (iPSC).
[0046] In certain embodiments, the cell further comprises one or more nucleic acid-guided nucleases, one or more guide nucleic acids, and / or one or more polynucleotides encoding one or more nucleic acid-guided nucleases and / or guide nucleic acids. In preferred embodiments, the cell comprises a nucleic acid-guided nuclease complexed with a gRNA. In certain embodiments, one or more nucleic acid-guided nucleases (see the Cas nucleases section below) are complexed with one or more guide nucleic acids (see the guide nucleic acids section below). In certain embodiments, the nuclease comprises a type V nuclease. In preferred embodiments, the nuclease comprises a type VA nuclease. In even more preferred embodiments, the nuclease comprises MAD7, e.g., MAD7 comprising one or more nuclear localization signals (NLSs), e.g., one to four NLSs, preferably four NLSs, more preferably one N-terminal NLS and three C-terminal NLSs. In certain embodiments, the cells further comprise a donor template, such as a donor template described herein, e.g., a donor template comprising a polynucleotide encoding one or more CARs or a portion thereof.
[0047] In certain embodiments, the cell further comprises a second genomic modification comprising a first transgene inserted into the genome. The first transgene can be inserted at any suitable location in the genome of the cell. In certain embodiments, the first transgene is inserted into a safe harbor site. The safe harbor site may be any suitable safe harbor site (see the Genomic Safe Harbor section below). In certain embodiments, the safe harbor site comprises the AAVS1 or Rosa 26 locus. In certain embodiments, the safe harbor site comprises any one of SEQ ID NOs: 2020-2043. In certain embodiments, the first transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In a preferred embodiment, the subunit is HLA-E. In a more preferred embodiment, the subunit is HLA-G. Additionally or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In a preferred embodiment, the transgene comprising a polynucleotide encoding a CAR or a portion thereof is inserted into a gene encoding a subunit of a TCR protein, e.g., the TRAC gene. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, the dual CAR comprises a first CAR or a portion thereof and a second CAR or a portion thereof, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. In certain embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In even more preferred embodiments, the CAR or a portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences in SEQ ID NOs: 86-104 or 116-124.
[0048] h. Surface proteins & CAR In certain embodiments, surface expression of cells containing genomic modifications in genes encoding subunits of HLA-1, HLA-2, and / or TCR proteins exhibits 90, 80, 70, 60, 50, 40, 30, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1%, preferably 20% or less, more preferably 10% or less, even more preferably 5% or less, and even more preferably 2% or less of active (immunogenic) protein compared to unmanipulated counterparts. In certain embodiments, endogenous surface-expressed HLA-1 protein can be measured using any suitable technique. In certain embodiments, the techniques include ELISA, proximity ligation assay, pull-down, and / or flow cytometry.
[0049] In certain embodiments, compositions comprising a CAR are provided herein. In certain embodiments, the CAR, or a portion thereof, comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, CD3 zeta, or a combination thereof. In certain embodiments, the CAR, or a portion thereof, comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 86-124. In preferred embodiments, the CAR, or a portion thereof, comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124. In certain embodiments, provided herein are compositions comprising dual CARs comprising a first CAR or portion thereof and a second CAR or portion thereof, separate or connected by one or more polypeptide linkers. In certain embodiments where the dual CARs are separate, the first CAR or portion thereof can be inserted into a first suitable location in the genome, and the second CAR or portion thereof can be inserted into a second suitable location in the genome, and / or a polycistronic gene can be introduced into a suitable location in the genome comprising two or more CARs or portions thereof, where each CAR is expressed on the surface of the cell. In certain embodiments, the dual CARs comprise the same CAR polypeptide sequence. In preferred embodiments, the dual CARs comprise different CAR polypeptide sequences.
[0050] [Table 1A]
[0051] [Table 1B]
[0052] [Table 1C]
[0053] [Table 1D]
[0054] [Table 1E]
[0055] 2. Cell populations containing genome modifications In certain embodiments, provided herein are compositions comprising one or more populations of cells having the genetic modifications described in the Cells Comprising Genomic Modifications section above. In certain embodiments, the composition comprises a single cell population, each cell comprising the same set of genomic modifications (1) through (3). In certain embodiments, provided herein are compositions comprising multiple cell populations, each cell population comprising a different set of genomic modifications. Generally, at least one cell population comprises cells comprising all of: (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of an HLA-1 protein; (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein; and (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding a subunit of a TCR protein, in addition to one or more additional cell populations that do not comprise all three genetic modifications. In certain embodiments, the one or more additional cell populations comprise cells comprising (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-1 protein, (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of an HLA-2 protein or transcription factors that regulate the expression of one or more subunits of an HLA-2 protein, and / or (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of a TCR protein, but not all of (1) through (3). In a preferred embodiment, the subunit of the HLA-1 protein comprises B2M. In a preferred embodiment, the transcription factor that regulates the expression of one or more subunits of an HLA-2 protein comprises CIITA. In certain embodiments, the subunit of the TCR protein is an alpha subunit or a beta subunit.In certain embodiments, the subunit of the TCR protein is a TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z protein. In a preferred embodiment, the gene encoding the subunit of the TCR protein is the TRAC gene. In a more preferred embodiment, at least one cell population comprising cells containing all three genomic modifications comprises: (1) one or more genomic modifications that partially or completely inactivate the B2M gene, (2) one or more genomic modifications that partially or completely inactivate the CIITA gene, and (3) one or more genomic modifications that partially or completely inactivate the TRAC gene. In certain embodiments, the plurality of cell populations comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or 45 and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 populations.
[0056] In certain embodiments, the first cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% of the total cells in the plurality of cell populations, e.g., 1-75%, preferably 5-75%, more preferably 10-75%, even more preferably 15-75%, and even more preferably 20-75% of the total cells in the plurality of cell populations. In certain embodiments, the second cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% of the total cells in the plurality of cell populations, e.g., 1-75%, preferably no more than 50%, more preferably no more than 30%, even more preferably no more than 20%, and even more preferably no more than 10% of the total cells in the plurality of cell populations. In certain embodiments, the third cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% of the total cells in the plurality of cell populations, e.g., 1-75%, preferably no more than 50%, more preferably no more than 30%, even more preferably no more than 20%, and even more preferably no more than 10% of the total cells in the plurality of cell populations. In certain embodiments, the fourth cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% of the total cells in the plurality of cell populations, e.g., 1-75%, preferably no more than 50%, more preferably no more than 30%, even more preferably no more than 20%, and even more preferably no more than 10% of the total cells in the plurality of cell populations.It is understood that the percentages for each of the multiple cell populations sum to 100%.
[0057] The number, relative abundance, and / or identity of a cell population in a plurality of cell populations can be measured by any suitable method. In certain embodiments, the number, relative abundance, and / or identity of a cell population in a plurality of cell populations can be measured by analyzing one or more nucleic acids in a sample using one or more methods, such as PCR, multiplex PCR, FISH, and / or sequencing. In certain embodiments, the number and / or identity of a cell population in a plurality of cell populations can be measured by analyzing one or more cell surface proteins and / or the absence thereof in a sample using one or more methods, such as immunostaining and microscopy, ELISA, pull-down, and / or flow cytometry.
[0058] 3. Complex of guide nucleic acid and nucleic acid-guided nuclease for generating genome modifications In certain embodiments, provided herein are compositions comprising a guide nucleic acid, a nucleic acid-guided nuclease, a nucleic acid-guided nuclease complex, and / or one or more polynucleotides encoding them. In certain embodiments, the nucleic acid-guided nuclease, the guide nucleic acid, and / or the complex thereof further comprise a donor template. In certain embodiments, the nucleic acid-guided nuclease, the guide nucleic acid, and / or the complex thereof further comprise an additive that stabilizes the nucleic acid-guided nuclease complex. In certain embodiments, the nucleic acid-guided nuclease and / or the guide nucleic acid are combined in the presence of an aqueous buffer. In certain embodiments, the nucleic acid-guided nuclease, the guide nucleic acid, and / or the complex thereof further comprise an excipient. In certain embodiments, the nucleic acid-guided nuclease, the guide nucleic acid, and / or the complex thereof are lyophilized, e.g., freeze-dried, along with one or more excipients.
[0059] a. A composition comprising a guide nucleic acid that includes a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-1 protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-1 protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the Guide Nucleic Acids section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0060] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided is a Type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11. * , and more preferably MAD7.
[0061] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0062] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0063] b. A composition comprising a guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-1 protein and / or a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-1 protein, as described above, and a second guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-2 protein or a transcription factor regulating the expression of one or more subunits of the HLA-2 protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the guide nucleic acid section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0064] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7.
[0065] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0066] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0067] c. A composition comprising a guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-1 protein, a gene encoding a subunit of an HLA-2 protein, or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein, and / or a gene encoding a subunit of a TCR protein. In certain embodiments, compositions are provided herein that include a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-1 protein, a second guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-2 protein or a transcription factor regulating the expression of one or more subunits of the HLA-2 protein, and a third guide nucleic acid directed to a target nucleotide sequence in a gene encoding a subunit of a TCR protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the Guide Nucleic Acid section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0068] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7.
[0069] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0070] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0071] d. A composition comprising a guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-1 protein and / or a gene encoding a subunit of a TCR protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of the HLA-1 protein, and a second guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of a TCR protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the guide nucleic acid section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0072] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7. In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0073] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0074] e. A composition comprising a guide nucleic acid that includes a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the Guide Nucleic Acids section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0075] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7.
[0076] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0077] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0078] f. A composition comprising a guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein, and / or a gene encoding a subunit of a TCR protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of an HLA-2 protein or a transcription factor that regulates the expression of one or more subunits of an HLA-2 protein, and a second guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of a TCR protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the Guide Nucleic Acids section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0079] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7.
[0080] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0081] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0082] g. A composition comprising a guide nucleic acid that includes a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of a TCR protein. In certain embodiments, provided herein are compositions comprising a first guide nucleic acid comprising a spacer sequence directed to a target nucleotide sequence in a gene encoding a subunit of a TCR protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, wherein the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, and the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. The spacer sequence may be any suitable sequence. In certain embodiments, the spacer sequence comprises any one of SEQ ID NOs: 125-2019. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In preferred embodiments, the guide nucleic acid comprises a dual guide nucleic acid (described in the Guide Nucleic Acids section below), in which the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at or near the 5' end, the 3' end, and / or both, as described in the gNA Modifications section below.
[0083] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. The guide nucleic acid can be combined and / or complexed with any suitable nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided comprises a type V CRISPR endonuclease, preferably MAD2, MAD7, ART2, ART11, and / or ART11*, more preferably MAD7.
[0084] In certain embodiments, the guide nucleic acid further comprises a nucleic acid-guided nuclease. Any suitable donor template can be combined with the guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a donor template described in the Donor Template section below. In certain embodiments, the donor template comprises a transgene. In preferred embodiments, the transgene comprises a polynucleotide encoding a B2M fusion protein, such as a B2M-HLA-1 subunit fusion protein. In certain embodiments, the HLA-1 subunit comprises HLA-C, HLA-E, or HLA-G, preferably HLA-E or HLA-G. In preferred embodiments, the subunit is HLA-E. In more preferred embodiments, the subunit is HLA-G. Additionally, or alternatively, the cell may comprise a transgene comprising a polynucleotide encoding a CAR or a portion thereof. In certain embodiments, the transgene comprises a polynucleotide encoding a dual CAR or a portion thereof, e.g., a fusion protein of a CAR or a portion thereof. In certain embodiments, a dual CAR comprises a first CAR or portion thereof and a second CAR or portion thereof, wherein the second CAR or portion thereof is different from the first CAR or portion thereof. In certain embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3 zeta. In preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124. In more preferred embodiments, the CAR or portion thereof comprises a polypeptide that binds to at least one of B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3 zeta.In even more preferred embodiments, the CAR or portion thereof comprises a polypeptide that is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104 or 116-124.
[0085] In certain embodiments, the guide nucleic acid, nucleic acid-guided nuclease, and / or donor template may further comprise a cell. The cell may be any suitable cell. In certain embodiments, the cell is a human cell, such as a human stem cell or a human immune cell, such as an immune cell comprising a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte. In a preferred embodiment, the human immune cell is a T cell. In certain embodiments, the T cell comprises a chimeric antigen receptor (CAR) T cell. In certain embodiments, the CAR T cell expresses multiple different CARs, e.g., two different CARs (dual CAR T cells). In certain embodiments, the human cell is a human stem cell, including a human pluripotent stem cell, a pluripotent stem cell, an embryonic stem cell, an induced pluripotent stem cell, a hematopoietic stem cell, or a CD34+ cell. In a preferred embodiment, the cell is a hematopoietic stem cell. In a more preferred embodiment, the cell is a CD34+ stem cell. In an even more preferred embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0086] [Table 2-1]
[0087] [Table 2-2]
[0088] [Table 2-3]
[0089] [Table 2-4]
[0090] Table 2-5
[0091] Table 2-6
[0092] Table 2-7
[0093] Table 2-8
[0094] Table 2-9
[0095] Table 2-10
[0096] Table 2-11
[0097] Table 2-12
[0098] Table 2-13
[0099] Table 2-14
[0100] Table 2-15
[0101] Table 2-16
[0102] Table 2-17
[0103] Table 2-18
[0104] Table 2-19
[0105] Table 2-20
[0106] Table 2-21
[0107] Table 2-22
[0108] Table 2-23
[0109] Table 2-24
[0110] Table 2-25
[0111] Table 2-26
[0112] Table 2-27
[0113] Table 2-28
[0114] Table 2-29
[0115] Table 2-30
[0116] Table 2-31
[0117] Table 2-32
[0118] Table 2-33
[0119] Table 2-34
[0120] Table 2-35
[0121] Table 2-36
[0122] Table 2-37
[0123] Table 2-38
[0124] Table 2-39
[0125] Table 2-40
[0126] Table 2-41
[0127] Table 2-42
[0128] Table 2-43
[0129] Table 2-44
[0130] Table 2-45
[0131] Table 2-46
[0132] Table 2-47
[0133] Table 2-48
[0134] Table 2-49
[0135] Table 2-50
[0136] Table 2-51
[0137] Table 2-52
[0138] Table 2-53
[0139] Table 2-54
[0140] Table 2-55
[0141] Table 2-56
[0142] Table 2-57
[0143] Table 2-58
[0144] Table 2-59
[0145] Table 2-60
[0146] Table 2-61
[0147] Table 2-62
[0148] Table 2-63
[0149] Table 2-64
[0150] Table 2-65
[0151] [Table 2-66]
[0152] [Table 2-67]
[0153] [Table 2-68]
[0154] B. Methods for reducing the immunogenicity of cells In certain embodiments, methods are provided herein. In certain embodiments, methods are provided herein for engineering cells, such as human cells. In certain embodiments, methods are provided herein for engineering cells to reduce the immunogenicity of the engineered cells. In certain embodiments, methods are provided herein for engineering cells to be introduced into a recipient that is allogeneic to the individual that was the source of the engineered cells (also referred to herein as "allogeneic cells") to reduce the immunogenicity of the engineered allogeneic cells.
[0155] In certain embodiments, methods are provided herein for generating one or more modifications in the genome of a target cell. In certain embodiments, the methods can simultaneously or sequentially generate at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and / or up to 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 100 genomic modifications, for example, 1 to 100 genomic modifications, preferably 1 to 20 genomic modifications (see the multiplexing section below). In certain embodiments, the first genomic modification is introduced into one or more target cells, where the target cells include wild-type cells or cells containing one or more genomic modifications (see the cells containing genomic modifications section above). In certain embodiments, the target cells comprise one or more modified cells described in the section on cells containing genomic modifications (above). In certain embodiments, the method involves generating one or more genomic modifications in one or more target cells, where the one or more genomic modifications are generated simultaneously in a single cell, e.g., by the introduction of all necessary components to produce the desired genomic modifications. In certain embodiments, the method involves generating one or more genomic modifications in one or more target cells, where the one or more genomic modifications are generated sequentially, e.g., when some of the desired genetic modifications are generated in a parent cell and the remaining desired genetic modifications are generated in one or more generations of progeny derived from the parent cell. In certain embodiments in which one or more genomic modifications are introduced sequentially, the one or more genomic modifications can be introduced in any suitable amount, order, and / or combination.For example, when three genomic modifications (A, B, and C) are introduced into one or more cells, the three genomic modifications can be introduced in any one of the following orders: (1) A, then B, then C; (2) A, then C, then B; (3) A and B, then C; (4) A, then B and C; (5) A and C, then B; (6) A, then C and B; (7) B, then A, then C; (8) B, then C, then A; (9) B and A, then C; (10) B, then A and C; (11) B and C, then A; (12) B, then C and A; (13) C, then A, then B; (14) C, then B, then C; (15) C and A, then B; (16) C, then A and B; (17) C, then B and A; (18) C and B, then A; or (19) A and B and C.
[0156] In certain embodiments, provided herein are methods for engineering one or more human cells. Any suitable human cell or cells can be used. In certain embodiments, the cells comprise one or more human stem cells or human immune cells. In certain embodiments, the cells comprise one or more human cells, including immune cells, including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, lymphocytes, or a combination thereof. In certain embodiments, the cells comprise one or more T cells. In certain embodiments, the cells comprise one or more chimeric antigen receptor (CAR)-T cells. In certain embodiments, the CAR T cells comprise a CAR polypeptide or a portion thereof. In certain embodiments, the CAR T cells comprise two or more CAR polypeptides or portions thereof. In certain embodiments, the CAR T cells comprise a dual CAR, wherein the dual CAR comprises a first CAR polypeptide or a portion thereof and a second CAR polypeptide or a portion thereof, wherein the second CAR polypeptide is different from the first CAR polypeptide, and the first and second CAR polypeptides are separate. In certain embodiments, the first and second CAR polypeptides are linked by a polypeptide linker. In certain embodiments, the cells comprise one or more human stem cells, including human pluripotent, pluripotent stem cells, embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells, CD34+ cells, or combinations thereof. In preferred embodiments, the cells comprise one or more hematopoietic stem cells. In more preferred embodiments, the cells comprise one or more CD34+ stem cells. In even more preferred embodiments, the cells comprise one or more induced pluripotent stem cells (iPSCs). In certain embodiments, the cells comprise allogeneic cells.
[0157] In certain embodiments, one or more cells containing one or more introduced genomic modifications are proliferated, e.g., expanded, or differentiated, e.g., iPSCs are differentiated into T cells. In certain embodiments in which two or more genomic modifications are introduced sequentially, one or more target cells are expanded after the introduction of a first set of genomic modifications, and a second set of genomic modifications is introduced into the progeny of the first set of cells. In certain embodiments, stem cells are differentiated before or after the introduction of one or more genomic modifications. In certain embodiments, stem cells are differentiated after the introduction of one or more genomic modifications.
[0158] In certain embodiments, one or more genomic modifications are introduced into a population of cells, where the resulting cell population comprises multiple cell populations each having a different set of genomic modifications (see the cell populations section above). For example, if three genomic modifications (A, B, C) are introduced sequentially and / or simultaneously into a population of cells, the resulting multiple cell populations may potentially comprise any number and / or combination of the following cell populations: (1) A, (2) AB, (3) AC, (4) ABC, (5) B, (6) BC, (7) C, and / or (8) no genomic modifications. In certain embodiments, each cell population in the multiple cell populations may be present in any percentage relative to the other cell populations, and the relative percentage of each population is influenced by several factors, including, but not limited to, the delivery efficiency of the editing components, the quality of the editing components, the concentration of the editing components, the relative efficiency and specificity of the editing events, the vitality of the cells, and / or the viability of the cells before or after the introduction of one or more genomic modifications.
[0159] In certain embodiments, provided herein are methods for manipulating cells, comprising delivering one or more site-specific nucleases to one or more target cells. In certain embodiments, the one or more site-specific nucleases are delivered to the target cells as polypeptides. In certain embodiments, the one or more site-specific nucleases are combined with a guide nucleic acid compatible with a nucleic acid-guided nuclease system, such as a CRISPR / cas system. In certain embodiments, one or more polynucleotides encoding one or more components of the nuclease system are delivered to the target cells. In preferred embodiments, the nucleic acid-guided nuclease system is a type V nuclease, more preferably a type VA nuclease, even more preferably MAD2, MAD7, ART2, ART11, ART11, or the like. * nuclease, even more preferably MAD7 nuclease.
[0160] In certain embodiments, one or more guide nucleic acids comprising a spacer sequence at least partially complementary to the target nucleotide sequence within the site where one or more genome modifications are to be introduced are delivered to the target cell. In certain embodiments, one or more nucleic acid-guided nucleases are delivered to the target cell. In certain embodiments, a combination of one or more guide nucleic acids and nucleic acid-guided nucleases is delivered to the target cell, where the one or more nucleic acid-guided nucleases are optionally complexed with the guide nucleic acid (see, e.g., the ribonucleoprotein (RNP) section below). In certain embodiments, one or more fully formed nucleic acid-guided nuclease complexes, e.g., RNPs, are delivered. In certain cases, any one of the embodiments described in the guide nucleic acid and donor template sections can be delivered to the target cell.
[0161]
[0013] In certain embodiments, provided herein are methods for producing non-immunogenic cells. In certain embodiments, provided herein are methods for producing non-immunogenic stem cells or immune cells. In certain embodiments, provided herein are methods for producing non-immunogenic CAR T cells. In certain embodiments, provided herein are methods for producing non-immunogenic CAR T cells, the methods comprising: (1) modifying the genome of a cell to reduce or eliminate cell surface expression of an active HLA-I protein in the cell and its progeny; (2) introducing into the genome of the cell or one or more of its progeny a first polynucleotide encoding surface expression of a first CAR, or a portion thereof, specific for a first antigen; and (3) introducing into the genome of the cell or one or more of its progeny a second polynucleotide encoding surface expression of a second CAR, or a portion thereof, specific for a second antigen. In certain embodiments, the method further comprises modifying the genome of the cell to reduce or eliminate cell surface expression of active HLA-1 protein, comprising introducing a genomic modification into the B2M gene that partially or completely inactivates the B2M gene. In certain embodiments, the B2M gene is completely inactivated. In certain embodiments in which the B2M gene is partially or completely inactivated, a first transgene encoding a B2M-HLA-1 subunit fusion protein is introduced. In certain embodiments, the B2M-HLA-1 subunit fusion protein comprising an HLA-1 subunit comprises HLA-C, -E, or -G. In preferred embodiments, the HLA-1 subunit comprises HLA-E or -G. In certain embodiments, the first and / or second CAR or portion thereof comprises any one of the CARs described in the Surface Proteins & CARs section above. In certain embodiments, the method further comprises modifying the genome of the cell or one of its progeny to reduce or eliminate surface expression of one or more subunits of the HLA-2 protein.In certain embodiments, one or more subunits of the HLA-2 protein are modified by introducing a genomic modification into a gene encoding a transcription factor for one or more genes encoding one or more subunits of the HLA-2 protein. In certain embodiments, the genomic modification in the transcription factor that regulates the expression of one or more subunits of the HLA-2 protein at least partially or completely inactivates the transcription factor. In certain embodiments, the transcription factor is completely inactivated. In a preferred embodiment, the transcription factor comprises CIITA. In certain embodiments, the method further comprises delivering to a cell one or more polynucleotides encoding a nucleic acid-guided nuclease system, or one or more portions of the system, comprising a nucleic acid-guided nuclease and a guide nucleic acid compatible with, capable of binding to, and activating the nucleic acid-guided nuclease, wherein the guide nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence, the spacer sequence being complementary to a target nucleotide sequence within a target polynucleotide in the genome of a human target cell, and a modulator nucleic acid comprising a modulator stem sequence complementary to the target stem sequence and, optionally, a 5' sequence, such that the nucleic acid-guided nuclease system targets and cleaves at least one strand of the target polynucleotide at or near the target nucleotide sequence. In certain embodiments, the nuclease comprises any suitable nuclease. In certain embodiments, the nuclease comprises any suitable nuclease described in the Cas protein section (below). In certain embodiments, the nuclease is a type V nuclease, preferably a type VA nuclease, ART2, ART11, ART11. *, MAD2, and / or MAD7 nuclease, even more preferably MAD7 nuclease. In certain embodiments, the nucleic acid-guided nuclease system comprises a guide nucleic acid comprising a single polynucleotide and / or a guide nucleic acid comprising one or more polynucleotides, e.g., a dual guide nucleic acid, preferably a dual guide nucleic acid capable of binding to and activating a nucleic acid-guided nuclease that, in a naturally occurring system, is activated by a single crRNA in the absence of tracrRNA. In certain embodiments, the guide nucleic acid comprises one or more chemical modifications described in the gNA Modifications section (below). In certain embodiments, the method further comprises delivering one or more donor templates described in the Donor Templates section below. In certain embodiments, at least a portion of the donor template is inserted via an innate cellular repair mechanism initiated by the generation of one or more strand breaks at or near the target nucleotide sequence by one or more nucleic acid-guided nucleases. In certain embodiments, delivery of one or more components for genome engineering is by electroporation.
[0162] In certain embodiments, provided herein is a method for producing a population of non-immunogenic CAR T cells, comprising: (1) modifying the genome of a first cell to reduce or eliminate cell surface expression of an HLA-1 protein in the first cell and its progeny, (2) introducing into the genome of the first cell a first polynucleotide encoding surface expression of a first CAR specific for a first antigen on the first cell, (3) modifying the genome of a second cell to reduce or eliminate cell surface expression of an HLA-1 protein in the second cell and its progeny, and (4) introducing into the genome of the second cell a second polynucleotide encoding surface expression of a second CAR specific for a second antigen on the second cell, wherein the first and second cells are the same cell, and the first cell is a descendant of the second cell, or the second cell is a descendant of the first cell. Steps (1) through (4) may be performed in any suitable order.
[0163] In certain embodiments, provided herein are methods for producing a population of non-immunogenic CAR T cells, comprising: (1) modifying the genome of a first cell to reduce or eliminate cell surface expression of an HLA-1 protein in the first cell and its progeny, (2) introducing into the genome of the first cell a first polynucleotide encoding surface expression of a first CAR specific for a first antigen on the first cell, (3) modifying the genome of a second cell to reduce or eliminate cell surface expression of an HLA-1 protein in the second cell and its progeny, and (4) introducing into the genome of the second cell a second polynucleotide encoding surface expression of a second CAR specific for a second antigen on the second cell, wherein the first and second cells are the same cell, and the first cell is a descendant of the second cell, or the second cell is a descendant of the first cell. In certain embodiments, when the first, second, third, and fourth cells are the same cell, steps (1)-(4) are performed simultaneously. In certain embodiments, one or more of steps (1)-(4) are performed sequentially, e.g., using any one of the following sequential permutations: ABCD, ABDC, ACBD, ACDB, ADBC, ADCB, BACD, BADC, BCAD, BCDA, BDAC, BDCA, CABD, CADB, CBAD, CBDA, CDAB, CDBA, DABC, DACB, DBAC, DBCA, DCAB, DCBA. In certain embodiments, when at least one step is performed sequentially, e.g., A, then BCD, or A and B, then C and D, one or more steps can be performed simultaneously.
[0164] In certain embodiments, provided herein is a method for modifying the genome of a human cell, comprising the steps of: (1) modifying a B2M gene in the genome to reduce or eliminate expression of the B2M gene; (2) modifying a T cell receptor (TCR) subunit gene in the genome to reduce or eliminate expression of the subunit; and (3) modifying a CIITA gene in the genome to reduce or eliminate expression of the CIITA gene, wherein at least two of (a) to (c) are performed sequentially rather than simultaneously, thereby producing a modified human cell.
[0165] II. Engineered, non-naturally occurring dual-guide CRISPR-cas systems A CRISPR-Cas system generally comprises a Cas protein and one or more guide nucleic acids (gNAs). The Cas protein can be directed to a specific location in a double-stranded DNA target by recognizing a protospacer adjacent motif (PAM) in the non-target strand of the DNA, and one or more guide nucleic acids can be directed to a specific location by hybridizing to a target nucleotide sequence, also referred to herein as a target sequence, in the target strand of a target polynucleotide. Typically, both PAM recognition and hybridization to the target nucleotide sequence are required for stable binding of the CRISPR-Cas complex to the DNA target and activation of the effector function (e.g., nuclease activity) if the Cas protein has an effector function. Consequently, when creating a CRISPR-Cas system, the guide nucleic acid can be designed to include a nucleotide sequence, referred to as a spacer sequence, that is at least partially complementary to and can hybridize with the target nucleotide sequence, located adjacent to the PAM in an orientation that allows the target nucleotide sequence to function with the Cas protein. It has been observed that not all CRISPR-Cas systems designed according to these criteria are equally effective.The larger polynucleotide in which the target nucleotide sequence is located; for example, chromosome or other genomic DNA, or a part thereof, or any other suitable polynucleotide in which the target nucleotide sequence is located can be referred to as target polynucleotide.The target polynucleotide in double-stranded DNA comprises two strands.The strand of the DNA duplex with which the spacer sequence is complementary is referred to herein as "target strand", while the strand with which the spacer sequence shares sequence identity is referred to herein as "non-target strand".
[0166] Two distinct classes of CRISPR-Cas systems have been identified. Class 1 CRISPR-Cas systems utilize multiprotein effector complexes, while class 2 CRISPR-Cas systems utilize single-protein effectors (see Makarova et al. (2017) Cell, 168: 328). Among the types of class 2 CRISPR-Cas systems, type II and type V systems typically target DNA, while type VI systems typically target RNA (ibid.). Naturally occurring type II effector complexes include Cas9, CRISPR RNA (crRNA), and trans-activating CRISPR RNA (tracrRNA), although the crRNA and tracrRNA can be fused as a single guide RNA in engineered systems for convenience (see Wang et al. (2016) Annu. Rev. Biochem., 85: 227). Certain naturally occurring type V systems, such as type VA, type VC, and type VD systems, do not require tracrRNA and use only crRNA as a guide for cleaving target DNA (see Zetsche et al. (2015) Cell 163: 759; Makarova et al. (2017) Cell 168: 328).
[0167] Naturally occurring type II CRISPR-Cas systems (e.g., CRISPR-Cas9 systems) generally contain two guide nucleic acids, called crRNA and tracrRNA, which form a complex through nucleotide hybridization. A single guide nucleic acid capable of activating type II Cas nucleases has been developed, for example, by linking crRNA and tracrRNA (see, e.g., U.S. Patent Nos. 10,266,850 and 8,906,616). Naturally occurring type II Cas proteins contain a RuvC-like nuclease domain and an HNH endonuclease domain, and recognize a 3' G-rich PAM located immediately downstream from the target nucleotide sequence, with its orientation determined using the non-target strand (i.e., the strand not hybridized with the spacer sequence) as a reference. CRISPR-Cas systems cleave double-stranded DNA to generate blunt ends. The cleavage site is generally 3-4 nucleotides upstream from the PAM on the non-target strand.
[0168] Naturally occurring VA-, VC-, and VD-type CRISPR-Cas systems lack a tracrRNA and rely on a single crRNA to guide the CRISPR-Cas complex to a target polynucleotide. Dual guide nucleic acids capable of activating VA-, VC-, or VD-type Cas nucleases have been developed, for example, by splitting a single crRNA into a targeter nucleic acid and a modulator nucleic acid (see, e.g., International (PCT) Application Publication No. WO 2021 / 067788). Naturally occurring VA-type Cas proteins contain a RuvC-like nuclease domain but lack an HNH endonuclease domain. They recognize a 5' T-rich PAM located immediately upstream from the target nucleotide sequence, and their orientation is determined using the non-target strand (i.e., the strand that did not hybridize with the spacer sequence) as a coordinate. These CRISPR-Cas systems cleave double-stranded DNA to generate staggered double-strand breaks rather than blunt ends. The cleavage site is distant from the PAM site (e.g., at least 10, 11, 12, 13, 14, or 15 nucleotides downstream from the PAM on the non-target strand and / or at least 15, 16, 17, 18, or 19 nucleotides upstream from the sequence complementary to the PAM on the target strand).
[0169] Elements in an exemplary single-guide CRISPR-Cas system, a VA-type CRISPR-Cas system, are shown in Figure 1A. A single gRNA, when present in the form of RNA, can also be referred to as a "crRNA" or "single gRNA." From 5' to 3', it can include an optional 5' sequence, e.g., a tail, a modulator stem sequence, a loop, a targeter stem sequence complementary to the modulator stem sequence, and a spacer sequence at least partially complementary to and capable of hybridizing with a target sequence in the target strand of a target polynucleotide. When a 5' tail is present, the sequence including the 5' tail and the modulator stem sequence can also be referred to herein as a "modulator sequence." The segment of the single-guide nucleic acid from the optional 5' tail to the targeter stem sequence, also referred to herein as a "scaffold sequence," binds to the Cas protein. Additionally, a PAM in the non-target strand of the target DNA binds to the Cas protein.
[0170] Elements of an exemplary dual-guide CRISPR-Cas system, e.g., a dual-guide VA-type CRISPR-Cas system, are shown in Figure 1B. The first guide nucleic acid, which may be referred to herein as the "modulator nucleic acid," comprises, from 5' to 3', an optional 5' tail and a modulator stem sequence. When a 5' tail is present, the sequence comprising the 5' tail and the modulator stem sequence may also be referred to herein as the "modulator sequence." The second guide nucleic acid, which may be referred to herein as the "targeter nucleic acid," comprises, from 5' to 3', a targeter stem sequence complementary to the modulator stem sequence and a spacer sequence at least partially complementary to and capable of hybridizing with a target sequence in the target strand of the target polynucleotide. The duplex of the modulator stem sequence and the targeter stem sequence, and the optional 5' tail, constitute a structure that binds to the Cas protein. Additionally, a PAM in the non-target strand of the target DNA binds to the Cas protein. It is understood that in a dual gNA, e.g., a dual gRNA, the targeter nucleic acid and the modulator nucleic acid are not in the same nucleic acid, i.e., are not joined end-to-end by a traditional internucleotide bond, but are covalently conjugated to each other by one or more chemical modifications introduced into these nucleic acids, which can increase the stability of the double-stranded complex and / or improve other characteristics of the system.
[0171] As used herein, the terms "targeter stem sequence" and "modulator stem sequence" may refer to a pair of nucleotide sequences in one or more guide nucleic acids that hybridize to each other. When the targeter stem sequence and the modulator stem sequence are contained in a single guide nucleic acid, the targeter stem sequence is proximal to a spacer sequence designed to hybridize with the target nucleotide sequence, and the modulator stem sequence is proximal to the targeter stem sequence. When the targeter stem sequence and the modulator stem sequence are in separate nucleic acids, the targeter stem sequence is in the same nucleic acid as the spacer sequence designed to hybridize with the target nucleotide sequence. In CRISPR-Cas systems that naturally contain separate crRNAs and tracrRNAs (e.g., Type II systems), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the duplex formed between the crRNA and tracrRNA. In CRISPR-Cas systems that naturally contain a single crRNA but do not contain a tracrRNA (e.g., VA-type systems), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the stem portion of the stem-loop structure in the scaffold sequence of the crRNA. It is understood that 100% complementarity between the targeter stem sequence and the modulator stem sequence is not required. However, in VA-type CRISPR-Cas systems, the targeter stem sequence is typically 100% complementary to the modulator stem sequence.
[0172] A. Cas proteins A guide nucleic acid, either as a single guide nucleic acid alone (the targeter and modulator nucleic acids are part of a single polynucleotide) or as a dual gNA comprising a separate targeter nucleic acid used with a cognate modulator nucleic acid, can bind to a CRISPR-associated (Cas) protein, e.g., a Cas nuclease. In certain embodiments, a guide nucleic acid, either as a single guide nucleic acid alone (the targeter and modulator nucleic acids are part of a single polynucleotide) or as a dual gNA comprising a separate targeter nucleic acid used with a cognate modulator nucleic acid, can activate a Cas nuclease. A gNA that can activate a particular Cas nuclease is said to be "compatible" with that Cas nuclease; a Cas nuclease that can be activated by a particular gNA is said to be "compatible" with that gNA.
[0173] The terms "CRISPR-associated protein," "Cas protein," and "Cas," used interchangeably herein, may refer to naturally occurring or engineered Cas proteins. Non-limiting examples of Cas protein engineering include, but are not limited to, mutations and modifications of Cas proteins that alter Cas activity, alter PAM specificity, expand the range of PAMs recognized, and / or decrease the ability to modify one or more off-target loci compared to the corresponding unmodified Cas. In certain embodiments, altered activity of an engineered Cas includes an altered ability (e.g., specificity or kinetics) to bind to a naturally occurring gNA, e.g., a gRNA, or an engineered gNA, e.g., a gRNA; an altered ability (e.g., specificity or kinetics) to bind to a target nucleotide sequence; an altered processivity of nucleic acid scanning; and / or an altered effector (e.g., nuclease) activity. A Cas protein with nuclease activity may be referred to as a "CRISPR-associated nuclease" or a "Cas nuclease" or simply a "nuclease," as used interchangeably herein.
[0174] In certain embodiments, the Cas protein is a type VA, type VC, or type VD Cas protein. In certain embodiments, the Cas protein is a type VA Cas protein. In other embodiments, the Cas protein is a type II Cas protein, such as a Cas9 protein.
[0175] In certain embodiments, the VA-type Cas nuclease comprises Cpf1. Cpf1 proteins are known in the art and are described, for example, in U.S. Patent Nos. 9,790,490 and 10,113,179. Cpf1 orthologs can be found in a variety of bacterial and archaeal genomes.For example, in certain embodiments, the Cpf1 protein is isolated from Francisella novicida U112 (Fn), Acidaminococcus sp. BV3L6 (As), Lachnospiraceae bacterium ND2006 (Lb), Lachnospiraceae bacterium MA2020 (Lb2), Candidatus Methanoplasma termitum (CMt), Moraxella bovoculi 237 (Mb), Porphyromonas crevioricanis (Pc), Prevotella disiens (Pd), Francisella tularensis, 1, Francisella tularensis subsp. Novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Eubacterium eligens, Leptospira inadai, Porphyromonas macacae, Prevotella brianti bryantii, Proteocatella sphenisci, Anaerovibrio sp. RM50, Moraxella caprae, Lachnospiraceae bacterium COE1, or Eubacterium coprostanoligenes.
[0176] In certain embodiments, the VA-type Cas nuclease comprises AsCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:3 of International (PCT) Application Publication WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO:3 of International (PCT) Application Publication WO 2021 / 158918.
[0177] In certain embodiments, the VA-type Cas nuclease comprises LbCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application Publication WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application Publication WO2021 / 158918.
[0178] In certain embodiments, the VA-type Cas nuclease comprises FnCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application Publication WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application Publication WO2021 / 158918.
[0179] In certain embodiments, the VA-type Cas nuclease comprises Prevotella briantii Cpf1 (PbCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:6 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO:6 of International (PCT) Application Publication No. WO 2021 / 158918.
[0180] In certain embodiments, the VA-type Cas nuclease comprises Proteocatella sphenisci Cpf1 (PsCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application Publication No. WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application Publication No. WO2021 / 158918.
[0181] In certain embodiments, the VA-type Cas nuclease comprises Cpf1 (As2Cpf1) of Anaerovibrio species RM50 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application Publication No. WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application Publication No. WO2021 / 158918.
[0182] In certain embodiments, the VA-type Cas nuclease comprises Moraxella caprae Cpf1 (McCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application Publication No. WO 2021 / 158918.
[0183] In certain embodiments, the VA-type Cas nuclease comprises Cpf1 (Lb3Cpf1) of Lachnospiraceae COE1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application Publication No. WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application Publication No. WO2021 / 158918.
[0184] In certain embodiments, the VA-type Cas nuclease comprises Eubacterium coprostanoligenes Cpf1 (EcCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application Publication No. WO2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application Publication No. WO2021 / 158918.
[0185] In certain embodiments, the VA-type Cas nuclease is not Cpf1. In certain embodiments, the VA-type Cas nuclease is not AsCpf1.
[0186] In certain embodiments, the VA-type Cas nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20, or a variant thereof, which are known in the art and described in U.S. Patent No. 9,982,279.
[0187] In certain embodiments, the VA-type Cas nuclease comprises MAD7 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 37. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 37.
[0188] MAD7 (SEQ ID NO: 37)
[0189] [ka]
[0190] In certain embodiments, the VA-type Cas nuclease comprises MAD2 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 38. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 38.
[0191] MAD2 (SEQ ID NO: 38)
[0192] [ka]
[0193] In certain embodiments, the VA-type Cas nuclease comprises Csm1. Csm1 proteins are known in the art and are described in U.S. Patent No. 9,896,696. Csm1 orthologs can be found in various bacterial and archaeal genomes. For example, in certain embodiments, the Csm1 protein is derived from Smithella sp. SCADC (Sm), Sulfuricurvum sp. (Ss), or Microgenomates (Roizmanbacteria) bacteria (Mb).
[0194] In certain embodiments, the VA-type Cas nuclease comprises SmCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 12 of International (PCT) Application Publication WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 12 of International (PCT) Application Publication WO 2021 / 158918.
[0195] In certain embodiments, the VA-type Cas nuclease comprises SsCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 13 of International (PCT) Application Publication WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 13 of International (PCT) Application Publication WO 2021 / 158918.
[0196] In certain embodiments, the VA-type Cas nuclease comprises MbCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 14 of International (PCT) Application Publication WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 14 of International (PCT) Application Publication WO 2021 / 158918.
[0197] In certain embodiments, the VA-type Cas nuclease comprises an ART nuclease or a variant thereof. Generally, such nuclease sequences have less than 60% AA sequence similarity with Cas12a, less than 60% AA sequence similarity with a positive control nuclease, and greater than 80% query coverage. In certain embodiments, the VA-type nuclease is selected from the group consisting of ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART28, ART30, ART31, ART32, ART33, ART34, ART35, and ART11, as shown in Table 3. *(i.e., ART11_L679F, i.e., ART11 in which the leucine (L) at amino acid position 679 is replaced with a phenylalanine (F)) nuclease. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence designated for an individual ART nuclease shown in Table 3. In certain embodiments, provided herein are nucleic acid-guided nucleases comprising a nucleic acid-guided nuclease polypeptide having at least 85% identity to the amino acid sequence represented by SEQ ID NOs: 1-36, or nucleic acids encoding a nucleic acid-guided nuclease polypeptide comprising at least 85% identity to a polynucleotide represented by SEQ ID NOs: 1-36. In certain embodiments, nucleic acid-guided nucleases are provided that comprise a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NOs: 1-36, wherein the polypeptide does not contain the peptide motif of YLFQIYNKDF (SEQ ID NO: 39). In certain embodiments, nucleic acid-guided nucleases are provided that comprise a nucleic acid encoding a polypeptide having at least 90% identity to a nucleic acid represented by SEQ ID NOs: 808-845, wherein the encoded polypeptide does not contain the peptide motif of YLFQIYNKDF (SEQ ID NO: 39). In certain embodiments, nucleic acid-guided nucleases are provided that comprise a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NOs: 1-9. In certain embodiments, nucleic acid-guided nucleases are provided that comprise a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NO: 2, 11, or 36.
[0198] [Table 3-1]
[0199] [Table 3-2]
[0200] Table 3-3
[0201] Table 3-4
[0202] Table 3-5
[0203] Table 3-6
[0204] Table 3-7
[0205] Table 3-8
[0206] Table 3-9
[0207] Table 3-10
[0208] Table 3-11
[0209] Table 3-12
[0210] Table 3-13
[0211] Table 3-14
[0212] Table 3-15
[0213] Table 3-16
[0214] Table 3-17
[0215] Table 3-18
[0216] Table 3-19
[0217] Table 3-20
[0218] Table 3-21
[0219] Table 3-22
[0220] Table 3-23
[0221] Table 3-24
[0222] Table 3-25
[0223] Table 3-26
[0224] Table 3-27
[0225] Table 3-28
[0226] Table 3-29
[0227] Table 3-30
[0228] Table 3-31
[0229] Table 3-32
[0230] Table 3-33
[0231] Table 3-34
[0232] Table 3-35
[0233] Table 3-36
[0234] In certain embodiments, the Cas nuclease is selected from the group consisting of ABW1 (SEQ ID NO: 3), ABW2 (SEQ ID NO: 16), ABW3 (SEQ ID NO: 29), ABW4 (SEQ ID NO: 42), ABW5 (SEQ ID NO: 55), ABW6 (SEQ ID NO: 68), ABW7 (SEQ ID NO: 81), ABW8 (SEQ ID NO: 94), and ABW9 (SEQ ID NO: 107) (all SEQ ID NOs for ABW1-9 and variants thereof from International (PCT) Application Publication No. WO 2021 / 108324), or any one of variants 1-10 of ABW1 (SEQ ID NOs: 4-13, respectively), any one of variants 1-10 of ABW2 (SEQ ID NOs: 17-26, respectively), any one of variants 1-10 of ABW3 (SEQ ID NOs: 30-39, respectively), any one of variants 1-10 of ABW4 (SEQ ID NOs: 43-52, respectively), and variants 1-10 of ABW5. (SEQ ID NOS: 56-65, respectively), any one of variants 1-10 of ABW6 (SEQ ID NOS: 69-78, respectively), any one of variants 1-10 of ABW7 (SEQ ID NOS: 82-91, respectively), any one of variants 1-10 of ABW8 (SEQ ID NOS: 95-104, respectively), any one of variants 1-10 of ABW9 (SEQ ID NOS: 108-117, respectively), and their variants. ABW1 to ABW9 and their variants are known in the art and are described in International (PCT) Application Publication No. WO2021 / 108324.
[0235] More VA-type Cas nucleases and their corresponding naturally occurring CRISPR-Cas systems can be identified by computational and experimental methods known in the art, for example, as described in U.S. Patent No. 9,790,490 and Shmakov et al. (2015) Mol. Cell, 60: 385. Exemplary computational methods include homology modeling, structural BLAST, PSI-BLAST, or HHPred analysis of predicted Cas proteins, and analysis of predicted CRISPR loci by identifying CRISPR arrays. Exemplary experimental methods include in vitro cleavage assays and intracellular nuclease assays (e.g., Surveyor assays) as described in Zetsche et al. (2015) Cell, 163: 759.
[0236] In certain embodiments, the Cas protein is a Cas nuclease that directs cleavage of one or both strands of the target locus, such as the target strand (i.e., the strand having a target nucleotide sequence that is at least partially complementary to and capable of hybridizing to a single guide nucleic acid or dual guide nucleic acid) and / or a non-target strand. In certain embodiments, the Cas nuclease directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more nucleotides of the first or last nucleotide of the target nucleotide sequence or its complement. In certain embodiments, the cleavage is staggered, i.e., generates sticky ends. In certain embodiments, the cleavage generates staggered cleavages with 5' overhangs. In certain embodiments, the cleavage generates staggered cleavages with 5' overhangs of 1 to 5 nucleotides, e.g., 4 or 5 nucleotides. In certain embodiments, the cleavage site is distant from the PAM, for example, cleavage occurs after the 18th nucleotide on the non-target strand and after the 23rd nucleotide on the target strand.
[0237] In certain embodiments, the compositions provided herein comprise a compatible guide nucleic acid (gNA), e.g., a Cas nuclease that can be activated by a gRNA. In certain embodiments, the compositions provided herein further comprise a Cas protein associated with a Cas nuclease that can be activated by a compatible guide nucleic acid (gNA), e.g., a gRNA. For example, in certain embodiments, the Cas protein comprises an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the Cas nuclease amino acid sequence. In certain embodiments, the Cas protein comprises a nuclease-inactive mutant of a Cas nuclease. In certain embodiments, the Cas protein further comprises an effector domain.
[0238] In certain embodiments, a Cas protein lacks substantially all DNA cleavage activity. Such a Cas protein can be generated, for example, by introducing one or more mutations into an active Cas nuclease (e.g., a naturally occurring Cas nuclease). A mutant Cas protein is considered to lack substantially all DNA cleavage activity if the DNA cleavage activity of the protein is about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less, e.g., zero or negligible, compared to the non-mutated form of the corresponding non-mutated form. Thus, a Cas protein may contain one or more mutations (e.g., a mutation in the RuvC domain of a VA-type Cas protein) and may be used as a genomic DNA-binding protein, with or without being fused to an effector domain. Exemplary mutations include D908A, E993A, and D1263A, with reference to the amino acid positions in AsCpf1; D832A, E925A, and D1180A, with reference to the amino acid positions in LbCpf1; and D917A, E1006A, and D1255A, with reference to the amino acid position numbers in FnCpf1. More mutations can be designed and generated according to the crystal structure described in Yamano et al. (2016) Cell, 165: 949.
[0239] Rather than losing nuclease activity to cleave all DNA, a Cas protein may lose the ability to cleave only the target strand or only the non-target strand of double-stranded DNA, thereby being understood to function as a nickase (see Gao et al. (2016) Cell Res., 26: 901). Thus, in certain embodiments, a Cas nuclease is a Cas nickase. In certain embodiments, a Cas nuclease has activity to cleave the non-target strand but substantially lacks activity to cleave the target strand, e.g., due to a mutation in the Nuc domain. In certain embodiments, a Cas nuclease has activity to cleave the target strand but substantially lacks activity to cleave the non-target strand.
[0240] In certain embodiments, the Cas nuclease has the activity to cleave double-stranded DNA, resulting in a double-strand break.
[0241] Cas proteins that lack substantially all DNA cleavage activity or have the ability to cleave only one strand can also be identified from naturally occurring systems. For example, certain naturally occurring CRISPR-Cas systems may retain the ability to bind to target nucleotide sequences but may have lost all or part of their DNA cleavage activity in eukaryotic (e.g., mammalian or human) cells. Such VA-type proteins are disclosed, for example, in Kim et al. (2017) ACS SYNTH. BIOL. 6(7): 1273-82 and Zhang et al. (2017) Cell Disov. 3:17018.
[0242] The activity of a Cas protein (e.g., a Cas nuclease) can be altered, for example, by creating an engineered Cas protein. In certain embodiments, the altered activity of the engineered Cas protein comprises increased targeting efficiency and / or reduced off-target binding. Without wishing to be bound by theory, it is hypothesized that off-target binding may be recognized by the Cas protein due to, for example, the presence of one or more mismatches between the spacer sequence and the target nucleotide sequence, which may affect the stability and / or conformation of the CRISPR-Cas complex. In certain embodiments, the altered activity comprises modified binding, e.g., increased binding to the target locus (e.g., the target strand or a non-target strand) and / or decreased binding to an off-target locus. In certain embodiments, the altered activity comprises altering the region of the protein that binds to a single guide nucleic acid or a dual guide nucleic acid. In certain embodiments, the altered activity of the engineered Cas protein comprises altering the region of the protein that binds to the target strand and / or a non-target strand. In certain embodiments, the altered activity of the engineered Cas protein comprises altering the region of the protein that binds to the off-target locus. The altered change may comprise a decrease in positive charge, a decrease in negative charge, an increase in positive charge, or an increase in negative charge. For example, a decrease in negative charge and an increase in positive charge may generally strengthen binding to nucleic acids, while a decrease in positive charge and an increase in negative charge may weaken binding to nucleic acids. In certain embodiments, the altered activity comprises an increase or decrease in steric hindrance between the protein and the single guide nucleic acid or dual guide nucleic acid. In certain embodiments, the altered activity comprises an increase or decrease in steric hindrance between the protein and the target strand and / or the non-target strand. In certain embodiments, the altered activity comprises an increase or decrease in steric hindrance between the protein and the off-target locus. In certain embodiments, the alteration or mutation comprises one or more substitutions of Lys, His, Arg, Glu, Asp, Ser, Gly, and / or Thr. In certain embodiments, the alteration or mutation comprises one or more substitutions of Gly, Ala, Ile, Glu, and / or Asp.In certain embodiments, the modification or mutation comprises one or more amino acid substitutions in the groove between the WED and RuvC domains of a Cas protein (e.g., a VA-type Cas protein).
[0243] In certain embodiments, the altered activity of the engineered Cas protein comprises an increase in the nuclease activity of cleaving the target locus. In certain embodiments, the altered activity of the engineered Cas protein comprises a decrease in the nuclease activity of cleaving the off-target locus. In certain embodiments, the altered activity of the engineered Cas protein comprises an alteration in helicase dynamics. In certain embodiments, the engineered Cas protein comprises a modification that alters the formation of a CRISPR complex.
[0244] In certain embodiments, a protospacer adjacent motif (PAM) or PAM-like motif directs binding of the Cas protein complex to the target locus. Many Cas proteins have PAM specificity. The exact sequence and length requirements of the PAM vary depending on the Cas protein used. The PAM sequence is typically 2-5 base pairs in length and is adjacent to (but located on a different strand of the target DNA from) the target nucleotide sequence. PAM sequences can be identified using any suitable method, such as testing the cleavage, targeting, or modification of the target nucleotide sequence and oligonucleotides with different PAM sequences.
[0245] Exemplary PAM sequences are provided in Table 2 and Table 3. In certain embodiments, the Cas protein comprises MAD7 and the PAM is TTTN (where N is A, C, G, or T). In certain embodiments, the Cas protein comprises MAD7 and the PAM is CTTN (where N is A, C, G, or T). In certain embodiments, the Cas protein comprises AsCpf1 and the PAM is TTTN (where N is A, C, G, or T). In certain embodiments, the Cas protein comprises FnCpf1 and the PAM is 5'TTN (where N is A, C, G, or T). PAM sequences for certain other VA-type Cas proteins are disclosed in Zetsche et al. (2015) Cell 163:759 and U.S. Patent No. 9,982,279. Furthermore, engineering the PAM-interacting (PI) domain of the Cas protein can allow programming of PAM specificity, improving target site recognition fidelity and / or increasing the versatility of engineered non-naturally occurring systems. Exemplary means of altering the PAM specificity of Cpfl are described in Gao et al. (2017) Nat. Biotechnol., 35:789.
[0246] In certain embodiments, engineered Cas proteins contain modifications that alter Cas protein specificity in concert with modifications to targeting scope. Cas mutants can be designed to have increased target specificity and coordinated modifications in PAM recognition, for example, by selecting mutations that alter PAM specificity (e.g., in the PI domain) and combining these mutations with groove mutations that increase (or, if desired, decrease) specificity for on-target versus off-target loci. The Cas modifications described herein can be used to counteract loss of specificity resulting from altered PAM recognition, enhance gain of specificity resulting from altered PAM recognition, counteract gain of specificity resulting from altered PAM recognition, or enhance loss of specificity resulting from altered PAM recognition.
[0247] In certain embodiments, the engineered Cas protein comprises one or more nuclear localization signal (NLS) motifs. In certain embodiments, the engineered Cas protein comprises at least two (e.g., at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs. Non-limiting examples of NLS motifs include the SV40 large T antigen NLS, having the amino acid sequence of PKKKRKV (SEQ ID NO: 40); an NLS derived from nucleoplasmin, e.g., the nucleoplasmin bisecting NLS, having the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 41); a c-myc NLS, having the amino acid sequence of PAAKRVKLD (SEQ ID NO: 42) or RQRRNELKRSP (SEQ ID NO: 43); hRNPA1 M9 NLS, having the amino acid sequence of NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 44); importin-α IBB domain NLS, having the amino acid sequence of RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 45); fibroid T protein NLS, having the amino acid sequence of VSRKRPRP (SEQ ID NO: 46) or PPKKARED (SEQ ID NO: 47); PQPKKKPL human p53 NLS having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO:48); mouse c-abl IV NLS having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO:49); influenza virus NS1 NLS having the amino acid sequence of DRLRR (SEQ ID NO:50) or PKQKKRK (SEQ ID NO:51); hepatitis virus delta antigen NLS having the amino acid sequence of RKLKKKIKKL (SEQ ID NO:52); mouse Mx1 protein NLS having the amino acid sequence of REKKKFLKRR (SEQ ID NO:53); human poly(ADP-ribose) polymerase NLS having the amino acid sequence of KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:54); human glucocorticoid receptor NLS having the amino acid sequence of RKCLQAGMNLEARKTKK (SEQ ID NO:55), and synthetic NLS motifs such as PAAKKKKLD (SEQ ID NO:56).
[0248] Generally, the one or more NLS motifs are of sufficient strength to drive the accumulation of detectable amounts of the Cas protein in the nucleus of a eukaryotic cell. The strength of the nuclear localization activity may derive from the number of NLS motifs in the Cas protein, the particular NLS motif used, the position of the NLS motif, or a combination of these and / or other factors. In certain embodiments, the engineered Cas protein contains at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the N-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the N-terminus). In certain embodiments, an engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the C-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more amino acids along the polypeptide chain from the C-terminus). In certain embodiments, an engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the C-terminus and at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the N-terminus. In certain embodiments, the engineered Cas protein comprises one, two, or three NLS motifs at or near the C-terminus, hi certain embodiments, the engineered Cas protein comprises one NLS motif at or near the N-terminus and one, two, or three NLS motifs at or near the C-terminus.In certain embodiments, the engineered Cas protein comprises a nucleoplasmin NLS motif at or near the C-terminus.
[0249] Detection of nuclear accumulation can be carried out by any suitable technique. For example, a detectable marker can be fused to the nucleic acid targeting protein so that its location within the cell can be visualized. Alternatively, after isolating the cell nucleus from the cell, its contents can be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. Nuclear accumulation can also be determined indirectly, such as by assays that detect the effect of nuclear transport of the Cas protein complex (e.g., assays for DNA cleavage or mutation at the target locus, or assays for altered gene expression activity), compared to controls that are not exposed to the Cas protein or that are exposed to a Cas protein lacking one or more NLS motifs.
[0250] The Cas protein may include a chimeric Cas protein, e.g., a Cas protein whose function is enhanced by being chimeric. A chimeric Cas protein may be a novel Cas protein containing fragments from more than one naturally occurring Cas protein or variant thereof. For example, fragments of multiple VA-type Cas homologs (e.g., orthologs) can be fused to form a chimeric Cas protein. In certain embodiments, the chimeric Cas protein comprises fragments of Cpf1 orthologs from multiple species and / or strains.
[0251] In certain embodiments, the Cas protein comprises one or more effector domains. The one or more effector domains may be located at or near the N-terminus of the Cas protein and / or at or near the C-terminus of the Cas protein. In certain embodiments, the effector domain comprised in the Cas protein is a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., a KRAB domain or a SID domain), an exogenous nuclease domain (e.g., FokI), a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), or a reverse transcriptase domain (e.g., a high-fidelity reverse transcriptase domain). Other activities of the effector domain include, but are not limited to, methylase activity, demethylase activity, transcription release factor activity, translation initiation activity, translation activation activity, translation repression activity, histone modification (e.g., acetylation or demethylation) activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity.
[0252] In certain embodiments, the Cas protein comprises one or more protein domains that enhance homology-directed repair (HDR) and / or inhibit non-homologous end joining (NHEJ). Exemplary protein domains with such functions are described in Jayavaradhan et al. (2019) Nat. Commun. 10(1): 2866 and Janssen et al. (2019) Mol. Ther. Nucleic Acids 16: 141-54. In certain embodiments, the Cas protein comprises a dominant-negative version of p53-binding protein 1 (53BP1), e.g., a fragment of 53BP1 containing a minimal focus-forming region (e.g., amino acids 1231-1644 of human 53BP1). In certain embodiments, the Cas protein comprises a motif targeted by APC-Cdh1, such as amino acids 1-110 of human geminin, thereby resulting in degradation of the fusion protein during the HDR-nonpermissive G1 phase of the cell cycle.
[0253] In certain embodiments, the Cas protein comprises an inducible or regulatable domain. Non-limiting examples of inducing or regulatable agents include light, hormones, and small molecule drugs. In certain embodiments, the Cas protein comprises a light-inducible or regulatable domain. In certain embodiments, the Cas protein comprises a chemical-inducible or regulatable domain.
[0254] In certain embodiments, the Cas protein comprises a tag protein or peptide to facilitate tracking and / or purification. Non-limiting examples of tag proteins and peptides include fluorescent proteins (e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato), HIS tags (e.g., 6xHis tag (SEQ ID NO: 2044), or gly-6xHis (SEQ ID NO: 2045); 8xHis (SEQ ID NO: 2046), or gly-8xHis (SEQ ID NO: 2047)), hemagglutinin (HA) tags, FLAG tags, 3xFLAG tags, and Myc tags.
[0255] In certain embodiments, the Cas protein is conjugated to a non-protein moiety, such as a fluorophore, useful for genomic imaging. In certain embodiments, the Cas protein is covalently conjugated to the non-protein moiety. As used herein, the terms "CRISPR-associated protein," "Cas protein," "Cas," "CRISPR-associated nuclease," and "Cas nuclease" include such conjugates despite the presence of one or more non-protein moieties.
[0256] B. Guide Nucleic Acid The guide nucleic acid may be a single gNA (sgNA, e.g., sgRNA), in which the gNA is a single polynucleotide, or a dual gNA (e.g., dual gRNA), in which the gNA comprises two separate polynucleotides (which may, in some cases, be covalently linked rather than by a conventional internucleotide linkage). In certain embodiments, a single guide nucleic acid is capable of activating only a Cas nuclease (e.g., in the absence of a tracrRNA).
[0257] Generally, gNAs comprise a modulator nucleic acid and a targeter nucleic acid. In sgNAs, the modulator and targeter nucleic acids are part of a single polynucleotide. In dual gNAs, the modulator and targeter nucleic acids are separate and not linked by a conventional nucleotide bond, e.g., not linked at all. The targeter nucleic acid comprises a spacer sequence and a targeter stem sequence. The modulator nucleic acid comprises the modulator stem sequence and generally additional nucleotides, such as nucleotides comprising the 5' tail. The modulator stem sequence and the targeter stem sequence may each comprise any suitable number of nucleotides and are sufficiently complementary to allow hybridization. In single gNAs, additional nucleotides may be present between the targeter stem sequence and the modulator stem sequence; these may form secondary structures, such as loops, in certain cases.
[0258] In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid that, together with the modulator nucleic acid, can bind to a Cas protein. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid that, together with the modulator nucleic acid, can activate a Cas protein. In certain embodiments, the system further comprises a Cas protein to which the targeter nucleic acid and modulator nucleic acid can bind, or a Cas nuclease that the targeter nucleic acid and modulator nucleic acid can activate.
[0259] It is contemplated that the single or dual guide nucleic acid must be compatible with a Cas protein (e.g., a Cas nuclease) to provide a workable CRISPR system. For example, the targeter stem sequence and modulator stem sequence may be derived from a naturally occurring crRNA that can activate a Cas nuclease in the absence of a tracrRNA. Alternatively, the targeter stem sequence and modulator stem sequence may be derived from a naturally occurring set of crRNAs and tracrRNAs, respectively, that can activate a Cas nuclease. In certain embodiments, the nucleotide sequences of the targeter stem sequence and modulator stem sequence are identical to the corresponding stem sequences of the stem-loop structure in such a naturally occurring crRNA.
[0260] Guide nucleic acid sequences operable with Type II or Type V Cas proteins are known in the art and are disclosed, for example, in U.S. Patent Nos. 9,790,490, 9,896,696, 10,113,179, and 10,266,850, and U.S. Patent Application Publication No. 2014 / 0242664. It is understood that these sequences are merely exemplary, and other guide nucleic acid sequences can also be used with these Cas proteins.
[0261] [Table 4A]
[0262] [Table 4B]
[0263] [Table 4C]
[0264] [Table 5A]
[0265] [Table 5B]
[0266] In certain embodiments, the guide nucleic acid, in the context of a VA-type CRISPR-Cas system, comprises a targeter stem sequence listed in Table 5. Targeter stem sequences that are the same as portions of the scaffold sequence are bold and underlined in Table 4.
[0267] In certain embodiments, the guide nucleic acid is a single guide nucleic acid comprising, from 5' to 3', a modulator stem sequence, a loop sequence, a targeter stem sequence, and a spacer sequence. In certain embodiments, the targeter stem sequence in the single guide nucleic acid is listed in Table 4 as a bold, underlined portion of the scaffold sequence, and the modulator stem sequence is complementary (e.g., 100% complementary) to the targeter stem sequence. In certain embodiments, the single guide nucleic acid comprises, from 5' to 3', a modulator sequence listed in Table 4 as an underlined portion of the scaffold sequence, a loop sequence, a targeter stem sequence as a bold, underlined portion of the same scaffold sequence, and a spacer sequence. In certain embodiments, the engineered, non-naturally occurring system comprises a single guide nucleic acid comprising a scaffold sequence listed in Table 4. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in a SEQ ID NO: listed in the same column of Table 4. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence set forth in a SEQ ID NO: listed in the same column of Table 4. In certain embodiments, the system is useful for targeting, editing, or modifying nucleic acids that comprise a target nucleotide sequence near or adjacent (e.g., immediately downstream) to a PAM listed in the same column of Table 4 when using a non-target strand (i.e., a strand that did not hybridize with the spacer sequence) as a coordinate.
[0268] In certain embodiments, the guide nucleic acid, e.g., a dual gNA, comprises a targeter guide nucleic acid comprising, from 5' to 3', a targeter stem sequence and a spacer sequence. In certain embodiments, the targeter stem sequence in the targeter nucleic acid is listed in Table 5. In certain embodiments, the engineered, non-naturally occurring system comprises a targeter nucleic acid and a modulator stem sequence that is complementary (e.g., 100% complementary) to the targeter stem sequence. In certain embodiments, the modulator nucleic acid comprises a modulator sequence listed in the same row of Table 5. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in a SEQ ID NO: listed in the same column of Table 5. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence set forth in a SEQ ID NO: listed in the same column of Table 5. In certain embodiments, the system is useful for targeting, editing, or modifying nucleic acids that comprise a target nucleotide sequence near or adjacent (e.g., immediately downstream) to a PAM listed in the same column of Table 5 when using a non-target strand (i.e., a strand that did not hybridize with a spacer sequence) as a coordinate.
[0269] The single guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid can be chemically synthesized or produced in a biological process (e.g., catalyzed by RNA polymerase in an in vitro reaction). Such a reaction or process may limit the length of the single guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid. In certain embodiments, the single guide nucleic acid is no more than 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides in length. In certain embodiments, the single guide nucleic acid is at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the single guide nucleic acid is 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 3 The targeter nucleic acid may be 0 to 50, 30 to 40, 40 to 100, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 100, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 100, 60 to 90, 60 to 80, 60 to 70, 70 to 100, 70 to 90, 70 to 80, 80 to 100, 80 to 90, or 90 to 100 nucleotides in length. In certain embodiments, the targeter nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides or less in length. In certain embodiments, the targeter nucleic acid is at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length.In certain embodiments, the targeter nucleic acid is 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, The length of the modulator nucleic acid may be 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60, 60-100, 60-90, 60-80, 60-70, 70-100, 70-90, 70-80, 80-100, 80-90, or 90-100 nucleotides. In certain embodiments, the modulator nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 20 nucleotides or less in length. In certain embodiments, the modulator nucleic acid is at least 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the modulator nucleic acid is selected from the group consisting of 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 15 to 20, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to The length is 60, 25-50, 25-40, 25-30, 30-100, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60, 60-100, 60-90, 60-80, 60-70, 70-100, 70-90, 70-80, 80-100, 80-90, or 90-100 nucleotides.
[0270] It is contemplated that the length of the duplex formed within a single guide nucleic acid or between a targeter nucleic acid and a modulator nucleic acid in a dual gNA can be a factor in providing an operable CRISPR system. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 10 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4, 5, 6, 7, 8, 9, or 10 nucleotides. It is understood that the composition of nucleotides in each sequence affects the stability of the duplex, with CG base pairs conferring greater stability than AU base pairs. In certain embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the base pairs are CG base pairs.
[0271] In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 5 nucleotides. Therefore, the targeter stem sequence and the modulator stem sequence form a 5-base pair duplex. In certain embodiments, 0 to 4, 0 to 3, 0 to 2, 0 to 1, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 5, 2 to 4, 2 to 3, 3 to 5, 3 to 4, or 4 to 5 of the 5 base pairs are CG base pairs. In certain embodiments, 0, 1, 2, 3, 4, or 5 of the 5 base pairs are CG base pairs. In certain embodiments, the targeter stem sequence consists of 5'-GUAGA-3' and the modulator stem sequence consists of 5'-UCUAC-3'. In certain embodiments, the targeter stem sequence consists of 5'-GUGGG-3' and the modulator stem sequence consists of 5'-CCCAC-3'.
[0272] In certain embodiments, in a VA-type system, the 3' end of the targeter stem sequence is linked to the 5' end of the spacer sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are adjacent to each other and directly linked by an internucleotide bond. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by a single nucleotide, such as a uridine. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by two or more nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.
[0273] In certain embodiments, the targeter nucleic acid further comprises an additional nucleotide sequence 5' to the targeter stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least 1 nucleotide (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50). In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In certain embodiments, the additional nucleotide sequence consists of 2 nucleotides. In certain embodiments, the additional nucleotide sequence mimics a loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 3' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It is understood that additional nucleotide sequences 5' to the targeter stem sequence may be optional. Thus, in certain embodiments, the targeter nucleic acid does not include any additional nucleotides 5' to the targeter stem sequence.
[0274] In certain embodiments, the targeter nucleic acid or single guide nucleic acid further comprises an additional nucleotide sequence containing one or more nucleotides at the 3' end that do not hybridize to the target nucleotide sequence. The additional nucleotide sequence can protect the targeter nucleic acid from degradation by 3'-5' exonucleases. In certain embodiments, the additional nucleotide sequence is 100 nucleotides or less in length. In certain embodiments, the additional nucleotide sequence is 90, 80, 70, 60, 50, 40, 30, 20, or 10 nucleotides or less in length. In certain embodiments, the additional nucleotide sequence is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides in length. In certain embodiments, the additional nucleotide sequence is between 5 and 100, 5 and 50, 5 and 40, 5 and 30, 5 and 25, 5 and 20, 5 and 15, 5 and 10, 10 and 100, 10 and 50, 10 and 40, 10 and 30, 10 and 25, 10 and 20, 10 and 15, 15 and 100, 15 and 50, 15 and 40, 15 and 30, 15 and 25, 15 and 20, 20 and 100, 20 and 50, 20 and 40, 20 and 30, 20 and 25, 25 and 100, 25 and 50, 25 and 40, 25 and 30, 30 and 100, 30 and 50, 30 and 40, 40 and 100, 40 and 50, or 50 and 100 nucleotides in length.
[0275] In certain embodiments, the additional nucleotide sequence forms a hairpin with the spacer sequence. Such secondary structures can increase the specificity of the guide nucleic acid or engineered, non-naturally occurring systems (see Kocak et al. (2019) Nat. Biotech. 37: 657-66). In certain embodiments, the free energy change during hairpin formation is greater than or equal to -20 kcal / mol, -15 kcal / mol, -14 kcal / mol, -13 kcal / mol, -12 kcal / mol, -11 kcal / mol, or -10 kcal / mol. In certain embodiments, the free energy change during hairpin formation is greater than or equal to -5 kcal / mol, -6 kcal / mol, -7 kcal / mol, -8 kcal / mol, -9 kcal / mol, -10 kcal / mol, -11 kcal / mol, -12 kcal / mol, -13 kcal / mol, -14 kcal / mol, or -15 kcal / mol. In certain embodiments, the free energy change during hairpin formation is greater than or equal to -20 to -10 kcal / mol, -20 to -11 kcal / mol, -20 to -12 kcal / mol, -20 to -13 kcal / mol, -20 to -14 kcal / mol, -20 to -15 kcal / mol, -15 to -10 kcal / mol, -15 to -11 kcal / mol, -15 to -12 kcal / mol, -15 to -13 kcal / mol. mol, -15 to -14 kcal / mol, -14 to -10 kcal / mol, -14 to -11 kcal / mol, -14 to -12 kcal / mol, -14 to -13 kcal / mol, -13 to -10 kcal / mol, -13 to -11 kcal / mol, -13 to -12 kcal / mol, -12 to -10 kcal / mol, -12 to -11 kcal / mol, or -11 to -10 kcal / mol. In other embodiments, the targeter nucleic acid or single guide nucleic acid does not contain any nucleotides 3' to the spacer sequence.
[0276] In certain embodiments, the modulator nucleic acid further comprises an additional nucleotide sequence 3' to the modulator stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least 1 nucleotide (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50). In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In certain embodiments, the additional nucleotide sequence consists of a single nucleotide (e.g., uridine). In certain embodiments, the additional nucleotide sequence consists of two nucleotides. In certain embodiments, the additional nucleotide sequence mimics the loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 5' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It is understood that the additional nucleotide sequence 3' to the modulator stem sequence may be optional. Thus, in certain embodiments, the modulator nucleic acid does not comprise any additional nucleotides 3' to the modulator stem sequence.
[0277] It is understood that the additional nucleotide sequence 5' to the targeter stem sequence and the additional nucleotide sequence 3' to the modulator stem sequence, if present, may interact with each other. For example, while the nucleotides immediately 5' to the targeter stem sequence and immediately 3' to the modulator stem sequence do not form Watson-Crick base pairs (which would otherwise constitute part of the targeter stem sequence and the modulator stem sequence, respectively), other nucleotides in the additional nucleotide sequence 5' to the targeter stem sequence and the additional nucleotide sequence 3' to the modulator stem sequence may form one, two, three, or more base pairs (e.g., Watson-Crick base pairs). Such interactions may affect the stability of a complex comprising the targeter nucleic acid and the modulator nucleic acid.
[0278] The stability of a complex containing a targeter nucleic acid and a modulator nucleic acid can be evaluated by the calculated or actually measured Gibbs free energy change (ΔG) during complex formation. When all predicted base pairing in the complex occurs between bases in the target nucleic acid and bases in the modulator nucleic acid, i.e., when there is no intrastrand secondary structure, ΔG during complex formation generally correlates with ΔG during secondary structure formation within the corresponding single guide nucleic acid. Methods for calculating or measuring ΔG are known in the art. An exemplary method is RNAfold (rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) disclosed in Gruber et al. (2008) Nucleic Acids Res., 36 (Web Server issue): W70-W74. Unless otherwise specified, ΔG in this disclosure is calculated by RNAfold for secondary structure formation within the corresponding single guide nucleic acid. In certain embodiments, ΔG is less than or equal to -1 kcal / mol, e.g., less than or equal to -2 kcal / mol, less than or equal to -3 kcal / mol, less than or equal to -4 kcal / mol, less than or equal to -5 kcal / mol, less than or equal to -6 kcal / mol, less than or equal to -7 kcal / mol, less than or equal to -7.5 kcal / mol, or less than or equal to -8 kcal / mol. In certain embodiments, ΔG is greater than or equal to -10 kcal / mol, e.g., greater than or equal to -9 kcal / mol, greater than or equal to -8.5 kcal / mol, or greater than or equal to -8 kcal / mol. In certain embodiments, ΔG is in the range of -10 to -4 kcal / mol.In certain embodiments, ΔG is in the range of -8 to -4 kcal / mol, -7 to -4 kcal / mol, -6 to -4 kcal / mol, -5 to -4 kcal / mol, -8 to -4.5 kcal / mol, -7 to -4.5 kcal / mol, -6 to -4.5 kcal / mol, or -5 to -4.5 kcal / mol. In certain embodiments, ΔG is about -8 kcal / mol, -7 kcal / mol, -6 kcal / mol, -5 kcal / mol, -4.9 kcal / mol, -4.8 kcal / mol, -4.7 kcal / mol, -4.6 kcal / mol, -4.5 kcal / mol, -4.4 kcal / mol, -4.3 kcal / mol, -4.2 kcal / mol, -4.1 kcal / mol, or -4 kcal / mol.
[0279] It is understood that ΔG can be affected by sequences in the targeter nucleic acid that are not within the targeter stem sequence and / or sequences in the modulator nucleic acid that are not within the modulator stem sequence. For example, one or more base pairs (e.g., Watson-Crick base pairs) between additional sequences 5' to the targeter stem sequence and additional sequences 3' to the modulator stem sequence can decrease ΔG, i.e., stabilize the nucleic acid complex. In certain embodiments, the nucleotide immediately 5' to the targeter stem sequence contains uracil or is a uridine, and the nucleotide immediately 3' to the modulator stem sequence contains uracil or is a uridine, thereby forming a non-conventional UU base pair.
[0280] In certain embodiments, the modulator nucleic acid or single guide nucleic acid comprises a nucleotide sequence referred to herein as the "5' tail," which is located 5' to the modulator stem sequence. In naturally occurring VA-type CRISPR-Cas systems, the 5' tail is the nucleotide sequence located 5' to the stem-loop structure of the crRNA. The 5' tail in an engineered VA-type CRISPR-Cas system, whether single-guide or dual-guide, can mimic the 5' tail in the corresponding naturally occurring VA-type CRISPR-Cas system.
[0281] Without being bound by theory, it is contemplated that the 5' tail may participate in the formation of a CRISPR-Cas complex. For example, in certain embodiments, the 5' tail forms a pseudoknot structure with the modulator stem sequence, which is recognized by the Cas protein (see Yamano et al. (2016) Cell, 165:949). In certain embodiments, the 5' tail is at least 3 (e.g., at least 4 or at least 5) nucleotides in length. In certain embodiments, the 5' tail is 3, 4, or 5 nucleotides in length. In certain embodiments, the 3'-terminal nucleotide of the 5' tail comprises uracil or is uridine. In certain embodiments, the second nucleotide in the 5' tail, counted from the 3' end, comprises uracil or is uridine. In certain embodiments, the third nucleotide in the 5' tail, counted from the 3' end, comprises adenine or is adenosine. This third nucleotide can form a base pair (e.g., a Watson-Crick base pair) with the nucleotide 5' to the modulator stem sequence. Thus, in certain embodiments, the modulator nucleic acid comprises a uridine- or uracil-containing nucleotide 5' to the modulator stem sequence. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-AUU-3'. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-AAUU-3'. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-UAAUU-3'. In certain embodiments, the 5' tail is located immediately 5' to the modulator stem sequence.
[0282] In certain embodiments, the single guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid are designed to reduce the degree of secondary structure other than hybridization between the targeter stem sequence and the modulator stem sequence. In certain embodiments, less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of a single guide nucleic acid other than the targeter stem sequence and the modulator stem sequence participate in self-complementary base pairing when optimally folded. In certain embodiments, less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of a target nucleic acid and / or modulator nucleic acid participate in self-complementary base pairing when optimally folded. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on minimum Gibbs free energy calculations. An example of one such algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), pp. 133-148). Another example of a folding algorithm is the online web server RNAfold, developed at the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., AR Gruber et al., 2008, Cell 106(1): pp. 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): pp. 1151-62).
[0283] The targeter nucleic acid is directed to a specific target nucleotide sequence, and the donor template can be designed to modify the target nucleotide sequence or an adjacent sequence. It is therefore understood that combining a single guide nucleic acid, targeter nucleic acid, or modulator nucleic acid with a donor template can increase editing efficiency and reduce off-targeting. Thus, in certain embodiments, the single guide nucleic acid or modulator nucleic acid further comprises a donor template recruitment sequence capable of hybridizing with the donor template (see Figure 2B). Donor templates are described in the "Donor Template" subsection of Section II below. The donor template and donor template recruitment sequence can be designed to have sequence complementarity. In certain embodiments, the donor template recruitment sequence is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) complementary to at least a portion of the donor template. In certain embodiments, the donor template recruitment sequence is 100% complementary to at least a portion of the donor template. In certain embodiments, if the donor template contains an engineered sequence that is not homologous to the sequence to be repaired, the donor template recruitment sequence can hybridize with the engineered sequence in the donor template. In certain embodiments, the donor template recruitment sequence is at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length. In certain embodiments, the donor template recruitment sequence is located at or near the 5' end of the single guide nucleic acid or the 5' end of the modulator nucleic acid. In certain embodiments, the donor template recruitment sequence is linked to the 5' tail, if present, or to the modulator stem sequence of the single guide nucleic acid or modulator nucleic acid by an internucleotide bond or nucleotide linker.
[0284] In certain embodiments, the single guide nucleic acid or modulator nucleic acid further comprises an editing enhancer sequence that increases the efficiency of gene editing and / or homology-directed repair (HDR) (see Figure 2C). Exemplary editing enhancer sequences are described in Park et al. (2018) Nat. Commun. 9:3313. In certain embodiments, the editing enhancer sequence, if present, is located 5' to the 5' tail or 5' to the single guide nucleic acid or modulator stem sequence. In certain embodiments, the editing enhancer sequence is 1-50, 4-50, 9-50, 15-50, 25-50, 1-25, 4-25, 9-25, 15-25, 1-15, 4-15, 9-15, 1-9, 4-9, or 1-4 nucleotides in length. In certain embodiments, the editing enhancer sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 55 nucleotides in length. The editing enhancer sequence is designed to minimize homology to the target nucleotide sequence or any other sequence that may bring the engineered, non-naturally occurring system into contact with, for example, the genomic sequence of the cell to which the engineered, non-naturally occurring system is delivered. In certain embodiments, the editing enhancer is designed to minimize the presence of hairpin structures. The editing enhancer may include one or more chemical modifications disclosed herein.
[0285] The single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid may further comprise a protective nucleotide sequence that prevents or reduces nucleic acid degradation. In certain embodiments, the protective nucleotide sequence is at least 5 (e.g., at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides in length. The length of the protective nucleotide sequence increases the time it takes for an exonuclease to reach the 5' tail, modulator stem sequence, targeter stem sequence, and / or spacer sequence, thereby protecting these portions of the single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid from exonucleolytic degradation. In certain embodiments, the protective nucleotide sequence forms a secondary structure, such as a hairpin or tRNA structure, to reduce the rate of exonucleolytic degradation (see, e.g., Wu et al. (2018) Cell. Mol. Life Sci., 75(19): 3593-3607). Secondary structure can be predicted by methods known in the art, such as the online web server RNAfold developed at the University of Vienna, which uses a centroid structure prediction algorithm (see Gruber et al. (2008) Nucleic Acids Res., 36:W70). Certain chemical modifications that may be present in the protected nucleotide sequence can also prevent or reduce nucleic acid degradation, as disclosed in the "RNA Modifications" subsection below.
[0286] The protective nucleotide sequence is typically located at the 5' or 3' end of the single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid. In certain embodiments, the single guide nucleic acid comprises a protective nucleotide sequence at the 5' end, the 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises a protective nucleotide sequence at the 5' end, the 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises a protective nucleotide sequence at the 5' end (see Figure 2A). In certain embodiments, the targeter nucleic acid comprises a protective nucleotide sequence at the 5' end, the 3' end, or both ends, optionally via a nucleotide linker.
[0287] As described above, various nucleotide sequences may be present in the 5' portion of a single nucleic acid or modulator nucleic acid, including, but not limited to, donor template recruitment sequences, editing enhancer sequences, protective nucleotide sequences, and linkers connecting such sequences, if present, to the 5' tail or to the modulator stem sequence. It is understood that the functions of donor template recruitment, editing enhancement, protection against degradation, and linking are not mutually exclusive, and a single nucleotide sequence may have one or more such functions. For example, in certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruitment sequence and an editing enhancer sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruitment sequence and a protective sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both an editing enhancer sequence and a protective sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is a donor template recruitment sequence, an editing enhancer sequence, and a protective sequence. In certain embodiments, if present, the nucleotide sequence 5' to the 5' tail or 5' to the modulator stem sequence is from 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 20 to 90, 20 to 80 , 20-70, 20-60, 20-50, 20-40, 20-30, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-90, 40-80, 40-70, 40-60, 40-50, 50-90, 50-80, 50-70, 50-60, 60-90, 60-80, 60-70, 70-90, 70-80, or 80-90 nucleotides in length.
[0288] In certain embodiments, the engineered, non-naturally occurring system further comprises one or more compounds (e.g., small molecule compounds) that enhance HDR and / or inhibit NHEJ. Exemplary compounds with such functionality are described in Maruyama et al. (2015) Nat Biotechnol. 33(5): 538-42; Chu et al. (2015) Nat Biotechnol. 33(5): 543-48; Yu et al. (2015) Cell Stem Cell 16(2): 142-47; Pinder et al. (2015) Nucleic Acids Res. 43(19): 9379-92; and Yagiz et al. (2019) Commun. Biol. 2: 198. In certain embodiments, the engineered, non-naturally occurring system further comprises one or more compounds selected from the group consisting of a DNA ligase IV antagonist (e.g., an SCR7 compound, an Ad4 E1B55K protein, and an Ad4 E4orf6 protein), a RAD51 agonist (e.g., RS-1), a DNA-dependent protein kinase (DNA-PK) antagonist (e.g., NU7441 and KU0060648), a β3-adrenergic receptor agonist (e.g., L755507), an inhibitor of intracellular protein transport from the ER to the Golgi apparatus (e.g., Brefeldin A), and any combination thereof.
[0289] In certain embodiments, engineered, naturally occurring systems comprising a targeter nucleic acid and a modulator nucleic acid are adjustable or inducible. For example, in certain embodiments, the targeter nucleic acid, modulator nucleic acid, and / or Cas protein can be introduced into the target nucleotide sequence at different times, and the system is active only when all components are present. In certain embodiments, the amount of targeter nucleic acid, modulator nucleic acid, and / or Cas protein can be titrated to achieve the desired efficiency and specificity. In certain embodiments, an excess amount of nucleic acid comprising a targeter stem sequence or a modulator stem sequence can be added to the system, thereby dissociating the complex between the targeter nucleic acid and the modulator nucleic acid and shutting down the system.
[0290] C. gNA Modification Guide nucleic acids, including single guide nucleic acids, targeter nucleic acids, and / or modulator nucleic acids, may comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, a single guide nucleic acid comprises DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, a targeter nucleic acid comprises DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, a modulator nucleic acid comprises DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. The spacer sequence can be presented as a DNA sequence by including thymidine (T) rather than uridine (U). It is understood that corresponding RNA sequences and DNA / RNA chimeric sequences are also contemplated. For example, if the spacer sequence is RNA, the sequence can be derived from the DNA sequences disclosed herein by replacing each T with U. Consequently, for purposes of describing nucleotide sequences, T and U are used interchangeably herein.
[0291] In certain embodiments, in a single guide nucleic acid, the targeter nucleic acid and modulator nucleic acid are part of a single polynucleotide, and in a dual guide nucleic acid, the targeter nucleic acid and modulator nucleic acid are separate nucleic acids, the engineered, non-naturally occurring system includes a target nucleic acid comprising a target nucleotide sequence and a spacer sequence designed to hybridize with the target stem sequence; and a modulator nucleic acid comprising a modulator stem sequence complementary to the target stem sequence, and optionally a 5' sequence, e.g., a tail sequence; the modifications may include one or more chemical modifications to one or more nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid (dual and single gNAs), at or near the 5' end of the targeter nucleic acid (dual gNAs), at or near the 3' end of the modulator nucleic acid (dual gNAs), at or near the 5' end of the modulator nucleic acid (single and dual gNAs), or, for single or dual gNAs, optionally in combination thereof. In certain embodiments, the Cas protein is a VA-type Cas nuclease. The modulator and / or targeter nucleic acid sequences may comprise additional sequences as detailed in the guide nucleic acid section, and modifications may be made, as needed, in these additional sequences, as will be apparent to those skilled in the art. In the embodiments described in this section below, in certain embodiments, the guide nucleic acid is from 5' at the modulator nucleic acid to 3' at the modulator stem sequence, and from 5' at the targeter stem sequence to 3' at the targeter sequence (see, e.g., Figures 1A and 1B); in certain embodiments, the guide nucleic acid is, as needed, from 3' at the modulator nucleic acid to 5' at the modulator stem sequence, and from 3' at the targeter stem sequence to 5' at the targeter sequence.
[0292] The targeter nucleic acid may comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. The modulator nucleic acid may comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, the targeter nucleic acid is RNA and the modulator nucleic acid is RNA. A targeter nucleic acid in the form of RNA is also referred to as a targeter RNA, and a modulator nucleic acid in the form of RNA is also referred to as a modulator RNA. The nucleotide sequences disclosed herein are presented as DNA sequences by including thymidine (T) and / or as RNA sequences by including uridine (U). It is understood that corresponding DNA sequences, RNA sequences, and DNA / RNA chimeric sequences are also contemplated. For example, if a spacer sequence is presented as a DNA sequence, a nucleic acid comprising this spacer sequence as RNA can be derived from the DNA sequence disclosed herein by replacing each T with U. Consequently, for purposes of describing nucleotide sequences, T and U are used interchangeably herein.
[0293] In certain embodiments, some or all of the gNAs are RNA, e.g., gRNAs. In certain embodiments, 5 to 100%, 10 to 100%, 20 to 100%, 30 to 100%, 40 to 100%, 50 to 100%, 60 to 100%, 70 to 100%, 80 to 100%, 90 to 100%, 95 to 100%, 99 to 100%, or 99.5 to 100% of the gNAs are gRNAs. In certain embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the gNAs are RNA. In certain embodiments, 50% of the gNAs are RNA. In certain embodiments, 70% of the gNAs are RNA. In certain embodiments, 90% of the gNAs are RNA. In certain embodiments, 100% of the gNAs are RNA, e.g., gRNA. In further embodiments, the remainder of the gNA that is not RNA comprises modified ribonucleotides, deoxyribonucleotides, modified deoxyribonucleotides, or synthetic, e.g., non-natural nucleotides, including, but not limited to, threose nucleic acids, locked nucleic acids, peptide nucleic acids, arabinonucleic acids, hexose nucleic acids, among others.
[0294] In certain embodiments, the targeter nucleic acid and / or modulator nucleic acid is an RNA having one or more modifications in the ribose group, one or more modifications in the phosphate group, one or more modifications in the nucleobase, one or more terminal modifications, or a combination thereof. Exemplary modifications are disclosed in U.S. Patent Nos. 10,900,034 and 10,767,175, U.S. Patent Application Publication No. 2018 / 0119140, Watts et al. (2008) Drug Discov. Today 13: 842-55, and Hendel et al. (2015) Nat. Biotechnol. 33: 985.
[0295] In certain embodiments, the targeter nucleic acid, e.g., RNA, comprises at least one nucleotide at or near the 3' end that comprises a ribose, a phosphate group, a modification to a nucleobase, or a terminal modification. In certain embodiments, the 3' end of the targeter nucleic acid comprises a spacer sequence. In certain embodiments, the 3' end of the targeter nucleic acid comprises a targeter stem sequence. Exemplary modifications are disclosed in Dang et al. (2015) Genome Biol. 16: 280, Kocaz et al. (2019) Nature Biotech. 37: 657-66, Liu et al. (2019) Nucleic Acids Res. 47(8): 4169-4180, Schubert et al. (2018) J. Cytokine Biol. 3(1): 121, Teng et al. (2019) Genome Biol. 20(1): 15, Watts et al. (2008) Drug Discovery Today 13(19-20): 842-55, and Wu et al. (2018) Cell Mol. Life. Sci. 75(19): 3593-607.
[0296] Modifications in the ribose group include, but are not limited to, modifications at the 2'-position or the 4'-position. For example, in certain embodiments, the ribose comprises a 2'-O-Ci_4 alkyl, such as 2'-O-methyl (2'-OMe, or M). In certain embodiments, the ribose comprises a 2'-O-Ci_3 alkyl-O-Ci_3 alkyl, such as 2'-methoxyethoxy (2'-O-CH2CHOCH3), also known as 2'-O-(2-methoxyethyl) or 2'-MOE. In certain embodiments, the ribose comprises a 2'-O-allyl. In certain embodiments, the ribose comprises a 2'-O-2,4-dinitrophenol (DNP). In certain embodiments, the ribose comprises a 2'-halo, such as 2'-F, 2'-Br, 2'-Cl, or 2'-I. In certain embodiments, the ribose comprises a 2'-NH2. In certain embodiments, the ribose comprises 2'-H (e.g., a deoxynucleotide). In certain embodiments, the ribose comprises 2'-arabino or 2'-F-arabino. In certain embodiments, the ribose comprises 2'-LNA or 2'-ULNA. In certain embodiments, the ribose comprises 4'-thioribosyl.
[0297] Modifications may include deoxy groups, for example, 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP).
[0298] Modifications of the internucleotide bond at the phosphate group include, but are not limited to, phosphorothioate (S), chiral phosphorothioate, phosphorodithioate, boranophosphonate, C 1~4Examples include alkyl phosphonates, such as methyl phosphonate, boranophosphonate, phosphonocarboxylates, such as phosphonoacetate (P), phosphonocarboxylate esters, such as phosphonoacetate ester, amides, thiophosphonocarboxylates, such as thiophosphonoacetate (SP), thiophosphonocarboxylate esters, such as thiophosphonoacetate ester, and phosphodiesters or 2',5'-linked phosphates having any of the above-mentioned modified phosphates. Various salts, mixed salts, and free acid forms are also included.
[0299] Modifications in the nucleobase include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine, 5-methyluracil, 5-hydroxymethylcytosine, 5-hydroxymethyl Examples include uracil, 5,6-dehydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil, 5-allylcytosine, 5-aminoallyluracil, 5-aminoallyl-cytosine, 5-bromouracil, 5-iodouracil, diaminopurine, difluorotoluene, dihydrouracil, abasic nucleotides, Z bases, P bases, unstructured nucleic acids, isoguanine, isocytosine (see Piccirilli et al. (1990) Nature, 343:33), 5-methyl-2-pyrimidine (see Rappaport (1993) Biochemistry, 32:3047), x(A,G,C,T), and y(A,G,C,T).
[0300] Terminal modifications include, but are not limited to, polyethylene glycol (PEG), hydrocarbon linkers (heteroatom (O, S, N)-substituted hydrocarbon spacers; halo-substituted hydrocarbon spacers; keto-, carboxyl-, amido-, thionyl-, carbamoyl-thiocarbamoyl-containing hydrocarbon spacers, propanediol, etc.), spermine linkers, dyes such as fluorescent dyes (e.g., fluorescein, rhodamine, cyanine), quenchers (e.g., dabcyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides, and / or proteins). In certain embodiments, terminal modifications include conjugation (or ligation) of the RNA to another molecule, including an oligonucleotide (such as a deoxyribonucleotide and / or a ribonucleotide), a peptide, a protein, a sugar, an oligosaccharide, a steroid, a lipid, folate, a vitamin, and / or other molecule. In certain embodiments, the terminal modification incorporated into the RNA is incorporated as a phosphodiester bond and is located within the RNA sequence via a linker, such as a 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker, which can be incorporated anywhere between two nucleotides in the RNA.
[0301] The modifications disclosed above can be combined in the targeter nucleic acid and / or modulator nucleic acid in the form of RNA. In certain embodiments, the modification in the RNA is selected from the group consisting of incorporation of 2'-O-methyl-3' phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-halo-3'-phosphorothioate (e.g., 2'-fluoro-3'-phosphorothioate), 2'-halo-3'-phosphonoacetate (e.g., 2'-fluoro-3'-phosphonoacetate), and 2'-halo-3'-thiophosphonoacetate (e.g., 2'-fluoro-3'-thiophosphonoacetate).
[0302] In certain embodiments, modifications may include 2'-O-methyl (M), phosphorothioate (S), phosphonoacetate (P), thiophosphonoacetate (SP), 2'-O-methyl-3'-phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP), or combinations thereof, at or near either the 3' or 5' end of either the targeter or modulator nucleic acid, for single or dual gNAs, as appropriate. In certain embodiments, modifications may include either a 5' or 3' propanediol or C3 linker modification.
[0303] In certain embodiments, the modification alters the stability of the RNA. In certain embodiments, the modification enhances the stability of the RNA, for example, by increasing the nuclease resistance of the RNA compared to the corresponding RNA that does not contain the modification. Stability-enhancing modifications include, but are not limited to, 2'-O-methyl, 2'-OC 1~4 Alkyl, 2'-halo (e.g., 2'-F, 2'-Br, 2'-Cl, or 2'-I), 2'MOE, 2'-OC 1~3 Alkyl-OC 1~3These modifications include the incorporation of alkyl, 2'-NH2, 2'-H (or 2'-deoxy), 2'-arabino, 2'-F-arabino, 4'-thioribosyl sugar moieties, 3'-phosphorothioates, 3'-phosphonoacetates, 3'-thiophosphonoacetates, 3'-methylphosphonates, 3'-boranophosphates, 3'-phosphorodithioates, locked nucleic acid ("LNA") nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring, and unlocked nucleic acid ("ULNA") nucleotides. Such modifications are suitable for use as protecting groups to prevent or reduce degradation of 5' sequences, such as tail sequences, modulator stem sequences (dual guide nucleic acids), targeter stem sequences (dual guide nucleic acids), and / or spacer sequences (see the "Targeter and Modulator Nucleic Acids" subsection).
[0304] In certain embodiments, the modification alters the specificity of the engineered, non-naturally occurring system. In certain embodiments, the modification enhances the specificity of the engineered, naturally occurring system, for example, by enhancing on-target binding and / or cleavage, or by reducing off-target binding and / or cleavage, or a combination thereof. Specificity-enhancing modifications include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, and pseudouracil. The 3'-terminus, e.g., within 10, 5, 4, 3, 2, or 1 nucleotide of the 3'-terminal nucleotide, is modified.
[0305] In certain embodiments, the modification alters the immunostimulatory effect of the RNA compared to the corresponding RNA that does not contain the modification, e.g., in certain embodiments, the modification reduces the ability of the RNA to activate TLR7, TLR8, TLR9, TLR3, RIG-I, and / or MDA5.
[0306] In certain embodiments, the targeter nucleic acid and / or modulator nucleic acid contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 modified nucleotides or internucleotide linkages. Modifications can be made at one or more positions in the targeter nucleic acid and / or modulator nucleic acid such that these nucleic acids retain function. For example, a modified nucleic acid can still direct a Cas protein to a target nucleotide sequence, allowing the Cas protein to exert its effector function. It is understood that a particular modification at a position can be selected based on the function of the nucleotide or internucleotide linkage at that position. For example, specificity-enhancing modifications may be suitable for nucleotides or internucleotide bonds in a spacer sequence, a targeter stem sequence, or a modulator stem sequence. Stability-enhancing modifications may be suitable for one or more terminal nucleotides or internucleotide bonds in a targeter nucleic acid and / or a modulator nucleic acid. In certain embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide bonds at or near the 5'-end and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide bonds at or near the 3'-end of a targeter nucleic acid are modified. In certain embodiments, no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 5' end and / or no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid are modified.In certain embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 5'-terminus and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 3'-terminus of a modulator nucleic acid are modified. In certain embodiments, no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 5'-terminus and / or no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 3'-terminus of a modulator nucleic acid are modified. Selection of positions for modification is described in U.S. Patent Nos. 10,900,034 and 10,767,175. As used in this paragraph, when the targeter or modulator nucleic acid is a combination of DNA and RNA, the nucleic acid as a whole is considered to be RNA, and the DNA nucleotides are considered to be modifications of the RNA, including 2'-H modifications of the ribose and, optionally, modifications of the nucleobases.
[0307] It is understood that in a dual guide nucleic acid system, the targeter nucleic acid and the modulator nucleic acid are not in the same nucleic acid, i.e., are not joined end-to-end by a traditional internucleotide bond, but are covalently conjugated to each other by one or more chemical modifications introduced into these nucleic acids, which can increase the stability of the double-stranded complex and / or improve other characteristics of the system.
[0308] III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA Engineered, non-naturally occurring systems such as those disclosed herein may be useful for targeting, editing, and / or modifying target nucleic acids, such as DNA (e.g., genomic DNA), in cells or organisms.
[0309] The present invention provides a method for cleaving a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, the method comprising contacting the target DNA with an engineered non-naturally occurring system disclosed herein, thereby effecting cleavage of the target DNA.
[0310] The present invention further provides a method for binding to a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, comprising contacting the target DNA with an engineered non-naturally occurring system disclosed herein, thereby causing binding of the system to the target DNA. This method may be useful, for example, for detecting the presence and / or location of a preselected target gene, e.g., when a component of the system (e.g., a Cas protein) comprises a detectable marker.
[0311] Also provided are methods for modifying a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, or a structure (e.g., a protein) associated with the target DNA (e.g., a histone protein in a chromosome), comprising contacting the target DNA with an engineered non-naturally occurring system disclosed herein, wherein the Cas protein comprises an effector domain or binds to an effector protein, thereby resulting in modification of the target DNA or the structure associated with the target DNA. The modification corresponds to the function of the effector domain or effector protein. The exemplary functions described in the "Cas Proteins" subsection in Section I above are applicable here.
[0312] The engineered, non-naturally occurring system can be contacted with a target nucleic acid as a complex. Thus, in certain embodiments, the method comprises contacting the target nucleic acid with a CRISPR-Cas complex comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein as disclosed herein. In certain embodiments, the Cas protein is a VA-type, VC-type, or VD-type Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a VA-type Cas protein (e.g., a Cas nuclease).
[0313] In certain embodiments, provided herein are methods for editing a human genome sequence at one of a group of preselected target loci, the method comprising delivering an engineered, non-naturally occurring system disclosed herein to a human cell, thereby resulting in editing of a genome sequence at the target locus in the human cell. In certain embodiments, provided herein are methods for detecting a human genome sequence at one of a group of preselected target loci, the method comprising delivering an engineered, non-naturally occurring system disclosed herein to a human cell, wherein a component of the system (e.g., a Cas protein) comprises a detectable marker, thereby detecting the target locus in the human cell. In certain embodiments, provided herein are methods for modifying a human chromosome at one of a group of preselected target loci, the method comprising delivering an engineered, non-naturally occurring system disclosed herein to a human cell, wherein the Cas protein comprises an effector domain or binds to an effector protein, thereby resulting in modification of a chromosome at the target locus in the human cell.
[0314] CRISPR-Cas complexes can be delivered to cells by introducing preformed ribonucleoprotein (RNP) complexes into cells. Alternatively, one or more components of the CRISPR-Cas complex can be expressed in cells. Exemplary delivery methods are known in the art and are described, for example, in U.S. Patent Nos. 8,697,359, 10,113,167, 10,570,418, 10,829,787, 11,118,194, and 11,125,739, and U.S. Patent Application Publication Nos. 2015 / 0344912, 2018 / 0119140, and 2018 / 0282763.
[0315] It is understood that contacting DNA (e.g., genomic DNA) in a cell with a CRISPR-Cas complex does not require delivery of all components of the complex to the cell. For example, one or more components may be naturally present in the cell. In certain embodiments, the cell (or its parent / ancestor cell) has been engineered to express a Cas protein, and a single guide nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a single guide nucleic acid), a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid), and / or a modulator nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid) are delivered to the cell. In certain embodiments, the cell (or its parent / ancestor cell) has been engineered to express a modulator nucleic acid, and a Cas protein (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a Cas protein) and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid) are delivered to the cell. In certain embodiments, a cell (or its parent / ancestor cell) has been engineered to express a Cas protein and a modulator nucleic acid, and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) is delivered to the cell.
[0316] In certain embodiments, the target DNA is in the genome of the target cell. Accordingly, the present invention also provides cells comprising the non-naturally occurring system or CRISPR expression system described herein. Furthermore, the present invention provides cells whose genomes have been modified by the CRISPR-Cas systems or complexes disclosed herein.
[0317] Target cells include bacterial cells (e.g., Escherichia coli (E. coli)), archaeal cells, cells of unicellular eukaryotes, plant cells, algal cells such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, etc., and fungal cells (e.g., Saccharomyces cerevisiae (S. The target cell may be a mitotic or post-mitotic cell from any organism, such as a yeast cell (e.g., Saccharomyces cervisiae), an animal cell, a cell from an invertebrate (e.g., a fruit fly, an enidarian, an echinoderm, a nematode, etc.), a cell from a vertebrate (e.g., a fish, an amphibian, a reptile, a bird, a mammal), a cell from a mammal, a cell from a rodent, or a cell from a human. Target cell types include, but are not limited to, stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells), somatic cells (e.g., fibroblasts, hematopoietic cells, T lymphocytes (e.g., CD8+ T lymphocytes), NK cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells), embryonic cells from any stage of embryo in vitro or in vivo (e.g., 1-cell, 2-cell, 4-cell, 8-cell stage zebrafish embryos). The cells may be derived from an established cell line or primary cells (i.e., cells and cell cultures derived from a subject and expanded in vitro for a limited number of culture passages). For example, a primary culture is a culture that may have been passaged 0, 1, 2, 4, 5, 10, or 15 times or less, but not enough times to survive a crisis stage. Typically, primary cell lines are maintained in vitro for fewer than 10 passages. When the cells are primary cells, they can be harvested from an individual by any method. For example, white blood cells can be harvested by apheresis, leukapheresis, or density gradient separation, while cells derived from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, or stomach can be harvested by biopsy.The harvested cells can be used directly or stored under frozen conditions using cryopreservation and thawed at a later time as is commonly known in the art.
[0318] A. Ribonucleoprotein (RNP) Delivery and "casRNA" Delivery The engineered, non-naturally occurring systems disclosed herein can be delivered to cells by any suitable method known in the art, including, but not limited to, ribonucleoprotein (RNP) delivery and "CasRNA" delivery, described below.
[0319] In certain embodiments, a CRISPR-Cas system comprising a single guide nucleic acid and a Cas protein, or a CRISPR-Cas system comprising a target nucleic acid, a modulator nucleic acid, and a Cas protein, can be combined into an RNP complex and then delivered to cells as a preformed complex.This method is suitable for actively modifying genetic or epigenetic information in cells for a limited period of time.For example, if a Cas protein has nuclease activity that modifies the genomic DNA of a cell, it only needs to maintain nuclease activity for a period that allows DNA cleavage, and prolonging nuclease activity may increase off-targeting.Similarly, once established, certain epigenetic modifications can be maintained in cells and inherited by daughter cells.
[0320] As used herein, "ribonucleoprotein" or "RNP" may refer to a complex comprising a nucleoprotein and a ribonucleic acid. As provided herein, "nucleoprotein" may refer to a protein capable of binding to nucleic acids (e.g., RNA, DNA). When a nucleoprotein binds to a ribonucleic acid, it can be referred to as a "ribonucleoprotein." The interaction between a ribonucleoprotein and a ribonucleic acid can be direct, for example, through a covalent bond, or indirect, for example, through a non-covalent bond (e.g., electrostatic interactions (e.g., ionic bonds, hydrogen bonds, halogen bonds), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion), ring stacking (pi effect), hydrophobic interactions, etc.). In certain embodiments, a ribonucleoprotein comprises an RNA-binding motif non-covalently bound to a ribonucleic acid. For example, a positively charged aromatic amino acid residue (e.g., lysine residue) in the RNA-binding motif can form an electrostatic interaction with the negative phosphate backbone of the RNA.
[0321] To ensure efficient loading of the Cas protein, a single guide nucleic acid, or a combination of targeter and modulator nucleic acids, can be provided in molar excess (e.g., at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold) relative to the Cas protein. In certain embodiments, the targeter and modulator nucleic acids are annealed under suitable conditions prior to complexing with the Cas protein. In other embodiments, the targeter nucleic acid, modulator nucleic acid, and Cas protein are mixed together directly to form the RNP.
[0322] Various delivery methods can be used to introduce the RNPs disclosed herein into cells. Exemplary delivery methods or vehicles include, but are not limited to, microinjection, liposomes (see, e.g., U.S. Pat. No. 10,829,787), e.g., molecular Trojan horse liposomes that deliver molecules across the blood-brain barrier (see, Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, microvesicles (e.g., exosomes and ARMMs), polycations, lipid:nucleic acid conjugates, electroporation, cell-penetrating peptides (see, U.S. Pat. No. 11,118,194), nanoparticles, nanowires (Shalek et al. (2012) Nano Letters, 12: Examples of RNP delivery methods include cell membrane perturbations (see U.S. Pat. No. 11,125,739; see U.S. Pat. No. 10,570,418; see U.S. Pat. No. 10,570,418). In certain embodiments, RNPs are delivered to cells by electroporation.
[0323] In certain embodiments, the CRISPR-Cas system is delivered to a cell by a "protocol," i.e., delivery of (a) a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, and (b) RNA (e.g., messenger RNA (mRNA)) encoding a Cas protein. The RNA encoding the Cas protein can be translated in the cell and form a complex with the single guide nucleic acid or the combination of a targeter nucleic acid and a modulator nucleic acid within the cell. As with the RNP approach, even if stability-enhancing modifications can be made in one or more RNAs, RNA has a limited half-life in the cell. Therefore, the "CasRNA" approach is suitable for active modification of genetic or epigenetic information in a cell for a limited period of time, such as DNA cleavage, and has the advantage of reducing off-targeting.
[0324] mRNA can be produced by transcription of DNA containing regulatory elements operably linked to a Cas coding sequence. Given that multiple copies of Cas protein can be generated from a single mRNA, a single guide nucleic acid, or targeter and modulator nucleic acids, are generally provided in molar excess (e.g., at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, or at least 100-fold) relative to the mRNA. In certain embodiments, the targeter and modulator nucleic acids are annealed under suitable conditions prior to delivery to cells. In other embodiments, the targeter and modulator nucleic acids are delivered to cells without being annealed in vitro.
[0325] A variety of delivery systems can be used to introduce the "CasRNA" system into cells. Non-limiting examples of delivery methods or vehicles include microinjection, biolistic particles, liposomes (see, e.g., U.S. Pat. No. 10,829,787), e.g., molecular Trojan horse liposomes that deliver molecules across the blood-brain barrier (see, e.g., Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, polycations, lipid:nucleic acid conjugates, electroporation, nanoparticles, nanowires (see, e.g., Shalek et al. (2012) Nano Letters, 12: 6498), exosomes, and cell membrane perturbation (e.g., by passing cells through a constriction in a microfluidic system, see, e.g., U.S. Pat. No. 11,125,739). A specific example of a "nucleic acid only" approach using electroporation is described in International (PCT) Publication No. WO2016 / 164356.
[0326] In certain embodiments, the CRISPR-Cas system is delivered to cells in the form of DNA containing (a) a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, and (b) a regulatory element operably linked to a Cas coding sequence. The DNA can be provided in a plasmid, a viral vector, or any other form described in the "CRISPR Expression System" subsection. Such delivery methods result in constitutive expression of the Cas protein in the target cell (e.g., when the DNA is maintained in the cell in an episomal vector or integrated into the genome), which may increase the risk of undesired off-targeting if the Cas protein has nuclease activity. Nevertheless, this approach is useful when the Cas protein contains a non-nuclease effector (e.g., a transcriptional activator or repressor). It is also useful for research purposes and plant genome editing.
[0327] B. CRISPR expression system Also provided herein is a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a guide nucleic acid disclosed herein.In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a single guide nucleic acid; this nucleic acid can constitute a CRISPR expression system by itself.In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid.In certain embodiments, the nucleic acid further comprises a nucleotide sequence encoding a modulator nucleic acid, wherein the nucleotide sequence encoding the modulator nucleic acid is operably linked to the same regulatory element as the nucleotide sequence encoding the targeter nucleic acid or a different regulatory element; this nucleic acid can constitute a CRISPR expression system by itself.
[0328] Additionally, the present invention provides a CRISPR expression system comprising: (a) a nucleic acid comprising a first regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid; and (b) a nucleic acid comprising a second regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid.
[0329] In certain embodiments, the CRISPR expression system further comprises a nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein, such as a Cas protein disclosed herein. In certain embodiments, the Cas protein is a VA-type, VC-type, or VD-type Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a VA-type Cas protein (e.g., a Cas nuclease).
[0330] As used in this context, the term "operably linked" may mean that the nucleotide sequence of interest is linked to regulatory elements in a manner that allows expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into a host cell).
[0331] The nucleic acids of the above-described CRISPR expression systems can be independently selected from a variety of nucleic acids, such as DNA (e.g., modified DNA) and RNA (e.g., modified RNA). In certain embodiments, the nucleic acid comprising a regulatory element operably linked to one or more nucleotide sequences encoding a guide nucleic acid is in the form of DNA. In certain embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of DNA. The third regulatory element may be a constitutive or inducible promoter that drives expression of the Cas protein. In other embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of RNA (e.g., mRNA).
[0332] The nucleic acid of the CRISPR expression system can be provided in one or more vectors. As used herein, the term "vector" may refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into cells, such as prokaryotic cells, eukaryotic cells, mammalian cells, or target tissues. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles, such as liposomes. Viral vector delivery systems include DNA and RNA viruses that have episomal or integrated genomes after delivery to cells. Gene therapy procedures are known in the art and are described in detail in Van Brunt (1988) Biotechnology, 6:1149; Anderson (1992) Science, 256:808; Nabel & Feigner (1993) TIBTECH, 11:211; Mitani & Caskey (1993) TIBTECH, 11:162; Dillon (1993) TIBTECH, 11:167; Miller (1992) Nature, 357:455; Vigne, (1995) Restorative Neurology and Neuroscience, 8:35; Kremer & Perricaudet (1995) British Medical Bulletin, 51:31; Haddada et al. (1995) Current Topics in Microbiology and Immunology, 199:297; Yu et al. (1994) Gene Therapy, 1: 13; and Doerfler and Bohm (eds.) (2012) The Molecular Repertoire of Adenoviruses II: Molecular Biology of Virus-Cell Interactions. In certain embodiments, at least one of the vectors is a DNA plasmid.In certain embodiments, at least one of the vectors is a viral vector (e.g., a retrovirus, adenovirus, or adeno-associated virus).
[0333] Certain vectors are capable of autonomous replication in host cells into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors and replication-defective viral vectors) do not autonomously replicate in host cells. However, certain vectors can integrate into the genome of a host cell, thereby replicating along with the host genome. Those skilled in the art will appreciate that various vectors may be suitable for various delivery methods and have various host tropisms, allowing them to select one or more vectors suitable for use.
[0334] As used herein, the term "regulatory element" may refer to transcriptional and / or translational control sequences, such as promoters, enhancers, transcription termination signals (e.g., polyadenylation signals), internal ribosome entry sites (IRES), proteolysis signals, and the like, that effect and / or regulate transcription of a non-coding sequence (e.g., a targeter nucleic acid or modulator nucleic acid) or a coding sequence (e.g., a Cas protein) and / or regulate translation of an encoded polypeptide. Such regulatory elements are described, for example, in Goeddel, Gene Expression Technology: Methods in Enzymology, 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). A tissue-specific promoter can direct expression primarily in a desired tissue of interest, such as muscle, neurons, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). Regulatory elements may or may not be tissue- or cell-type-specific and can also direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner. In certain embodiments, the vector includes one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters.Examples of Pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally containing the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally containing the CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. Also encompassed by the term "regulatory element" are the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (see Takebe et al. (1988) Mol. Cell. Biol., 8:466); enhancer elements such as the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (see O'Hare et al. (1981) Proc. Natl. Acad. Sci. USA, 78:1527). It will be appreciated by those skilled in the art that the design of the expression vector can depend on factors such as the choice of the host cell to be transformed, the level of expression desired, etc. The vectors can be introduced into host cells to produce the transcripts, proteins, or peptides (e.g., CRISPR transcripts, proteins, enzymes, mutant forms thereof, or fusion proteins thereof), including fusion proteins or peptides, encoded by the nucleic acids described herein.
[0335] In certain embodiments, the nucleotide sequence encoding the Cas protein is codon-optimized for expression in prokaryotic cells, such as Escherichia coli, eukaryotic host cells, such as yeast cells (e.g., Saccharomyces cerevisiae), mammalian cells (e.g., mouse cells, rat cells, or human cells), or plant cells. Various species exhibit specific biases for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often associated with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, inter alia, on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The preference of a selected tRNA in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the "Codon Usage Database" available at kazusa.or.jp / codon / , and these tables can be adapted in several ways (see Nakamura et al. (2000) Nucl. Acids Res., 28:292). Computer algorithms are also available for codon-optimizing particular sequences for expression in particular host cells, such as GeneForge (Aptagen; Jacobus, Pa.). In certain embodiments, codon optimization facilitates or improves expression of the Cas protein in the host cell.
[0336] C. Donor template Cleavage of a target nucleotide sequence in a cell's genome by a CRISPR-Cas system or complex can activate DNA damage pathways, allowing the cleaved DNA fragments to be rejoined by NHEJ or HDR. HDR requires an endogenous or exogenous repair template to transfer sequence information from the repair template to the target.
[0337] In certain embodiments, the engineered non-naturally occurring system or CRISPR expression system further comprises a donor template. As used herein, the term "donor template" may refer to a nucleic acid designed to serve as a repair template at or near a target nucleotide sequence when introduced into a cell or organism. In certain embodiments, the donor template is complementary to a polynucleotide comprising the target nucleotide sequence or a portion thereof. When optimally aligned, the donor template may overlap with one or more nucleotides (e.g., about 1, 5, 10, 15, 20, 25, 30, 35, 40 or more nucleotides) of the target nucleotide sequence. The nucleotide sequence of the donor template is typically not identical to the genomic sequence it replaces. Rather, the donor template may contain one or more substitutions, insertions, deletions, inversions, or rearrangements with respect to the genomic sequence, as long as there is sufficient homology to support homology-directed repair. In certain embodiments, the donor template comprises a non-homologous sequence flanked by two homologous regions (i.e., homology arms) such that homology-directed repair between the target DNA region and the two flanking sequences results in insertion of the non-homologous sequence at the target region. In certain embodiments, the donor template comprises a non-homologous sequence located between the two homology arms, the non-homologous sequence being 10 to 100 nucleotides, 50 to 500 nucleotides, 100 to 1,000 nucleotides, 200 to 2,000 nucleotides, or 500 to 5,000 nucleotides in length.
[0338] Generally, the homologous region of the donor template has at least 50% sequence identity to the genomic sequence with which recombination is desired. The homology arms are designed or selected so that they can recombine with nucleotide sequences flanking the target nucleotide sequence under intracellular conditions. In certain embodiments, when HDR of the non-target strand is desired, the donor template comprises a first homology arm homologous to a sequence 5' to the target nucleotide sequence and a second homology arm homologous to a sequence 3' to the target nucleotide sequence. In certain embodiments, the first homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 5' to the target nucleotide sequence. In certain embodiments, the second homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 3' to the target nucleotide sequence. In certain embodiments, the polynucleotides comprising the donor template sequence and the target nucleotide sequence are optimally aligned, and the nearest nucleotide of the donor template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, or more nucleotides of the target nucleotide sequence.
[0339] In certain embodiments, the donor template further comprises an engineered sequence that is not homologous to the sequence to be repaired. Such engineered sequence may carry a barcode and / or sequence that can hybridize to the donor template recruitment sequence disclosed herein.
[0340] In certain embodiments, the donor template further comprises one or more mutations in the genome sequence, wherein the one or more mutations reduce or prevent cleavage by the same CRISPR-Cas system of the donor template or of a modified genome sequence containing at least a portion of the integrated donor template sequence. In certain embodiments, in the donor template, a PAM adjacent to the target nucleotide sequence and recognized by a Cas nuclease is mutated to a sequence not recognized by the same Cas nuclease. In certain embodiments, in the donor template, the target nucleotide sequence (e.g., a seed region) is mutated. In certain embodiments, the one or more mutations are silent with respect to the reading frame of the protein-coding sequence encompassing the mutation site.
[0341] Donor template can be provided to cell as single-stranded DNA, single-stranded RNA, double-stranded DNA or double-stranded RNA.It is understood that CRISPR-Cas system such as the system disclosed herein can have the nuclease activity of cutting target strand, non-target strand or both.When HDR of target strand is desired, donor template also contemplates having the nucleic acid sequence complementary to target strand.
[0342] The donor template can be introduced into cells in a linear or circular form. When introduced in a linear form, the ends of the donor template can be protected (e.g., from exonuclease degradation) by methods known to those skilled in the art. For example, one or more dideoxynucleotide residues can be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides can be ligated to one or both ends (see, e.g., Chang et al. (1987) Proc. Natl. Acad Sci USA, 84: 4959; Nehls et al. (1996) Science, 272: 886; also see the chemical modifications for increasing RNA stability and / or specificity disclosed above). Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino groups and the use of modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. As an alternative to protecting the ends of the linear donor template, additional lengths of sequence outside the homologous regions can be included that can be degraded without affecting recombination.
[0343] The donor template may be a component of a vector described herein, contained in a separate vector, or provided as a separate polynucleotide, such as an oligonucleotide, linear polynucleotide, or synthetic polynucleotide. In certain embodiments, the donor template is DNA. In certain embodiments, the donor template is in the same nucleic acid as sequences encoding a single guide nucleic acid, a targeter nucleic acid, a modulator nucleic acid, and / or a Cas protein, as appropriate. In certain embodiments, the donor template is provided in a separate nucleic acid. The donor template polynucleotide may be of any suitable length, such as about 50, 75, 100, 150, 200, 500, 1000, 2000, 3000, 4000 nucleotides or more, or at least about 50, 75, 100, 150, 200, 500, 1000, 2000, 3000, 4000 nucleotides or more in length.
[0344] The donor template can be introduced into cells as an isolated nucleic acid. Alternatively, the donor template can be introduced into cells as part of a vector (e.g., a plasmid) that has additional sequences, such as an origin of replication, a promoter, and a gene encoding antibiotic resistance, that are not intended for insertion into the DNA region of interest. Alternatively, the donor template can be delivered by a virus (e.g., adenovirus, adeno-associated virus (AAV)). In certain embodiments, the donor template is introduced as an AAV, e.g., a pseudotyped AAV. One of skill in the art can select the capsid protein of the AAV based on the tropism and target cell type of the AAV. For example, in certain embodiments, the donor template is introduced into hepatocytes as AAV8 or AAV9. In certain embodiments, the donor template is introduced into hematopoietic stem cells, hematopoietic progenitor cells, or T lymphocytes (e.g., CD8) as AAV6 or AAVHSC (see U.S. Pat. No. 9,890,396). + It is understood that the sequence of the capsid protein (VP1, VP2, or VP3) can be modified from the wild-type AAV capsid protein, for example, by having at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the wild-type AAV capsid sequence.
[0345] The donor template can be delivered to cells (e.g., primary cells) by various delivery methods, such as viral or non-viral methods disclosed herein. In certain embodiments, a non-viral donor template is introduced into a target cell as naked nucleic acid or in a complex with liposomes or poloxamers. In certain embodiments, a non-viral donor template is introduced into a target cell by electroporation. In other embodiments, a viral donor template is introduced into a target cell by infection. The engineered, non-naturally occurring system can be delivered before, after, or simultaneously with the donor template (see International (PCT) Application Publication No. WO 2017 / 053729). One of skill in the art can select the appropriate timing based on the form of delivery (e.g., taking into account the time required for transcription and translation of the RNA and protein components) and the half-life of the molecule in the cell. In certain embodiments, when the CRISPR-Cas system including the Cas proteins is delivered by electroporation (e.g., as an RNP), the donor template (e.g., as an AAV) is introduced into the cell within 4 hours (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 90, 120, 150, 180, 210, or 240 minutes) after introduction of the engineered, non-naturally occurring system.
[0346] In certain embodiments, the donor template is covalently conjugated to the modulator nucleic acid. Suitable covalent bonds for this conjugation are known in the art and are described, for example, in U.S. Pat. No. 9,982,278 and Savic et al. (2018) eLife 7:e33761. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., at the 5' end of the modulator nucleic acid) via an internucleotide bond. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., at the 5' end of the modulator nucleic acid) via a linker.
[0347] In certain embodiments, the donor template may comprise any nucleic acid chemistry. In certain embodiments, the donor template may comprise DNA and / or RNA nucleotides. In certain embodiments, the donor template may comprise single-stranded DNA, linear single-stranded RNA, linear double-stranded DNA, linear double-stranded RNA, circular single-stranded DNA, circular single-stranded RNA, circular double-stranded DNA, or circular double-stranded RNA. In certain embodiments, the donor template comprises a mutation in the PAM sequence that partially or completely abolishes RNP binding to DNA. In certain embodiments, the donor template comprises at least 0.05, 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, or 4, and / or 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, 4, or 5 μg μL -1 For example, 0.01 to 5 μg μL -1 In certain embodiments, the donor template comprises one or more promoters. In certain embodiments, the donor template comprises a promoter that shares at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99.5% sequence identity with any one of SEQ ID NOs: 78-85 in Table 6.
[0348] [Table 6A]
[0349] [Table 6B]
[0350] [Table 6C]
[0351] [Table 6D]
[0352] D. Efficiency and Specificity Engineered, non-naturally occurring systems can be evaluated in terms of efficiency and / or specificity in nucleic acid targeting, cleavage, or modification.
[0353] In certain embodiments, the engineered, non-naturally occurring system has high efficiency. For example, in certain embodiments, at least 1, 1.5, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of a population of nucleic acids having a target nucleotide sequence and a cognate PAM are targeted, cleaved, or modified when contacted with the engineered, non-naturally occurring system. In certain embodiments, when the engineered, non-naturally occurring system is delivered to cells, the genomes of at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of the cells' population are targeted, ablated, or modified.
[0354] For a given spacer sequence, the occurrence of on-target events and the occurrence of off-target events have generally been observed to be correlated. For certain therapeutic purposes, a lower on-target efficiency may be acceptable, and a low off-target frequency is more desirable. For example, when editing or modifying the proliferation of cells delivered to a subject and grown in vivo, there is a low tolerance for off-target events. Prior to delivery, evaluation of on-target and off-target events allows for the selection of one or more colonies with desired editing or modification and lacking undesired editing or modification. Nevertheless, the on-target efficiency may need to meet certain criteria to be suitable for therapeutic use. High editing efficiency in standard CRISPR-Cas systems allows for tuning of the system, for example, by reducing binding of guide nucleic acids to Cas proteins without losing therapeutic applicability.
[0355] In certain embodiments, when a population of nucleic acids having a target nucleotide sequence and a cognate PAM is contacted with an engineered, naturally occurring system disclosed herein, the frequency of off-target events (e.g., targeting, cleavage, or modification, depending on the function of the CRISPR-Cas system) is reduced. Methods for assessing off-target events are summarized in Lazzarotto et al. (2018) Nat. Protoc. 13(11): 2615-42 and include discovery of Cas off-targets in situ by sequencing (DISCOVER-seq), disclosed in Wienert et al. (2019) Science 364(6437): 286-89; genome-wide unbiased identification of double-strand breaks (DSBs) enabled by sequencing (GUIDE-seq), disclosed in Kleinstiver et al. (2016) Nat. Biotech. 34: 869-74; and circularization for in vitro reporting of cleavage efficacy by sequencing (CIRCLE-seq), described in Kocak et al. (2019) Nat. Biotech. 37: 657-66. In certain embodiments, an off-target event comprises targeting, cleavage, or modification at a given off-target locus (e.g., the locus that exhibits the highest incidence of detected off-target events). In certain embodiments, an off-target event comprises targeting, cleavage, or modification at all loci that collectively exhibit detectable off-target events.
[0356] In certain embodiments, genomic mutations are 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.00 Detected in less than 7%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% of the cells (in total). In certain embodiments, the ratio of the percentage of cells having an on-target event to the percentage of cells having any off-target event (e.g., the ratio of the percentage of cells having an on-target editing event to the percentage of cells having a mutation at any off-target locus) is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000. It is understood that genetic alterations may be present in a population of cells, for example, due to spontaneous mutations, and such mutations are not included as off-target events.
[0357] E. Multiplexing The methods for targeting, editing, and / or modifying genomic DNA disclosed herein can be performed in a variety of ways. For example, a library of targeter nucleic acids can be used to target multiple loci; a library of donor templates can be used to generate multiple insertions, deletions, and / or substitutions. Multiplex assays can be performed in screening methods, where each separate cell culture (e.g., in the wells of a 96-well or 384-well plate) is exposed to a different guide nucleic acid with a different targeter stem sequence and / or a different donor template. Multiplex assays can also be performed in selection methods, where cell cultures are exposed to a mixed population of different guide nucleic acids and / or donor templates, and cells with desired properties (e.g., functionality) are enriched or selected by advantageous survival or growth, resistance to a particular drug, expression of a detectable protein (e.g., a fluorescent protein detectable by flow cytometry), etc.
[0358] In certain embodiments, multiple guide nucleic acids and / or multiple donor templates are designed for saturation editing.For example, in certain embodiments, each nucleotide position in the target sequence is systematically modified with all four traditional bases, A, T, G, and C.In other embodiments, at least one sequence in each gene from a pool of target genes is modified, for example, according to a CRISPR design algorithm.In certain embodiments, each sequence from a pool of target exogenous elements (e.g., protein-coding sequences, non-protein-coding sequences, regulatory elements) is inserted into one or more given loci of the genome.
[0359] It is understood that multiplexing methods suitable for performing screening or selection methods, typically performed for research purposes, may differ from methods suitable for therapeutic purposes. For example, constitutive expression of certain elements (e.g., Cas nucleases and / or guide nucleic acids) may be undesirable for therapeutic purposes due to the increased potential for off-targeting. Conversely, constitutive expression of Cas nucleases and / or guide nucleic acids may be desirable for research purposes. For example, constitutive expression provides a greater opportunity to introduce other elements. If a stable cell line is established for constitutive expression, the number of exogenous elements that need to be simultaneously delivered to a single cell is also reduced. Thus, constitutive expression of certain elements can increase the efficiency and reduce the complexity of the screening or selection process. Inducible expression of certain elements of the systems disclosed herein can also be used for research purposes, with similar advantages. Expression can be induced by exogenous agents (e.g., small molecules) or by endogenous molecules or complexes present in a particular cell type (e.g., at a particular differentiation stage). Methods known in the art, such as those described herein, can be used to constitutively or inducibly express one or more elements. For example, the specificity of a CRISPR nuclease is determined at least in part by the uniqueness of the spacer (along with the proximity of the spacer sequence to the required PAM), and its off-target score can be calculated using algorithms such as crispr.mit.edu (Hsu et al. (2013) Nat. Biotech. 31: 827-832). The highest possible score is 100, indicating high specificity and a low probability of off-targeting. Because our SHS library targets intergenic regions, algorithms for gRNA prediction should be able to align with repetitive regions and low-complexity sequences.
[0360] It is further understood that although it may be necessary to introduce multiple elements—a single guide nucleic acid and Cas protein; or a targeter nucleic acid, a modulator nucleic acid, and a Cas protein—these elements can be delivered to a cell as a single complex of preformed RNPs. Thus, efficiency of the screening or selection process can also be achieved by preassembling multiple RNP complexes in a multiplexed manner.
[0361] In certain embodiments, the methods disclosed herein further comprise identifying a guide nucleic acid, a Cas protein, a donor template, or a combination of two or more of these elements from the screening or selection process. To facilitate identification, for example, a set of barcodes can be used in the donor template between two homology arms. In certain embodiments, the methods further comprise harvesting the population of cells; selectively amplifying genomic DNA or RNA containing the target nucleotide sequence and / or barcode; and / or sequencing the selectively amplified genomic DNA or RNA sample and / or barcode.
[0362] Additionally, the present invention provides libraries comprising a plurality of guide nucleic acids, such as a plurality of guide nucleic acids disclosed herein. In another aspect, the present invention provides libraries comprising a plurality of nucleic acids, each comprising a regulatory element operably linked to a different guide nucleic acid, such as a different guide nucleic acid disclosed herein. These libraries can be used with one or more Cas proteins or Cas-encoding nucleic acids, such as those disclosed herein, and / or one or more donor templates, such as those disclosed herein, for screening or selection methods.
[0363] F. Genomic Safe Harbor Genome engineering is a research field that aims to modify the genes of living organisms, particularly to improve our understanding of gene function and develop methods for genome engineering to treat genetic or acquired diseases. To modify the genome of a target cell, those skilled in the art use one or more available means to introduce changes into the genome at a target position to modify the sequence of a target polynucleotide, such as a target gene, in a desired manner, for example, by modulating gene expression, modulating gene sequence, removing gene sequence, or introducing exogenous DNA, such as a transgene. Efficient transgene insertion can be achieved by non-precision methods including, but not limited to, viral vectors such as retroviral vectors, e.g., adeno-associated virus (AAV), or precision methods including, but not limited to, inducible nucleases such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), homing endonucleases, e.g., restriction endonucleases, or nucleic acid-guided nucleases, e.g., CRISPR-cas, e.g., Cas9 and Cas12a and engineered versions thereof.
[0364] For example, exogenous genes, e.g., transgenes, inserted into the genome of target human cells randomly by retroviral vectors or in a targeted manner by the action of nucleic acid-guided nucleases, e.g., Cas, can interact with other genomic elements in unpredictable ways. Due to the complex transcriptional regulation of genes in mammalian cells through a network of cis- and trans-regulatory elements, such as proximal and distal enhancers and multiple transcription factors, attempts to alter the default genome structure by integrating exogenous DNA, e.g., transgenes, or synthetic sequences, can affect the expression of the transgene itself, resulting in complete attenuation or complete silencing, and / or expression of both nearby and distant endogenous genes that can alter cellular behavior in dramatic ways, i.e., impairing safety checkpoints possessed by healthy cells, including, for example, dysregulation of the expression of critical genes, such as oncogenes and tumor suppressor genes, which can promote clonal amplification or malignant transformation of the host.
[0365] It has been shown that gene integration next to the regulatory element of proto-oncogene causes oncogenic transformation, which is particularly important when manipulating cells for therapeutic applications.Therefore, it is desirable to identify suitable target polynucleotides, including target nucleotide sequences in the human genome, where the insertion of transgene leads to the appropriate expression of the transgene without disrupting neighboring genes.In particular, for gene and cell therapy applications, it is desirable to identify suitable target polynucleotides, including target nucleotide sequences in the human genome, where the insertion of transgene leads to the sufficient expression of the transgene in therapeutic cells, such as T cells, for example, CAR T cells; or progenitor cells, such as stem cells, for example, hematopoietic stem cells, without causing malignant transformation or any other disruption that would be harmful to individuals after implantation.
[0366] Expression of an exogenous gene, e.g., a transgene, in a desired cell type and / or developmental / differentiation stage relies on integration from a candidate locus into a suitable target polynucleotide containing a target nucleotide sequence that results in sufficient expression to a sufficient degree for the intended purpose. Expression from a specific genomic site can be influenced by many factors, including, but not limited to, cell type and differentiation stage, as well as changes in chromatin structure, since one or more components of the target polynucleotide are activated during differentiation while others are silenced. Therefore, it is desirable to identify a suitable target polynucleotide containing a target nucleotide sequence in the human genome where insertion of exogenous DNA, e.g., a transgene, results in sufficient expression in target human cells, and in the case of stem cells, expression is maintained at a sufficient level through (1) differentiation and (2) clonal expansion. The present disclosure provides a significant advance in the ability to manipulate the human genome by providing compositions and methods for targeting and delivering exogenous genes, e.g., transgenes, to a suitable target polynucleotide containing a target nucleotide sequence.
[0367] Compositions and methods for genome manipulation are provided herein. Certain embodiments include compositions. Certain embodiments include compositions for editing a genome. Embodiments disclosed herein relate to novel guide nucleic acids (gNAs), e.g., gRNAs, that are complementary to a target nucleotide sequence in a target polynucleotide. As used herein, "target polynucleotide" includes a polynucleotide in which a target nucleotide sequence is located. As used herein, "target nucleotide sequence" includes a sequence to which a guide sequence can bind, e.g., is complementary, where binding between the target nucleotide sequence and the guide sequence can enable the activity of a nucleic acid-guided nuclease complex. Further embodiments disclosed herein relate to novel gNAs, e.g., gRNAs, that are complementary to a target nucleotide sequence in a target polynucleotide, where insertion of exogenous DNA, e.g., a transgene, does not negatively affect a cell, e.g., does not significantly affect the expression of one or more endogenous genes or result in malignant transformation of the cell. In further embodiments disclosed herein, gene expression exhibited in a human target cell is maintained by differentiation of the human target cell and / or propagation in one or more progeny cells at a level sufficient for the cell's ultimate use. Certain embodiments disclosed herein relate to novel nucleic acid-guided nuclease complexes, e.g., RNPs, such as Cas bound to gNAs, which are complementary to a target nucleotide sequence within a target polynucleotide and hydrolyze (also referred to as cleave or cut) the phosphodiester backbone at at least one position on at least one strand of the target polynucleotide. Certain embodiments disclosed herein relate to methods for selecting and using gNAs, e.g., gRNAs, for genome engineering. Certain embodiments relate to methods using gNAs that are complementary to a target nucleotide sequence within a target polynucleotide, synthesizing gNAs and nucleic acid-guided nucleases, and / or combining nucleic acid-guided nucleases with gNAs to form nucleic acid-guided nuclease complexes, e.g., RNPs. Certain embodiments disclosed herein relate to methods.Certain embodiments disclosed herein relate to methods for genome engineering, including methods in which a nucleic acid-guided nuclease complex, e.g., an RNP, is introduced, e.g., transfected, into a human target cell along with a donor template, e.g., exogenous DNA, e.g., a transgene, wherein the nucleic acid-guided nuclease makes a backbone cleavage at least one position in at least one of the strands of the target polynucleotide, and the donor template is used to repair the cleaved target polynucleotide, thereby introducing at least a portion of the donor template into the target polynucleotide. As used herein, "exogenous DNA" or "transgene" includes any gene, natural or synthetic, introduced into the genome of a non-endogenous organism or cell. The transgene may or may not retain the ability to express and / or produce RNA or protein in the human target cell. The transgene may or may not alter the resulting phenotype of the human target cell. Certain embodiments include human target cells, e.g., eukaryotic cells, e.g., mammalian cells such as human cells, e.g., stem cells or immune cells, produced by a method in which a nucleic acid-guided nu...
Claims
1. A composition comprising modified human cells, wherein the cells are a. A genome modification comprising a first portion of a first polynucleotide, wherein the first portion comprises a transgene inserted into a site containing a TCR subunit gene, thereby partially or completely inactivating the TCR subunit gene and expressing the transgene. A composition containing the following:
2. Modified human cells, b. A genome modification comprising a second polynucleotide encoding a fusion protein of B2M and HLA-E or HLA-G inserted into the B2M gene, thereby partially or completely inactivating endogenous B2M and expressing the fusion protein; and / or c. Genomic modification of the CIITA gene, wherein the CIITA gene is partially or completely inactivated. The composition according to claim 1, further comprising:
3. The composition according to claim 1 or 2, wherein the TCR subunit gene is completely inactivated.
4. The composition according to claim 2, wherein the endogenous B2M gene and / or CIITA gene are completely inactivated.
5. The composition according to claim 1 or 2, wherein the TCR subunit gene comprises the TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z gene.
6. The composition according to claim 1 or 2, wherein the TCR subunit gene comprises the CD3E gene.
7. The composition according to claim 1 or 2, wherein the introduced gene comprises a polynucleotide encoding a chimeric antigen receptor (CAR), a dual CAR protein, or a portion thereof.
8. The composition according to claim 7, wherein CAR or a portion thereof comprises a polypeptide bound to at least one of B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, and CD3 zeta.
9. The composition according to claim 8, comprising a polypeptide in which CAR or a portion thereof is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% identical to any one of the amino acid sequences of SEQ ID NOs. 86 to 124.
10. The composition according to claim 1 or 2, wherein the modified human cells are stem cells or immune cells.
11. The composition according to claim 10, wherein the modified human cells are stem cells selected from hematopoietic stem cells, CD34+ stem cells, and induced pluripotent stem cells (iPSCs).
12. The composition according to claim 10, wherein the modified human cells are immune cells selected from neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, and lymphocytes.
13. The composition according to claim 1 or 2, wherein the modified human cell is a T cell, preferably a CAR-containing T cell.
14. A pharmaceutical product comprising the composition described in claim 1 or 2.
15. A method for manipulating human cells, a. A step of introducing a genome modification, which includes inserting a transgene into a TCR subunit gene, wherein the modification partially or completely inactivates the TCR subunit gene. Methods that include...
16. b. A step of introducing a genome modification, comprising introducing a fusion protein of B2M and HLA-E or HLA-G into the B2M gene, wherein the modification partially or completely inactivates endogenous B2M; and / or c. A process of introducing a genome modification that partially or completely inactivates the CIITA gene. The method according to claim 15, further comprising:
17. The method according to claim 15 or 16, wherein the TCR subunit gene comprises the TRAC, TRBC, CD3E, CD3D, CD3G, or CD3Z gene.
18. The method according to claim 15 or 16, wherein the TCR subunit gene comprises the CD3E gene.
19. The cells are manipulated by delivering a composition to them that includes a VA-type Cas nuclease or a polynucleotide encoding the nuclease, and a guide nucleic acid, wherein the guide nucleic acid is: (1) Targeter nucleic acids including a targeter stem sequence and a spacer sequence; and (2) A modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence, and optionally a 5' sequence. The method according to claim 15 or 16, including the method described in claim 15 or 16.