Compositions and methods for genome editing
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- セリントラ·セラピューティクス·エスア
- Filing Date
- 2023-07-25
- Publication Date
- 2026-07-21
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Statement of Federally Funded Research none.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application Nos. 63 / 392,041, filed July 25, 2022, 63 / 412,772, filed October 3, 2022, and 63 / 521,084, filed June 14, 2023, the disclosures of which are incorporated herein by reference in their entireties for all purposes. [Background technology]
[0003] CRISPR-Cas systems have been engineered for a variety of purposes, including genomic DNA cleavage, base editing, epigenome editing, and genome imaging. Although considerable development has been achieved, there remains a need for novel and useful CRISPR-Cas systems as powerful and precise genome targeting tools. The invention disclosed herein includes CRISPR-Cas-based compositions and methods for high transgene integration and expression efficiency, along with high post-transfection cell viability in eukaryotic cells.
[0004] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Summary of the Invention [Means for solving the problem]
[0005] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which: [Brief explanation of the drawings]
[0006] [Figure 1A] FIG. 1 shows a schematic diagram illustrating the structure of an exemplary single-guide VA-type CRISPR system. [Figure 1B] FIG. 1 is a schematic diagram showing the structure of an exemplary dual-guide VA-CRISPR system. [Figure 2A] A series of schematic diagrams are shown illustrating the incorporation of protecting groups (e.g., protective nucleotide sequences or chemical modifications) (Figure 2A), donor template recruitment sequences (Figure 2B), and editing enhancers (Figure 2C) into a Type VA CRISPR-Cas system. While these additional elements are shown with respect to a dual-guide Type VA CRISPR system, it is understood that they may also be present in other CRISPR systems, including single-guide Type VA CRISPR systems, single-guide Type II CRISPR systems, or dual-guide Type II CRISPR systems. [Figure 2B] A series of schematic diagrams are shown illustrating the incorporation of protecting groups (e.g., protective nucleotide sequences or chemical modifications) (Figure 2A), donor template recruitment sequences (Figure 2B), and editing enhancers (Figure 2C) into a Type VA CRISPR-Cas system. While these additional elements are shown with respect to a dual-guide Type VA CRISPR system, it is understood that they may also be present in other CRISPR systems, including single-guide Type VA CRISPR systems, single-guide Type II CRISPR systems, or dual-guide Type II CRISPR systems. [Figure 2C]A series of schematic diagrams are shown illustrating the incorporation of protecting groups (e.g., protective nucleotide sequences or chemical modifications) (Figure 2A), donor template recruitment sequences (Figure 2B), and editing enhancers (Figure 2C) into a Type VA CRISPR-Cas system. While these additional elements are shown with respect to a dual-guide Type VA CRISPR system, it is understood that they may also be present in other CRISPR systems, including single-guide Type VA CRISPR systems, single-guide Type II CRISPR systems, or dual-guide Type II CRISPR systems. [Figure 3] 1 shows a schematic diagram of a VA-type nucleic acid-guided nuclease containing dual guide nucleic acids. [Figure 4] A schematic diagram of a polynucleotide containing upstream and downstream PAMs and a target nucleotide sequence in a P1 + T1 + D + T2 - P2 - configuration (401) flanking a donor template is shown. The polynucleotide can be circular, e.g., a plasmid, or linear double-stranded DNA with or without covalently closed ends. The diagram also illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (402). [Figure 5] A schematic diagram of a polynucleotide containing upstream and downstream PAMs and a target nucleotide sequence flanking a donor template in a T1-P1-D+T2-P2-configuration (501) is shown. The polynucleotide can be circular, e.g., a plasmid, or linear double-stranded DNA with or without covalently closed ends. The diagram also illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (502). The processed donor template (502) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 6]A schematic diagram of a polynucleotide containing upstream and downstream PAMs and a target nucleotide sequence flanking a donor template in a P1 + T1 + D + P2 + T2 + configuration (601) is shown. The polynucleotide can be circular, e.g., a plasmid, or linear double-stranded DNA with or without covalently closed ends. The diagram also illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (602). The processed donor template (602) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 7] A schematic diagram of a polynucleotide containing upstream and downstream PAMs and a target nucleotide sequence flanking a donor template in a T1 -P1 -D+P2 +T2 + configuration (701) is shown. The polynucleotide can be circular, e.g., a plasmid, or linear double-stranded DNA with or without covalently closed ends. The diagram also illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (702). The processed donor template (702) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 8] An exemplary plasmid is shown. [Figure 9] Cell viability (y-axis) data after knock-in of heterologous DNA using various donor templates is shown. [Figure 10] Gene editing data measured by % edits with complete HDR as a percentage of total edits (y-axis) at the TGFBR2 locus are shown using MAD7 complexed with gR007 and various donor templates. [Figure 11]Gene editing data measured by % edits with complete HDR as a percentage of total edits (y-axis) at the TGFBR2 locus are shown using MAD7 complexed with gR008 and various donor templates. [Figure 12] 1 shows gene editing data measured by % edits with complete HDR as a percentage of total edits at the FAS locus using MAD7 complexed with gR94 and various donor templates. [Figure 13] Gene editing data measured by GFP fluorescence and cell viability after treatment with various homology-dependent repair templates are shown. [Figure 14] A schematic diagram of a linear double-stranded polynucleotide with covalently closed ends, including upstream and downstream PAMs and a target nucleotide sequence in a P1 + T1 + D + T2 - P2 - configuration (1401) flanking a donor template, is shown. The diagram further illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (1402). The processed donor template (1402) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 15] A schematic diagram of a linear double-stranded polynucleotide with covalently closed ends, including upstream and downstream PAMs and a target nucleotide sequence in a T1-P1-D+T2-P2- configuration (1501) flanking a donor template, is shown. The diagram further illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (1502). The processed donor template (1502) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 16]A schematic diagram of a linear double-stranded polynucleotide with covalently closed ends, including upstream and downstream PAMs and a target nucleotide sequence flanking a donor template in a P1 + T1 + D + P2 + T2 + configuration (1601). The diagram further illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (1602). The processed donor template (1602) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 17] A schematic diagram of a linear double-stranded polynucleotide with covalently closed ends, including upstream and downstream PAMs and a target nucleotide sequence flanking a donor template in a T1-P1-D+P2+T2+ configuration (1701). The diagram further illustrates a method for cleaving a polynucleotide using a nucleic acid-guided nuclease. The nuclease can cleave at one or more target sites on the polynucleotide. The diagram also illustrates a processed donor template (1702). The processed donor template (1702) can include a bound nucleic acid-guided nuclease containing a functional domain, such as one or more nuclear localization signals (NLSs). The NLS can improve transport of the polynucleotide into the nucleus. [Figure 18] 1 shows a schematic diagram of a polynucleotide comprising one or more additional PAMs and a target site. [Figure 19] Figure 1 shows gene editing efficiency for miniplasmids and linear double-stranded DNA with covalently closed ends, measured by % CAR expression (y-axis), as a function of ssODN concentration, donor template concentration, gRNA:nuclease ratio, and nuclease concentration. [Figure 20] Figure 1 shows gene editing efficiency for miniplasmids and linear double-stranded DNA with covalently closed ends, as measured by % CAR expression (y-axis) as a function of template concentration, nuclease concentration, and ssODN concentration. [Figure 21]Figure 1 shows cell viability for miniplasmids and linear double-stranded DNA with covalently closed ends, as measured by the % viable cells (y-axis) in treated cell populations as a function of template concentration, nuclease concentration, and ssODN concentration. DETAILED DESCRIPTION OF THE INVENTION
[0007] overview I. Compositions and Methods of Use of Polynucleotides and Their Binding Polypeptides A. Polynucleotides B. Polynucleotides and their associated polypeptides C. Methods of Use of Polynucleotides and Binding Polypeptides II. Engineered non-native dual-guide CRISPR-cas systems A. Cas proteins B. Guide Nucleic Acid C.gNA modification III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA A. Ribonucleoprotein (RNP) Delivery and "cas RNA" Delivery B. CRISPR Expression System C. Donor Template D. Efficiency and Specificity E. Multiplex Method F. Genomic Safe Harbor IV. Pharmaceutical Compositions V. Therapeutic Use A. Gene Therapy B. Immune cell manipulation VI. Kit VII. Embodiments VIII. Working Examples IX. Equivalents
[0008] I. Compositions and Methods of Use of Polynucleotides and Their Binding Polypeptides Genome targeting technologies have advanced in recent years. For example, specific loci in genomic DNA can be targeted, edited, or otherwise modified by designer meganucleases, zinc finger nucleases, or transcription activator-like effectors (TALEs). Furthermore, the CRISPR-Cas system of bacterial and archaeal adaptive immunity has been adapted for precise targeting of genomic DNA in eukaryotic cells. Compared to earlier generations of genome editing tools, the CRISPR-Cas system is easy to set up, scalable, and applicable to targeting multiple locations within eukaryotic genomes, thereby providing a major resource for new applications in genome engineering. In certain embodiments, provided herein are compositions, methods, and / or kits for genome engineering. In certain embodiments, provided herein are compositions, methods, and / or kits for genome engineering of eukaryotic cells. In certain embodiments, provided herein are compositions, methods, and / or kits for genome engineering of human cells. In certain embodiments, provided herein are compositions, methods, and / or kits for genome engineering of human immune cells or stem cells. In certain embodiments, provided herein are compositions, methods, and / or kits for efficient genome engineering. In certain embodiments, provided herein are compositions, methods, and / or kits for genome engineering that result in improved survival of cells, e.g., T cells, treated with a composition, method, and / or kit disclosed herein. In certain embodiments, provided herein are compositions, methods, and / or kits comprising a polynucleotide and / or a polynucleotide and a polypeptide bound thereto. In certain embodiments, provided herein are compositions and methods for improving the editing efficiency of one or more cells treated with a composition comprising a polynucleotide disclosed herein, e.g., increasing the editing efficiency by at least 5, 10, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, or 500%.
[0009] A. Polynucleotides In certain embodiments, compositions comprising a polynucleotide are provided herein. The polynucleotide can be in any suitable form, such as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA, preferably double-stranded DNA, e.g., a plasmid. In certain embodiments, the linear double-stranded DNA comprises covalently closed ends (e.g., doggybone DNA, dbDNA). Exemplary linear double-stranded DNA polypeptides with covalently closed ends are shown in Figures 14-17. Linear double-stranded DNA with covalently closed ends can be generated in any suitable manner, for example, by treating the double-stranded polynucleotide with a suitable enzyme, such as a protelomerase enzyme, e.g., TelN, or a suitable alternative, such as ligating a hairpin loop to the linear double-stranded DNA. In certain embodiments, the polynucleotide is a single-stranded DNA or RNA, a single-stranded DNA or RNA with a 5' hairpin, a single-stranded DNA or RNA with a 3' hairpin, a single-stranded DNA or RNA with a 5' hairpin and a 3' hairpin, a half-loop, a single-stranded DNA or RNA with a linked 5' oligo, a single-stranded DNA or RNA with a linked 5' oligo and a 3' hairpin, a single-stranded DNA or RNA with a linked 5' oligo and a linked 3' oligo, a single-stranded DNA or RNA with a linked 3' oligo, a double hairpin, a mixed loop, a mixed strand, or a combination thereof. Exemplary polynucleotide designs can be found in Shy et al. (2023) Nature Biotechnology. In certain embodiments, the ligated hairpin loop comprises a nuclear localization signal attached to the hairpin loop. In certain embodiments, the polynucleotide further comprises one or more of: (1) a donor template (D); (2) a first PAM (P1) and a first target nucleotide sequence (T1); and, optionally, (3) a second PAM (P2) and a second target nucleotide sequence (T2).A polynucleotide can include any suitable number of PAMs and target nucleotide sequences, e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or no more than 10, 9, 8, 7, 6, 5, 4, 3, or 2, e.g., 1 to 10, preferably 1 to 6, more preferably 1 to 3, and even more preferably 2 PAMs and target nucleotide sequences. As used herein, the term "target nucleotide sequence" includes a sequence in a polynucleotide, e.g., a protospacer, to which a nucleic acid-guided nuclease can bind. The target nucleotide sequence comprises at least partial complementarity to the spacer sequence of the nucleic acid-guided nuclease such that the nucleic acid-guided nuclease can bind to and optionally cleave one or more strands of the polynucleotide at or near the target nucleotide sequence. Typically, the target nucleotide sequence is present within a target site in a target polynucleotide, and upon binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence, the nuclease can optionally generate strand cleavage in one or more strands of the polynucleotide. The donor template can include any suitable sequence, including any suitable number and combination of components, e.g., transgenes, as needed for the present application. For example, the donor template can include a regulatory element, such as a promoter (any suitable regulatory element can be used, e.g., one or more of the sequences shown in Table 5), a payload, e.g., a heterologous gene, e.g., a transgene or payload, e.g., a chemokine receptor, a chimeric antigen receptor (CAR, e.g., a sequence shown in Table 1), or a chimeric autoantibody receptor (CAAR), and / or a terminator. In certain embodiments, the transgene encodes a self-cleaving peptide, such as a 2A peptide. In certain embodiments, the transgene encodes a reporter protein, e.g., a fluorescent protein, an antibiotic resistance marker, or the like. In certain embodiments, the heterologous gene comprises a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof that binds to a binding partner, e.g., B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ, or a portion thereof.In certain embodiments, the heterologous gene comprises a polynucleotide encoding a CAR or a portion thereof comprising a polypeptide at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124 shown in Table 1. In other embodiments, the donor template comprises a sequence that, when inserted into the target site, generates one or more mutations in the gene at the target site, such as SNPs and INDELs or missense mutations. In certain embodiments, the polynucleotide may further comprise one or more homology arms, e.g., upstream and downstream homology arms, to the donor template. In certain embodiments, the polynucleotide comprises a first PAM (P1) and a first target nucleotide sequence (T1) either upstream or downstream of the donor template (D). In certain embodiments, the first PAM (P1) and the first target nucleotide sequence (T1) are adjacent to, but not within, the donor template. The first PAM (P1) and the first target nucleotide sequence (T1) should be oriented so that the respective nucleic acid-guided nucleases can recognize the PAM and target nucleotide sequence. For example, when using a VA-type nucleic acid-guided nuclease, such as a nuclease disclosed herein, such as cpf1 and / or MAD7, the first PAM (P1) and the first target nucleotide sequence (T1) should be oriented so that the first PAM (P1) is 5' of the first target nucleotide sequence (T1). In certain embodiments, for a first PAM (P1) and a first target nucleotide sequence (T1) upstream of a VA-type nucleic acid-guided nuclease, the first PAM (P1) and the first target nucleotide sequence (T1) can be oriented as follows: (1) 5'P1. + T1 + D + 3' or (2) 5' T1 - P1 - D + 3' (where D +(represents the sense strand of the donor template). In certain embodiments, for the first PAM (P1) and first target nucleotide sequence (T1) downstream of the VA-type nucleic acid-guided nuclease, the first PAM (P1) and first target nucleotide sequence (T1) can be oriented as follows: (1) 5'D + P1 + T1 + 3' or (2) 5'D + T1 - P1 - 3' (where D + (where ' denotes the sense strand of the donor template). In certain embodiments, the polynucleotide comprises a first PAM (P1) and a first target nucleotide sequence (T1) upstream of the donor template, and a second PAM (P2) and a second target nucleotide sequence (T2) downstream of the donor template. In certain embodiments, the first PAM (P1) and the first target nucleotide sequence (T1) may be oriented as follows: (1) 5'P1 + T1 + D + 3' or (2) 5' T1 - P1 - D + The 3' and second PAM (P2) and second target nucleotide sequence (T2) can be oriented as follows: (1) 5'D + P2 + T2 + 3' or (2) 5'D + T2 - P2 - 3'. In certain embodiments, the first and second PAM target nucleotide sequences are 5'T1 - P1 - D + P2 + T2 + 3' (as shown in Figure 7, the sense strand of the donor template is marked 701), 5'T1 - P1 - D + T2 - P2 - 3' (as shown in Figure 5, the sense strand of the donor template is marked 501), 5'P1 + T1 + D+ T2 - P2 - 3' (as shown in Figure 4, the sense strand of the donor template is marked 401) or 5' P1 + T1 + D + P2 + T2 + 3' (as shown in Figure 6, the sense strand of the donor template is marked 601), preferably oriented 5'T1 - P1 - D + P2 + T2 + 3' or 5' P1 + T1 + D + T2 - P2 - It is oriented 3'.
[0010] In certain embodiments, the PAM comprises a sequence suitable for the nucleic acid-guided nuclease being used. For example, the PAM will comprise a sequence suitable for recognition by a type V nucleic acid-guided nuclease for applications using the type V nucleic acid-guided nucleases disclosed herein (e.g., Tables 3 and 4). In certain embodiments, the PAM comprises the sequence CTTN or TTTN. It should be understood that any suitable PAM sequence may be used for each nucleic acid-guided nuclease. For example, engineered nucleic acid-guided nucleases in which the PAM specificity is adjusted, modified, suppressed, etc., will have a first PAM comprising the respective sequence recognized by the engineered nucleic acid-guided nuclease.
[0011] The target nucleotide sequence may comprise any suitable sequence required for the intended use.
[0012] A polynucleotide can contain any suitable number of binding sites comprising a PAM suitable for a nucleic acid-guided nuclease and a target nucleotide sequence, and thus a composition can comprise a polynucleotide having any suitable number of nucleic acid-guided nuclease complexes bound thereto. In certain embodiments, a polynucleotide contains 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer binding sites, e.g., 1 to 10, preferably 1 or 2 binding sites. Illustratively, a composition comprising a polynucleotide containing two binding sites can contain up to two nucleic acid-guided nuclease complexes bound to the polynucleotide, and a composition comprising a polynucleotide containing three binding sites can contain up to three nucleic acid-guided nuclease complexes bound to the polynucleotide.
[0013] In certain embodiments, provided herein are compositions comprising a plurality of polynucleotides. In certain embodiments, the compositions comprise a plurality of polynucleotides as described above. In certain embodiments, the plurality of polynucleotides comprises: (1) a donor template (D); x (2) a first preferred PAM (P1) x and a first target nucleotide sequence (T1) x , and (3) a second suitable PAM (P2) x and a second target nucleotide sequence (T2) xwhere for each integer x, the polynucleotide comprises a different donor template sequence. The donor template can comprise any sequence suitable for the intended use. In certain embodiments, the donor template comprises a sequence encoding a heterologous gene such as a CAR or CAAR, e.g., a sequence encoding a polypeptide comprising a CAR or a portion thereof that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ, or a portion thereof, preferably a sequence encoding a CAR or a portion thereof comprising a polypeptide at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of the amino acid sequences set forth in SEQ ID NOs: 86-124. The plurality of polynucleotides can comprise any suitable number of different polynucleotides, i.e., the integer x. In certain embodiments, the number of different integers x is at least 2, 3, 4, 5, 6, 7, 8, or 9 and / or no more than 10, 9, 8, 6, 5, 4, or 3, such as 2 to 10, preferably 2 to 5. In certain embodiments, the number of different integers x is at least 10, 20, 30, 40, 50, 60, 70, 80, or 90 and / or no more than 100, 90, 80, 60, 50, 40, 30, or 20, such as 10 to 100, preferably 10 to 50.
[0014] In certain embodiments, provided herein are compositions comprising cells comprising the above-described polynucleotides and / or nucleotides. In certain embodiments, the cells are human cells. In certain embodiments, the cells are human cells or stem cells. In certain embodiments, the human cells are immune cells, including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, or lymphocytes, preferably T cells. In certain embodiments, the human cells are human pluripotent or pluripotent stem cells, embryonic stem cells, induced pluripotent stem cells (iPSCs), human hematopoietic stem cells, CD34+ cells, preferably iPSCs. In certain embodiments, the cells are cells that exhibit reduced immunogenicity when placed in an allogeneic host. In certain embodiments, the cells are non-immunogenic when placed in an allogeneic host.
[0015] [Table 1]
[0016] [Table 2]
[0017] [Table 3]
[0018] [Table 4]
[0019] B. Polynucleotides and their associated polypeptides In certain embodiments, provided herein are compositions comprising a polynucleotide and a polypeptide bound to the polynucleotide. In certain embodiments, the polynucleotide can comprise any of the polynucleotides disclosed herein. In preferred embodiments, the polypeptide further comprises one or more functional moieties, such as a nuclear localization sequence (NLS) or an affinity tag. In more preferred embodiments, the composition comprises a polynucleotide and a polypeptide bound to the polynucleotide, wherein the polypeptide comprises an NLS. In certain embodiments, the polypeptide comprises at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or no more than 10, 9, 8, 7, 6, 5, 4, 3, or 2 functional moieties on the N-terminus and / or C-terminus, e.g., 1 to 10 functional domains. In certain embodiments, the polypeptide comprises 3 to 7 to 4 to 6 functional domains on the N-terminus. Any suitable polypeptide can be used, e.g., a DNA-binding protein, a homing endonuclease, or a nucleic acid-guided nuclease such as a Class I or Class II nucleic acid-guided nuclease, e.g., a Type V nucleic acid-guided nuclease. In preferred embodiments, the polypeptide comprises a nucleic acid-guided nuclease, i.e., a nucleic acid-guided nuclease complex, e.g., a ribonucleoprotein (RNP), complexed with a compatible guide nucleic acid (gNA). The polynucleotide can comprise any suitable number of binding sites comprising a PAM and target nucleotide sequence suitable for the nucleic acid-guided nuclease, and thus the composition can comprise a polynucleotide bound to any suitable number of nucleic acid-guided nuclease complexes. In certain embodiments, the polynucleotide comprises 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer binding sites, e.g., 1 to 10, preferably 1 or 2 binding sites. Illustratively, a composition comprising a polynucleotide comprising two binding sites can comprise up to two nucleic acid-guided nuclease complexes bound to the polynucleotide, and a composition comprising a polynucleotide comprising three binding sites can comprise up to three nucleic acid-guided nuclease complexes bound to the polynucleotide.In certain embodiments, one or more bound nucleic acid-guided nuclease complexes are the same. In other embodiments, they are different. In preferred embodiments, one or more nucleic acid-guided nuclease complexes are the same, such that the first, second, and / or additional target sequences share sufficient complementarity with a spacer sequence in the guide nucleic acid of the nucleic acid-guided nuclease complex so that the complexes can bind to the target nucleotide sequence. In certain embodiments, at least one strand of a polynucleotide comprising a target nucleotide sequence is cleaved at or near the target nucleotide sequence upon binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence. In preferred embodiments, the polynucleotide comprises double-stranded DNA, and both strands of the polynucleotide are cleaved at or near the target nucleotide sequence upon binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence. In other embodiments, neither strand of a polynucleotide comprising a target nucleotide sequence is cleaved at or near the target nucleotide sequence upon binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence. In certain cases, engineered nucleic acid-guided nuclease complexes can be used that lack the ability to cleave one or more strands of a polynucleotide. In other cases, the gNA contains a spacer sequence that is sufficiently complementary to the target nucleotide sequence to bind to the target nucleotide sequence, but not sufficiently complementary to cleave the polynucleotide at or near the target nucleotide sequence.
[0020] The nucleic acid-guided nuclease complex may comprise any suitable nucleic acid-guided nuclease and a compatible gNA as disclosed herein.
[0021] In certain embodiments, the nucleic acid-guided nuclease comprises an engineered non-naturally occurring nuclease. In certain embodiments, the nucleic acid-guided nuclease comprises a Class 1 or Class 2 nucleic acid-guided nuclease, such as a Type II or Type V, e.g., a Type VA, Type VB, Type VC, Type VD, or Type VE nucleic acid-guided nuclease, e.g., a Type VA (Va) nucleic acid-guided nuclease. In certain embodiments, the nucleic acid-guided nuclease comprises a MAD, ART, or ABW nuclease. In preferred embodiments, the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99 or 100% identical to the amino acid sequence of a MAD, ART or ABW nucleic acid-guided nuclease, preferably an amino acid sequence that is at least 80, 85, 90, 95, 99 or 100% identical to the amino acid sequence of a MAD nuclease, more preferably an amino acid sequence that is at least 80, 85, 90, 95, 99 or 100% identical to the amino acid sequence of a MAD7 nuclease, and even more preferably an amino acid sequence that is at least 80, 85, 90, 95, 99 or 100% identical to the amino acid sequence of SEQ ID NO: 37.
[0022] In certain embodiments, the nucleic acid-guided nuclease complex comprises at least four NLSs, preferably at least five NLSs, arranged in any suitable orientation. In certain embodiments, the nucleic acid-guided nuclease comprises one N-terminal NLS and three C-terminal NLSs. In other embodiments, the nucleic acid-guided nuclease comprises five or more N-terminal NLSs. The NLSs can comprise any suitable sequence as disclosed herein, preferably any one of SEQ ID NOS: 40-56, more preferably SEQ ID NOS: 40, 51, and 56.
[0023] In certain embodiments, the gNA comprises an engineered non-naturally occurring nuclease. In certain embodiments, the gNA comprises a spacer sequence heterologous to any naturally occurring spacer sequence for the respective nucleic acid-guided nuclease. In certain embodiments, the gNA comprises a single polynucleotide. In preferred embodiments, the gNA comprises a dual gNA as disclosed herein, e.g., a guide nucleic acid comprising a targeter nucleic acid and a modulator nucleic acid that can bind to and activate a nucleic acid-guided nuclease that is activated by a single crRNA in the absence of tracrRNA in a naturally occurring system.
[0024] In certain embodiments, compositions comprising a polynucleotide and a polypeptide bound thereto are provided herein. In preferred embodiments, the polypeptide bound thereto comprises one or more NLSs. The polynucleotide may comprise single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA, preferably double-stranded DNA, e.g., a plasmid. In certain embodiments, the linear double-stranded DNA comprises covalently closed ends, such as those generated by a telomerase enzyme, e.g., TelN, or a suitable alternative. In certain embodiments, the composition comprises a polynucleotide comprising one or more of: (1) a donor template (D); (2) a first PAM (P1) and a first target nucleotide sequence (T1) and a bound polypeptide, optionally comprising an NLS; and / or (3) a second PAM (P2) and a second target nucleotide sequence (T2) and a bound polypeptide, optionally comprising an NLS. The donor template may comprise any suitable sequence, including any suitable number and combination of components, as required for the present application. For example, the donor template can include sequences encoding a promoter (e.g., a sequence shown in Table 5), a heterologous gene such as a chimeric antigen receptor (CAR) or chimeric autoantibody receptor (CAAR), and / or a terminator. In certain embodiments, the heterologous gene includes a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof that binds to a binding partner including B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ, or a portion thereof. In certain embodiments, the heterologous gene includes a polynucleotide encoding a CAR or a portion thereof that comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124 shown in Table 1.
[0025] In certain embodiments, the composition comprises a polynucleotide comprising a first PAM (P1) and a first target nucleotide sequence (T1) either upstream or downstream of a donor template (D), and a polypeptide comprising an NLS bound thereto. In certain embodiments, the first PAM (P1) and the first target nucleotide sequence (T1) are adjacent to, but not within, the donor template. The first PAM (P1) and the first target nucleotide sequence (T1) should be oriented in a manner that is functional for the respective nucleic acid-guided nuclease. For example, when using a Va-type nucleic acid-guided nuclease such as cpf1 or MAD7, the first PAM (P1) and the first target nucleotide sequence (T1) should be oriented so that the first PAM (P1) is upstream of the first target nucleotide sequence (T1). In certain embodiments, for a first PAM (P1) and a first target nucleotide sequence (T1) upstream of a VA-type nucleic acid-guided nuclease, the first PAM (P1) and the first target nucleotide sequence (T1) may be oriented as follows: (1) 5'P1 + T1 + D + 3' or (2) 5' T1 - P1 - D + 3' (where D + (represents the sense strand of the donor template). In certain embodiments, for the first PAM (P1) and first target nucleotide sequence (T1) downstream of the VA-type nucleic acid-guided nuclease, the first PAM (P1) and first target nucleotide sequence (T1) can be oriented as follows: (1) 5'D + P1 + T1 + 3' or (2) 5'D + T1 - P1 - 3' (where D +(represents the sense strand of the donor template). In certain embodiments, the polynucleotide comprises a polypeptide comprising a first PAM (P1) and a first target nucleotide sequence (T1) with an NLS bound thereto upstream of the donor template (D), and a polypeptide comprising a second PAM (P2) and a second target nucleotide sequence (T2) with an NLS bound thereto downstream of the donor template (D). In certain embodiments, the first PAM (P1) and the first target nucleotide sequence (T1) may be oriented as follows: (1) 5'P1 + T1 + D + 3' or (2) 5' T1 - P1 - D + The 3' and second PAM (P2) and second target nucleotide sequence (T2) can be oriented as follows: (1) 5'D + P2 + T2 + 3' or (2) 5'D + T2 - P2 - 3'. In certain embodiments, the first and second PAM target nucleotide sequences are 5'T1 - P1 - D + P2 + T2 + 3' (as shown in Figure 7, the sense strand of the donor template is marked 701), 5'T1 - P1 - D + T2 - P2 - 3' (as shown in Figure 5, the sense strand of the donor template is marked 501), 5'P1 + T1 + D + T2 - P2 - 3' (as shown in Figure 4, the sense strand of the donor template is marked 401) or 5' P1 + T1 + D + P2 + T2 +3' (as shown in Figure 6, the sense strand of the donor template is marked 601), preferably oriented 5'T1 - P1 - D + P2 + T2 + 3' or 5' P1 + T1 + D + T2 - P2 - It is oriented 3'.
[0026] Illustrated in Figure 4 is the 5'P1 + T1 + D + T2 - P2 - 1 is an exemplary composition comprising a polynucleotide comprising a 3'-oriented donor template, first and second PAMs, and first and second target nucleotide sequences, further comprising a first nucleic acid-guided nuclease complex comprising an NLS bound to the first PAM and the target nucleotide sequence, and a second nucleic acid-guided nuclease comprising an NLS bound to the second PAM and the target nucleotide sequence. The sense strand of the donor template is represented as 401. In certain embodiments, the nucleic acid-guided nuclease complex is not released after cleavage, and the cleavage product comprises the donor template (402) and the remainder of the polynucleotide with the bound nucleic acid-guided nuclease complex. The resulting donor template (402) is free of the bound nucleic acid-guided nuclease complex comprising the NLS.
[0027] Illustrated in Figure 5 is the 5'T1 - P1 - D + T2 - P2 -1 is an exemplary composition comprising a polynucleotide comprising a 3'-oriented donor template, first and second PAMs, and first and second target nucleotide sequences, further comprising a first nucleic acid-guided nuclease complex comprising an NLS bound to the first PAM and the target nucleotide sequence, and a second nucleic acid-guided nuclease complex comprising an NLS bound to the second PAM and the target nucleotide sequence. The sense strand of the donor template is represented as 501. In certain embodiments, the nucleic acid-guided nuclease complex is not released after cleavage, and the cleavage product comprises the donor template and the upstream PAM and target nucleotide sequence (502), as well as a bound nucleic acid-guided nuclease complex comprising an NLS for the remainder of the polynucleotide. The resulting donor template (502) comprises one bound nucleic acid-guided nuclease complex comprising an NLS.
[0028] Illustrated in Figure 6 is the 5'P1 + T1 + D + P2 + T2 + 6 is an exemplary composition comprising a polynucleotide comprising a 3'-oriented donor template, first and second PAMs, and first and second target nucleotide sequences, further comprising a first nucleic acid-guided nuclease complex comprising an NLS bound to the first PAM and the target nucleotide sequence, and a second nucleic acid-guided nuclease complex comprising an NLS bound to the second PAM and the target nucleotide sequence. The sense strand of the donor template is represented as 601. In certain embodiments, the nucleic acid-guided nuclease complex is not released after cleavage, and the cleavage products comprise the donor template and the downstream PAM and target nucleotide sequence (602), as well as a bound nucleic acid-guided nuclease complex comprising an NLS to the remainder of the polynucleotide. The resulting donor template (602) comprises one nucleic acid-guided nuclease complex comprising an NLS.
[0029] Illustrated in Figure 7 is the 5'T1 - P1 - D + P2+ T2 + 7A and 7B are exemplary compositions comprising a polynucleotide comprising a 3'-oriented donor template, first and second PAMs, and first and second target nucleotide sequences, further comprising a first nucleic acid-guided nuclease complex comprising an NLS bound to the first PAM and the target nucleotide sequence, and a second nucleic acid-guided nuclease complex comprising an NLS bound to the second PAM and the target nucleotide sequence. The sense strand of the donor template is represented as 701. In certain embodiments, the nucleic acid-guided nuclease complex is not released after cleavage, and the cleavage product comprises the donor template and two bound nucleic acid-guided nuclease complexes comprising both the upstream and downstream PAMs and target nucleotide sequences (702) and an NLS relative to the remainder of the polynucleotide. The resulting donor template (702) comprises two nucleic acid-guided nuclease complexes comprising an NLS.
[0030] Without wishing to be bound by theory, an NLS operably linked to a donor template can facilitate functional uptake of the donor template into the nucleus of a target cell, further increasing the likelihood that at least a portion of the donor template will be introduced at or near the target site in the genome of the target cell upon cleavage of the genome by a nucleic acid-guided nuclease complex via homology-directed repair (HDR). This efficiency of HDR of the donor template can be improved by increasing the local concentration of the donor template (HDR template), facilitated by increased nuclear localization by one or more bound polypeptides containing an NLS. Surprisingly, the ability of the nucleic acid-guided nuclease complex to remain bound to the polynucleotide after cleavage improves the nuclear uptake of the cleaved product.
[0031] In certain embodiments, the polynucleotide further comprises one or more additional PAMs and target sites between the donor template and the first and / or second PAM and target site. In certain cases, the one or more additional PAMs and target sites are modified so that the nucleic acid-guided nuclease complex can bind to the one or more additional PAMs and target sites and does not produce one or more cleavages in the polynucleotide at or near the site. Any suitable modification can be used to modify the polynucleotide to prevent the production of one or more cleavages in the polynucleotide at or near the site, such as by designing the guide nucleic acid to have a spacer sequence with fewer complementary nucleotides or to have one or more mismatches in the spacer sequence, for example, within the seed sequence. In certain embodiments, the length of complementary nucleotides between the spacer sequence and the target site contains at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, and up to 10, 9, 8, 7, 6, 5, 4, 3, or 2 fewer complementary nucleotides, for example, 1 to 10 fewer complementary nucleotides, compared to a spacer sequence with perfect complementarity. In other words, the spacer sequence can contain at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, and / or no more than 25 or 20, nucleotides complementary to the target sequence, e.g., 5 to 20 complementary nucleotides. In certain embodiments, the spacer sequence contains at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, and / or no more than 10, 9, 8, 7, 6, 5, 4, 3, or 2, mismatches, e.g., 1 to 10 mismatches, compared to the target sequence. Any suitable number of mismatches can be present in the seed sequence, e.g., a total of 1, 2, 3, 4, or 5 mismatches can be present in the seed sequence.
[0032] An exemplary polynucleotide comprising one or more additional PAMs and a target site is shown in Figure 18. Specifically, Figure 18 shows a polynucleotide comprising a donor template (1801) flanked by a first PAM (1804) and a second PAM (1805) and an upstream homology arm (1802) and a downstream homology arm (1803) further adjacent to the target site, preferably designed such that one or more strand cleavages occur upon binding of a nucleic acid-guided nuclease complex, and one or more additional PAMs (1806 and 1807) and further adjacent to the target site, preferably designed such that one or more strand cleavages cannot occur as disclosed herein upon binding of a nucleic acid-guided nuclease complex. A polynucleotide comprising a donor template (1801) flanked by an upstream homology arm (1802) and a downstream homology arm (1803) can include any number of suitable additional elements, such as one or more of 1804, 1805, 1806, and / or 1807, such as one of 1804, 1805, 1806, and / or 1807, two of 1804, 1805, 1806, and / or 1807, three of 1804, 1805, 1806, and / or 1807, or all of 1804, 1805, 1806, and / or 1807.
[0033] In certain embodiments, the composition further comprises an additive that stabilizes the nucleic acid-guided nuclease complex. In certain embodiments, one or more additives that stabilize the nucleic acid-guided nuclease complex are combined with the nuclease and guide nucleic acid. In certain embodiments, one or more additives that stabilize the nucleic acid-guided nuclease complex are combined with the guide nucleic acid prior to combination with the nuclease. In certain embodiments, one or more additives that stabilize the nucleic acid-guided nuclease complex are combined with the nuclease prior to combination with the guide nucleic acid. In certain embodiments, one or more additives that stabilize the nucleic acid-guided nuclease complex are combined with a preformed nucleic acid-guided nuclease complex comprising one or more nucleases and a guide nucleic acid. In certain embodiments, one or more additives that stabilize the nucleic acid-guided nuclease complex prevent aggregation and / or support dispersion of the nucleic acid-guided nuclease complex in a population of nucleic acid-guided nuclease complexes. In certain embodiments, the RNP stabilizer can comprise any suitable protein stabilizer, such as protein stabilizers known in the art. In certain embodiments, the RNP stabilizer is 1,2,3-heptanetriol, 2-amino-2-(hydroxymethyl)-1,3-propanediol (tris), 3-(1-pyridino)-1-propanesulfonate (NDSB 201), 3-[(3-cholamidopropyl)dimethylammonio]-1-propanesulfonate (CHAPS), 6-aminocaproic acid, adenosine diphosphate (ADP), adenosine triphosphate (ATP), alpha-cyclodextrin, amidosulfobetaine-14 (ASB-14), ammonium acetate, ammonium nitrate, ammonium sulfate, arginine, arginine ethyl ester, barium chloride, barium iodide, benzamidine HCl, beta-cyclodextrin, beta-mercaptoethanol (BME), biotin, calcium chloride, cesium chloride, cesium sulfate, cetyltrimethylammonium bromide (CTAB), choline chloride, citric acid, cobalt chloride, copper(II) chloride, cyclohexanol, D-sorbitol, dimethylethylammonium propanesulfonate (NDSB)195), dithiothreitol (DTT), erythritol, ethanol, ethylene glycol, ethylene glycol-bis(β-beta-aminoethyl ether)-N,N,N',N'-tetraacetic acid (EGTA), ethylenediaminetetraacetic acid (EDTA), formamide, gadolinium bromide, gamma-butyrolactone, glucose, glutamic acid, glutamine, glycerol, glycine, glycine betaine, glycine-glycine-glycine, guanidine HCl, guanosine triphosphate (GTP), holmium chloride, imidazole, iron(III) chloride, Jeffamine M-600, lanthanum acetate, lauryl sulfobetaine, lauryldimethylamine N-oxide (LDAO), lithium sulfate, magnesium chloride, magnesium sulfate, manganese chloride, mannitol, N-(2-hydroxyethyl)piperazine-N'-(3-propanesulfonic acid) (EPPS), N-dodecyl beta-D-maltoside (DDM), N-ethyl urea, n-hexanol, N-lauryl sarcoside, N-lauryl sarcosine, N-methylformamide, N-methyl urea, n-octyl-b-D-glucoside (OG: octyl glucoside), n-pentanol, nickel chloride, non-surfactant sulfobetaine (NDSB), Nonidet P40 (NP40), octyl beta-D-glucopyranoside, poly-L-glutamic acid, polyethylene glycol (e.g., PEG 300, PEG 3350, PEG 4000), polyethylene glycol lauryl ether (Brij 35), polyoxyethylene (2) oleyl ether (Brij 93), polyoxyethylene cetyl ether (Brij56), polyvinylpyrrolidone 40 (PVP40), potassium chloride, potassium citrate, potassium nitrate, proline, putrescine, spermidine, spermine, riboflavin, samarium bromide, sarcosine, sodium acetate, sodium chloride, sodium dodecyl sulfate (SDS), sodium fluoride, sodium iodide, sodium lauryl sarcosinate (sarcosyl), sodium malonate, sodium molybdate, sodium selenite, sodium sulfate, sodium thiocyanate, sucrose, taurine, trehalose, tricine, triethylamine, trimethylamine N-oxide (TMAO), tris(2-carboxyethyl)phosphine (TCEP), Triton X-100, Tween 20, Tween 60, Tween 80, urea, vitamin B12, xylitol, yttrium chloride, yttrium nitrate, zinc chloride, Zwittergent 3-08, Zwittergent 3-14, or a combination thereof. In certain embodiments, the RNP stabilizer comprises a negatively charged polymer. In certain embodiments, the RNP stabilizer comprises poly-L-glutamic acid (PGA) or a suitable alternative. In certain embodiments, provided herein are compositions, methods, and / or kits comprising poly-L-glutamic acid.
[0034] In certain embodiments, provided herein are compositions comprising cells comprising a polynucleotide and / or a plurality of nucleotides and polypeptides bound thereto, as described above. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a human cell or a stem cell. In certain embodiments, the human cell is an immune cell, including a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, or lymphocyte, preferably a T cell. In certain embodiments, the human cell is a human pluripotent or pluripotent stem cell, embryonic stem cell, induced pluripotent stem cell (iPSC), human hematopoietic stem cell, or stem cell that is a CD34+ cell, preferably an iPSC. In certain embodiments, the cell is a cell that exhibits reduced immunogenicity when placed in an allogeneic host. In certain embodiments, the cell is non-immunogenic when placed in an allogeneic host.
[0035] C. Methods of Use of Polynucleotides and Binding Polypeptides In certain embodiments, methods are provided herein. In certain embodiments, methods of using any one of the compositions disclosed herein are provided herein. In certain embodiments, the methods provided are suitable for one or more genome engineering applications.
[0036] In certain embodiments, provided herein are methods for cleaving one or more polynucleotides comprising a PAM and a target nucleotide sequence adjacent to, but not within, a donor template. In certain embodiments, the polynucleotide comprises a circular polynucleotide, which binds to a target nucleic acid sequence and, upon cleavage by a nucleic acid-guided nuclease complex at or near the target nucleotide sequence, generates at least one strand break, resulting in linearization of the circular polynucleotide. In certain embodiments, a method for preparing a linearized polynucleotide comprises contacting a polynucleotide as disclosed herein with a nucleic acid-guided nuclease complex, wherein the nucleic acid-guided nuclease complex binds to a target site on the polynucleotide and generates at least one strand break at or near the target nucleotide sequence. In certain embodiments, the polynucleotide comprises two target sites, and at least one strand break is generated at or near each of the target nucleotide sequences by one or more nucleic acid-guided nuclease complexes. In certain embodiments, the nucleic acid-guided nuclease complex remains bound to the polynucleotide. In certain embodiments, a composition comprising the linearized product is delivered to a cell.
[0037] In certain embodiments, provided herein are methods of manipulating the genome of a cell, comprising delivering to the cell a composition comprising a polynucleotide as described in the Polynucleotides section above, and a nucleic acid-guided nuclease system comprising a nucleic acid-guided nuclease and a gNA. In certain embodiments, the method of manipulating the genome of a cell comprises delivering to the cell a composition comprising any one of the compositions as disclosed herein.
[0038] In certain embodiments, the method of genome engineering comprises delivering multiple exogenous nucleic acids into the genome of a target cell, which converts the target cell into a donor template (D), each of which is a donor template (D). x and multiple nucleic acid-induced nuclease complexes (N) x wherein for each integer x, the donor template comprises a different sequence. In certain embodiments, the polynucleotide comprises a first suitable PAM (P1) adjacent to the 5' end of the donor template but not within the donor template. x and a first target nucleotide sequence (T1) x and a second suitable PAM (P2) adjacent to the 3' end of the donor template but not within the donor template. x and a second target nucleotide sequence (T2) x In certain embodiments, the polynucleotide further comprises (P1) x (T1) x and (D) x The first homology arm (HA1) between x And, (P2) x (T2) x and (D) x The second homology arm (HA2) between x and further including (HA1) x and (HA2) x is a target site (TS) selected from multiple target sites in the genome of the target cell. x (D) x In certain embodiments, each target site (TS) can initiate host cell-mediated recombination of at least a portion of the target site (TS). x About (TS) x and (T1) x and (T2) x a plurality of nucleic acid-induced nuclease complexes (N) capable of cleaving at least one of x Including (TS) x and (T1) x and (T2) x At least one cleavage of (TS) x (D) in xIn a preferred embodiment, at least one of the plurality of donor templates comprises a polynucleotide encoding a CAR or CAAR.
[0039] In certain embodiments, the method comprises combining a nucleic acid-guided nuclease and a gNA in vitro in a suitable buffer to produce a composition comprising the nucleic acid-guided nuclease and the gNA, and optionally allowing the nucleic acid-guided nuclease and the gNA to form a nucleic acid-guided nuclease complex. In certain embodiments, one or more polynucleotides are then added to the composition. In preferred embodiments, the composition comprises a buffer in which at least one component necessary for the activity of the nucleic acid-guided nuclease complex is at an activity-limiting concentration such that the complex cannot generate strand breaks when bound to the polynucleotide. Exemplary components include magnesium (Mg 2+ ) such Mg 2+ The restriction buffer contains Mg 2+ This can be achieved by the addition of EDTA, which chelates Mg and makes it unavailable to the complex. 2+ The limiting composition is delivered to the target cells and inhibits intracellular Mg 2+ is no longer limiting. Such compositions would prevent cleavage of one or more polynucleotides prior to delivery to the intended target cell.
[0040] In certain embodiments, the method further comprises treating the cells with an HDR enhancer (NHEJ inhibitor). In certain embodiments, the one or more additives that inhibit NHEJ are introduced into the target cells prior to delivery of the nucleic acid-guided nuclease, guide nucleic acid and / or donor template, or one or more polynucleotides encoding the nucleic acid-guided nuclease, guide nucleic acid and / or donor template. In certain embodiments, the one or more additives that inhibit NHEJ are introduced into the target cells after delivery of the nucleic acid-guided nuclease, guide nucleic acid and / or donor template, or one or more polynucleotides encoding the nucleic acid-guided nuclease, guide nucleic acid and / or donor template. In certain embodiments, the one or more additives that inhibit NHEJ are introduced into the target cells both before and after delivery of the nucleic acid-guided nuclease, guide nucleic acid and / or donor template, or one or more polynucleotides encoding the nucleic acid-guided nuclease, guide nucleic acid and / or donor template. In certain embodiments, the one or more additives that inhibit NHEJ are introduced into the cell culture medium, allowing the one or more NHEJ inhibitors to enter the cells.
[0041] In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that indirectly or directly affects the interaction of p53-binding protein 1 (53BP1) with ubiquitinated histones at double-strand breaks, such as iP53. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the interaction of Ku protein with DNA, such as STL127705. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the activity of DNA-dependent protein kinase, such as M3814, KU-0060648, or NU7026. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the activity of ATM-Rad3-related (ATR) protein, such as VE-822. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the activity of a ligase, such as ligase IV, such as SCR7. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the activity of RAD51 binding to ssDNA, such as RS-1. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects active cell cycle progression, such as aphidicolin, mimosine, thymidine, hydroxyurea, nocodazole, ABT-751, XL413, etc. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects active beta-3-adrenergic receptors, such as L755507. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects the activity of intracellular transport from the endoplasmic reticulum (ER) to the Golgi, such as brefeldin A. In certain embodiments, the one or more additives that inhibit NHEJ include a molecule that directly or indirectly affects active histone deacetylase, such as valproic acid (VPA). In certain embodiments, the one or more additives that inhibit NHEJ include M3814.
[0042] II. Engineered non-native dual-guide CRISPR-cas systems CRISPR-Cas systems generally include a Cas protein and one or more guide nucleic acids (gNAs). The Cas protein can be directed to a specific location in a double-stranded DNA target by recognizing a protospacer adjacent motif (PAM) in the non-target strand of the DNA, and one or more guide nucleic acids can be directed to a specific location by hybridizing to a target nucleotide sequence, also referred to herein as a target sequence, in the target strand of the target polynucleotide. Typically, both PAM recognition and target nucleotide sequence hybridization are required for stable binding of the CRISPR-Cas complex to the DNA target and activation of the effector function (e.g., nuclease activity) if the Cas protein has such a function. Consequently, when generating a CRISPR-Cas system, the guide nucleic acid can be designed to include a nucleotide sequence, referred to as a spacer sequence, that is at least partially complementary to and can hybridize with the target nucleotide sequence, with the target nucleotide sequence positioned adjacent to the PAM in an orientation that allows the Cas protein to function. It has been observed that not all CRISPR-Cas systems designed according to these criteria are equally effective. The larger polynucleotide in which the target nucleotide sequence is located can be called a target polynucleotide, such as a chromosome or other genomic DNA or a portion thereof, or any other suitable polynucleotide in which the target nucleotide sequence is located. In double-stranded DNA, the target polynucleotide comprises two strands. The strand of the DNA duplex to which the spacer sequence is complementary is referred to herein as the "target strand," while the strand with which the spacer sequence shares sequence identity is referred to herein as the "non-target strand."
[0043] Two distinct classes of CRISPR-Cas systems have been identified. Class 1 CRISPR-Cas systems use multiprotein effector complexes, while class 2 CRISPR-Cas systems use single-protein effectors (see Makarova et al. (2017) CELL, 168:328). Among the types of class 2 CRISPR-Cas systems, type II and type V systems typically target DNA, while type VI systems typically target RNA (id.). The native type II effector complex contains Cas9, CRISPR RNA (crRNA), and trans-activating CRISPR RNA (tracrRNA), although the crRNA and tracrRNA can be fused as a single guide RNA in engineered systems for simplicity (see Wang et al. (2016) ANNU. REV. BIOCHEM., 85:227). Some natural type V systems, such as type VA, type VC, and type VD systems, do not require tracrRNA and use a single crRNA as a guide for target DNA cleavage (see Zetsche et al. (2015) CELL, 163:759; Makarova et al. (2017) CELL, 168:328).
[0044] Natural type II CRISPR-Cas systems (e.g., CRISPR-Cas9 systems) generally contain two guide nucleic acids, called crRNA and tracrRNA, which form a complex through nucleotide hybridization. Single guide nucleic acids capable of activating type II Cas nucleases have been developed, for example, by linking crRNA and tracrRNA (see, e.g., U.S. Pat. Nos. 10,266,850 and 8,906,616). Natural type II Cas proteins contain a RuvC-like nuclease domain and an HNH endonuclease domain and recognize a 3' G-rich PAM located immediately downstream from the target nucleotide sequence, with the orientation determined using the non-target strand (i.e., the strand not hybridized with the spacer sequence) as a coordinate. CRISPR-Cas systems cleave double-stranded DNA to generate blunt ends. The cleavage site is generally 3 to 4 nucleotides upstream from the PAM in the non-target strand.
[0045] Naturally occurring VA-, VC-, and VD-type CRISPR-Cas systems lack a tracrRNA and rely on a single crRNA to guide the CRISPR-Cas complex to the target polynucleotide. Dual guide nucleic acids capable of activating VA-, VC-, or VD-type Cas nucleases have been developed, for example, by splitting a single crRNA into a targeter nucleic acid and a modulator nucleic acid (see, for example, International PCT Application Publication No. WO 2021 / 067788). Naturally occurring VA-type Cas proteins contain a RuvC-like nuclease domain but lack an HNH endonuclease domain and recognize a 5' T-rich PAM located immediately upstream from the target nucleotide sequence (i.e., the strand not hybridized with the spacer sequence) as their orientation, determined using the non-target strand. These CRISPR-Cas systems cleave double-stranded DNA to generate staggered double-strand breaks rather than blunt ends. The cleavage site is distant from the PAM site (e.g., at least 10, 11, 12, 13, 14, or 15 nucleotides downstream from the PAM on the non-target strand and / or at least 15, 16, 17, 18, or 19 nucleotides upstream from the sequence complementary to the PAM on the target strand).
[0046] The elements of an exemplary single-guide CRISPR-Cas system, e.g., a VA-type CRISPR-Cas system, are shown in Figure 1A. A single gRNA, when present in the form of RNA, may also be referred to as a "crRNA" or "single gRNA." From 5' to 3', it may include any 5' sequence, such as a tail, a modulator stem sequence, a loop, a targeter stem sequence complementary to the modulator stem sequence, and a spacer sequence at least partially complementary to and capable of hybridizing with a target sequence in the target strand of a target polynucleotide. When a 5' tail is present, the sequence comprising the 5' tail and the modulator stem sequence is also referred to herein as a "modulator sequence." The segment of the single-guide nucleic acid from the optional 5' tail to the targeter stem sequence, also referred to herein as a "scaffold sequence," binds to the Cas protein. Additionally, a PAM in the non-target strand of the target DNA binds to the Cas protein.
[0047] Elements in an exemplary dual-guide CRISPR-Cas system, such as a dual-guide VA-type CRISPR-Cas system, are shown in Figure 1B. The first guide nucleic acid, which may be referred to herein as a "modulator nucleic acid," comprises, from 5' to 3', an optional 5' tail and a modulator stem sequence. If a 5' tail is present, the sequence comprising the 5' tail and the modulator stem sequence may also be referred to herein as a "modulator sequence." The second guide nucleic acid, which may be referred to herein as a "targeter nucleic acid," comprises, from 5' to 3', a targeter stem sequence complementary to the modulator stem sequence and a spacer sequence at least partially complementary to and capable of hybridizing with a target sequence in the target strand of a target polynucleotide. The duplex between the modulator stem sequence and the targeter stem sequence and the optional 5' tail constitute a structure that binds to a Cas protein. Additionally, a PAM in the non-target strand of the target DNA binds to a Cas protein. It is understood that in dual gNAs, e.g., dual gRNAs, the targeter nucleic acid and modulator nucleic acid are not in the same nucleic acid, i.e., are not linked end-to-end via traditional inter-polynucleotide bonds, but may be covalently conjugated to each other via one or more chemical modifications introduced into these nucleic acids, thereby increasing the stability of the double-stranded complex and / or improving other characteristics of the system.
[0048] The terms "targeter stem sequence" and "modulator stem sequence," as used herein, may refer to a pair of nucleotide sequences in one or more guide nucleic acids that hybridize to each other. When the targeter stem sequence and modulator stem sequence are contained in a single guide nucleic acid, the targeter stem sequence is proximal to a spacer sequence designed to hybridize with the target nucleotide sequence, and the modulator stem sequence is proximal to the target stem sequence. When the targeter stem sequence and modulator stem sequence are in separate nucleic acids, the targeter stem sequence is in the same nucleic acid as the spacer sequence designed to hybridize with the target nucleotide sequence. In CRISPR-Cas systems that naturally contain separate crRNAs and tracrRNAs (e.g., Type II systems), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the duplex formed between the crRNA and tracrRNA. In CRISPR-Cas systems that naturally contain a single crRNA but do not contain a tracrRNA (e.g., a VA-type system), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the stem portion of the stem-loop structure in the scaffold sequence of the crRNA. It is understood that 100% complementarity is not required between the targeter stem sequence and the modulator stem sequence. However, in VA-type CRISPR-Cas systems, the targeter stem sequence is typically 100% complementary to the modulator stem sequence.
[0049] An illustrative example of a nucleic acid-guided nuclease complex is shown in Figure 3. Specifically, Figure 3 shows a VA-type nucleic acid-guided nuclease (301) complexed with a dual gNA comprising a modulator nucleic acid (306) and a targeter nucleic acid (307), where the modulator nucleic acid and targeter nucleic acid are hybridized through the stem. The targeter nucleic acid further comprises a spacer sequence (305), i.e., a protospacer, in the target polynucleotide (302) adjacent to a suitable PAM (303), that is at least partially complementary to the target nucleotide sequence (304). Upon binding to the target nucleotide sequence, the nucleic acid-guided nuclease complex can generate one or more strand breaks (308) in the target polynucleotide at or near the target nucleotide sequence.
[0050] A. Cas proteins Guide nucleic acids, either alone (where the targeter and modulator nucleic acids are part of a single polynucleotide) or as dual gNAs with separate targeter nucleic acids used in combination with the cognate modulator nucleic acid, can bind to CRISPR-associated (Cas) proteins, such as Cas nucleases. In certain embodiments, guide nucleic acids, either alone (where the targeter and modulator nucleic acids are part of a single polynucleotide) or as dual gNAs with separate targeter nucleic acids used in combination with the cognate modulator nucleic acid, can activate Cas nucleases. A gNA that can activate a particular Cas nuclease is said to be "compatible" with that Cas nuclease, and a Cas nuclease that can be activated by a particular gNA is said to be "compatible" with that gNA.
[0051] The terms "CRISPR-associated protein," "Cas protein," and "Cas," used interchangeably herein, may refer to a native Cas protein or an engineered Cas protein. Non-limiting examples of Cas protein engineering include, but are not limited to, mutations and modifications of Cas proteins that alter Cas activity, alter PAM specificity, expand the range of recognized PAMs, and / or reduce its ability to modify one or more off-target loci compared to the corresponding unmodified Cas. In certain embodiments, the altered activity of an engineered Cas includes an altered ability (e.g., specificity or kinetics) to bind to a native gNA, e.g., gRNA, or an engineered gNA, e.g., gRNA; an altered ability (e.g., specificity or kinetics) to bind to a target nucleotide sequence; an altered processivity of nucleic acid scanning; and / or an altered effector (e.g., nuclease) activity. Cas proteins with nuclease activity may be referred to interchangeably herein as "CRISPR-associated nucleases," or "Cas nucleases," or simply "nucleases."
[0052] In certain embodiments, the Cas protein is a type VA, type VC, or type VD Cas protein. In certain embodiments, the Cas protein is a type VA Cas protein. In other embodiments, the Cas protein is a type II Cas protein, such as a Cas9 protein.
[0053] In certain embodiments, the VA-type Cas nuclease comprises Cpf1. Cpf1 proteins are known in the art and are described, for example, in U.S. Patent Nos. 9,790,490 and 10,113,179. Cpf1 orthologs are found in a variety of bacterial and archaeal genomes. For example, in certain embodiments, the Cpf1 protein is isolated from Francisella novicida U112 (Fn), Acidaminococcus sp. BV3L6 (As), Lachnospiraceae bacterium ND2006 (Lb), Lachnospiraceae bacterium MA2020 (Lb2), Candidatus Methanoplasma termitum (CMt), Moraxella bovoculi 237 (Mb), Porphyromonas crevioricanis (Pc), Prevotella dysenteriae (Pc), or any of the following bacteria: disiens (Pd), Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.The gene is derived from SCADC, Eubacterium eligens, Leptospira inadai, Porphyromonas macacae, Prevotella bryantii, Proteocatella sphenisci, Anaerovibrio sp. RM50, Moraxella caprae, Lachnospiraceae bacterium COE1, or Eubacterium coprostanoligenes.
[0054] In certain embodiments, the VA-type Cas nuclease comprises AsCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 3 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 3 of International (PCT) Application Publication No. WO 2021 / 158918.
[0055] In certain embodiments, the VA-type Cas nuclease comprises LbCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 4 of International (PCT) Application Publication No. WO 2021 / 158918.
[0056] In certain embodiments, the VA-type Cas nuclease comprises FnCpf1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 5 of International (PCT) Application Publication No. WO 2021 / 158918.
[0057] In certain embodiments, the VA-type Cas nuclease comprises Prevotella bryantii Cpf1 (PbCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 6 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 6 of International (PCT) Application Publication No. WO 2021 / 158918.
[0058] In certain embodiments, the VA-type Cas nuclease comprises Proteocatella sphenisci Cpf1 (PsCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 7 of International (PCT) Application Publication No. WO 2021 / 158918.
[0059] In certain embodiments, the VA-type Cas nuclease comprises Anaerovibrio sp. RM50 Cpf1 (As2Cpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 8 of International (PCT) Application Publication No. WO 2021 / 158918.
[0060] In certain embodiments, the VA-type Cas nuclease comprises Moraxella caprae Cpf1 (McCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO:9 of International (PCT) Application Publication No. WO 2021 / 158918.
[0061] In certain embodiments, the VA-type Cas nuclease comprises Lachnospiraceae bacterium COE1 Cpf1 (Lb3Cpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 10 of International (PCT) Application Publication No. WO 2021 / 158918.
[0062] In certain embodiments, the VA-type Cas nuclease comprises the Eubacterium coprostanoligenes gene Cpf1 (EcCpf1) or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 11 of International (PCT) Application Publication No. WO 2021 / 158918.
[0063] In certain embodiments, the VA-type Cas nuclease is not Cpf1. In certain embodiments, the VA-type Cas nuclease is not AsCpf1.
[0064] In certain embodiments, the VA-type Cas nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20, or a variant thereof. MAD1-MAD20 are known in the art and are described in U.S. Patent No. 9,982,279.
[0065] In certain embodiments, the VA-type Cas nuclease comprises MAD7 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 37. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 37. MAD7 (SEQ ID NO: 37)
[0066] In certain embodiments, the VA-type Cas nuclease comprises MAD2 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identical to the amino acid sequence set forth in SEQ ID NO: 38. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 38. MAD2 (SEQ ID NO: 38)
[0067] In certain embodiments, the VA-type Cas nuclease comprises Csm1. Csm1 proteins are known in the art and are described in U.S. Patent No. 9,896,696. Csm1 orthologs are found in various bacterial and archaeal genomes. For example, in certain embodiments, the Csm1 protein is derived from Smithella sp. SCADC (Sm), Sulfuricurvum sp. (Ss), or Microgenomates (Roizmanbacteria) bacterium (Mb).
[0068] In certain embodiments, the VA-type Cas nuclease comprises SmCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 12 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 12 of International (PCT) Application Publication No. WO 2021 / 158918.
[0069] In certain embodiments, the VA-type Cas nuclease comprises SsCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 13 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 13 of International (PCT) Application Publication No. WO 2021 / 158918.
[0070] In certain embodiments, the VA-type Cas nuclease comprises MbCsm1 or a variant thereof. In certain embodiments, the VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 14 of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the VA-type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 14 of International (PCT) Application Publication No. WO 2021 / 158918.
[0071] In certain embodiments, the VA-type Cas nuclease comprises an ART nuclease or a variant thereof. Generally, such nuclease sequences have <60% AA sequence similarity with Cas12a, <60% AA sequence similarity with a positive control nuclease, and >80% query coverage. In certain embodiments, the VA type nuclease comprises an ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART28, ART30, ART31, ART32, ART33, ART34, ART35, or ART11* (i.e., ART11_L679F, i.e., ART11, wherein the leucine (L) at amino acid position 679 is replaced with a phenylalanine (F)) nuclease, as shown in Table 2. In certain embodiments, a VA-type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence designated for an individual ART nuclease, as set forth in Table 2. In certain embodiments, a nucleic acid-guided nuclease is provided that comprises a nucleic acid-guided nuclease polypeptide having at least 85% identity to an amino acid sequence represented by SEQ ID NOs: 1-36, or a nucleic acid encoding a nucleic acid-guided nuclease polypeptide comprising at least 85% identity to a polynucleotide represented by SEQ ID NOs: 1-36. In certain embodiments, a nucleic acid-guided nuclease is provided that comprises a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NOs: 1-36, wherein the polypeptide does not contain the peptide motif YLFQIYNKDF (SEQ ID NO: 39).In certain embodiments, nucleic acid-guided nucleases are provided that include a nucleic acid encoding a polypeptide having at least 90% identity to a nucleic acid represented by SEQ ID NOs:808-845, wherein the encoded polypeptide does not contain the peptide motif YLFQIYNKDF (SEQ ID NO:39). In certain embodiments, nucleic acid-guided nucleases are provided that include a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NOs:1-9. In certain embodiments, nucleic acid-guided nucleases are provided that include a polypeptide having at least 90% identity to an amino acid sequence represented by SEQ ID NOs:2, 11, or 36.
[0072] [Table 5]
[0073] [Table 6]
[0074] [Table 7]
[0075] [Table 8]
[0076] [Table 9]
[0077] [Table 10]
[0078] [Table 11]
[0079]
Table 12
[0080]
Table 13
[0081]
Table 14
[0082]
Table 15
[0083] Table 16
[0084] Table 17
[0085] Table 18
[0086] Table 19
[0087] Table 20
[0088] Table 21
[0089] Table 22
[0090] Table 23
[0091] Table 24
[0092] Table 25
[0093] Table 26
[0094] Table 27
[0095] Table 28
[0096] Table 29
[0097]
Table 30
[0098] Table 31
[0099] Table 32
[0100] [Table 33]
[0101] [Table 34]
[0102] In certain embodiments, the Cas nuclease is ABW1 (SEQ ID NO: 3), ABW2 (SEQ ID NO: 16), ABW3 (SEQ ID NO: 29), ABW4 (SEQ ID NO: 42), ABW5 (SEQ ID NO: 55), ABW6 (SEQ ID NO: 68), ABW7 (SEQ ID NO: 81), ABW8 (SEQ ID NO: 94), or ABW9 (SEQ ID NO: 107) (all SEQ ID NOs for ABW1-9 from International (PCT) Application Publication WO 2021 / 108324 and variants thereof) or variants thereof, such as any one of variants 1-10 of ABW1 (SEQ ID NOs: 4-13, respectively), variants 1-10 of ABW2 (SEQ ID NOs: 5-16, respectively), variants 1-10 of ABW3 (SEQ ID NOs: 5-18, respectively), variants 1-10 of ABW4 (SEQ ID NOs: 5-19, respectively), variants 1-10 of ABW5 (SEQ ID NOs: 5-19, respectively), variants 1-10 of ABW6 (SEQ ID NOs: 5-18, respectively), variants 1-10 of ABW7 (SEQ ID NOs: 5-19, respectively), variants 1-10 of ABW8 (SEQ ID NOs: 5-19, respectively), variants 1-10 of ABW9 (SEQ ID NOs: 5-19, respectively), variants 1-10 of ABW1 ... Any one of ABW3 variants 1 to 10 (SEQ ID NOS: 30 to 39, respectively), any one of ABW4 variants 1 to 10 (SEQ ID NOS: 43 to 52, respectively), any one of ABW5 variants 1 to 10 (SEQ ID NOS: 56 to 65, respectively), any one of ABW6 variants 1 to 10 (SEQ ID NOS: 69 to 78, respectively), any one of ABW7 variants 1 to 10 (SEQ ID NOS: 82 to 91, respectively), any one of ABW8 variants 1 to 10 (SEQ ID NOS: 95 to 104, respectively), and any one of ABW9 variants 1 to 10 (SEQ ID NOS: 108 to 117, respectively). ABW1 to ABW9 and variants thereof are known in the art and are described in International PCT Publication WO 2021 / 108324.
[0103] Additional VA-type Cas nucleases and their corresponding native CRISPR-Cas systems can be identified by computational and experimental methods known in the art, such as those described in U.S. Patent No. 9,790,490 and Shmakov et al. (2015) MOL. CELL, 60:385. Exemplary computational methods include homology modeling, structural BLAST, PSI-BLAST, or HHPred analysis of putative Cas proteins, and CRISPR array identification of putative CRISPR loci. Exemplary experimental methods include in vitro cleavage assays and intracellular nuclease assays (e.g., Surveyor assays), such as those described in Zetsche et al. (2015) CELL, 163:759.
[0104] In certain embodiments, the Cas protein is a Cas nuclease that directs cleavage of one or both strands at the target locus, e.g., the target strand (i.e., the strand having a target nucleotide sequence at least partially complementary to and capable of hybridizing to a single guide nucleic acid or dual guide nucleic acid) and / or a non-target strand. In certain embodiments, the Cas nuclease directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more nucleotides from the first or last nucleotide of the target nucleotide sequence or its complement. In certain embodiments, the cleavage is staggered, i.e., generates sticky ends. In certain embodiments, the cleavage generates staggered cleavages with 5' overhangs. In certain embodiments, the cleavage generates staggered cleavages with 5' overhangs of 1 to 5 nucleotides, e.g., 4 or 5 nucleotides. In certain embodiments, the cleavage site is distant from the PAM, for example, cleavage occurs after the 18th nucleotide on the non-target strand and after the 23rd nucleotide on the target strand.
[0105] In certain embodiments, the compositions provided herein comprise a Cas nuclease that can be activated by a compatible guide nucleic acid (gNA), e.g., a gRNA. In certain embodiments, the compositions provided herein further comprise a Cas protein associated with the Cas nuclease that can be activated by a compatible guide nucleic acid (gNA), e.g., a gRNA. For example, in certain embodiments, the Cas protein comprises an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the Cas nuclease amino acid sequence. In certain embodiments, the Cas protein comprises a nuclease-inactive mutant of a Cas nuclease. In certain embodiments, the Cas protein further comprises an effector domain.
[0106] In certain embodiments, the Cas protein lacks substantially all DNA cleavage activity. Such Cas proteins can be generated, for example, by introducing one or more mutations into an active Cas nuclease (e.g., a naturally occurring Cas nuclease). A mutated Cas protein is considered to lack substantially all DNA cleavage activity if the DNA cleavage activity of the protein is about 25%, 10%, 5%, 1%, 0.1%, 0.01% or less of the DNA cleavage activity of the corresponding non-mutated form, e.g., none or negligible compared to the non-mutated form. Thus, Cas proteins can contain one or more mutations (e.g., mutations in the RuvC domain of a VA-type Cas protein) and can be used as general DNA-binding proteins with or without fusion to an effector domain. Exemplary mutations include D908A, E993A, and D1263A relative to amino acid positions in AsCpf1; D832A, E925A, and D1180A relative to amino acid positions in LbCpf1; and D917A, E1006A, and D1255A relative to amino acid position numbering in FnCpf1. Further mutations can be designed and generated according to the crystal structure described in Yamano et al. (2016) CELL, 165:949.
[0107] It is understood that a Cas protein does not lose nuclease activity to cleave all DNA, but may lose the ability to cleave only the target strand or only the non-target strand of double-stranded DNA, thereby functioning as a nickase (see Gao et al. (2016) CELL RES., 26:901). Thus, in certain embodiments, a Cas nuclease is a Cas nickase. In certain embodiments, a Cas nuclease has activity to cleave the non-target strand but substantially lacks activity to cleave the target strand, e.g., due to a mutation in the Nuc domain. In certain embodiments, a Cas nuclease has cleavage activity to cleave the target strand but substantially lacks activity to cleave the non-target strand.
[0108] In certain embodiments, the Cas nuclease has the activity to cleave double-stranded DNA, resulting in a double-strand break.
[0109] Cas proteins that lack substantially all DNA cleavage activity or have the ability to cleave only one strand can also be identified from natural systems. For example, certain natural CRISPR-Cas systems may retain the ability to bind to target nucleotide sequences but lose all or partial DNA cleavage activity in eukaryotic (e.g., mammalian or human) cells. Such VA-type proteins are disclosed, for example, in Kim et al. (2017) ACS SYNTH.BIOL.6(7):1273-82 and Zhang et al. (2017) CELL DISCOV.3:17018.
[0110] The activity of a Cas protein (e.g., a Cas nuclease) can be altered, for example, by generating an engineered Cas protein. In certain embodiments, the altered activity of the engineered Cas protein includes increased targeting efficiency and / or decreased off-target binding. Without wishing to be bound by theory, it is hypothesized that off-target binding may be recognized by the Cas protein due to, for example, the presence of one or more mismatches between the spacer sequence and the target nucleotide sequence, which may affect the stability and / or conformation of the CRISPR-Cas complex. In certain embodiments, the altered activity includes modified binding, such as increased binding to the target locus (e.g., the target strand or a non-target strand) and / or decreased binding to an off-target locus. In certain embodiments, the altered activity includes altered charge in a region of the protein associated with a single guide nucleic acid or a dual guide nucleic acid. In certain embodiments, the altered activity of the engineered Cas protein includes altered charge in a region of the protein associated with the target strand and / or a non-target strand. In certain embodiments, the altered activity of the engineered Cas protein comprises an altered charge in a region of the protein associated with an off-target locus. The altered charge can comprise a decreased positive charge, a decreased negative charge, an increased positive charge, or an increased negative charge. For example, a decreased negative charge and an increased positive charge can generally strengthen binding to a nucleic acid, while a decreased positive charge and an increased negative charge can weaken binding to a nucleic acid. In certain embodiments, the altered activity comprises increased or decreased steric hindrance between the protein and a single-guide nucleic acid or a dual-guide nucleic acid. In certain embodiments, the altered activity comprises increased or decreased steric hindrance between the protein and a target strand and / or a non-target strand. In certain embodiments, the altered activity comprises increased or decreased steric hindrance between the protein and an off-target locus. In certain embodiments, the modification or mutation comprises one or more substitutions of Lys, His, Arg, Glu, Asp, Ser, Gly, and / or Thr. In certain embodiments, the modification or mutation comprises one or more substitutions of Gly, Ala, Ile, Glu, and / or Asp.In certain embodiments, the modification or mutation comprises one or more amino acid substitutions in the groove between the WED and RuvC domains of a Cas protein (e.g., a VA-type Cas protein).
[0111] In certain embodiments, the altered activity of the engineered Cas protein comprises increased nuclease activity for cleaving the target locus. In certain embodiments, the altered activity of the engineered Cas protein comprises decreased nuclease activity for cleaving the off-target locus. In certain embodiments, the altered activity of the engineered Cas protein comprises altered helicase kinetics. In certain embodiments, the engineered Cas protein comprises a modification that alters the formation of a CRISPR complex.
[0112] In certain embodiments, a protospacer adjacent motif (PAM) or PAM-like motif directs binding of the Cas protein complex to the target locus. Many Cas proteins have PAM specificity. The exact sequence and length requirements of the PAM vary depending on the Cas protein used. The PAM sequence is typically 2-5 base pairs in length and is adjacent to the target nucleotide sequence (but located on a different strand of the target DNA than the target nucleotide sequence). PAM sequences can be identified using any suitable method, such as cleavage assays, targeting, or modification of the target nucleotide sequence and an oligonucleotide with a different PAM sequence.
[0113] Exemplary PAM sequences are shown in Tables 2 and 3. In certain embodiments, the Cas protein comprises MAD7, the PAM is TTTN, and N is A, C, G, or T. In certain embodiments, the Cas protein comprises MAD7, the PAM is CTTN, and N is A, C, G, or T. In certain embodiments, the Cas protein comprises AsCpf1, the PAM is TTTN, and N is A, C, G, or T. In certain embodiments, the Cas protein comprises FnCpf1, the PAM is 5'TTN, and N is A, C, G, or T. PAM sequences for certain other VA-type Cas proteins are disclosed in Zetsche et al. (2015) CELL, 163:759 and U.S. Patent No. 9,982,279. Furthermore, engineering the PAM-interacting (PI) domain of a Cas protein may allow programming of PAM specificity, improving target site recognition fidelity and / or increasing the versatility of engineered non-natural systems. An exemplary approach for modifying the PAM specificity of Cpf1 is described in Gao et al. (2017) NAT. BIOTECHNOL., 35:789.
[0114] In certain embodiments, engineered Cas proteins contain modifications that alter Cas protein specificity in conjunction with modifications to targeting scope. Cas mutants can be designed to have increased target specificity and corresponding modifications in PAM recognition, for example, by selecting mutations that alter PAM specificity (e.g., in the PI domain) and combining those mutations with groove mutations that increase (or, optionally, decrease) specificity for on-target loci relative to off-target loci. The Cas modifications described herein can be used to reduce the loss of specificity resulting from altered PAM recognition, enhance the increase in specificity resulting from altered PAM recognition, reduce the increase in specificity resulting from altered PAM recognition, or enhance the loss of specificity resulting from altered PAM recognition.
[0115] In certain embodiments, the engineered Cas protein comprises one or more nuclear localization signal (NLS) motifs, hi certain embodiments, the engineered Cas protein comprises at least two (e.g., at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs. Non-limiting examples of NLS motifs include the SV40 large T antigen NLS having the amino acid sequence of PKKKRKV (SEQ ID NO: 40); an NLS from nucleoplasmin, such as the nucleoplasmin bipartite NLS having the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 41); a c-myc NLS having the amino acid sequence of PAAKRVKLD (SEQ ID NO: 42) or RQRRNELKRSP (SEQ ID NO: 43); an hRNPA1 M9 NLS having the amino acid sequence of NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 44); an importin-α NLS having the amino acid sequence of RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 45). IBB domain NLS; fibroid T protein NLS having the amino acid sequence of VSRKRPRP (SEQ ID NO: 46) or PPKKARED (SEQ ID NO: 47); human p53 NLS having the amino acid sequence of PQPKKKPL (SEQ ID NO: 48); mouse c-abl IV NLS having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO: 49); influenza virus NS1 NLS having the amino acid sequence of DRLRR (SEQ ID NO: 50) or PKQKKRK (SEQ ID NO: 51); hepatitis virus delta antigen NLS having the amino acid sequence of RKLKKKIKKL (SEQ ID NO: 52); mouse Mx1 protein NLS having the amino acid sequence of REKKKFLKRR (SEQ ID NO: 53); human poly(ADP-ribose) polymerase NLS having the amino acid sequence of KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 54); human glucocorticoid receptor NLS having the amino acid sequence of RKCLQAGMNLEARKTKK (SEQ ID NO: 55), as well as synthetic NLS motifs such as PAAKKKKLD (SEQ ID NO: 56).
[0116] Generally, the one or more NLS motifs are strong enough to promote the accumulation of detectable amounts of the Cas protein in the nucleus of a eukaryotic cell. The strength of the nuclear localization activity can depend on the NLS motifs in the Cas protein, the number of specific NLS motifs used, the location of the NLS motifs, or a combination of these and / or other factors. In certain embodiments, the engineered Cas protein contains at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the N-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N-terminus). In certain embodiments, the engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the C-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the C-terminus). In certain embodiments, the engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the C-terminus and at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motif at or near the N-terminus. In certain embodiments, the engineered Cas protein comprises one, two, or three NLS motifs at or near the C-terminus. In certain embodiments, the engineered Cas protein comprises one NLS motif at or near the N-terminus and one, two, or three NLS motifs at or near the C-terminus. In certain embodiments, the engineered Cas protein comprises a nucleoplasmic NLS at or near the C-terminus.
[0117] Detection of nuclear accumulation can be performed by any suitable technique. For example, a detectable marker can be fused to the nucleic acid targeting protein so that its location within the cell can be visualized. Cell nuclei can also be isolated from cells, and their contents can then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assays. Nuclear accumulation can also be determined indirectly, such as by assays that detect the effect of transport of the Cas protein complex into the nucleus (e.g., assays for DNA breaks or mutations at the target locus or assays for altered gene expression activity) compared to controls not exposed to the Cas protein or exposed to a Cas protein lacking one or more of the NLS motifs.
[0118] The Cas protein may include a chimeric Cas protein, e.g., a Cas protein with enhanced function due to its chimeric nature. A chimeric Cas protein may be a novel Cas protein containing fragments from two or more naturally occurring Cas proteins or variants thereof. For example, fragments of multiple VA-type Cas homologs (e.g., orthologs) may be fused to form a chimeric Cas protein. In certain embodiments, the chimeric Cas protein comprises fragments of Cpf1 orthologs from multiple species and / or strains.
[0119] In certain embodiments, a Cas protein comprises one or more effector domains. The one or more effector domains may be located at or near the N-terminus of the Cas protein and / or at or near the C-terminus of the Cas protein. In certain embodiments, the effector domain comprised in the Cas protein is a transcription activation domain (e.g., VP64), a transcription repression domain (e.g., a KRAB domain or a SID domain), an exogenous nuclease domain (e.g., FokI), a deaminase domain (e.g., a cytidine deaminase or an adenine deaminase), or a reverse transcriptase domain (e.g., a high-fidelity reverse transcriptase domain). Other activities of effector domains include, but are not limited to, methylase activity, demethylase activity, transcription release factor activity, translation initiation activity, translation activation activity, translation repression activity, histone modification (e.g., acetylation or demethylation) activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity.
[0120] In certain embodiments, the Cas protein comprises one or more protein domains that enhance homology-directed repair (HDR) and / or inhibit non-homologous end joining (NHEJ). Exemplary protein domains with such functionality are described in Jayavaradhan et al. (2019) NAT.COMMUN.10(1):2866 and Janssen et al. (2019) MOL.THER.NUCLEIC ACIDS 16:141-54. In certain embodiments, the Cas protein comprises a dominant-negative form of p53-binding protein 1 (53BP1), such as a fragment of 53BP1 that contains a minimal focus-forming region (e.g., amino acids 1231-1644 of human 53BP1). In certain embodiments, the Cas protein comprises a motif targeted by APC-Cdh1, such as amino acids 1-110 of human geminin, thereby resulting in degradation of the fusion protein during the HDR-nonpermissive G1 phase of the cell cycle.
[0121] In certain embodiments, the Cas protein comprises an inducible or regulatory domain. Non-limiting examples of inducers or regulators include light, hormones, and small molecule drugs. In certain embodiments, the Cas protein comprises a light-inducible or regulatory domain. In certain embodiments, the Cas protein comprises a chemical-inducible or regulatory domain.
[0122] In certain embodiments, the Cas protein comprises a tag protein or peptide to facilitate tracking and / or purification. Non-limiting examples of tag proteins and peptides include fluorescent proteins (e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato), HIS tags (e.g., 6xHis tag or gly-6xHis; 8xHis or gly-8xHis), hemagglutinin (HA) tags, FLAG tags, 3xFLAG tags, and Myc tags.
[0123] In certain embodiments, the Cas protein is conjugated to a non-protein moiety, such as a fluorophore useful for genomic imaging. In certain embodiments, the Cas protein is covalently conjugated to the non-protein moiety. The terms "CRISPR-associated protein," "Cas protein," "Cas," "CRISPR-associated nuclease," and "Cas nuclease" are used herein to include such conjugates, regardless of the presence of one or more non-protein moieties.
[0124] B. Guide Nucleic Acid A guide nucleic acid can be a single gNA (sgNA, e.g., sgRNA), in which the gNA is a single polynucleotide, or a dual gNA (e.g., dual gRNA), in which the gNA comprises two separate polynucleotides (which may in some cases be covalently linked, but not via a conventional internucleotide linkage). In certain embodiments, a single guide nucleic acid can activate a Cas nuclease by itself (e.g., in the absence of a tracrRNA).
[0125] Generally, a gNA comprises a modulator nucleic acid and a targeter nucleic acid. In a single gNA, the modulator and targeter nucleic acid are part of a single polynucleotide. In a dual gNA, the modulator and targeter nucleic acids are separate, for example, not linked by a conventional nucleotide bond, for example, not linked at all. The targeter nucleic acid comprises a spacer sequence and a targeter stem sequence. The modulator nucleic acid comprises a modulator stem sequence and generally further nucleotides, for example, nucleotides including a 5' tail. The modulator stem sequence and the targeter stem sequence may each comprise any suitable number of nucleotides and have sufficient complementarity to allow them to hybridize. In a single gNA, additional nucleotides may be present between the targeter stem sequence and the modulator stem sequence, which may form a secondary structure such as a loop in certain cases.
[0126] In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid capable of binding to a Cas protein in combination with a modulator nucleic acid. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid capable of activating a Cas nuclease in combination with a modulator nucleic acid. In certain embodiments, the system further comprises a Cas protein to which the targeter nucleic acid and modulator nucleic acid can bind, or a Cas nuclease that the targeter nucleic acid and modulator nucleic acid can activate.
[0127] The single or dual guide nucleic acid may need to be compatible with a Cas protein (e.g., a Cas nuclease) to provide an operational CRISPR system. For example, the targeter stem sequence and modulator stem sequence may be derived from a naturally occurring crRNA that can activate a Cas nuclease in the absence of a tracrRNA. Alternatively, the targeter stem sequence and modulator stem sequence may be derived from a naturally occurring pair of crRNA and tracrRNA, respectively, that can activate a Cas nuclease. In certain embodiments, the nucleotide sequences of the targeter stem sequence and modulator stem sequence are identical to the corresponding stem sequences of the stem-loop structure in such a naturally occurring crRNA.
[0128] Guide nucleic acid sequences that work with Type II or Type V Cas proteins are known in the art and are disclosed, for example, in U.S. Patent Nos. 9,790,490, 9,896,696, 10,113,179, and 10,266,850, and U.S. Patent Application Publication No. 2014 / 0242664. It is understood that these sequences are exemplary only, and that other guide nucleic acid sequences can also be used with these Cas proteins.
[0129] [Table 35]
[0130] [Table 36]
[0131] [Table 37]
[0132] [Table 38]
[0133] [Table 39]
[0134] In certain embodiments, the guide nucleic acid comprises a targeter stem sequence listed in Table 4 for a type VA CRISPR-Cas system. Targeter stem sequences that are the same as portions of the scaffold sequence are bolded and underlined in Table 3.
[0135] In certain embodiments, the guide nucleic acid is a single guide nucleic acid comprising, from 5' to 3', a modulator stem sequence, a loop sequence, a targeter stem sequence, and a spacer sequence. In certain embodiments, the targeter stem sequence in the single guide nucleic acid is listed in Table 3 as the bold, underlined portion of the scaffold sequence, and the modulator stem sequence is complementary (e.g., 100% complementary) to the targeter stem sequence. In certain embodiments, the single guide nucleic acid comprises, from 5' to 3', a modulator sequence listed in Table 3 as the underlined portion of the scaffold sequence, a loop sequence, a targeter stem sequence that is the bold, underlined portion of the same scaffold sequence, and a spacer sequence. In certain embodiments, the engineered non-natural system comprises a single guide nucleic acid comprising a scaffold sequence listed in Table 3. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identical to the amino acid sequence set forth in a SEQ ID NO: listed in the same row of Table 3. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising the amino acid sequence set forth in a SEQ ID NO: listed in the same row of Table 3. In certain embodiments, the system is useful for targeting, editing or modifying nucleic acids comprising a target nucleotide sequence near or adjacent to (e.g., immediately downstream of) a PAM listed in the same row of Table 3 when using a non-target strand (i.e., a strand that does not hybridize to the spacer sequence) as a coordinate.
[0136] In certain embodiments, the guide nucleic acid, e.g., a dual gNA, comprises a targeter guide nucleic acid comprising, from 5' to 3', a targeter stem sequence and a spacer sequence. In certain embodiments, the targeter stem sequence in the targeter nucleic acid is listed in Table 4. In certain embodiments, the engineered non-natural system comprises a targeter nucleic acid and a modulator stem sequence that is complementary (e.g., 100% complementary) to the targeter stem sequence. In certain embodiments, the modulator nucleic acid comprises a modulator sequence listed in the same row of Table 4. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in a SEQ ID NO: listed in the same row of Table 4. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) comprising the amino acid sequence set forth in a SEQ ID NO: listed in the same row of Table 4. In certain embodiments, the system is useful for targeting, editing, or modifying a nucleic acid comprising a target nucleotide sequence near or adjacent to (e.g., immediately downstream of) a PAM listed in the same row of Table 4 when using a non-target strand (i.e., a strand that does not hybridize to the spacer sequence) as a coordinate.
[0137] Single guide nucleic acids, targeter nucleic acids, and / or modulator nucleic acids can be chemically synthesized or generated in biological processes (e.g., catalyzed by RNA polymerase in an in vitro reaction). Such reactions or processes can limit the length of the single guide nucleic acids, targeter nucleic acids, and / or modulator nucleic acids. In certain embodiments, single guide nucleic acids are no longer than 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides in length. In certain embodiments, single guide nucleic acids are at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the single guide nucleic acid is 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60 , 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60, 60-100, 60-90, 60-80, 60-70, 70-100, 70-90, 70-80, 80-100, 80-90, or 90-100 nucleotides in length. In certain embodiments, the targeter nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides in length or less. In certain embodiments, the targeter nucleic acid is at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length.In certain embodiments, the targeter nucleic acid is 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, The modulator nucleic acid may be 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60, 60-100, 60-90, 60-80, 60-70, 70-100, 70-90, 70-80, 80-100, 80-90, or 90-100 nucleotides in length. In certain embodiments, the modulator nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 20 nucleotides or less in length. In certain embodiments, the modulator nucleic acid is at least 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the modulator nucleic acid is selected from the group consisting of 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 15 to 20, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 The length is up to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 100, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 100, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 100, 60 to 90, 60 to 80, 60 to 70, 70 to 100, 70 to 90, 70 to 80, 80 to 100, 80 to 90, or 90 to 100 nucleotides.
[0138] It is believed that the length of the duplex formed within a single guide nucleic acid or between a targeter nucleic acid and a modulator nucleic acid, e.g., in a dual guide nucleic acid, can be a factor in providing an operational CRISPR system. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 10 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4, 5, 6, 7, 8, 9, or 10 nucleotides. It is understood that the composition of the nucleotides in each sequence affects the stability of the duplex, with CG base pairs conferring greater stability than AU base pairs. In certain embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the base pairs are CG base pairs.
[0139] In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of five nucleotides. Thus, the targeter stem sequence and the modulator stem sequence form a five-base pair duplex. In certain embodiments, 0 to 4, 0 to 3, 0 to 2, 0 to 1, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 5, 2 to 4, 2 to 3, 3 to 5, 3 to 4, or 4 to 5 of the five base pairs are CG base pairs. In certain embodiments, 0, 1, 2, 3, 4, or 5 of the five base pairs are CG base pairs. In certain embodiments, the targeter stem sequence consists of 5'-GUAGA-3' and the modulator stem sequence consists of 5'-UCUAC-3'. In certain embodiments, the targeter stem sequence consists of 5'-GUGGG-3' and the modulator stem sequence consists of 5'-CCCAC-3'.
[0140] In certain embodiments, in a VA-type system, the 3' end of the targeter stem sequence is linked to the 5' end of the spacer sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are adjacent to each other and directly linked by an inter-polynucleotide bond. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by a single nucleotide, such as a uridine. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by two or more nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.
[0141] In certain embodiments, the targeter nucleic acid further comprises an additional nucleotide sequence 5' to the targeter stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least one nucleotide (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50). In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In certain embodiments, the additional nucleotide sequence consists of two nucleotides. In certain embodiments, the additional nucleotide sequence is similar to a loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 3' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It can be appreciated that additional nucleotide sequences 5' to the targeter stem sequence may not be essential. Thus, in certain embodiments, the targeter nucleic acid does not include any additional nucleotides 5' to the targeter stem sequence.
[0142] In certain embodiments, the targeter nucleic acid or single guide nucleic acid further comprises an additional nucleotide sequence containing one or more nucleotides at the 3' end that do not hybridize to the target nucleotide sequence. The additional nucleotide sequence may protect the targeter nucleic acid from degradation by 3'-5' exonucleases. In certain embodiments, the additional nucleotide sequence is 100 nucleotides or less in length. In certain embodiments, the additional nucleotide sequence is 90, 80, 70, 60, 50, 40, 30, 20, or 10 nucleotides or less in length. In certain embodiments, the additional nucleotide sequence is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides in length. In certain embodiments, the further nucleotide sequence is 5 to 100, 5 to 50, 5 to 40, 5 to 30, 5 to 25, 5 to 20, 5 to 15, 5 to 10, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 10 to 25, 10 to 20, 10 to 15, 15 to 100, 15 to 50, 15 to 40, 15 to 30, 15 to 25, 15 to 20, 20 to 100, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 50, 30 to 40, 40 to 100, 40 to 50 or 50 to 100 nucleotides in length.
[0143] In certain embodiments, the additional nucleotide sequence forms a hairpin together with the spacer sequence. Such secondary structures can increase the specificity of the guide nucleic acid or engineered non-natural system (see Kocak et al. (2019) Nat. Biotech. 37:657-66). In certain embodiments, the free energy change during hairpin formation is -20 kcal / mol, -15 kcal / mol, -14 kcal / mol, -13 kcal / mol, -12 kcal / mol, -11 kcal / mol, or -10 kcal / mol or greater. In certain embodiments, the free energy change during hairpin formation is -5 kcal / mol, -6 kcal / mol, -7 kcal / mol, -8 kcal / mol, -9 kcal / mol, -10 kcal / mol, -11 kcal / mol, -12 kcal / mol, -13 kcal / mol, -14 kcal / mol, or -15 kcal / mol or greater. In certain embodiments, the free energy change during hairpin formation is between -20 and -10 kcal / mol, -20 and -11 kcal / mol, -20 and -12 kcal / mol, -20 and -13 kcal / mol, -20 and -14 kcal / mol, -20 and -15 kcal / mol, -15 and -10 kcal / mol, -15 and -11 kcal / mol, -15 and -12 kcal / mol, -15 and -13 kcal / mol ol, -15 to -14 kcal / mol, -14 to -10 kcal / mol, -14 to -11 kcal / mol, -14 to -12 kcal / mol, -14 to -13 kcal / mol, -13 to -10 kcal / mol, -13 to -11 kcal / mol, -13 to -12 kcal / mol, -12 to -10 kcal / mol, -12 to -11 kcal / mol, or -11 to -10 kcal / mol. In other embodiments, the targeter nucleic acid or single guide nucleic acid does not contain any nucleotides 3' to the spacer sequence.
[0144] In certain embodiments, the modulator nucleic acid further comprises an additional nucleotide sequence 3' to the modulator stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotide. In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. The additional nucleotide sequence consists of one nucleotide (e.g., uridine). In certain embodiments, the additional nucleotide sequence consists of two nucleotides. In certain embodiments, the additional nucleotide sequence is similar to a loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 5' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It is understood that additional nucleotide sequences 3' to the modulator stem sequence may not be required. Thus, in certain embodiments, the modulator nucleic acid does not include any additional nucleotides 3' to the modulator stem sequence.
[0145] It is understood that the additional nucleotide sequence 5' to the targeter stem sequence and the additional nucleotide sequence 3' to the modulator stem sequence, if present, may interact with each other. For example, the nucleotide immediately 5' to the targeter stem sequence and the nucleotide immediately 3' to the modulator stem sequence do not form Watson-Crick base pairs (indeed, they may constitute part of the targeter stem sequence and part of the modulator stem sequence, respectively), but other nucleotides in the additional nucleotide sequence 5' to the targeter stem sequence and the additional nucleotide sequence 3' to the modulator stem sequence may form one, two, three, or more base pairs (e.g., Watson-Crick base pairs). Such interactions may affect the stability of a complex comprising the targeter nucleic acid and the modulator nucleic acid.
[0146] The stability of a complex containing a target nucleic acid and a modulator nucleic acid can be evaluated by the Gibbs free energy change (ΔG) during complex formation, either calculated or actually measured. When all predicted base pairings of the complex occur between bases in the target nucleic acid and bases in the modulator nucleic acid, i.e., when no intrastrand secondary structure is present, ΔG during complex formation generally correlates with ΔG during the formation of a secondary structure within the corresponding single-guide nucleic acid. Methods for calculating or measuring ΔG are known in the art. An exemplary method is RNAfold (rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi), as disclosed in Gruber et al. (2008) Nucleic Acids Res., 36 (Web Server issue): W70-W74. Unless otherwise indicated, ΔG values in the present disclosure are calculated by RNAfold for the formation of a secondary structure within the corresponding single-guide nucleic acid. In certain embodiments, ΔG is -1 kcal / mol or less, e.g., -2 kcal / mol or less, -3 kcal / mol or less, -4 kcal / mol or less, -5 kcal / mol or less, -6 kcal / mol or less, -7 kcal / mol or less, -7.5 kcal / mol or less, or -8 kcal / mol or less. In certain embodiments, ΔG is -10 kcal / mol or more, e.g., -9 kcal / mol or more, -8.5 kcal / mol or more, or -8 kcal / mol or more. In certain embodiments, ΔG is in the range of -10 to -4 kcal / mol. In certain embodiments, ΔG is in the range of −8 to −4 kcal / mol, −7 to −4 kcal / mol, −6 to −4 kcal / mol, −5 to −4 kcal / mol, −8 to −4.5 kcal / mol, −7 to −4.5 kcal / mol, −6 to −4.5 kcal / mol, or −5 to −4.5 kcal / mol.In certain embodiments, ΔG is about −8 kcal / mol, −7 kcal / mol, −6 kcal / mol, −5 kcal / mol, −4.9 kcal / mol, −4.8 kcal / mol, −4.7 kcal / mol, −4.6 kcal / mol, −4.5 kcal / mol, −4.4 kcal / mol, −4.3 kcal / mol, −4.2 kcal / mol, −4.1 kcal / mol, or −4 kcal / mol.
[0147] It is understood that ΔG can be affected by sequences in the targeter nucleic acid that are not within the targeter stem sequence and / or sequences in the modulator nucleic acid that are not within the modulator stem sequence. For example, one or more base pairs (e.g., Watson-Crick base pairs) between additional sequences 5' to the targeter stem sequence and additional sequences 3' to the modulator stem sequence can decrease ΔG, i.e., stabilize the nucleic acid complex. In certain embodiments, the nucleotide immediately 5' to the targeter stem sequence contains uracil or is a uridine, and the nucleotide immediately 3' to the modulator stem sequence contains uracil or is a uridine, thereby forming a non-conventional UU base pair.
[0148] In certain embodiments, the modulator nucleic acid or single-guide nucleic acid comprises a nucleotide sequence located 5' to the modulator stem sequence, referred to herein as the "5' tail." In native type VA CRISPR-Cas systems, the 5' tail is the nucleotide sequence located 5' to the stem-loop structure of the crRNA. The 5' tail in an engineered type VA CRISPR-Cas system can resemble the 5' tail in the corresponding native type VA CRISPR-Cas system, whether single-guide or dual-guide.
[0149] Without being bound by theory, it is believed that the 5' tail may be involved in the formation of a CRISPR-Cas complex. For example, in certain embodiments, the 5' tail forms a pseudoknot structure with the modulator stem sequence, which is recognized by the Cas protein (see Yamano et al. (2016) Cell, 165:949). In certain embodiments, the 5' tail is at least 3 (e.g., at least 4 or at least 5) nucleotides in length. In certain embodiments, the 5' tail is 3, 4, or 5 nucleotides in length. In certain embodiments, the nucleotide at the 3' end of the 5' tail comprises uracil or is uridine. In certain embodiments, the nucleotide at the second position from the 3' end of the 5' tail comprises uracil or is uridine. In certain embodiments, the nucleotide at the third position from the 3' end of the 5' tail comprises adenine or is adenosine. This third nucleotide can form a base pair (e.g., a Watson-Crick base pair) with the nucleotide 5' to the modulator stem sequence. Thus, in certain embodiments, the modulator nucleic acid comprises a uridine- or uracil-containing nucleotide 5' to the modulator stem sequence. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-AUU-3'. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-AAUU-3'. In certain embodiments, the 5' tail comprises the nucleotide sequence 5'-UAAUU-3'. In certain embodiments, the 5' tail is located immediately 5' to the modulator stem sequence.
[0150] In certain embodiments, the single guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid are designed to reduce the degree of secondary structure other than hybridization between the targeter stem sequence and the modulator stem sequence. In certain embodiments, less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of the single guide nucleic acid, other than the targeter stem sequence and the modulator stem sequence, participate in self-complementary base pairing when optimally folded. In certain embodiments, less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of the targeter nucleic acid and / or modulator nucleic acid participate in self-complementary base pairing when optimally folded. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on minimum Gibbs free energy calculations. One example of such an algorithm is mFold, as described in Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example of a folding algorithm is the online web server RNAfold, developed by the Institute for Theoretical Chemistry at the University of Vienna, which uses a centroid structure prediction algorithm (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and P.A. Carr and G.M. Church, 2009, Nature Biotechnology 27(12):1151-62).
[0151] The targeter nucleic acid is directed to a specific target nucleotide sequence, and the donor template can be designed to modify the target nucleotide sequence or a sequence nearby. It is therefore understood that binding a single guide nucleic acid, targeter nucleic acid, or modulator nucleic acid to a donor template can increase editing efficiency and reduce off-target effects. Thus, in certain embodiments, the single guide nucleic acid or modulator nucleic acid further comprises a donor template recruitment sequence capable of hybridizing with the donor template (see Figure 2B). Donor templates are described in the "Donor Template" subsection of Section II below. The donor template and donor template recruitment sequence can be designed so that they have sequence complementarity. In certain embodiments, the donor template recruitment sequence is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) complementary to at least a portion of the donor template. In certain embodiments, the donor template recruitment sequence is 100% complementary to at least a portion of the donor template. In certain embodiments, if the donor template contains an engineered sequence that is not homologous to the sequence to be repaired, the donor template recruitment sequence can hybridize to the engineered sequence in the donor template. In certain embodiments, the donor template recruitment sequence is at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length. In certain embodiments, the donor template recruitment sequence is located at or near the 5' end of the single guide nucleic acid or the 5' end of the modulator nucleic acid. In certain embodiments, the donor template recruitment sequence is linked to the 5' tail or to the modulator stem sequence, if present, of the single guide nucleic acid or the modulator nucleic acid via an interpolynucleotide bond or a nucleotide linker.
[0152] In certain embodiments, the single guide nucleic acid or modulator nucleic acid further comprises an editing enhancer sequence, which enhances the efficiency of gene editing and / or homology-directed repair (HDR) (see Figure 2C). Exemplary editing enhancer sequences are described in Park et al. (2018) Nat. Commun. 9:3313. In certain embodiments, the editing enhancer sequence, if present, is located 5' to the 5' tail or 5' to the single guide nucleic acid or modulator stem sequence. In certain embodiments, the editing enhancer sequence is 1-50, 4-50, 9-50, 15-50, 25-50, 1-25, 4-25, 9-25, 15-25, 1-15, 4-15, 9-15, 1-9, 4-9, or 1-4 nucleotides in length. In certain embodiments, the editing enhancer sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 55 nucleotides in length. The editing enhancer sequence is designed to minimize homology with the target nucleotide sequence or any other sequences that the engineered non-native system may contact, such as the genomic sequence of the cell to which the engineered non-native system is delivered. In certain embodiments, the editing enhancer is designed to minimize the presence of hairpin structures. The editing enhancer may include one or more of the chemical modifications disclosed herein.
[0153] The single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid may further comprise a protective nucleotide sequence that prevents or reduces nucleic acid degradation. In certain embodiments, the protective nucleotide sequence is at least 5 (e.g., at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides in length. The length of the protective nucleotide sequence increases the time it takes exonucleases to reach the 5' tail, modulator stem sequence, targeter stem sequence, and / or spacer sequence, thereby protecting these portions of the single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid from exonucleolytic degradation. In certain embodiments, the protective nucleotide sequence forms a secondary structure, such as a hairpin or tRNA structure, to reduce the rate of exonucleolytic degradation (see, e.g., Wu et al. (2018) Cell. Mol. Life Sci., 75(19):3593-3607). Secondary structure can be predicted by methods known in the art, such as the online web server RNAfold, developed by the University of Vienna and using a centroid structure prediction algorithm (see Gruber et al. (2008) Nucleic Acids Res., 36:W70). Certain chemical modifications that may be present in the protected nucleotide sequence can also prevent or reduce nucleic acid degradation, as disclosed in the "RNA Modifications" subsection below.
[0154] The protected nucleotide sequence is typically located at the 5' or 3' end of the single guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid. In certain embodiments, the single guide nucleic acid comprises a protected nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises a protected nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises a protected nucleotide sequence at the 5' end (see Figure 2A). In certain embodiments, the targeter nucleic acid comprises a protected nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker.
[0155] As described above, various nucleotide sequences can be present in the 5' portion of a single nucleic acid or modulator nucleic acid, including, but not limited to, donor template recruitment sequences, editing enhancer sequences, protective nucleotide sequences, and linkers connecting such sequences to the 5' tail or modulator stem sequence, if present. It is understood that the functions of donor template recruitment, editing enhancement, protection against degradation, and binding are not mutually exclusive, and a single nucleotide sequence can have one or more of such functions. For example, in certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruitment sequence and an editing enhancer sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruitment sequence and a protective sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is both an editing enhancer sequence and a protective sequence. In certain embodiments, a single guide nucleic acid or modulator nucleic acid comprises a nucleotide sequence that is a donor template recruitment sequence, an editing enhancer sequence, and a protective sequence. In certain embodiments, if present, the nucleotide sequence 5' to the 5' tail or 5' to the modulator stem sequence is from 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 20 to 90, 20 to 80, The length is 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 90, 60 to 80, 60 to 70, 70 to 90, 70 to 80, or 80 to 90 nucleotides.
[0156] In certain embodiments, the engineered non-natural system further comprises one or more compounds (e.g., small molecule compounds) that enhance HDR and / or inhibit NHEJ. Exemplary compounds with such functionality are described in Maruyama et al. (2015) Nat Biotechnol. 33(5):538-42; Chu et al. (2015) Nat Biotechnol. 33(5):543-48; Yu et al. (2015) Cell Stem Cell 16(2):142-47; Pinder et al. (2015) Nucleic Acids Res. 43(19):9379-92; and Yagiz et al. (2019) Commun. Biol. 2:198. In certain embodiments, the engineered non-natural system further comprises one or more compounds selected from the group consisting of DNA ligase IV antagonists (e.g., SCR7 compounds, Ad4 E1B55K protein, and Ad4 E4orf6 protein), RAD51 agonists (e.g., RS-1), DNA-dependent protein kinase (DNA-PK) antagonists (e.g., NU7441 and KU0060648), beta3 adrenergic receptor agonists (e.g., L755507), inhibitors of intracellular protein transport from the ER to the Golgi apparatus (e.g., Brefeldin A), and any combination thereof.
[0157] In certain embodiments, an engineered non-natural system comprising a targeter nucleic acid and a modulator nucleic acid is regulatable or inducible. For example, in certain embodiments, the targeter nucleic acid, modulator nucleic acid, and / or Cas protein can be introduced into the target nucleotide sequence at different times, and the system is active only when all components are present. In certain embodiments, the amount of targeter nucleic acid, modulator nucleic acid, and / or Cas protein can be adjusted to achieve desired efficiency and specificity. In certain embodiments, an excess amount of nucleic acid comprising a targeter stem sequence or a modulator stem sequence can be added to the system, thereby dissociating the targeter nucleic acid and modulator nucleic acid complex and turning off the system.
[0158] C.gNA modification Guide nucleic acids, including single guide nucleic acids, targeter nucleic acids, and / or modulator nucleic acids, can comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, single guide nucleic acids comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, targeter nucleic acids comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, modulator nucleic acids comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. A spacer sequence may be designated as a DNA sequence by including thymidine (T) rather than uridine (U). It is understood that corresponding RNA sequences and DNA / RNA chimeric sequences are also contemplated. For example, if the spacer sequence is RNA, the sequence may be derived from a DNA sequence disclosed herein by replacing each T with U. Consequently, T and U are used interchangeably herein to describe nucleotide sequences.
[0159] In certain embodiments, the engineered non-natural system comprises a targeter nucleic acid comprising: a target nucleotide sequence and a spacer sequence designed to hybridize with the targeter stem sequence; and a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence, and optionally a 5' sequence, e.g., a tail sequence, wherein in a single guide nucleic acid, the targeter nucleic acid and modulator nucleic acid are part of a single polynucleotide; in a dual guide nucleic acid, the targeter nucleic acid and modulator nucleic acid are separate nucleic acids, and the modifications may comprise one or more chemical modifications to one or more nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid (dual and single gNA), at or near the 5' end of the targeter nucleic acid (dual gNA), at or near the 3' end of the modulator nucleic acid (dual gNA), at or near the 5' end of the modulator nucleic acid (single and dual gNA), or combinations thereof, as appropriate for single or dual gNA. In certain embodiments, the Cas nuclease is a VA-type Cas nuclease. The modulator and / or targeter nucleic acid sequences can include additional sequences, as detailed in the guide nucleic acid section, and modifications may be made in these additional sequences, as desired and as would be apparent to one of skill in the art. In the embodiments described in this section below, in certain embodiments, the guide nucleic acid is oriented 5' to 3' of the modulator stem sequence and 5' to 3' of the targeter stem sequence (see, e.g., Figures 1A and 1B), and in certain embodiments, optionally, the guide nucleic acid is oriented 3' to 5' of the modulator stem sequence and 3' to 5' of the targeter stem sequence.
[0160] The targeter nucleic acid may comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. The modulator nucleic acid may comprise DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, the targeter nucleic acid is RNA and the modulator nucleic acid is RNA. A targeter nucleic acid in the form of RNA is also referred to as a targeter RNA, and a modulator nucleic acid in the form of RNA is also referred to as a modulator RNA. The nucleotide sequences disclosed herein are represented as DNA sequences by including RNA sequences that include thymidine (T) and / or uridine (U). It is understood that corresponding DNA sequences, RNA sequences, and DNA / RNA chimeric sequences are also contemplated. For example, if a spacer sequence is represented as a DNA sequence, a nucleic acid comprising this spacer sequence as RNA can be derived from the DNA sequence disclosed herein by replacing each T with U. Consequently, T and U are used interchangeably herein to describe nucleotide sequences.
[0161] In certain embodiments, some or all of the gNAs are RNA, e.g., gRNAs. In certain embodiments, 5-100%, 10-100%, 20-100%, 30-100%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, 95-100%, 99-100%, or 99.5-100% of the gNAs are gRNAs. In certain embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the gNAs are RNA. In certain embodiments, 50% of the gNAs are RNA. In certain embodiments, 70% of the gNAs are RNA. In certain embodiments, 90% of the gNAs are RNA. In certain embodiments, 100% of the gNAs are RNA, e.g., gRNA. In further embodiments, the remainder of the gNA that is not RNA comprises modified ribonucleotides, deoxyribonucleotides, modified deoxyribonucleotides or synthetic, e.g., non-natural nucleotides, including, but not limited to, threose nucleic acids, locked nucleic acids, peptide nucleic acids, arabinonucleic acids, hexose nucleic acids, among others.
[0162] In certain embodiments, the targeter nucleic acid and / or modulator nucleic acid is an RNA having one or more modifications in the ribose group, one or more modifications in the phosphate group, one or more modifications in the nucleobase, one or more terminal modifications, or a combination thereof. Exemplary modifications are disclosed in U.S. Patent Nos. 10,900,034 and 10,767,175, U.S. Patent Application Publication No. 2018 / 0119140, Watts et al. (2008) Drug Discov. Today 13:842-55, and Hendel et al. (2015) Nat. Biotechnol. 33:985.
[0163] In certain embodiments, the targeter nucleic acid, e.g., RNA, comprises at least one nucleotide at or near the 3' end that comprises a ribose, a phosphate group, a modification to the nucleobase, or a terminal modification. In certain embodiments, the 3' end of the targeter nucleic acid comprises a spacer sequence. In certain embodiments, the 3' end of the targeter nucleic acid comprises a targeter stem sequence. Exemplary modifications are disclosed in Dang et al. (2015) Genome Biol. 16:280, Kocaz et al. (2019) Nature Biotech. 37:657-66, Liu et al. (2019) Nucleic Acids Res. 47(8):4169-4180, Schubert et al. (2018) J. Cytokine Biol. 3(1):121, Teng et al. (2019) Genome Biol. 20(1):15, Watts et al. (2008) Drug Discov. Today 13(19-20):842-55 and Wu et al. (2018) Cell Mol. Life. Sci. 75(19):3593-607.
[0164] Modifications in the ribose group include, but are not limited to, modifications at the 2'-position or modifications at the 4'-position. For example, in certain embodiments, the ribose comprises a 2'-O-Ci_4 alkyl, such as 2'-O-methyl (2'-OMe or M). In certain embodiments, the ribose comprises a 2'-O-Ci_3 alkyl-O-Ci_3 alkyl, such as 2'-O-(2-methoxyethyl) or 2'-methoxyethoxy (2'-O-CH2CHOCH3), also known as 2'-MOE. In certain embodiments, the ribose comprises a 2'-O-allyl. In certain embodiments, the ribose comprises a 2'-O-2,4-dinitrophenol (DNP). In certain embodiments, the ribose comprises a 2'-halo, such as 2'-F, 2'-Br, 2'-Cl, or 2'-I. In certain embodiments, the ribose comprises a 2'-NH2. In certain embodiments, the ribose comprises 2'-H (e.g., a deoxynucleotide). In certain embodiments, the ribose comprises 2'-arabino or 2'-F-arabino. In certain embodiments, the ribose comprises 2'-LNA or 2'-ULNA. In certain embodiments, the ribose comprises 4'-thioribosyl.
[0165] Modifications may also include deoxy groups, such as 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP).
[0166] Internucleotide linkage modifications at the phosphate group include, but are not limited to, phosphorothioate (S), chiral phosphorothioate, phosphorodithioate, boranophosphonate, C 1~4 These include alkyl phosphonates, such as methyl phosphonates, boranophosphonates, phosphonocarboxylates, such as phosphonoacetates (P), phosphonocarboxylic acid esters, such as phosphonoacetate esters, amides, thiophosphonocarboxylates, such as thiophosphonoacetates (SP), thiophosphonocarboxylic acid esters, such as thiophosphonoacetate esters, and phosphodiesters or modified phosphates having a 2',5'-linkage, as described above. Various salts, mixed salts, and free acid forms are also included.
[0167] Modifications in the nucleobase include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine, 5-methyluracil, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6 5-dehydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil, 5-allylcytosine, 5-aminoallyluracil, 5-aminoallyl-cytosine, 5-bromouracil, 5-iodouracil, diaminopurine, difluorotoluene, dihydrouracil, abasic nucleotides, Z bases, P bases, unstructured nucleic acids, isoguanine, isocytosine (see Piccirilli et al. (1990) NATURE, 343:33), 5-methyl-2-pyrimidine (see Rappaport (1993) BIOCHEMISTRY, 32:3047), x(A, G, C, T), and y(A, G, C, T).
[0168] Terminal modifications include, but are not limited to, polyethylene glycol (PEG), hydrocarbon linkers (heteroatom (O,S,N)-substituted hydrocarbon spacers; halo-substituted hydrocarbon spacers; keto-, carboxyl-, amido-, thionyl-, carbamoyl-, thionocarbamoyl-containing hydrocarbon spacers, propanediol), spermine linkers, dyes, such as fluorescent dyes (e.g., fluorescein, rhodamine, cyanine), quenchers (e.g., dabcyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides, and / or proteins, etc.). In certain embodiments, terminal modifications include conjugation (or ligation) of the RNA to another molecule, including oligonucleotides (such as deoxyribonucleotides and / or ribonucleotides), peptides, proteins, sugars, oligosaccharides, steroids, lipids, folate, vitamins, and / or other molecules. In certain embodiments, the terminal modification incorporated into the RNA is incorporated within the RNA sequence as a phosphodiester bond and is positioned via a linker such as a 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker, which can be incorporated anywhere between two nucleotides in the RNA.
[0169] The modifications disclosed above can be combined in targeter and / or modulator nucleic acids in the form of RNA. In certain embodiments, the modification in the RNA is selected from the group consisting of incorporation of 2'-O-methyl-3' phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-halo-3'-phosphorothioate (e.g., 2'-fluoro-3'-phosphorothioate), 2'-halo-3'-phosphonoacetate (e.g., 2'-fluoro-3'-phosphonoacetate), and 2'-halo-3'-thiophosphonoacetate (e.g., 2'-fluoro-3'-thiophosphonoacetate).
[0170] In certain embodiments, modifications may include 2'-O-methyl (M), phosphorothioate (S), phosphonoacetate (P), thiophosphonoacetate (SP), 2'-O-methyl-3'-phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP), or combinations thereof, at or near the 3' or 5' end of either the targeter or modulator nucleic acid, as appropriate for single or dual gNAs. In certain embodiments, modifications may include either a 5' or 3' propanediol or C3 linker modification.
[0171] In certain embodiments, the modifications alter the stability of the RNA. In certain embodiments, the modifications improve the stability of the RNA, for example, by increasing the nuclease resistance of the RNA compared to the corresponding RNA without the modification. Stability-enhancing modifications include, but are not limited to, 2'-O-methyl, 2'-OC 1~4 Alkyl, 2'-halo (e.g., 2'-F, 2'-Br, 2'-Cl, or 2'-I), 2'MOE, 2'-OC 1~3 These modifications include the incorporation of alkyl-O-Ci_3 alkyl, 2'-NH2, 2'-H (or 2'-deoxy), 2'-arabino, 2'-F-arabino, 4'-thioribosyl sugar moieties, 3'-phosphorothioate, 3'-phosphonoacetate, 3'-thiophosphonoacetate, 3'-methylphosphonate, 3'-boranophosphate, 3'-phosphorodithioate, locked nucleic acid ("LNA") nucleotides and unlocked nucleic acid ("ULNA") nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring. Such modifications are suitable for use as protecting groups to prevent or reduce degradation of 5' sequences, such as tail sequences, modulator stem sequences (dual guide nucleic acids), targeter stem sequences (dual guide nucleic acids), and / or spacer sequences (see the "Targeter and Modulator Nucleic Acids" subsection).
[0172] In certain embodiments, the modifications alter the specificity of the engineered non-natural system. In certain embodiments, the modifications improve the specificity of the engineered non-natural system, for example, by improving off-target binding and / or cleavage, or by reducing off-target binding and / or cleavage, or a combination thereof. Specificity-enhancing modifications include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, and pseudouracil. Within 10, 5, 4, 3, 2, or 1 nucleotide of the 3'-terminus, for example, the 3'-terminal nucleotide, is modified.
[0173] In certain embodiments, the modification alters the immunostimulatory activity of the RNA compared to the corresponding RNA without the modification, e.g., in certain embodiments, the modification reduces the ability of the RNA to activate TLR7, TLR8, TLR9, TLR3, RIG-I and / or MDA5.
[0174] In certain embodiments, the targeter nucleic acid and / or modulator nucleic acid contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 modified nucleotides or internucleotide linkages. Modifications can be made at one or more positions in the targeter nucleic acid and / or modulator nucleic acid such that these nucleic acids retain functionality. For example, a modified nucleic acid can still direct a Cas protein to a target nucleotide sequence, allowing the Cas protein to exert its effector function. It is understood that a particular modification at a given position can be selected based on the functionality of the nucleotide or internucleotide linkage at that position. For example, specificity-enhancing modifications may be suitable for nucleotides or internucleotide linkages in a spacer sequence, a targeter stem sequence, or a modulator stem sequence. Stability-enhancing modifications may be suitable for one or more terminal nucleotides or internucleotide linkages in a targeter nucleic acid and / or a modulator nucleic acid. In certain embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 5'-end and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 3'-end of a targeter nucleic acid are modified. In certain embodiments, five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides or internucleotide linkages at or near the 5'-end and / or five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides or internucleotide linkages at or near the 3'-end of a targeter nucleic acid are modified.In certain embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 5'-terminus and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 3'-terminus of the modulator nucleic acid are modified. In certain embodiments, no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 5'-terminus and / or no more than five (e.g., no more than one, no more than two, no more than three, or no more than four) terminal nucleotides or internucleotide linkages at or near the 3'-terminus of the modulator nucleic acid are modified. Selection of positions for modification is described in U.S. Patent Nos. 10,900,034 and 10,767,175. As used in this paragraph, when the targeter or modulator nucleic acid is a combination of DNA and RNA, the nucleic acid as a whole is considered to be RNA, and the DNA nucleotides are considered to be modifications of the RNA, including 2'-H modifications of the ribose and optionally modifications of the nucleobases.
[0175] It is understood that in a dual guide nucleic acid system, the targeter nucleic acid and the modulator nucleic acid are not in the same nucleic acid, i.e., are not linked end-to-end via traditional inter-polynucleotide bonds, but may be covalently conjugated to each other via one or more chemical modifications introduced into these nucleic acids, thereby increasing the stability of the double-stranded complex and / or improving other features of the system.
[0176] III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA Engineered non-natural systems such as those disclosed herein can be useful for targeting, editing, and / or modifying target nucleic acids, such as DNA (e.g., genomic DNA), within a cell or organism.
[0177] The present invention provides a method for cleaving a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, the method comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in cleavage of the target DNA.
[0178] Additionally, the present invention provides a method for binding a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in binding of the system to the target DNA. This method is useful, for example, for detecting the presence and / or location of a preselected target gene, e.g., when a component of the system (e.g., a Cas protein) comprises a detectable marker.
[0179] Further provided are methods for modifying a target nucleic acid (e.g., DNA) comprising a preselected target sequence or a portion thereof, or a structure (e.g., a protein) associated with the target DNA (e.g., a histone protein in a chromosome), comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in modification of the target DNA or a structure associated with the target DNA, wherein the Cas protein comprises an effector domain or is associated with an effector protein. The modification corresponds to the function of the effector domain or effector protein; the exemplary functions described in the "Cas Proteins" subsection in Section I above are applicable hereto.
[0180] The engineered non-natural system can be contacted with the target nucleic acid as a complex. Thus, in certain embodiments, the method comprises contacting the target nucleic acid with a CRISPR-Cas complex comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein disclosed herein. In certain embodiments, the Cas protein is a VA-type, VC-type, or VD-type Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a VA-type Cas protein (e.g., a Cas nuclease).
[0181] In certain embodiments, provided herein are methods for editing a human genome sequence at one of a group of preselected target loci, the method comprising delivering an engineered non-native system disclosed herein into a human cell, thereby resulting in editing of the genome sequence at the target locus in the human cell. In certain embodiments, provided herein are methods for detecting a human genome sequence at one of a group of preselected target loci, the method comprising delivering an engineered non-native system disclosed herein into a human cell, thereby resulting in detection of the target locus in the human cell, wherein a component of the system (e.g., a Cas protein) comprises a detectable marker. In certain embodiments, provided herein are methods for modifying a human chromosome at one of a group of preselected target loci, the method comprising delivering an engineered non-native system disclosed herein into a human cell, thereby resulting in modification of the chromosome at the target locus in the human cell, wherein the Cas protein comprises an effector domain or is associated with an effector protein.
[0182] CRISPR-Cas complex can be delivered to cells by introducing preformed ribonucleoprotein (RNP) complex into cells.Alternatively, one or more components of CRISPR-Cas complex can be expressed in cells.Exemplary delivery methods are known in the art, and are described in, for example, U.S. Patent Nos. 8,697,359, 10,113,167, 10,570,418, 10,829,787, 11,118,194 and 11,125,739 and U.S. Patent Application Publication Nos. 2015 / 0344912, 2018 / 0119140 and 2018 / 0282763.
[0183] It is understood that contacting DNA (e.g., genomic DNA) within a cell with a CRISPR-Cas complex does not require delivery of all components of the complex into the cell. For example, one or more of the components may already be present in the cell. In certain embodiments, the cell (or its parent / ancestor cell) has been engineered to express a Cas protein, and a single guide nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a single guide nucleic acid), a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid), and / or a modulator nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid) is delivered into the cell. In certain embodiments, the cell (or its parent / ancestor cell) has been engineered to express a modulator nucleic acid, and a Cas protein (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a Cas protein) and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid) are delivered into the cell. In certain embodiments, the cell (or its parent / ancestor cell) has been engineered to express a Cas protein, and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) is delivered into the cell.
[0184] In certain embodiments, the target DNA is within the genome of the target cell. Accordingly, the present invention also provides cells comprising a non-native system or a CRISPR expression system described herein. Additionally, the present invention provides cells whose genomes have been modified by a CRISPR-Cas system or complex disclosed herein.
[0185] Target cells can be mitotic or post-mitotic cells from any organism, such as bacterial cells (e.g., E. coli), archaeal cells, cells of unicellular eukaryotes, plant cells, algae cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc. The target cell may be a fungal cell (e.g., yeast cell, e.g., S. cerevisiae), animal cell, cell from an invertebrate (e.g., fruit fly, enidarian, echinoderm, nematode, etc.), cell from a vertebrate (e.g., fish, amphibian, reptile, bird, mammal), cell from a mammal, cell from a rodent, or cell from a human. Target cell types include, but are not limited to, stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, embryonic cells), somatic cells (e.g., fibroblasts, hematopoietic cells, T lymphocytes (e.g., CD8+ T lymphocytes), NK cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells), in vitro or in vivo embryonic cells of any stage of embryo (e.g., 1-cell, 2-cell, 4-cell, 8-cell zebrafish embryos). The cells can be derived from an established cell line or primary cells (i.e., cells and cell cultures derived from a subject and propagated in vitro for a limited number of passages). For example, a primary culture is a culture that can be passaged 0, 1, 2, 4, 5, 10, or 15 times, but not enough times to reach the transition phase. Typically, a primary cell line is maintained in vitro for fewer than 10 passages. When the cells are primary cells, they can be collected from an individual by any suitable method. For example, white blood cells can be collected by apheresis, leukapheresis, or density gradient separation, while cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, or stomach can be collected by biopsy.The harvested cells can be used immediately or can be stored under frozen conditions with a cryoprotectant and thawed at a later time by methods commonly known in the art.
[0186] A. Ribonucleoprotein (RNP) Delivery and "casRNA" Delivery The engineered non-natural systems disclosed herein can be delivered into cells by any suitable method known in the art, including but not limited to ribonucleoprotein (RNP) delivery and "CasRNA" delivery, as described below.
[0187] In certain embodiments, a CRISPR-Cas system comprising a single guide nucleic acid and a Cas protein, or a CRISPR-Cas system comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein, can be assembled into an RNP complex and then delivered into cells as a preformed complex. This method is suitable for actively modifying the genetic or epigenetic information of cells for a limited period of time. For example, if a Cas protein has nuclease activity to modify the genomic DNA of a cell, the nuclease activity only needs to be maintained for a period of time to allow DNA cleavage, and prolonged nuclease activity may increase off-target effects. Similarly, some specific epigenetic modifications can be maintained in cells once established and inherited by daughter cells.
[0188] As used herein, the term "ribonucleoprotein" or "RNP" can refer to a complex comprising a nucleoprotein and a ribonucleic acid. As provided herein, a "nucleoprotein" can refer to a protein capable of binding to nucleic acids (e.g., RNA, DNA). When a nucleoprotein binds to a ribonucleic acid, it can be referred to as a "ribonucleoprotein." The interaction between a ribonucleoprotein and a ribonucleic acid can be direct, e.g., by a covalent bond, or indirect, e.g., by a non-covalent bond (e.g., electrostatic interactions (e.g., ionic bonds, hydrogen bonds, halogen bonds), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion forces), ring stacking (π effect), hydrophobic interactions, etc.). In certain embodiments, a ribonucleoprotein comprises an RNA-binding motif non-covalently bound to a ribonucleic acid. For example, a positively charged aromatic amino acid residue (e.g., lysine residue) in the RNA-binding motif can form an electrostatic interaction with the negative nucleic acid phosphate backbone of the RNA.
[0189] To ensure efficient loading of the Cas protein, the single guide nucleic acid or the combination of targeter and modulator nucleic acids can be provided in molar excess (e.g., at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold) relative to the Cas protein. In certain embodiments, the targeter and modulator nucleic acids are annealed under suitable conditions before forming a complex with the Cas protein. In other embodiments, the targeter nucleic acid, modulator nucleic acid, and Cas protein are mixed together directly to form the RNP.
[0190] Various delivery methods can be used to introduce the RNPs disclosed herein into cells. Exemplary delivery methods or vehicles include, but are not limited to, microinjection, liposomes (see, e.g., U.S. Pat. No. 10,829,787), such as molecular trojan horse liposomes that deliver molecules across the blood-brain barrier (see, Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, microvesicles (e.g., exosomes and ARMMs), polycations, lipid:nucleic acid conjugates, electroporation, cell-penetrating peptides (see, U.S. Pat. No. 11,118,194), nanoparticles, nanowires (Shalek et al. (2012) Nano Letters, 12:6498), exosomes, and perturbation of the cell membrane (e.g., by passing cells through a constriction in a microfluidic system, see U.S. Pat. No. 11,125,739). When the target cell is a proliferating cell, the efficiency of RNP delivery can be improved by synchronization of the cell cycle (see U.S. Pat. No. 10,570,418). In certain embodiments, RNPs are delivered into cells by electroporation.
[0191] In certain embodiments, the CRISPR-Cas system is delivered into cells by the following approach: (a) delivery of a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, and (b) delivery of RNA (e.g., messenger RNA (mRNA)) encoding a Cas protein. The RNA encoding the Cas protein can be translated within the cell and form a complex within the cell with the single guide nucleic acid or the combination of the targeter nucleic acid and the modulator nucleic acid. Similar to RNP approaches, RNA has a limited half-life within the cell, even if stability-enhancing modifications are made to one or more of the RNAs. Therefore, the "Cas RNA" approach is suitable for active modification of a cell's genetic or epigenetic information, e.g., DNA cleavage, for a limited period of time, with the advantage of reduced off-target effects.
[0192] The mRNA can be produced by transcription of DNA comprising a regulatory element operably linked to a Cas coding sequence. Given that multiple copies of the Cas protein can be generated from a single mRNA, the single guide nucleic acid or targeter nucleic acid and modulator nucleic acid are generally provided in molar excess (e.g., at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, or at least 100-fold) relative to the mRNA. In certain embodiments, the targeter nucleic acid and modulator nucleic acid are annealed under suitable conditions before delivery into the cell. In other embodiments, the targeter nucleic acid and modulator nucleic acid are delivered into the cell without in vitro annealing.
[0193] The "Cas RNA" system can be introduced into cells using a variety of delivery systems. Non-limiting examples of delivery methods or vehicles include microinjection, particle guns, liposomes (see, e.g., U.S. Pat. No. 10,829,787), molecular Trojan horse liposomes that deliver molecules across the blood-brain barrier (Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, polycations, lipid:nucleic acid conjugates, electroporation, nanoparticles, nanowires (see, e.g., Shalek et al. (2012) Nano Letters, 12:6498), exosomes, and cell membrane perturbation (e.g., by passing cells through a constriction in a microfluidic system, see, U.S. Pat. No. 11,125,739). A specific example of a "nucleic acid only" approach using electroporation is described in International (PCT) Publication No. WO 2016 / 164356.
[0194] In certain embodiments, the CRISPR-Cas system is delivered into cells in the form of DNA containing (a) a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, and (b) a regulatory element operably linked to the Cas coding sequence. The DNA can be provided in a plasmid, a viral vector, or any other form described in the "CRISPR Expression Systems" subsection. Such delivery methods may result in constitutive expression of the Cas protein in the target cell (e.g., if the DNA is maintained in the cell on an episomal vector or integrated into the genome), which may increase the risk of undesired off-target effects if the Cas protein has nuclease activity. Nevertheless, this approach is useful when the Cas protein contains a non-nuclease effector (e.g., a transcriptional activator or repressor). It is also useful for research purposes and plant genome editing.
[0195] B. CRISPR Expression System Also provided herein are nucleic acids comprising a regulatory element operably linked to a nucleotide sequence encoding a guide nucleic acid disclosed herein. In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a single guide nucleic acid, and this nucleic acid can constitute a CRISPR expression system by itself. In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid. In certain embodiments, the nucleic acid further comprises a nucleotide sequence encoding a modulator nucleic acid, and the nucleotide sequence encoding the modulator nucleic acid is operably linked to the same or a different regulatory element as the nucleotide sequence encoding the targeter nucleic acid, and this nucleic acid can constitute a CRISPR expression system by itself.
[0196] Additionally, the present invention provides a CRISPR expression system comprising: (a) a nucleic acid comprising a first regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid; and (b) a nucleic acid comprising a second regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid.
[0197] In certain embodiments, the CRISPR expression system further comprises a nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein, e.g., a Cas protein disclosed herein. In certain embodiments, the Cas protein is a VA-type, VC-type, or VD-type Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a VA-type Cas protein (e.g., a Cas nuclease).
[0198] As used in this context, the term "operably linked" means that the nucleotide sequence of interest is linked to regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when a vector is introduced into a host cell).
[0199] The nucleic acids of the above CRISPR expression systems can be independently selected from a variety of nucleic acids, such as DNA (e.g., modified DNA) and RNA (e.g., modified RNA). In certain embodiments, the nucleic acid comprising a regulatory element operably linked to one or more nucleotide sequences encoding a guide nucleic acid is in the form of DNA. In certain embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of DNA. The third regulatory element can be a constitutive or inducible promoter that drives expression of the Cas protein. In other embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of RNA (e.g., mRNA).
[0200] The nucleic acid of the CRISPR expression system can be provided in one or more vectors. As used herein, the term "vector" can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into cells, such as prokaryotic cells, eukaryotic cells, mammalian cells, or target tissues. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcription products of the vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery vehicle, such as liposomes. Viral vector delivery systems include DNA and RNA viruses that have either episomal or integrated genomes after delivery to cells. Gene therapy procedures are known in the art and are described in Van Brunt (1988) BIOTECHNOLOGY, 6:1149; Anderson (1992) SCIENCE, 256:808; Nabel & Feigner (1993) TIBTECH, 11:211; Mitani & Caskey (1993) TIBTECH, 11:162; Dillon (1993) TIBTECH, 11:167; Miller (1992) NATURE, 357:455; Vigne, (1995) RESTORATIVE NEUROLOGY AND NEUROSCIENCE, 8:35; Kremer & Perricaudet (1995) BRITISH MEDICAL BULLETIN, 51:31; Haddada et al. (1995) CURRENT TOPICS IN MICROBIOLOGY AND IMMUNOLOGY, 199:297; Yu et al. (1994) GENE THERAPY, 1:13; and Doerfler and Bohm (Eds.) (2012) The Molecular Repertoire of Adenoviruses II: Molecular Biology of Virus-Cell Interactions. In certain embodiments, at least one of the vectors is a DNA plasmid. In certain embodiments, at least one of the vectors is a viral vector (e.g., a retrovirus, an adenovirus, or an adeno-associated virus).
[0201] Certain vectors are capable of autonomous replication in host cells into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors and replication-deficient viral vectors) do not autonomously replicate in host cells. However, certain vectors can be integrated into the genome of the host cell, thereby replicating along with the host genome. One skilled in the art will understand that different vectors may be suitable for different delivery methods and have different host tropisms, and may select one or more vectors suitable for use.
[0202] The term "regulatory element," as used herein, can refer to transcriptional and / or translational control sequences, such as promoters, enhancers, transcription termination signals (e.g., polyadenylation signals), internal ribosome entry sites (IRES), proteolysis signals, and the like, that effect and / or regulate transcription of a non-coding sequence (e.g., a targeter nucleic acid or modulator nucleic acid) or a coding sequence (e.g., a Cas protein) and / or regulate translation of an encoded polypeptide. Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY, 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). A tissue-specific promoter may direct expression primarily in a desired tissue of interest, such as muscle, neurons, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). Regulatory elements may also direct expression in a time-dependent manner, such as a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell-type-specific. In certain embodiments, the vector comprises one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and H1 promoters.Examples of Pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" also encompasses enhancer elements, such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (see Takebe et al. (1988) MOL. CELL. BIOL., 8:466); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (see O'Hare et al. (1981) PROC. NATL. ACAD. SCI. USA., 78:1527). It will be understood by those skilled in the art that the design of the expression vector can depend on factors such as the choice of the host cell to be transformed, the level of expression desired, etc. The vectors can be introduced into host cells to produce transcripts, proteins, or peptides (including fusion proteins or peptides) encoded by the nucleic acids described herein (e.g., CRISPR transcripts, proteins, enzymes, mutant forms thereof, or fusion proteins thereof).
[0203] In certain embodiments, nucleotide sequences encoding Cas proteins are codon-optimized for expression in prokaryotic cells, such as Escherichia coli (E. coli), eukaryotic host cells, such as yeast cells (e.g., S. cerevisiae), mammalian cells (e.g., mouse, rat, or human cells), or plant cells. Various species exhibit specific biases toward some codons for specific amino acids. Codon bias (differences in codon usage among organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which is therefore thought to depend, inter alia, on the properties of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example, at the "Codon Usage Database," available at kazusa.or.jp / codon / , and these tables can be adapted in several ways (see Nakamura et al. (2000) NUCL. ACIDS RES., 28:292). Computer algorithms are also available, such as Gene Forge (Aptagen; Jacobus, Pa.), for codon-optimizing particular sequences for expression in particular host cells. In certain embodiments, codon optimization facilitates or improves expression of the Cas protein in the host cell.
[0204] C. Donor Template Cleavage of a target nucleotide sequence in a cell's genome by a CRISPR-Cas system or complex can activate DNA damage pathways, allowing the cleaved DNA fragments to be rejoined by NHEJ or HDR, which requires either an endogenous or exogenous repair template to transfer sequence information from the repair template to the target.
[0205] In certain embodiments, the engineered non-natural system or CRISPR expression system further comprises a donor template. As used herein, the term "donor template" may refer to a nucleic acid designed to serve as a repair template at or near a target nucleotide sequence when introduced into a cell or organism. In certain embodiments, the donor template is complementary to a polynucleotide comprising a target nucleotide sequence or a portion thereof. When optimally aligned, the donor template may overlap with one or more nucleotides of the target nucleotide sequence (e.g., about 1, 5, 10, 15, 20, 25, 30, 35, 40 or more nucleotides) of the target nucleotide sequence. The nucleotide sequence of the donor template is typically not identical to the genomic sequence it replaces. Instead, the donor template may contain one or more substitutions, insertions, deletions, inversions, or rearrangements relative to the genomic sequence, as long as there is sufficient homology to support homologous recombination repair. In certain embodiments, the donor template comprises a non-homologous sequence flanked by two homologous regions (i.e., homology arms), such that homology-directed repair between the target DNA region and the two flanking sequences results in insertion of the non-homologous sequence in the target region. In certain embodiments, the donor template comprises a non-homologous sequence 10-100 nucleotides, 50-500 nucleotides, 100-1,000 nucleotides, 200-2,000 nucleotides, or 500-5,000 nucleotides in length located between the two homology arms.
[0206] Generally, the homologous region of the donor template has at least 50% sequence identity with the genomic sequence with which recombination is desired. The homologous arms are designed or selected so that they are capable of recombining with the nucleotide sequence adjacent to the target nucleotide sequence under intracellular conditions. In certain embodiments, when HDR of the non-target strand is desired, the donor template comprises a first homologous arm homologous to a sequence 5' of the target nucleotide sequence and a second homologous arm homologous to a sequence 3' of the target nucleotide sequence. In certain embodiments, the first homologous arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 5' of the target nucleotide sequence. In certain embodiments, the second homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 3' of the target nucleotide sequence. In certain embodiments, when polynucleotides comprising the donor template sequence and the target nucleotide sequence are optimally aligned, the nearest nucleotide of the donor template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, or more nucleotides of the target nucleotide sequence.
[0207] In certain embodiments, the donor template further comprises an engineered sequence that is not homologous to the sequence to be repaired. Such engineered sequence may have a barcode and / or sequence that can hybridize to the donor template recruitment sequence disclosed herein.
[0208] In certain embodiments, the donor template further comprises one or more mutations in the genomic sequence, which reduce or prevent cleavage of the donor template or cleavage of a modified genomic sequence incorporating at least a portion of the donor template sequence by the same CRISPR-Cas system. In certain embodiments, in the donor template, a PAM adjacent to the target nucleotide sequence and recognized by a Cas nuclease is mutated to a sequence not recognized by the same Cas nuclease. In certain embodiments, in the donor template, the target nucleotide sequence (e.g., a seed region) is mutated. In certain embodiments, the one or more mutations are silent with respect to the reading frame of the protein-coding sequence containing the mutation site.
[0209] The donor template can be provided to cells as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. In certain embodiments, the linear double-stranded DNA comprises covalently closed ends, such as those generated by telomerase enzymes such as TelN or suitable alternatives. It is understood that the CRISPR-Cas system, such as the system disclosed herein, can have nuclease activity to cleave the target strand, the non-target strand, or both. When HDR of the target strand is desired, the donor template can also be considered to have a nucleic acid sequence complementary to the target strand.
[0210] The donor template can be introduced into cells in linear or circular form. If introduced in linear form, the ends of the donor template can be protected (e.g., from exonuclease degradation) by methods known to those skilled in the art. For example, one or more dideoxynucleotide residues can be added to the 3' end of the linear molecule, and / or self-complementary oligonucleotides can be ligated to one or both ends (see, for example, Chang et al. (1987) PROC.NATL.ACAD SCI USA, 84:4959; Nehls et al. (1996) SCIENCE, 272:886, and also see the chemical modifications for increasing RNA stability and / or specificity disclosed above). Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino groups and the use of modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methylribose or deoxyribose residues. For example, linear covalently closed DNA, such as that produced by a telomerase enzyme such as TelN, can be used as the donor template. As an alternative to protecting the ends of the linear donor template, additional lengths of sequence can be included outside the homologous regions that can be degraded without affecting recombination.
[0211] The donor template can be a component of a vector described herein, can be included in a separate vector, or can be provided as a separate polynucleotide, such as an oligonucleotide, a linear polynucleotide, or a synthetic polynucleotide. In certain embodiments, the donor template is DNA. In certain embodiments, the donor template is in the same nucleic acid as the sequence encoding the single guide nucleic acid, the sequence encoding the targeter nucleic acid, the sequence encoding the modulator nucleic acid, and / or the sequence encoding the Cas protein, if applicable. In certain embodiments, the donor template is provided in a separate nucleic acid. The donor template polynucleotide can be of any suitable length, for example, about or at least about 50, 75, 100, 150, 200, 500, 1000, 2000, 3000, 4000, or more nucleotides in length.
[0212] The donor template can be introduced into cells as an isolated nucleic acid. Alternatively, the donor template can be introduced into cells as part of a vector (e.g., a plasmid) that contains additional sequences not intended for insertion into the DNA region of interest, such as an origin of replication, a promoter, and a gene encoding antibiotic resistance. Alternatively, the donor template can be delivered by a virus (e.g., an adenovirus or an adeno-associated virus (AAV)). In certain embodiments, the donor template is introduced as an AAV, e.g., a pseudotyped AAV. The capsid protein of the AAV can be selected by one skilled in the art based on the tropism of the AAV and the target cell type. For example, in certain embodiments, the donor template is introduced into hepatocytes as AAV8 or AAV9. In certain embodiments, the donor template is introduced into hematopoietic stem cells, hematopoietic progenitor cells, or T lymphocytes (e.g., CD8+ T lymphocytes) as AAV6 or AAVHSC (see U.S. Pat. No. 9,890,396). It is understood that the sequence of the capsid protein (VP1, VP2, or VP3) can be modified from a wild-type AAV capsid protein, for example, having at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the wild-type AAV capsid sequence.
[0213] The donor template can be delivered to cells (e.g., primary cells) by a variety of delivery methods, including viral or non-viral methods disclosed herein. In certain embodiments, the non-viral donor template is introduced into the target cell as naked nucleic acid or complexed with liposomes or poloxamers. In certain embodiments, the non-viral donor template is introduced into the target cell by electroporation. In other embodiments, the viral donor template is introduced into the target cell by infection. The engineered non-natural system can be delivered before, after, or simultaneously with the donor template (see International (PCT) Application Publication No. WO 2017 / 053729). One of skill in the art will be able to select the appropriate timing based on the form of delivery (e.g., taking into account the time required for transcription and translation of the RNA and protein components) and the half-life of the molecule within the cell. In certain embodiments, when the CRISPR-Cas system comprising the Cas proteins is delivered by electroporation (e.g., as an RNP), the donor template (e.g., as an AAV) is introduced into the cell within 4 hours (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 90, 120, 150, 180, 210, or 240 minutes) after introduction of the engineered non-native system.
[0214] In certain embodiments, the donor template is covalently conjugated to the modulator nucleic acid. Suitable covalent bonds for this conjugation are known in the art and are described, for example, in U.S. Pat. No. 9,982,278 and Savic et al. (2018) ELIFE 7:e33761. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., at the 5' end of the modulator nucleic acid) via an inter-polynucleotide bond. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., at the 5' end of the modulator nucleic acid) via a linker.
[0215] In certain embodiments, the donor template may comprise any nucleic acid chemistry. In certain embodiments, the donor template may comprise DNA and / or RNA nucleotides. In certain embodiments, the donor template may comprise single-stranded DNA, linear single-stranded RNA, linear double-stranded DNA, linear double-stranded RNA, circular single-stranded DNA, circular single-stranded RNA, circular double-stranded DNA, or circular double-stranded RNA. In certain embodiments, the linear double-stranded DNA comprises covalently closed ends, such as those generated by a telomerase enzyme, such as TelN, or a suitable alternative. In certain embodiments, the donor template comprises a mutation in the PAM sequence to partially or completely abolish RNP binding to DNA. In certain embodiments, the donor template comprises at least 0.05, 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, or 4 and / or 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, 4, or 5 μg μL -1 For example, 0.01 to 5 μg μL -1 In certain embodiments, the donor template comprises one or more promoters. In certain embodiments, the donor template comprises a promoter that shares at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99.5% sequence identity with any one of SEQ ID NOs: 78-85 in Table 5.
[0216] [Table 40]
[0217] [Table 41]
[0218] [Table 42]
[0219] [Table 43]
[0220] D. Efficiency and Specificity Engineered non-natural systems can be evaluated for efficiency and / or specificity in nucleic acid targeting, cleavage or modification.
[0221] In certain embodiments, the engineered non-natural system has high efficiency. For example, in certain embodiments, at least 1, 1.5, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of a population of nucleic acids having a target nucleotide sequence and a cognate PAM are targeted, cleaved, or modified when contacted with the engineered non-natural system. In certain embodiments, at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of the genomes of a population of cells are targeted, cleaved, or modified when the engineered non-natural system is delivered into the cells.
[0222] For a given spacer sequence, the occurrence of on-target events and the occurrence of off-target events are generally observed to be correlated. For certain therapeutic purposes, a lower on-target efficiency may be acceptable, and a low off-target frequency is more desirable. For example, when editing or modifying proliferative cells delivered to a subject and growing in vivo, the tolerance for off-target events is low. Prior to delivery, on-target and off-target events can be evaluated, thereby selecting one or more colonies with the desired editing or modification and no undesired editing or modification. Nevertheless, the on-target efficiency may need to meet certain criteria to be suitable for therapeutic use. The high editing efficiency of standard CRISPR-Cas systems allows for tuning of the system, for example, by reducing the binding of guide nucleic acids to Cas proteins, without losing therapeutic applicability.
[0223] In certain embodiments, when a population of nucleic acids having a target nucleotide sequence and a cognate PAM is contacted with an engineered non-natural system disclosed herein, the frequency of off-target events (e.g., targeting, cleavage, or modification, depending on the function of the CRISPR-Cas system) is reduced. Methods for assessing off-target events are summarized in Lazzarotto et al. (2018) Nat. Protoc. 13(11):2615-42 and include in situ Cas off-target discovery and validation by sequencing (DISCOVER-seq), as disclosed in Wienert et al. (2019) Science 364(6437):286-89; genome-wide unbiased identification of double-strand breaks (DSBs) enabled by sequencing (GUIDE-seq), as disclosed in Kleinstiver et al. (2016) Nat. Biotech. 34:869-74; and circularization for in vitro reporting of cleavage efficacy by sequencing (CIRCLE-seq), as described in Kocak et al. (2019) Nat. Biotech. 37:657-66. In certain embodiments, an off-target event comprises targeting, cleavage, or modification at a given off-target locus (e.g., the locus at which the most occurrences of the off-target event were detected). In certain embodiments, an off-target event comprises targeting, cleavage, or modification at all loci collectively with detectable off-target events.
[0224] In certain embodiments, genomic mutations occur in 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.00 Detected in less than 6%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4% or 5% (total). In certain embodiments, the ratio of the percentage of cells having on-target events to the percentage of cells having any off-target events (e.g., the ratio of the percentage of cells having on-target editing events to the percentage of cells having mutations at any off-target loci) is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000. It is understood that genetic variation may be present within a population of cells due to, for example, spontaneous mutations, and such mutations are not included as off-target events.
[0225] E. Multiplex Method The methods for targeting, editing, and / or modifying genomic DNA disclosed herein can be performed in a multiplexed manner. For example, a library of targeter nucleic acids can be used to target multiple genomic loci, and a library of donor templates can also be used to generate multiple insertions, deletions, and / or substitutions. Multiplex assays can be performed in a screening method, where each separate cell culture (e.g., in the wells of a 96-well or 384-well plate) is exposed to a different guide nucleic acid with a different targeter stem sequence and / or a different donor template. Multiplex assays can also be performed in a selection method, where cell cultures are exposed to a mixed population of different guide nucleic acids and / or donor templates, and cells with desired characteristics (e.g., functionality) are enriched or selected by advantageous survival or growth, resistance to a specific drug, expression of a detectable protein (e.g., a fluorescent protein detectable by flow cytometry), etc.
[0226] In certain embodiments, multiple guide nucleic acids and / or multiple donor templates are designed for saturation editing. For example, in certain embodiments, each nucleotide position in a target sequence is systematically modified with each of all four conventional bases, A, T, G, and C. In other embodiments, at least one sequence in each gene from a pool of target genes is modified, for example, according to a CRISPR design algorithm. In certain embodiments, each sequence from a pool of target exogenous elements (e.g., protein-coding sequences, non-protein-coding genes, regulatory elements) is inserted into one or more given loci of a genome.
[0227] It is understood that multiplex methods suitable for performing screening or selection methods typically performed for research purposes may differ from methods suitable for therapeutic purposes. For example, constitutive expression of certain elements (e.g., Cas nuclease and / or guide nucleic acid) may be undesirable for therapeutic purposes due to the potential for increased off-target effects. Conversely, constitutive expression of Cas nuclease and / or guide nucleic acid may be desirable for research purposes. For example, constitutive expression provides a broader time frame in which other elements can be introduced. When stable cell lines are established for constitutive expression, the number of exogenous elements that need to be co-delivered into a single cell is also reduced. Thus, constitutive expression of certain elements may increase the efficiency and reduce the complexity of the screening or selection process. Inducible expression of certain elements of the systems disclosed herein may also be used for research purposes, considering similar advantages. Expression can be induced by exogenous factors (e.g., small molecules) or endogenous molecules or complexes present in a particular cell type (e.g., at a particular differentiation stage). Methods known in the art, such as those described herein, can be used to constitutively or inducibly express one or more elements. For example, the specificity of a CRISPR nuclease is determined at least in part by the uniqueness of the spacer (combined with the proximity of the spacer sequence to the required PAM), and its off-target score can be calculated using an algorithm, such as crispr.mit.edu (Hsu et al. (2013) Nat. Biotech. 31:827-832). The highest possible score is 100, indicating high specificity and little off-target potential. Because our SHS library targets intergenic regions, algorithms for gRNA prediction should be able to align using repeated regions and low-complexity sequences.
[0228] It is further understood that despite the need to introduce multiple elements—a single guide nucleic acid and a Cas protein, or a targeter nucleic acid, a modulator nucleic acid, and a Cas protein—these elements can be delivered into the cell as a single complex of preformed RNPs. Thus, efficiency of the screening or selection process can also be achieved by preassembling multiple RNP complexes in a multiplexed manner.
[0229] In certain embodiments, the methods disclosed herein further comprise identifying the guide nucleic acid, Cas protein, donor template, or a combination of two or more of these elements by a screening or selection process. A set of barcodes can be used on the donor template, for example, between two homology arms, to facilitate identification. In certain embodiments, the methods further comprise harvesting a population of cells, selectively amplifying a genomic DNA or RNA sample and / or barcodes comprising the target nucleotide sequence, and / or sequencing the selectively amplified genomic DNA or RNA sample and / or barcodes.
[0230] Additionally, the present invention provides libraries comprising a plurality of guide nucleic acids, e.g., a plurality of guide nucleic acids disclosed herein. In another aspect, the present invention provides libraries comprising a plurality of nucleic acids, each comprising a regulatory element operably linked to a different guide nucleic acid, e.g., a different guide nucleic acid disclosed herein. These libraries can be used in combination with one or more Cas proteins or Cas-encoding nucleic acids, e.g., those disclosed herein, and / or one or more donor templates, e.g., those disclosed herein, for screening or selection methods.
[0231] F. Genomic Safe Harbor Genome engineering is a field of research that seeks to modify the genes of organisms in order to improve understanding of gene function and, inter alia, develop methods of genome engineering to treat inherited or acquired diseases. To modify the genome of a target cell, those skilled in the art use one or more available means to introduce changes into the genome at targeted locations to modify the sequence of a target polynucleotide, e.g., a target gene, in a desired manner, for example, to regulate gene expression, regulate a gene sequence, delete a gene sequence, introduce a gene, e.g., exogenous DNA, e.g., a transgene, etc. Efficient transgene insertion can be achieved by non-precision methods including, but not limited to, viral vectors, such as retroviral vectors, e.g., adeno-associated virus (AAV), or precision methods including, but not limited to, inducible nucleases, such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), homing endonucleases, e.g., restriction endonucleases or nucleic acid-guided nucleases, such as CRISPR-cas, e.g., Cas9 and Cas12a, and engineered forms thereof.
[0232] Exogenous genes, e.g., transgenes, inserted randomly, e.g., via retroviral vectors, or in a targeted manner, e.g., by the action of nucleic acid-guided nucleases, e.g., Cas, into the genome of target human cells can interact with other genomic elements in unpredictable ways. Due to the complex transcriptional regulation of genes in mammalian cells by networks of cis and trans regulatory elements, e.g., proximal and distal enhancers and multiple transcription factors, attempts to modify the default genome structure by incorporating exogenous DNA, e.g., transgenes or synthetic sequences, can affect the expression of the transgene itself, which can, for example, impair safety checkpoints possessed by healthy cells and dramatically alter cellular behavior, i.e., completely attenuate or completely silence and / or express both proximal and distal endogenous genes, including down-regulation of the expression of key genes, e.g., oncogenes and tumor suppressor genes, which can promote clonal proliferation or malignant transformation of the host.
[0233] Gene integration adjacent to the regulatory elements of protooncogenes has been shown to cause oncogenic transformation, which is particularly important when manipulating cells for therapeutic applications.Therefore, it is desired to identify a suitable target polynucleotide comprising a target nucleotide sequence in the human genome, such that the insertion of a transgene leads to the appropriate expression of the transgene without disrupting adjacent genes.In particular, for gene and cell therapy applications, it is desired to identify a suitable target polynucleotide comprising a target nucleotide sequence in the human genome, such that the insertion of a transgene leads to sufficient expression of the transgene in therapeutic cells, such as T cells, for example, CAR T cells, or progenitor cells, for example, stem cells, for example, hematopoietic stem cells, without causing malignant transformation or any other disruption that would be harmful to individuals after transplantation.
[0234] Expression of an exogenous gene, e.g., a transgene, in a desired cell type and / or developmental / differentiation stage depends on integration from a candidate locus into a suitable target polynucleotide containing a target nucleotide sequence that confers sufficient expression for the intended purpose. Expression from a particular genomic site can be influenced by many factors, including, but not limited to, cell type and differentiation stage, as well as changes in chromatin structure, when one or more components of the target polynucleotide are activated during differentiation while other components are silenced. Therefore, it is desirable to identify a suitable target polynucleotide containing a target nucleotide sequence in the human genome where insertion of exogenous DNA, e.g., a transgene, results in sufficient expression in target human cells, and in the case of stem cells, expression is maintained at a sufficient level through (1) differentiation and (2) clonal expansion. The present disclosure provides a significant advance in the ability to manipulate the human genome by providing compositions and methods for targeting and offsetting exogenous genes, e.g., transgenes, to a suitable target polynucleotide containing a target nucleotide sequence.
[0235] Compositions and methods for genome manipulation are provided herein. Certain embodiments include compositions. Certain embodiments include compositions for genome editing. Embodiments disclosed herein relate to novel guide nucleic acids (gNAs), e.g., gRNAs, complementary to a target nucleotide sequence in a target polynucleotide. As used herein, "target polynucleotide" includes a polynucleotide in which a target nucleotide sequence is located. As used herein, "target nucleotide sequence" includes a sequence to which a guide sequence can bind, e.g., has complementarity, and binding between the target nucleotide sequence and the guide sequence can enable activity of a nucleic acid-guided nuclease complex. Further embodiments disclosed herein relate to novel gNAs, e.g., gRNAs, complementary to a target nucleotide sequence in a target polynucleotide, such that insertion of exogenous DNA, e.g., a transgene, does not adversely affect the cell, e.g., does not significantly affect the expression of one or more endogenous genes or result in malignant transformation of the cell. In further embodiments disclosed herein, gene expression exhibited in a human target cell is maintained by differentiation of the human target cell and / or proliferation of one or more progeny cells at a level sufficient for the end use of the cell. Certain embodiments disclosed herein relate to novel nucleic acid-guided nuclease complexes, e.g., Cas bound to RNPs, e.g., gNAs, which are complementary to a target nucleotide sequence within a target polynucleotide and hybridize (also referred to as cleavage or cutting) their phosphodiester backbones at at least one position in at least one strand of the target polynucleotide. Certain embodiments disclosed herein relate to methods of selecting and using gNAs, e.g., gRNAs, for genome engineering. Certain embodiments relate to methods using gNAs complementary to a target nucleotide sequence within a target polynucleotide, synthesizing gNAs and nucleic acid-guided nucleases, and / or combining nucleic acid-guided nucleases with gNAs to form nucleic acid-guided nuclease complexes, e.g., RNPs. Certain embodiments disclosed herein relate to methods. Certain embodiments disclosed herein relate to methods for genome engineering.Certain embodiments disclosed herein relate to methods in which a nucleic acid-guided nuclease complex, e.g., an RNP, is introduced, e.g., transfected, into a human target cell along with a donor template, e.g., exogenous DNA, e.g., a transgene, where the nucleic acid-guided nuclease cleaves the backbone at least one position in at least one of the strands of the target polynucleotide, repairs the cleaved target polynucleotide using the donor template, and introduces at least a portion of the donor template into the target polynucleotide. As used herein, "exogenous DNA" or "transgene" includes any gene, natural or synthetic, introduced into the genome of an organism or cell to which it is not endogenous. The transgene may or may not retain the ability to be expressed and / or produce RNA or protein in the human target cell. The transgene may or may not alter the resulting phenotype of the human target cell. Certain embodiments include human target cells, e.g., eukaryotic cells, e.g., mammalian cells, e.g., human cells, e.g., stem cells or immune cells, generated by a method in which a nucleic acid-guided nuclease complex, e.g., RNP, is introduced, e.g., transfected, into a human target cell along with a donor template, e.g., exogenous DNA or a transgene, e.g., a chimeric antigen receptor (CAR), where the nucleic acid-guided nuclease cleaves at or near a target sequence in the target polynucleotide, and the donor template is used to introduce at least a portion of the donor template into the target polynucleotide to repair the cleaved target polynucleotide. Certain embodiments disclosed herein include promoter sequences adjacent to the exogenous gene, e.g., a transgene, and in certain cases, constructs containing the promoter, when introduced into a target polynucleotide of a human target cell, e.g., an immune cell or stem cell, maintain sufficient gene expression in the edited human target cell for the intended purpose of the cell or its progeny. In certain embodiments, the human target cell survives after introduction of the exogenous DNA.
[0236] As used herein, "human target cells" include cells into which an exogenous product, such as a protein, a nucleic acid, or a combination thereof, has been introduced. In certain cases, human target cells can be used to produce a gene product from exogenous DNA, such as a transgene, such as an exogenous protein, such as a CAR. In certain cases, human target cells can contain a target nucleotide sequence within a target polynucleotide, and a nucleic acid-guided nuclease hybridizes and cleaves at cleavage sites at one or more positions in one or more strands of the target polynucleotide at or near the target nucleotide sequence.
[0237] As used herein, a "cleavage site" refers to one or more locations where a nucleic acid-guided nuclease complex hydrolyzes the phosphodiester backbone of a single-stranded or double-stranded target polynucleotide after binding to a target nucleotide sequence in the target polynucleotide. In certain cases where the target polynucleotide of a nucleic acid-guided nuclease complex is double-stranded, binding of the nucleic acid-guided nuclease complex to a target nucleotide sequence within the target polynucleotide can result in hydrolysis of one of the strands of the target polynucleotide at or near the target nucleotide sequence, resulting in strand cleavage. In such cases, the nucleic acid-guided nuclease complex can cleave either strand of the target polynucleotide. In certain cases, binding of the nucleic acid-guided nuclease complex to a target nucleotide sequence within the target polynucleotide can result in hydrolysis of both strands of the target polynucleotide at or near the target nucleotide sequence, resulting in cleavage of both strands. The cleavage site can be the same for both strands, resulting in a blunt end, or the cleavage sites for each strand can be offset, resulting in a single-stranded overhang, e.g., a sticky end. In certain cases, a mismatch at or near the cleavage site may or may not affect the cleavage efficiency of the nucleic acid-guided nuclease complex.
[0238] In certain cases, uncontrolled gene integration adjacent to regulatory elements of proto-oncogenes is known to cause oncogenic transformation, which is particularly important when engineering cells for therapeutic applications. Therefore, it is desirable to identify suitable target polynucleotides containing target nucleotide sequences that result in safe and stable integration of exogenous DNA with sufficient expression in human target cells and their resulting progeny.
[0239] Exemplary characteristics of target nucleotide sequences that may exhibit predictable function without potentially deleterious changes in human target cell genome activity are: (1) >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from known cancer-associated genes; (2) >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from any miRNA / other functional small RNA; (3) >10 kb, e.g., >20 kb, away from any 5' gene end; , e.g., >30 and in some cases >50 kb away, (4) >10 kb, e.g., >20, e.g., >30 and in some cases >50 kb away from any replication origin, (5) >10 kb, e.g., >20, e.g., >30 and in some cases >50 kb away from any ultraconserved element, (6) exhibits low transcriptional activity, (7) is outside a copy number variable region, (8) is located in open chromatin, and (9) is unique, i.e., contains one or more copies per genome.
[0240] In certain embodiments, provided herein are compositions. In certain embodiments, provided herein are compositions for engineering a human target cell with a preferred target nucleotide sequence within a target polynucleotide of the human target cell.
[0241] In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least one of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least two of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least three of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least four of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least five of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least six of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least seven of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least eight of the exemplary characteristics. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has all of the exemplary characteristics.
[0242] In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least one additional exemplary feature. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least two additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least three additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, from any 5' gene end and further comprises at least four additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, from any 5' gene end and further comprises at least five additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, from any 5' gene end and further comprises at least six additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, from any 5' gene end and further comprises at least seven additional exemplary features. In certain embodiments, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises all eight additional exemplary features.
[0243] In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, and further comprises at least one additional exemplary feature. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, and further comprises at least two additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, and further comprises at least three additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene and further comprises at least four additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene and further comprises at least five additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene and further comprises at least six additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene and further comprises at least seven additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, and further comprises all eight additional exemplary characteristics.
[0244] In certain embodiments, a suitable target polynucleotide is >150 kb, such as >200, for example >250, and in some cases >300 kb, from a known cancer-associated gene, and >10 kb, such as >20, for example >30, and in some cases >50 kb, from any 5' gene end. In certain embodiments, a suitable target polynucleotide is >150 kb, such as >200, for example >250, and in some cases >300 kb, from a known cancer-associated gene, and >10 kb, such as >20, for example >30, and in some cases >50 kb, from any 5' gene end, and further comprises at least one additional exemplary feature. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least two additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least three additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least four additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least five additional exemplary features.In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises at least six additional exemplary features. In certain embodiments, a suitable target polynucleotide is >150 kb, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene, >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and further comprises all seven additional exemplary features.
[0245] In a preferred embodiment, a suitable target polynucleotide is >10 kb, e.g., >20, e.g., >30, and in some cases >50 kb, away from any 5' gene end, and >150, e.g., >200, e.g., >250, and in some cases >300 kb, away from a known cancer-associated gene.
[0246] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020-2043 in Table 6. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or completely identical to any one of SEQ ID NOs: 2020-2043. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020-2043. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020-2043.
[0247] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020-2042 in Table 6. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or completely identical to any one of SEQ ID NOs: 2020-2042. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020-2042. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020-2042.
[0248] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOS: 2020-2041 and 2043 in Table 6. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or completely identical to any one of SEQ ID NOS: 2020-2041 and 2043. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOS: 2020-2041 and 2043. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020-2041 and 2043.
[0249] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020-2041 in Table 6. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or completely identical to any one of SEQ ID NOs: 2020-2041. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020-2041. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020-2041.
[0250] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise at least a portion of any one of SEQ ID NOs: 2020-2030 of Table 6, e.g., nucleotides 1-495, 1-490, 1-485, 1-480, 1-475, 1-470, 1-465, 1-460, 1-455, 1-450, 1-445, 1-440, 1-435, 1-430, 1-425, 1-420, 1-415, 1-410, 1-405, or 1-400. In certain embodiments, suitable target polynucleotides comprising a target nucleotide sequence are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or completely identical to a portion of any one of SEQ ID NOs: 2020-2030.
[0251] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise at least a portion of any one of SEQ ID NOs: 2031-2041 in Table 6, e.g., nucleotides 5-500, 10-500, 15-500, 20-500, 25-500, 30-500, 35-500, 40-500, 45-500, 50-500, 55-500, 60-500, 65-500, 70-500, 75-500, 80-500, 85-500, 90-500, 95-500, or 100-500. In certain embodiments, suitable target polynucleotides comprising a target nucleotide sequence are at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or completely identical to a portion of any one of SEQ ID NOs: 2031-2041.
[0252] [Table 44]
[0253] [Table 45]
[0254] [Table 46]
[0255] [Table 47]
[0256] [Table 48]
[0257] Table 49
[0258]
Table 50
[0259] Table 51
[0260] Table 52
[0261] Table 53
[0262] Table 54
[0263] Table 55
[0264] Table 56
[0265] Table 57
[0266] Table 58
[0267] In certain cases, expression of exogenous DNA, e.g., a transgene, inserted into a target polynucleotide at or near a target nucleotide sequence may depend on the cell type and differentiation stage, where one or more components of the target polynucleotide are activated during differentiation while other components are silenced, which may or may not correlate with rearrangements in chromatin structure reorganization during differentiation. To overcome this, in certain embodiments, in addition to the exemplary characteristics described above, a suitable target polynucleotide comprising a target nucleotide sequence demonstrates suitable expression of the inserted exogenous DNA, e.g., a transgene, throughout differentiation and clonal expansion.
[0268] IV. Pharmaceutical Compositions Provided herein are compositions (e.g., pharmaceutical compositions) comprising a guide nucleic acid, an engineered non-natural system, or a eukaryotic cell, such as a guide nucleic acid disclosed herein, an engineered non-natural system, or a eukaryotic cell. In certain embodiments, the composition comprises an RNP comprising a guide nucleic acid, such as a guide nucleic acid disclosed herein, and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition comprises a single guide nucleic acid, such as a single guide nucleic acid disclosed herein. In certain embodiments, the composition comprises an RNP comprising a single guide nucleic acid and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition comprises an RNP comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition comprises a complex of a targeter nucleic acid and a modulator nucleic acid, such as a complex of a targeter nucleic acid and a modulator nucleic acid disclosed herein. In certain embodiments, the composition comprises an RNP comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein (e.g., a Cas nuclease).
[0269] In certain embodiments, provided herein are methods of producing a composition, the method comprising incubating a single guide nucleic acid, such as a single guide nucleic acid disclosed herein, with a Cas protein, thereby producing a complex of the single guide nucleic acid and the Cas protein (e.g., an RNP). In certain embodiments, the method further comprises purifying the complex (e.g., an RNP).
[0270] In certain embodiments, provided herein are methods of producing a composition, the method comprising incubating a targeter nucleic acid and a modulator nucleic acid, such as a targeter nucleic acid and a modulator nucleic acid disclosed herein, under suitable conditions, thereby producing a composition (e.g., a pharmaceutical composition) comprising a complex of the targeter nucleic acid and the modulator nucleic acid. In certain embodiments, the method further comprises incubating the targeter nucleic acid and the modulator nucleic acid with a Cas protein (e.g., a Cas nuclease or related Cas protein that the targeter nucleic acid and the modulator nucleic acid can activate), thereby producing a complex of the targeter nucleic acid, the modulator nucleic acid, and the Cas protein (e.g., an RNP). In certain embodiments, the method further comprises purifying the complex (e.g., an RNP).
[0271] For therapeutic use, the guide nucleic acid, engineered non-native system, CRISPR expression system, or cells containing such systems or modified by such systems disclosed herein may be combined with a pharmaceutically acceptable carrier. As used herein, the term "pharmaceutically acceptable" may refer to compounds, materials, compositions, and / or dosage forms that are suitable for use in contact with the tissues of human beings and animals, within the scope of sound medical judgment, at a reasonable benefit / risk ratio, and without undue toxicity, irritation, allergic response, or other problem or complication.
[0272] As used herein, the term "pharmaceutically acceptable carrier" includes buffers, carriers, and excipients that are suitable for use in contact with human and animal tissues without undue toxicity, irritation, allergic response, or other problems or complications, commensurate with a reasonable benefit / risk ratio. Pharmaceutically acceptable carriers include any of the standard pharmaceutical carriers, such as phosphate-buffered saline, water, emulsions (e.g., oil-in-water or water-in-oil emulsions), and various types of wetting agents. The compositions may also contain stabilizers and preservatives. For examples of carriers, stabilizers, and adjuvants, see, for example, Martin, Remington's Pharmaceutical Sciences, 15th Ed., Mack Publ. Co., Easton, PA (1975). Pharmaceutically acceptable carriers include buffers, solvents, dispersion media, coatings, isotonic and absorption delaying agents, and the like, that are compatible with pharmaceutical administration. The use of such media and agents for pharmacologically active substances is well known in the art.
[0273] In certain embodiments, the pharmaceutical compositions disclosed herein comprise salts such as NaCl, MgCl, KCl, MgSO, etc.; buffers such as Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), MES sodium salt, 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.; solubilizing agents; detergents such as non-ionic detergents, such as Tween-20; nuclease inhibitors, etc. For example, in certain embodiments, the compositions of the invention comprise a DNA-targeting RNA, e.g., a gRNA, and a buffer for stabilizing the nucleic acid.
[0274] In certain embodiments, pharmaceutical compositions may contain formulatory materials to adjust, maintain, or preserve, for example, the pH, osmolality, viscosity, clarity, color, isotonicity, odor, sterility, stability, rate of dissolution or release, adsorption, or permeability of the composition. In such embodiments, suitable formulation materials include, but are not limited to, amino acids (such as glycine, glutamine, asparagine, arginine, or lysine); antimicrobial agents; antioxidants (such as ascorbic acid, sodium sulfite, or sodium bisulfite); buffers (such as borate, bicarbonate, Tris-HCl, citrate, phosphate, or other organic acids); bulking agents (such as mannitol or glycine); chelating agents (such as ethylenediaminetetraacetic acid (EDTA)); complexing agents (such as caffeine, polyvinylpyrrolidone, β-cyclodextrin, or hydroxypropyl-β-cyclodextrin); fillers; monosaccharides; disaccharides; and other carbohydrates (such as glucose, mannose, or dextrin); proteins (such as serum albumin, gelatin, or immunoglobulins); colorants, flavoring agents, and diluents; emulsifiers; hydrophilic polymers (such as polyvinylpyrrolidone); low molecular weight polypeptides; salt-forming agents preservatives (such as benzalkonium chloride, benzoic acid, salicylic acid, thimerosal, phenethyl alcohol, methylparaben, propylparaben, chlorhexidine, sorbic acid, or hydrogen peroxide); solvents (such as glycerin, propylene glycol, or polyethylene glycol); sugar alcohols (such as mannitol or sorbitol); suspending agents; surfactants or wetting agents (such as pluronics, PEG, sorbitan esters, polysorbates such as polysorbate 20, polysorbate, triton, tromethamine, lecithin, cholesterol, tyloxapal, etc.); stability enhancers (such as sucrose or sorbitol); tonicity enhancers (such as alkali metal halides, preferably sodium or potassium chloride, mannitol, sorbitol, etc.); delivery vehicles; diluents; excipients and / or pharmaceutical adjuvants (see Remington's Pharmaceutical Sciences, 18th ed. (Mack Publishing Company, 1990)).
[0275] In certain embodiments, the pharmaceutical composition may contain nanoparticles, such as polymeric nanoparticles, liposomes, or micelles (see Anselmo et al. (2016) Bioeng. Transl. Med. 1:10-29). In certain embodiments, the pharmaceutical composition comprises inorganic nanoparticles. Exemplary inorganic nanoparticles include, for example, magnetic nanoparticles (e.g., Fe3MnO2) or silica. The outer surface of the nanoparticle may be conjugated with a positively charged polymer (e.g., polyethyleneimine, polylysine, polyserine) that allows for attachment (e.g., conjugation or encapsulation) of a payload. In certain embodiments, the pharmaceutical composition comprises organic nanoparticles (e.g., encapsulation of a payload within the nanoparticle). Exemplary organic nanoparticles include, for example, SNALP liposomes coated with polyethylene glycol (PEG) and containing cationic lipids together with neutral helper lipids, and protamine and nucleic acid complexes coated with a lipid coating. In certain embodiments, the pharmaceutical composition comprises a liposome, such as a liposome disclosed in International (PCT) Application Publication No. WO 2015 / 148863.
[0276] In certain embodiments, the pharmaceutical compositions comprise a targeting moiety to enhance target cell binding or renewal of nanoparticles and liposomes. Exemplary targeting moieties include cell-specific antigens, monoclonal antibodies, single-chain antibodies, aptamers, polymers, saccharides, and cell membrane-penetrating peptides. In certain embodiments, the pharmaceutical compositions comprise fusogenic or endosome-destabilizing peptides or polymers.
[0277] In certain embodiments, the pharmaceutical composition may contain a sustained- or controlled-delivery formulation. Techniques for formulating sustained- or controlled-delivery means, such as liposome carriers, biodegradable microparticles or porous beads, and depot injections, are also known to those skilled in the art. Sustained-release formulations may include, for example, porous polymeric microparticles or semipermeable polymer matrices in the form of shaped articles, such as films or microcapsules. Sustained-release matrices may include polyesters, hydrogels, polylactides, copolymers of L-glutamic acid and gamma-ethyl-L-glutamate, poly(2-hydroxyethyl-inethacrylate), ethylene vinyl acetate, or poly-D(-)-3-hydroxybutyric acid. Sustained-release compositions may also include liposomes, which may be prepared by any of several methods known in the art.
[0278] The pharmaceutical compositions of the present invention can be administered by various methods known in the art. The route and / or method of administration vary depending on the desired results. Administration can be intravenous, intramuscular, intraperitoneal, or subcutaneous, or can be administered proximal to the target site. The pharmaceutically acceptable carrier should be suitable for intravenous, intramuscular, subcutaneous, parenteral, spinal, or epidermal administration (e.g., by injection or infusion). Depending on the route of administration, the active compound (e.g., the guide nucleic acid, engineered non-natural system, or CRISPR expression system disclosed herein) can be coated with a material to protect the compound from the action of acids and other natural conditions that may inactivate the compound.
[0279] Suitable formulation components for parenteral administration include sterile diluents such as water for injection, saline solution, fixed oils, polyethylene glycol, glycerin, propylene glycol, or other synthetic solvents; antibacterial agents such as benzyl alcohol or methyl parabens; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as EDTA; buffers such as acetate, citrate, or phosphate; and agents for the adjustment of tonicity such as sodium chloride or dextrose.
[0280] For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor EL™ (BASF, Parsippany, NJ), or phosphate buffered saline (PBS). The carrier should be stable under the conditions of manufacture and storage and should be preserved against microorganisms. The carrier can be, for example, a solvent or dispersion medium containing water, ethanol, polyol (e.g., glycerol, propylene glycol, and liquid polyethylene glycol), and suitable mixtures thereof.
[0281] Pharmaceutical preparations are preferably sterilized.Sterilization can be carried out by any suitable method, for example, by filtration through a sterile filtration membrane.When the composition is lyophilized, sterilization by filtration can be carried out before or after lyophilization and reconstitution.In certain embodiments, pharmaceutical compositions are lyophilized and then reconstituted with buffered saline at the time of administration.
[0282] The pharmaceutical compositions of the present invention can be prepared according to methods well known and routinely practiced in the art. See, for example, Remington: The Science and Practice of Pharmacy, Mack Publishing Co., 20th ed., 2000; and Sustained and Controlled Release Drug Delivery Systems, JR Robinson, ed., Marcel Dekker, Inc., New York, 1978. Pharmaceutical compositions are preferably manufactured under GMP conditions. Typically, a therapeutically effective dose or an effective dose of the guide nucleic acid, engineered non-natural system, or CRISPR expression system disclosed herein is used in the pharmaceutical compositions of the present invention. The compositions disclosed herein are formulated into pharmaceutically acceptable dosage forms by conventional methods known to those skilled in the art. The dosage regimen is adjusted to provide the optimal desired response (e.g., therapeutic response). For example, a single bolus can be administered, or several divided doses can be administered over time, or the dose can be proportionally reduced or increased as indicated by the exigencies of the therapeutic situation. It is particularly advantageous to formulate parenteral compositions into dosage unit form for ease of administration and uniformity of dosage.As used herein, dosage unit form refers to physically discrete units suitable as unitary dosages for the subject to be treated, each unit containing a predetermined amount of active compound calculated to produce desired therapeutic effect together with necessary pharmaceutical carriers.
[0283] The actual dosage level of the active ingredient in the pharmaceutical compositions of the present invention may be varied to obtain an amount of the active ingredient that is effective to obtain the desired therapeutic response for a particular patient, composition, and method of administration, and that is not toxic to the patient. The selected dosage level will depend on various pharmacokinetic factors, including the activity of the particular composition disclosed herein or its ester, salt, or amide employed, the route of administration, the time of administration, the excretion rate of the particular compound employed, the duration of treatment, other drugs, compounds, and / or materials used in combination with the particular composition employed, and factors such as the age, sex, weight, condition, overall health, and medical history of the patient being treated.
[0284] V. Therapeutic Use For example, the guide nucleic acids, engineered non-natural systems, and CRISPR expression systems disclosed herein are useful for targeting, editing, and / or modifying genomic DNA in a cell or organism. These guide nucleic acids and systems, and cells comprising one of the systems or whose genomes have been modified by one of the systems, can be used to treat diseases or disorders in which modification of genetic or epigenetic information is desirable. Accordingly, provided herein are methods for treating a disease or disorder, comprising administering to a subject in need thereof a guide nucleic acid, non-natural system, CRISPR expression system, or cell disclosed herein.
[0285] The term "subject" includes human and non-human animals. Non-human animals include all vertebrates, e.g., mammals and non-mammals, such as non-human primates, sheep, dogs, cows, chickens, amphibians, and reptiles. Except where noted, the terms "patient" and "subject" are used interchangeably herein.
[0286] The terms "treatment," "treating," "treat," "treated," and the like, as used herein, may refer to obtaining a desired pharmacological and / or physiological effect. The effect may be therapeutic in terms of a partial or complete cure of a disease and / or adverse effects resulting from a disease, or a delay in the progression of a disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal, e.g., a human, and includes (a) inhibiting the disease, i.e., halting its progression, and (b) relieving the disease, i.e., causing regression of the disease. It is understood that diseases or disorders may be identified by genetic methods and treated prior to the manifestation of any medical symptoms.
[0287] To minimize toxicity and off-target effects, it may be important to control the concentration of the delivered CRISPR-Cas system. The optimal concentration can be determined by testing various concentrations in cell models, tissue models, or non-human eukaryotic animal models and analyzing the degree of modification at potential off-target genomic loci using deep sequencing. The concentration that achieves the highest level of on-target modification while minimizing the level of off-target modification is generally selected for ex vivo or in vivo delivery.
[0288] It is understood that the guide nucleic acids, engineered non-natural systems, and CRISPR expression systems disclosed herein can be used to treat any suitable disease or disorder that can be ameliorated by an intracellular system.
[0289] For therapeutic purposes, certain methods disclosed herein are particularly suitable for editing or modifying proliferative cells, such as stem cells (e.g., hematopoietic stem cells), progenitor cells (e.g., hematopoietic progenitor cells or lymphoid progenitor cells), or memory cells (e.g., memory T cells). Given that such cells are delivered to a subject and expanded in vivo, there is a low tolerance for off-target events. However, on-target and off-target events can be assessed prior to delivery, thereby selecting one or more colonies that have the desired editing or modification and not undesired editing or modification. Thus, lower editing or modification efficiency can be tolerated for such cells. The engineered non-natural system of the present invention has the advantage of increasing or decreasing the efficiency of nucleic acid cleavage, for example, by adjusting the hybridization of dual guide nucleic acids. As a result, it can be used to minimize off-target events when generating genetically modified proliferative cells.
[0290] In certain embodiments, the guide nucleic acids, engineered non-native systems, and / or CRISPR expression systems disclosed herein can be used to engineer immune cells. Immune cells include, but are not limited to, lymphocytes (e.g., B lymphocytes or B cells, T lymphocytes or T cells, and natural killer cells), myeloid cells (e.g., monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes), and stem and progenitor cells capable of differentiating into these cell types (e.g., hematopoietic stem cells, hematopoietic progenitor cells, and lymphoid progenitor cells). Cells can include autologous cells from the subject being treated or allogeneic cells from a donor.
[0291] In certain embodiments, the immune cells are T cells, which can be, for example, cultured T cells, primary T cells, T cells from a cultured T cell line (e.g., Jurkat, SupTi), or T cells obtained from a mammal, e.g., the subject to be treated. If obtained from a mammal, the T cells can be obtained from many sources, including, but not limited to, blood, bone marrow, lymph nodes, thymus, or other tissues or fluids. The T cells can also be enriched or purified. The T cells can be any type of T cell, including, but not limited to, CD4 + / CD8 + Double positive T cells, CD4 + Helper T cells (e.g., Th1 and Th2 cells), CD8 + The T cells may be at any stage of development, including T cells (e.g., cytotoxic T cells), tumor-infiltrating lymphocytes (TILs), memory T cells (e.g., central memory T cells and effector memory T cells), regulatory T cells, naive T cells, etc.
[0292] In certain embodiments, immune cells, such as T cells, are engineered to express exogenous gene.For example, in certain embodiments, the engineered CRISPR system disclosed herein can catalyze the DNA cleavage in locus, and enable the site-specific integration of exogenous gene in locus by HDR.
[0293] In certain embodiments, immune cells, e.g., T cells, are engineered to express a chimeric antigen receptor (CAR), i.e., the T cells contain an exogenous nucleotide sequence encoding the CAR. As used herein, the term "chimeric antigen receptor" or "CAR" includes any artificial receptor that contains an antigen-specific binding moiety and one or more signaling chains derived from an immune receptor. A CAR may contain a single-chain variable fragment (scFv) of an antigen-specific antibody linked via a hinge and transmembrane region to the cytoplasmic domain of a T cell signaling molecule, e.g., a T cell triggering domain (e.g., from CD3ζ) and a tandem T cell costimulatory domain (e.g., from CD28, CD137, OX40, ICOS, or CD27). T cells that express a chimeric antigen receptor are referred to as CAR T cells. Exemplary CAR T cells include CD19-targeted CTL019 cells (see Grupp et al. (2015) BLOOD, 126:4983), 19-28z cells (see Park et al. (2015) J. CLIN. ONCOL., 33:7010), and KTE-C19 cells (see Locke et al. (2015) BLOOD, 126:3991). Further exemplary CAR T cells are described in U.S. Pat. Nos. 7,446,190, 8,399,645, 8,906,682, 9,181,527, 9,272,002, 9,266,960, 10,253,086, 10640569, and 10,808,035 and International (PCT) Publication Nos. WO 2013 / 142034, WO 2015 / 120180, WO 2015 / 188141, WO 2016 / 120220, and WO 2017 / 040945.Exemplary techniques for expressing CARs using the CRISPR system are described in Hale et al. (2017) MOL THER METHODS CLIN DEV., 4:192, MacLeod et al. (2017) MOL THER, 25:949, and Eyquem et al. (2017) NATURE, 543:113.
[0294] In certain embodiments, immune cells, e.g., T cells, bind to an antigen, e.g., a cancer antigen, via an endogenous T cell receptor (TCR). In certain embodiments, immune cells, e.g., T cells, are engineered to express an exogenous TCR, e.g., an exogenous natural TCR or an exogenous engineered TCR. T cell receptors comprise two chains, termed α and β chains, which combine on the surface of T cells to form a heterodimeric receptor capable of recognizing MHC-restricted antigens. Each of the α and β chains comprises a constant region and a variable region. Each variable region of the α and β chains defines three loops, termed complementarity-determining regions (CDRs), known as CDR1, CDR2, and CDR3, which confer antigen-binding activity and binding specificity to the T cell receptor.
[0295] In certain embodiments, the CAR or TCR is selected from the group consisting of B-cell maturation antigen (BCMA), mesothelin, prostate-specific membrane antigen (PSMA), prostate stem cell antigen (PSCA), carbonic anhydrase IX (CAIX), carcinoembryonic antigen (CEA), CD5, CD7, CD10, CD19, CD20, CD22, CD30, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD70, CD74, CD123, CD133, CD138, epithelial glycoprotein 2 (EGP2), epithelial glycoprotein- 40 (EGP-40), epithelial cell adhesion molecule (EpCAM), receptor tyrosine-protein kinase (FLT3), folate-binding protein (FBP), fetal acetylcholine receptor (AChR), folate receptor-α and β (FRa and β), ganglioside G2 (GD2), ganglioside G3 (GD3), epidermal growth factor receptor 2 (HER-2 / ERB2), epidermal growth factor receptor vIII (EGFRvIII), ERB3, ERB4, human telomerase reverse transcriptase (hTERT), interleukin (IL-1) Interleukin-13 receptor subunit alpha-2 (IL-13Ra2), K-light chain, kinase insert domain receptor (KDR), Lewis A (CA19.9), Lewis Y (LeY), LI cell adhesion molecule (LICAM), melanoma-associated antigen 1 (melanoma antigen family A1, MAGE-A1), mucin 16 (MUC-16), mucin 1 (MUC-1, e.g., truncated MUC-1), KG2D ligand, cancer-testis antigen NY-ESO-1, oncofetal antigen (h5T4), tumor-associated glycoprotein 72 (TAG-72), The antibody binds to a cancer antigen selected from vascular endothelial growth factor R2 (VEGF-R2), Wilms' tumor protein (WT-1), tyrosine-protein kinase type 1 transmembrane receptor (ROR1), B7-H3 (CD276), B7-H6 (Nkp30), chondroitin sulfate proteoglycan-4 (CSPG4), DNAX accessory molecule (DNAM-1), ephrin type A receptor 2 (EpHA2), fibroblast-associated protein (FAP), Gpl00 / HLA-A2, glypican 3 (GPC3), HA-IH, HERK-V, IL-1 IRa, latency membrane protein 1 (LMP1), neural cell adhesion molecule (N-CAM / CD56), and TRAIL receptor (TRAIL-R).
[0296] Suitable loci for insertion of a CAR coding sequence or an exogenous TCR coding sequence include, but are not limited to, safe harbor loci (e.g., the AAVS1 locus) and TCR subunit loci (e.g., the TCR alpha constant region (TRAC) locus, the TCR beta constant region 1 (TRBC1) locus, and the TCR beta constant region 2 (TRBC2) locus). It is understood that insertion at the TRAC locus reduces tonic CAR signaling and improves T cell potency (see Eyquem et al. (2017) NATURE, 543:113). Furthermore, inactivation of the endogenous TRAC, TRBC1, or TRBC2 gene may reduce graft-versus-host disease (GVHD) responses, thereby enabling the use of allogeneic T cells as starting material for the preparation of CAR T cells. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of endogenous TCRs or TCR subunits, e.g., TRAC, TRBC1, and / or TRBC2. Cells can be engineered to have partially reduced or no expression of endogenous TCRs or TCR subunits. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) expression of endogenous TCRs or TCR subunits compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of endogenous TCRs or TCR subunits. Exemplary techniques for reducing TCR expression using the CRISPR system are described in U.S. Pat. No. 9,181,527, Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, Cooper et al. (2018) LEUKEMIA, 32:1970, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0297] It is understood that certain immune cells, e.g., T cells, also express major histocompatibility complex (MHC) or human leukocyte antigen (HLA) genes, and inactivation of these endogenous genes may reduce the immune response, thereby enabling the use of allogeneic T cells as starting material for the preparation of CAR T cells. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of one or more endogenous class I or class II MHC or HLA (e.g., β2-microglobulin (B2M), class II major histocompatibility complex transactivator (CIITA)). Cells can be engineered to have partially reduced or no expression of endogenous MHC or HLA. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) expression of endogenous MHC (e.g., B2M, CIITA) compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of endogenous MHC (e.g., B2M, CIITA). In certain cases, cells may be engineered to have expression of, for example, HLA-E and / or HLA-G to avoid attack by natural killer (NK) cells. Exemplary techniques for reducing MHC expression using the CRISPR system are described in Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0298] Other genes that may be inactivated include, but are not limited to, CD3, CD52, and deoxycytidine kinase (DCK). For example, inactivation of DCK may render immune cells (e.g., T cells) resistant to purine nucleotide analog (PNA) compounds, which are often used to suppress the host immune system to reduce GVHD responses during immune cell therapy. In certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) of endogenous CD52 or DCK expression compared to corresponding unmodified or parental cells.
[0299] It is understood that the activity of immune cells (e.g., T cells) can be enhanced by inactivating or reducing the expression of immunosuppressants, such as immune checkpoint proteins. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of immune checkpoint proteins. Exemplary immune checkpoint proteins expressed by wild-type T cells include, but are not limited to, PDCD1 (PD-1), CTLA4, ADORA2A (A2AR), B7-H3, B7-H4, BTLA, KIR, LAG3, HAVCR2 (TIM3), TIGIT, VISTA, PTPN6 (SHP-1), and FAS. Cells can be modified to have partially reduced or no expression of immune checkpoint proteins. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%) expression of immune checkpoint proteins compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of immune checkpoint proteins. Exemplary approaches for reducing immune checkpoint protein expression using CRISPR systems are described in International (PCT) Publication No. WO 2017 / 017184, Cooper et al. (2018) LEUKEMIA, 32:1970, Su et al. (2016) ONCOIMMUNOLOGY, 6:e1249558, and Zhang et al. (2017) FRONT MED, 11:554.
[0300] Immune cells can be engineered by gene editing or modification to have reduced expression of endogenous genes, such as the endogenous genes described above. For example, in certain embodiments, the engineered CRISPR system disclosed herein can cause DNA breaks at the locus, thereby inactivating the targeted gene. In other embodiments, the engineered CRISPR system disclosed herein can be fused to an effector domain (e.g., a transcriptional repressor or histone methylase) to reduce the expression of the target gene.
[0301] Immune cells can also be engineered to express exogenous proteins (in addition to the antigen binding proteins listed above) at the locus of the human ADORA2A, B2M, CD52, CIITA, CTLA4, DCK, FAS, HAVCR2, LAG3, PDCD1, PTPN6, TIGIT, TRAC, TRBC1, TRBC2, CARD11, CD247, IL7R, LCK, or PLCG1 genes.
[0302] In certain embodiments, immune cells, e.g., T cells, are modified to express a dominant-negative form of an immune checkpoint protein. In certain embodiments, the dominant-negative form of the checkpoint inhibitor may act as a decoy receptor that binds to or sequesters the natural ligand that normally binds and activates the wild-type immune checkpoint protein. Examples of engineered immune cells, e.g., T cells, containing dominant-negative forms of immunosuppressants are described, for example, in International (PCT) Publication No. WO 2017 / 040945.
[0303] In certain embodiments, immune cells, e.g., T cells, are modified to express genes (e.g., transcription factors, cytokines, or enzymes) that regulate immune cell survival, proliferation, activity, or differentiation (e.g., into memory cells). In certain embodiments, immune cells are modified to express TET2, FOXO1, IL-12, IL-15, IL-18, IL-21, IL-7, GLUT1, GLUT3, HK1, HK2, GAPDH, LDHA, PDK1, PKM2, PFKFB3, PGK1, ENO1, GYS1, and / or ALDOA. In certain embodiments, the modification is an insertion of a nucleotide sequence encoding a protein operably linked to a regulatory element. In certain embodiments, the modification is a substitution of a single nucleotide polymorphism (SNP) site in the endogenous gene. In certain embodiments, immune cells, e.g., T cells, are modified to express a variant of a gene, e.g., a variant having higher activity than the respective wild-type gene. In certain embodiments, immune cells are modified to express mutants of CARD11, CD247, IL7R, LCK, or PLCG1. For example, several gain-of-function mutants of IL7R are disclosed in Zenatti et al., (2011) NAT. GENET. 43(10):932-39. The mutants can be expressed from the native locus of the respective wild-type gene by delivering the engineered systems described herein to target the native locus in combination with a donor template bearing the mutant or a portion thereof.
[0304] In certain embodiments, immune cells, e.g., T cells, are modified to express proteins (e.g., cytokines or enzymes) that regulate the microenvironment (e.g., the tumor microenvironment) into which the immune cells are designed to migrate. In certain embodiments, the immune cells are modified to express CA9, CA12, V-ATPase subunit, NHE1, and / or MCT-1.
[0305] A. Gene Therapy For example, it is understood that the engineered non-natural systems and CRISPR expression systems disclosed herein can be used to treat genetic diseases or disorders, i.e., diseases or disorders associated with or mediated by unwanted mutations in a subject's genome.
[0306] Exemplary genetic diseases or disorders include age-related macular degeneration, adrenoleukodystrophy (ALD), Alagille syndrome, alpha-1-antitrypsin deficiency, argininemia, argininosuccinic aciduria, ataxia (e.g., Friedreich's ataxia, spinocerebellar ataxia, ataxia-telangiectasia, essential tremor, spastic paraplegia), autism, biliary atresia, biotinidase deficiency, carbamoyl phosphate synthase I deficiency, glycoprotein glycosylation syndrome (CDGS), central nervous system (CNS)-related diseases (e.g., Alzheimer's disease, amyotrophic lateral sclerosis), and disorders of the retina (e.g., retinal ventricles, retinal plexus, retinal plexus, retinal plexus) associated with retinal plexus syndrome (RVS). amyotrophic lateral sclerosis (ALS), Canavan disease (CD), ischemia, multiple sclerosis (MS), neuropathic pain, Parkinson's disease), Bloom's syndrome, cancer, Charcot-Marie-Tooth disease (e.g., peroneal muscular atrophy, hereditary motor and sensory neuropathy), congenital hepatic porphyria, citrullinemia, Crigler-Najjar syndrome, cystic fibrosis (CF), dentatorubral-pallidoluysian atrophy (DRPLA), diabetes insipidus, Fabry disease, familial hypercholesterolemia (LDL receptor deficiency), Fanconi anemia, fragile X syndrome, fatty acid metabolism disorders, galactosemia, Glucose-6-phosphate dehydrogenase (G6PD), glycogen storage diseases (e.g., types I (glucose-6-phosphatase deficiency, von Gierke disease), II (α-glucosidase deficiency, Pompe disease), III (debranching enzyme deficiency, Corey disease), IV (branching enzyme deficiency, Anderson disease), V (muscle glycogen phosphorylase deficiency, McArdle disease), VII (muscle phosphofructokinase deficiency, Tauri disease), VI (hepatic phosphorylase deficiency, Yale disease), IX (hepatic glycogen phosphorylase kinase deficiency)), hemophilia Hemophilia A (associated with factor VIII deficiency), hemophilia B (associated with factor IX deficiency), Huntington's disease, glutaric aciduria, hypophosphatemia, Krabbe disease, lactic acidosis, Lafora disease, Leber congenital amaurosis, Lesch-Nyhan syndrome, lysosomal storage diseases, metachromatic leukodystrophy disease (MLD), mucopolysaccharidoses (MPS) (e.g., Hunter syndrome, Hurler syndrome, Maroteaux-Lamy syndrome, Sanfilippo syndrome, Scheie syndrome, Morquio syndrome, others), MPS I, MPS II, MPS III, MS IV, MPS7), muscular / skeletal disorders (e.g., muscular dystrophy, Duchenne muscular dystrophy), myotonic dystrophy (DM), neoplasia, N-acetylglutamate synthetase deficiency, ornithine transcarbamylase deficiency, phenylketonuria, primary open-angle glaucoma, retinitis pigmentosa, schizophrenia, severe combined immunodeficiency (SCID), spinal-bulbar muscular atrophy (SBMA), sickle cell anemia, Usher syndrome, Tay-Sachs disease, thalassemia (e.g., β-thalassemia), trinucleotide repeat diseases, tyrosinemia, Wilson's disease, Wiskott-Aldrich syndrome, X-linked chronic granulomatous disease (CGD), X-linked severe combined immunodeficiency, and xeroderma pigmentosum.
[0307] Further exemplary genetic diseases or disorders and related information are available on the World Wide Web at kumc.edu / gec / support, genome.gov / 10001200 and ncbi.nlm.nih.gov / books / NBK22183 / . Further exemplary genetic diseases or disorders, associated gene mutations, and gene therapy approaches for treating genetic diseases or disorders are described in International (PCT) Publication Nos. WO 2013 / 126794, WO 2013 / 163628, WO 2015 / 048577, WO 2015 / 070083, WO 2015 / 089354, WO 2015 / 134812, WO 2015 / 138510, WO 2015 / 148670, WO 2015 / 148860, WO 2015 / 148863, WO 2015 / 153780, WO 2015 / 153789, and WO 2015 / 153791 pamphlet, U.S. Patent Nos. 8,383,604, 8,859,597, 8,956,828, 9,255,130 and 9,273,296, and U.S. Patent Application Publication Nos. 2009 / 0222937, 2009 / 0271881, 2010 / 0229252 and 2010 / 0311124 and the like, as described in the patent documents 2011 / 0016540, 2011 / 0023139, 2011 / 0023144, 2011 / 0023145, 2011 / 0023146, 2011 / 0023153, 2011 / 0091441, 2012 / 0159653 and 2013 / 0145487.
[0308] B. Immune cell manipulation It is understood that the engineered non-natural system comprising the ssODNs disclosed herein can be used to engineer immune cells. Immune cells include, but are not limited to, lymphocytes (e.g., B lymphocytes or B cells, T lymphocytes or T cells, and natural killer cells), myeloid cells (e.g., monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes), and stem and progenitor cells that can differentiate into these cell types (e.g., hematopoietic stem cells, hematopoietic progenitor cells, and lymphoid progenitor cells). Cells can include autologous cells derived from the subject to be treated or allogeneic cells derived from a donor.
[0309] It is understood that the CRISPR systems comprising the ssODNs disclosed herein can be used to treat any disease or disorder that can be ameliorated by editing or modifying a target sequence, and exemplary genes containing target sequences that can be modified for therapeutic purposes include ADORA2A, ALPNR, B2M, BBS1, CALR, CARD11, CD3E, CD3G, CD38, CD40LG, CD52, CD58, CD247, CIITA, COL17A1, CSF1R, CSF2, CTLA4, DCK, DEFB134, DHODH, ERAP1, ERAP2, FAS, mir-101-2, and the like in cells. , HAVCR2 (also called TIM3), IFNGR1, IFNGR2, IL7R, JAK1, JAK2, LAG3, LCK, LCK1, MLANA, MVD, PDCD1 (also called PD-1), PLCG1, PLK1, PSMB5, PSMB8, PSMB9, PTCD2, PTPN1, PTPN6, PTPN11, RFX5, RFXAP, RPL23, RXANK, SOX10, SRP54, STAT1, Tap1, TAP2, TAPBP, TGFBR2, TIGIT, TIM3, TRAC, TRBC1, TRBC1+2, TRBC2, TUBB, TWF1 and / or U6 genes.
[0310] In certain embodiments, the immune cells are T cells, which can be, for example, cultured T cells, primary T cells, T cells from a cultured T cell line (e.g., Jurkat, SupTi), or T cells obtained from a mammal, e.g., the subject to be treated. If obtained from a mammal, the T cells can be obtained from many sources, including, but not limited to, blood, bone marrow, lymph nodes, thymus, or other tissues or fluids. The T cells can also be enriched or purified. The T cells can be any type of T cell and at any stage of development, including, but not limited to, CD4+ / CD8+ double positive T cells, CD4+ helper T cells (e.g., Th1 and Th2 cells), CD8+ T cells (e.g., cytotoxic T cells), tumor-infiltrating lymphocytes (TILs), memory T cells (e.g., central memory T cells and effector memory T cells), regulatory T cells, naive T cells, etc.
[0311] In certain embodiments, immune cells, such as T cells, are engineered to express an exogenous gene. For example, in certain embodiments, the engineered CRISPR systems disclosed herein can be used to engineer immune cells to express an exogenous gene. For example, in certain embodiments, the guide nucleic acids, engineered non-native systems, and CRISPR expression systems disclosed herein can be used to engineer immune cells to express an exogenous gene, such as human ADORA2A, ALPNR, B2M, BBS1, CALR, CARD11, CD3E, CD3G, CD38, CD40LG, CD52, CD58, CD247, CIITA, COL17A1, CSF1R, CSF2, CTLA4, DCK, DEFB134, DHODH, ERAP1, ERAP2, FAS, mir-101-2, HAVCR2 (also known as TIM3), IFNGR1, IFNGR2, IL7R, JAK1, JAK2 , LAG3, LCK, LCK1, MLANA, MVD, PDCD1 (also known as PD-1), PLCG1, PLK1, PSMB5, PSMB8, PSMB9, PTCD2, PTPN1, PTPN6, PTPN11, RFX5, RFXAP, RPL23, RXANK, SOX10, SRP54, STAT1, Tap1, TAP2, TAPBP, TGFBR2, TIGIT, TIM3, TRAC, TRBC1, TRBC1+2, TRBC2, TUBB, TWF1, and / or U6 gene loci. For example, in certain embodiments, an engineered CRISPR system comprising the ssODNs disclosed herein can catalyze DNA cleavage at the locus, allowing for site-specific integration of the exogenous gene at the locus via HDR, while reducing off-target effects by integrating the wild-type gene back into the off-target cleavage site via HDR.
[0312] In certain embodiments, immune cells, e.g., T cells, are engineered to express a chimeric antigen receptor (CAR), i.e., the T cells contain an exogenous nucleotide sequence encoding the CAR. As used herein, the term "chimeric antigen receptor" or "CAR" includes any artificial receptor that contains an antigen-specific binding moiety and one or more signaling chains derived from an immune receptor. A CAR may contain a single-chain variable fragment (scFv) of an antigen-specific antibody linked via a hinge and transmembrane region to the cytoplasmic domain of a T cell signaling molecule, e.g., a T cell triggering domain (e.g., from CD3≒i, aC) and a tandem T cell costimulatory domain (e.g., from CD28, CD137, OX40, ICOS, or CD27). T cells expressing a chimeric antigen receptor are referred to as CAR T cells. Exemplary CAR T cells include CD19-targeted CTL019 cells (see Grupp et al. (2015) BLOOD, 126:4983), 19-28z cells (see Park et al. (2015) J. CLIN. ONCOL., 33:7010), and KTE-C19 cells (see Locke et al. (2015) BLOOD, 126:3991). Further exemplary CAR T cells are described in U.S. Pat. Nos. 8,399,645, 8,906,682, 7,446,190, 9,181,527, 9,272,002, and 9,266,960, U.S. Patent Application Publication Nos. 2016 / 0362472, 2016 / 0200824, and 2016 / 0311917, and International (PCT) Publication Nos. WO 2013 / 142034, WO 2015 / 120180, WO 2015 / 188141, WO 2016 / 120220, and WO 2017 / 040945.Exemplary techniques for expressing CARs using the CRISPR system are described in Hale et al. (2017) MOL THER METHODS CLIN DEV., 4:192, MacLeod et al. (2017) MOL THER, 25:949, and Eyquem et al. (2017) NATURE, 543:113.
[0313] In certain embodiments, immune cells, e.g., T cells, bind to an antigen, e.g., a cancer antigen, via an endogenous T cell receptor (TCR). In certain embodiments, immune cells, e.g., T cells, are engineered to express an exogenous TCR, e.g., an exogenous natural TCR or an exogenous engineered TCR. T cell receptors comprise two chains, termed the OE± chain and the OE≦ chain, which combine on the surface of T cells to form a heterodimeric receptor capable of recognizing MHC-restricted antigens. Each of the OE± chain and the OE≦ chain comprises a constant region and a variable region. Each variable region of the OE± chain and the OE≦ chain defines three loops, termed complementarity-determining regions (CDRs), known as CDR1, CDR2, and CDR3, which confer antigen-binding activity and binding specificity to the T cell receptor.
[0314] In certain embodiments, the CAR or TCR is selected from the group consisting of B-cell maturation antigen (BCMA), mesothelin, prostate-specific membrane antigen (PSMA), prostate stem cell antigen (PCSA), carbonic anhydrase IX (CAIX), carcinoembryonic antigen (CEA), CD5, CD7, CD10, CD19, CD20, CD22, CD30, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD70, CD74, CD123, CD133, CD138, epithelial glycoprotein 2 (EGP2), epithelial glycoprotein-40 (EGP-40), epithelial cell adhesion molecule (EpCAM), receptor tyrosine-protein kinase (FLT3), folate binding protein (FBP), fetal acetylcholine receptor (AChR), folate receptor-OE± and OE≦(FR OE± and OE≦), ganglioside G2 (GD2), ganglioside G3 (GD3), epidermal growth factor receptor 2 (HER-2 / ERB2), epidermal growth factor receptor vIII (EGFRvIII), ERB3, ERB4, human telomerase reverse transcriptase (hTERT), interleukin-13 receptor subunit alpha-2 (IL-13Ra2), K-light chain, kinase insert domain receptor (KDR), Lewis A (CA19.9), Lewis Y (LeY), LI cell adhesion molecule (LICAM), melanoma-associated antigen 1 (melanoma antigen family A1, MAGE-A1), mucin 16 (MUC-16), mucin 1 (MUC-1, e.g., truncated MUC -1), KG2D ligand, cancer-testis antigen NY-ESO-1, oncofetal antigen (h5T4), tumor-associated glycoprotein 72 (TAG-72), vascular endothelial growth factor R2 (VEGF-R2), Wilms tumor protein (WT-1), type 1 tyrosine-protein kinase transmembrane receptor (ROR1), B7-H3 (CD276), B7-H6 (Nkp30), chondroitin sulfate proteoglycan-4 (CSPG4), DNAX accessory molecule (DNAM-1), ephrin type A receptor 2 (EpHA2), fibroblast-associated protein (FAP), Gpl00 / HLA-A2, glypican 3 (GPC3), HA-IH, HERK-V, IL-1 IRa, latency membrane protein 1 (LMP1), neural cell adhesion molecule (N-CAM / CD56), and TRAIL receptor (TRAIL-R).
[0315] Suitable loci for insertion of a CAR coding sequence or an exogenous TCR coding sequence include, but are not limited to, safe harbor loci (e.g., the AAVS1 locus), TCR subunit loci (e.g., the TCR OE±constant (TRAC) locus), and other loci associated with particular advantages (e.g., the CCR5 locus, whose inactivation can prevent or reduce HIV infection). Insertion at the TRAC locus is understood to reduce tonic CAR signaling and improve T cell potency (see Eyquem et al. (2017) NATURE, 543:113). Furthermore, inactivation of the endogenous TRAC gene may reduce graft-versus-host disease (GVHD) responses, thereby enabling the use of allogeneic T cells as starting material for the preparation of CAR-T cells. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of an endogenous TCR or TCR subunit, e.g., the TCR OE±subunit constant (TRAC). Cells can be engineered to have partially reduced or no expression of endogenous TCRs or TCR subunits. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) expression of endogenous TCRs or TCR subunits compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of endogenous TCRs or TCR subunits. Exemplary techniques for reducing TCR expression using the CRISPR system are described in U.S. Pat. No. 9,181,527, Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, Cooper et al. (2018) LEUKEMIA, 32:1970, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0316] It is understood that certain immune cells, e.g., T cells, also express major histocompatibility complex (MHC) or human leukocyte antigen (HLA) genes, and inactivation of these endogenous genes may reduce GVHD responses, thereby enabling the use of allogeneic T cells as starting material for the preparation of CAR-T cells. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of one or more endogenous class I or class II MHC or HLA (e.g., β2-microglobulin (B2M), class II major histocompatibility complex transactivator (CIITA), HLA-E, and / or HLA-G). Cells can be engineered to have partially reduced or no expression of endogenous MHC or HLA. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) expression of endogenous MHC (e.g., B2M, CIITA, HLA-E, or HLA-G) compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of endogenous MHC (e.g., B2M, CIITA, HLA-E, or HLA-G). Exemplary approaches for reducing MHC expression using CRISPR systems are described in Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0317] Other genes that can be inactivated to reduce GVHD responses include, but are not limited to, CD3, CD52, and deoxycytidine kinase (DCK). For example, inactivation of DCK can render immune cells (e.g., T cells) resistant to purine nucleotide analog (PNA) compounds, which are often used to suppress the host immune system to reduce GVHD responses during immune cell therapy. In certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) of endogenous CD52 or DCK expression compared to corresponding unmodified or parental cells.
[0318] In certain embodiments, immune cells, such as T cells, are engineered to have reduced expression of exogenous genes. For example, in certain embodiments, the engineered CRISPR system disclosed herein can be used to engineer immune cells to have reduced expression of endogenous genes. For example, in certain embodiments, the engineered CRISPR system disclosed herein can cause DNA breaks at the locus, thereby inactivating the targeted gene. In other embodiments, the engineered CRISPR system disclosed herein can be fused to an effector domain (e.g., a transcriptional repressor or histone methylase) to reduce the expression of the target gene.
[0319] It is understood that the activity of immune cells (e.g., T cells) can be enhanced by inactivating or reducing the expression of immunosuppressants, such as immune checkpoint proteins. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have reduced expression of immune checkpoint proteins. Exemplary immune checkpoint proteins expressed by wild-type T cells include, but are not limited to, PDCD1 (PD-1), CTLA4, ADORA2A (A2AR), B7-H3, B7-H4, BTLA, KIR, LAG3, HAVCR2 (TIM3), TIGIT, VISTA, PTPN6 (SHP-1), and FAS. Cells can be modified to have partially reduced or no expression of immune checkpoint proteins. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5%) expression of immune checkpoint proteins compared to corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of immune checkpoint proteins. Exemplary approaches for reducing immune checkpoint protein expression using CRISPR systems are described in International (PCT) Publication No. WO 2017 / 017184, Cooper et al. (2018) LEUKEMIA, 32:1970, Su et al. (2016) ONCOIMMUNOLOGY, 6:e1249558, and Zhang et al. (2017) FRONT MED, 11:554.
[0320] Immune cells express human ADORA2A, ALPNR, B2M, BBS1, CALR, CARD11, CD3E, CD3G, CD38, CD40LG, CD52, CD58, CD247, CIITA, COL17A1, CSF1R, CSF2, CTLA4, DCK, DEFB134, DHODH, ERAP1, ERAP2, FAS, mir-101-2, HAVCR2 (also known as TIM3), IFNGR1, IFNGR2, IL7R, JAK1, JAK2, LAG3, LCK, LCK1, MLANA, MVD, and PDCD1. (also known as PD-1), PLCG1, PLK1, PSMB5, PSMB8, PSMB9, PTCD2, PTPN1, PTPN6, PTPN11, RFX5, RFXAP, RPL23, RXANK, SOX10, SRP54, STAT1, Tap1, TAP2, TAPBP, TGFBR2, TIGIT, TIM3, TRAC, TRBC1, TRBC1+2, TRBC2, TUBB, TWF1, and / or U6 gene loci can also be engineered to express exogenous proteins (in addition to the antigen binding proteins listed above).
[0321] In certain embodiments, immune cells, e.g., T cells, are modified to express a dominant-negative form of an immune checkpoint protein. In certain embodiments, the dominant-negative form of the checkpoint inhibitor may act as a decoy receptor that binds to or sequesters the natural ligand that would normally bind to and activate the wild-type immune checkpoint protein. Examples of engineered immune cells, e.g., T cells, containing dominant-negative forms of immunosuppressants are described, for example, in International (PCT) Publication No. WO 2017 / 040945.
[0322] In certain embodiments, immune cells, e.g., T cells, are modified to express genes (e.g., transcription factors, cytokines, or enzymes) that regulate immune cell survival, proliferation, activity, or differentiation (e.g., into memory cells). In certain embodiments, immune cells are modified to express TET2, FOXO1, IL-12, IL-15, IL-18, IL-21, IL-7, GLUT1, GLUT3, HK1, HK2, GAPDH, LDHA, PDK1, PKM2, PFKFB3, PGK1, ENO1, GYS1, and / or ALDOA. In certain embodiments, the modification is an insertion of a nucleotide sequence encoding a protein operably linked to a regulatory element. In certain embodiments, the modification is a substitution of a single nucleotide polymorphism (SNP) site in the endogenous gene. In certain embodiments, immune cells, e.g., T cells, are modified to express a variant of a gene, e.g., a variant having higher activity than the respective wild-type gene. In certain embodiments, immune cells are modified to express mutants of CARD11, CD247, IL7R, LCK, or PLCG1. For example, several gain-of-function mutants of IL7R are disclosed in Zenatti et al., (2011) NAT. GENET. 43(10):932-39. The mutants can be expressed from the native locus of the respective wild-type gene by delivering the engineered systems described herein to target the native locus in combination with a donor template bearing the mutant or a portion thereof.
[0323] In certain embodiments, immune cells, e.g., T cells, are modified to express proteins (e.g., cytokines or enzymes) that regulate the microenvironment (e.g., the tumor microenvironment) into which the immune cells are designed to migrate. In certain embodiments, the immune cells are modified to express CA9, CA12, V-ATPase subunit, NHE1, and / or MCT-1.
[0324] In certain embodiments, methods are provided for treating a disease, e.g., cancer, by administering to a subject suffering from the disease an effective amount of T cells that have been modified to express a disease-specific CAR using the modified guide nucleic acid and CRISPR-Cas system described herein, e.g., in Sections IA, IA1, IB, IC, and IVB. In certain embodiments, the T cells are autologous cells removed from the subject, treated to modify their genomic DNA to express a CAR, expanded, and administered to the subject; in certain embodiments, the T cells are allogeneic T cells that have been treated to modify their genomic DNA to express a CAR. In certain embodiments, the disease is a blood cancer such as leukemia or lymphoma; in certain embodiments, the disease is a solid tumor cancer.
[0325] VI. Kit It is understood that the guide nucleic acids, engineered non-natural systems, CRISPR expression systems, and / or libraries disclosed herein can be packaged into kits suitable for use by healthcare providers. Accordingly, in another aspect, the present invention provides kits comprising any one or more of the elements disclosed in the above systems, libraries, methods, and compositions. In certain embodiments, the kits include instructions for using the engineered non-natural systems and kits disclosed herein. The instructions can be specific to the applications and methods described herein. In certain embodiments, one or more of the system elements are provided in solution. In certain embodiments, one or more of the system elements are provided in lyophilized form, and the kit further comprises a diluent. The elements can be provided individually or in combination and can be provided in any suitable container, such as a vial, bottle, tube, or immobilized on the surface of a solid substrate (e.g., a chip or microarray). In certain embodiments, the kits comprise one or more of the nucleic acids and / or proteins described herein. In certain embodiments, the kits include all of the elements of the systems of the present invention.
[0326] In certain embodiments of kits comprising an engineered non-natural dual guide system, the targeter nucleic acid and modulator nucleic acid are provided in separate containers, hi other embodiments, the targeter nucleic acid and modulator nucleic acid are pre-complexed and the complex is provided in a single container.
[0327] In certain embodiments, the kit comprises a Cas protein or a nucleic acid comprising a regulatory element operably linked to a nucleic acid encoding a Cas protein provided in a separate container, hi other embodiments, the kit comprises a Cas protein pre-complexed with a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, the complexes being provided in a single container.
[0328] In certain embodiments, the kit further comprises one or more donor templates provided in one or more separate containers. In certain embodiments, the kit comprises a plurality of donor templates disclosed herein (e.g., in separate tubes or immobilized on the surface of a solid substrate such as a chip or microarray), one or more guide nucleic acids disclosed herein, and optionally a regulatory element operably linked to a Cas protein or a nucleic acid encoding a Cas protein disclosed herein. Such a kit is useful for identifying a donor template that introduces optimal genetic modifications in a multiplex assay. The CRISPR expression system disclosed herein is also suitable for use in the kit.
[0329] In certain embodiments, the kit further comprises one or more reagents and / or buffers for use in a process employing one or more of the components described herein. The reagents may be provided in any suitable container and may be provided in a form ready for use in a particular assay or in a form requiring the addition of one or more other components prior to use (e.g., a concentrate or lyophilized form). The buffer may be a reaction or storage buffer, including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In certain embodiments, the buffer has a pH of about 7 to about 10. In certain embodiments, the kit further comprises a pharmaceutically acceptable carrier. In certain embodiments, the kit further comprises one or more devices or other materials for administration to a subject.
[0330] VII. Embodiments In embodiment 1, provided herein is a composition comprising: (A) a double-stranded DNA polynucleotide; and (B) a polypeptide comprising a nuclear localization signal (NLS) linked to the polynucleotide.
[0331] In embodiment 2, provided herein is a composition according to embodiment 1, wherein the double-stranded DNA polynucleotide comprises a plasmid.
[0332] In embodiment 3, provided herein is a composition according to embodiment 1, wherein the double-stranded DNA polynucleotide comprises a linear double-stranded DNA polynucleotide having covalently closed ends.
[0333] In embodiment 4, there is provided herein a composition according to any one of embodiments 1 to 3, wherein the double-stranded DNA polynucleotide comprises a donor template.
[0334] In embodiment 5, provided herein is a composition according to any one of embodiments 1 to 4, wherein the polypeptide comprises a nucleic acid-guided nuclease complex comprising (1) a nucleic acid-guided nuclease and (2) a guide nucleic acid (gNA).
[0335] In embodiment 6, provided herein is a composition according to embodiment 5, wherein the double-stranded DNA polynucleotide comprises a first PAM (P1) recognized by a nucleic acid-guided nuclease and a first target nucleotide sequence (T1) adjacent to, but not within, the donor template (D).
[0336] In embodiment 7, provided herein is a composition according to embodiment 6, wherein the double-stranded DNA polynucleotide further comprises a second PAM (P2) and a second target nucleotide sequence (T2) recognized by the nucleic acid-guided nuclease adjacent to, but not within, the donor template (D).
[0337] In embodiment 8, there is provided herein a composition according to embodiment 6 or 7, wherein the first PAM and the first target nucleotide sequence are oriented 5'P1+T1+D+3'.
[0338] In embodiment 9, provided herein is a composition according to embodiment 6 or 7, wherein the first PAM and the first target nucleotide sequence are oriented 5'T1-P1-D+3'.
[0339] In embodiment 10, there is provided herein a composition according to any one of embodiments 7 to 9, wherein the second PAM and the second target nucleotide sequence are oriented 5'D+P2+T2+3'.
[0340] In embodiment 11, there is provided herein a composition according to any one of embodiments 7 to 9, wherein the second suitable PAM and the second target nucleotide sequence are oriented 5'D+T2-P2-3'.
[0341] In embodiment 12, provided herein is a composition according to embodiment 7, wherein the nucleotide sequences of the first and second PAM targets are oriented 5'T1-P1-D+P2+T2+3'.
[0342] In embodiment 13, there is provided herein a composition according to embodiment 7, wherein the first and second PAM target nucleotide sequences are oriented 5'P1+T1+D+T2-P2-3'.
[0343] In embodiment 14, provided herein is a composition according to embodiment 5, wherein the nucleic acid-guided nuclease comprises an engineered non-naturally occurring nuclease.
[0344] In embodiment 15, provided herein is a composition according to any one of embodiments 5 to 14, wherein the nucleic acid-guided nuclease complex comprises a class 1 or class 2 nucleic acid-guided nuclease complex.
[0345] In embodiment 16, provided herein is a composition according to embodiment 15, wherein the nucleic acid-guided nuclease complex comprises a type II or type V nucleic acid-guided nuclease.
[0346] In embodiment 17, provided herein is a composition according to embodiment 16, wherein the nucleic acid-guided nuclease comprises a type VA, type VB, type VC, type VD, or type VE nucleic acid-guided nuclease.
[0347] In embodiment 18, provided herein is a composition according to embodiment 17, wherein the nucleic acid-guided nuclease comprises a type VA nucleic acid-guided nuclease.
[0348] In embodiment 19, provided herein is a composition of embodiment 18, wherein the nucleic acid-guided nuclease comprises a MAD nuclease, an ART nuclease, or an ABW nucleic acid-guided nuclease.
[0349] In embodiment 20, provided herein is a composition of embodiment 19, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of a MAD, ART, or ABW nucleic acid-guided nuclease.
[0350] In embodiment 21, provided herein is a composition of embodiment 19, wherein the nucleic acid-guided nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20.
[0351] In embodiment 22, provided herein is a composition of embodiment 19, wherein the nucleic acid-guided nuclease comprises ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART11*, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, or ART35.
[0352] In embodiment 23, provided herein is a composition of embodiment 19, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of MAD2, MAD7, ART2, ART11, or ART11*.
[0353] In embodiment 24, provided herein is a composition according to embodiment 19, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of SEQ ID NO: 37.
[0354] In embodiment 25, provided herein is a composition according to any one of embodiments 6 to 24, wherein the first PAM (P1) is a PAM recognized by a type V nucleic acid-guided nuclease.
[0355] In embodiment 26, provided herein is a composition of embodiment 25, wherein the PAM comprises a sequence of CTTN.
[0356] In embodiment 27, provided herein is a composition according to any one of embodiments 7 to 26, wherein the second PAM (P2) is a PAM recognized by a type V nucleic acid-guided nuclease.
[0357] In embodiment 28, provided herein is a composition according to any one of embodiments 5 to 27, wherein the gNA is an engineered, non-naturally occurring gNA.
[0358] In embodiment 29, there is provided herein a composition according to any one of embodiments 5 to 28, wherein the gNA comprises a single polynucleotide.
[0359] In embodiment 30, the present specification provides a composition described in any one of embodiments 5 to 28, wherein the gNA comprises a dual gNA comprising a targeter nucleic acid and a modulator nucleic acid, and the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides.
[0360] In embodiment 31, provided herein is a composition described in embodiment 30, wherein the dual gNAs can bind to and activate a nucleic acid-guided nuclease that is activated by a single crRNA in the absence of tracrRNA in a natural system.
[0361] In embodiment 32, provided herein is a composition described in any one of embodiments 28 to 31, wherein the gNA comprises a heterologous spacer sequence that shares complementarity with a first target sequence in the double-stranded DNA polynucleotide and a second target sequence in the human genome.
[0362] In embodiment 33, provided herein is a composition described in any one of embodiments 28 to 31, wherein the gNA comprises a heterologous spacer sequence that does not share complementarity with a target sequence in the human genome.
[0363] In embodiment 34, provided herein is a composition according to any one of embodiments 1 to 31, wherein the nucleic acid-guided nuclease comprises at least four NLSs.
[0364] In embodiment 35, provided herein is a composition according to embodiment 34, wherein the nucleic acid-guided nuclease comprises one N-terminal NLS and three C-terminal NLSs.
[0365] In embodiment 36, provided herein is a composition according to embodiment 34, wherein the nucleic acid-guided nuclease comprises five or more N-terminal NLSs.
[0366] In embodiment 37, provided herein is a composition according to any one of embodiments 1 to 35, wherein the NLS comprises any one of SEQ ID NOs: 40 to 56.
[0367] In embodiment 38, provided herein is a composition according to embodiment 37, wherein the NLS comprises SEQ ID NOs: 40, 51 and 56.
[0368] In embodiment 39, provided herein is a composition according to any one of embodiments 1 to 38, further comprising at least one of: (1) a buffer; (2) an RNP stabilizer.
[0369] In embodiment 40, provided herein is a composition according to embodiment 39, wherein the buffer lacks magnesium.
[0370] In embodiment 41, provided herein is a composition according to embodiment 39 or 40, wherein the RNP stabilizer comprises a peptide, poly-L-glutamic acid (PGA), or a single-stranded oligodeoxynucleotide (ssODN).
[0371] In embodiment 42, provided herein is a composition comprising: (A) a donor template (D); and (B) a polynucleotide comprising a first PAM (P1) and a first target nucleotide sequence (T1) adjacent to, but not within, the donor template (D), the first PAM being recognized by a type V nucleic acid-guided nuclease.
[0372] In embodiment 43, provided herein is a composition according to embodiment 42, wherein the polynucleotide comprises double-stranded DNA.
[0373] In embodiment 44, provided herein is a composition according to embodiment 42 or 43, wherein the polynucleotide comprises circular DNA.
[0374] In embodiment 45, there is provided herein a composition according to any one of embodiments 42 to 44, wherein the polynucleotide is a plasmid.
[0375] In embodiment 46, provided herein is a composition according to any one of embodiments 42 to 45, wherein the polynucleotide further comprises a second PAM (P2) and a second target nucleotide sequence (T2) recognized by a V-type nucleic acid-guided nuclease adjacent to, but not within, the donor template (D).
[0376] In embodiment 47, there is provided herein a composition according to any one of embodiments 42 to 46, wherein the first PAM and the first target nucleotide sequence are oriented 5'P1+T1+D+3'.
[0377] In embodiment 48, there is provided herein a composition according to any one of embodiments 42 to 46, wherein the first PAM and the first target nucleotide sequence are oriented 5'T1-P1-D+3'.
[0378] In embodiment 49, there is provided herein a composition according to any one of embodiments 46 to 48, wherein the second PAM and the second target nucleotide sequence are oriented 5'D+P2+T2+3'.
[0379] In embodiment 50, provided herein is a composition according to any one of embodiments 46 to 48, wherein the second suitable PAM and the second target nucleotide sequence are oriented 5'D+T2-P2-3'.
[0380] In embodiment 51, provided herein is a composition according to embodiment 46, wherein the nucleotide sequences of the first and second PAM targets are oriented 5'T1-P1-D+P2+T2+3'.
[0381] In embodiment 52, provided herein is a composition according to any one of embodiments 42 to 51, wherein the polynucleotide comprises a selectable marker and / or an...
Claims
1. (A) a donor template (D), a double-stranded DNA polynucleotide comprising a first PAM (P1) and a first target nucleotide sequence (T1) that are adjacent to but not within the donor template (D) and recognized by a nucleic acid-inducible nuclease, (B) A polypeptide comprising a nucleic acid-inducible nuclease complex comprising (1) a nucleic acid-inducible nuclease and (2) a guide nucleic acid (gNA), and a nuclear localization signal (NLS) bound to the polynucleotide. A composition containing the following:
2. The composition according to claim 1, wherein the double-stranded DNA polynucleotide comprises a plasmid.
3. The composition according to claim 1, wherein the double-stranded DNA polynucleotide comprises a linear double-stranded DNA polynucleotide having covalently closed ends.
4. The aforementioned double-stranded DNA polynucleotide is adjacent to, but not within, the donor template (D), and is recognized by a nucleic acid-inducible nuclease as a second PAM (P 2 ) and the second target nucleotide sequence (T 2 The composition according to claim 1, further comprising )
5. The first PAM and the first target nucleotide sequence are oriented 5’T 1 - P 1 - D + 3’, and / or the second PAM and the second target nucleotide sequence are oriented 5’D + P 2 + T 2 + 3’, the composition according to claim 1 or 4
6. The first and second PAM target nucleotide sequences are 5'T 1 - P 1 - D + P 2 + T 2 + The composition according to claim 5, which is oriented to 3'.
7. The composition according to claim 1, wherein the nucleic acid-inducible nuclease comprises a V-A type nucleic acid-inducible nuclease.
8. The composition according to claim 7, wherein the nucleic acid-induced nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of MAD2 (SEQ ID NO: 38), MAD7 (SEQ ID NO: 37), ART2 (SEQ ID NO: 2), ART11 (SEQ ID NO: 11), or ART11* (SEQ ID NO: 36).
9. The composition according to claim 8, wherein the nucleic acid-induced nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of SEQ ID NO:
37.
10. The composition according to claim 1, wherein the first PAM (P1) comprises a sequence of CTTN.
11. The second PAM (P 2 The composition according to claim 4, comprising a sequence of CTTN.
12. The composition according to claim 1, wherein the gNA comprises a dual gNA comprising a targeter nucleic acid and a modulator nucleic acid, the targeter nucleic acid and the modulator nucleic acid being separate polynucleotides.
13. The composition according to claim 1, wherein the gNA includes heterogeneous spacer sequences that share complementarity with a first target sequence in the double-stranded DNA polynucleotide and a second target sequence in the human genome, or the gNA includes heterogeneous spacer sequences that do not share complementarity with the target sequence in the human genome.
14. (1) Buffer solution, (2) RNP stabilizer The composition according to claim 1, further comprising at least one of the following.
15. The composition according to claim 14, wherein the buffer solution lacks magnesium, and / or the RNP stabilizer comprises a peptide, poly-L-glutamic acid (PGA), or single-stranded oligodeoxynucleotide (ssODN).
16. The composition according to claim 1, wherein the donor template comprises a first sequence encoding a first polypeptide including a first CAR or a part thereof.
17. The composition according to claim 16, wherein the first CAR or a part thereof is bonded to a bonding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a part thereof.
18. (A) A linear double-stranded DNA donor template (D) having covalently closed ends, (B) A first PAM (P) recognized by a V-type nucleic acid-inducible nuclease, adjacent to but not within the donor template (D) 1 ) and the first target nucleotide sequence (T 1 ) and, by choice, (C) A second PAM (P) that is adjacent to, but not within, the donor template (D), and is recognized by a V-type nucleic acid-inducible nuclease. 2 ) and the second target nucleotide sequence (T 2 )and A composition containing a polynucleotide.
19. The first PAM and the first target nucleotide sequence are 5'T 1 - P 1 - D + The 3'-oriented and / or the second PAM and the second target nucleotide sequence are 5'D + P 2 + T 2 + The composition according to claim 18, which is oriented to 3'.
20. The first and second PAM target nucleotide sequences are 5'T 1 - P 1 - D + P 2 + T 2 + The composition according to claim 19, which is oriented to 3'.
21. A method comprising contacting cells with the composition described in claim 18.