Compositions and methods for improving efficacy of cas9-based knock-in strategies
Patent Information
- Application Number
- JP2025065457
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-07-03
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-29
AI Technical Summary
Existing CRISPR-Cas systems face challenges in efficiently targeting and modifying sequences in eukaryotic cells without off-target effects and in generating sticky ends for seamless integration of sequences.
A non-naturally occurring CRISPR-Cas system, comprising a Cas9 effector protein (stiCas9) capable of generating sticky ends and a guide polynucleotide that forms a complex with stiCas9, specifically hybridizing to eukaryotic sequences, is developed to address these challenges.
The system enables precise, site-specific modification of eukaryotic cell sequences with reduced off-target effects and facilitates seamless integration of nucleotide sequences.
Smart Images

Figure 00000104_0000 
Figure 00000104_0001 
Figure 00000104_0002
Abstract
Description
[Technical Field]
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The ASCII copy was created on November 16, 2018, is named 0098-0002WO1_SL.txt, and is 1,105,014 bytes in size.
[0002] The present disclosure provides a non-naturally occurring CRISPR-Cas system that includes a Cas9 effector protein (stiCas9) that can generate sticky ends, and a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell, and the complex does not occur in nature. [Background technology]
[0003] The clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) system is a prokaryotic immune system first discovered by Ishino in Escherichia coli (E. coli) (Non-Patent Document 1, incorporated herein by reference in its entirety). This immune system provides immunity to viruses and plasmids by sequence-specifically targeting their nucleic acids. See also Non-Patent Document 2, incorporated herein by reference in its entirety. CRISPR-Cas systems have been classified into three major types: Type I, Type II, and Type III. The main characteristics that define the different types are the different cas genes used and the respective proteins they encode. The cas1 and cas2 genes appear to be universal across the three major types, while cas3, cas9, and cas10 appear to be specific to Type I, Type II, and Type III systems, respectively. See, e.g., Non-Patent Document 3, incorporated herein by reference in its entirety.
[0004] There are two main stages involved in this immune system: the first stage is acquisition, and the second stage is interference. The first stage involves cleaving the genome of invading viruses and plasmids and integrating this segment into the CRISPR locus of the organism. The segment integrated into the genome is known as a protospacer and helps protect the organism from subsequent attacks by the same virus or plasmid. The second stage involves attacking the invading virus or plasmid. This stage relies on the protospacer being transcribed into RNA, which, after some processing, hybridizes with a complementary sequence in the DNA of the invading virus or plasmid, while also associating with a protein or protein complex that effectively cleaves the DNA.
[0005] CRISPR RNA processing proceeds differently depending on the bacterial species. For example, in the Type II system first described in the bacterium Streptococcus pyogenes, the transcribed RNA pairs with a transactivating RNA (tracrRNA) and is then cleaved by RNase III to form individual CRISPR-RNAs (crRNAs). The crRNA is further processed after binding by the Cas9 nuclease to produce mature crRNAs. The crRNA / Cas9 complex then binds to DNA containing a sequence complementary to the capture region (called the protospacer). The Cas9 protein then site-specifically cleaves both strands of DNA, forming double-strand breaks (DSBs). This provides a DNA-based "memory" and leads to rapid degradation of viral or plasmid DNA upon repeated exposure and / or infection. Natural CRISPR systems have been comprehensively reviewed (see, for example, Non-Patent Document 3).
[0006] Since its initial discovery, multiple groups have conducted intensive research into the potential applications of CRISPR systems in genetic engineering, e.g., gene editing (Non-Patent Document 4; Non-Patent Document 5; and Non-Patent Document 6; each of which is incorporated herein by reference in its entirety). One major development has been the use of chimeric RNAs to target Cas9 proteins engineered around individual units from CRISPR arrays fused to tracrRNA. This creates a single RNA species called a small guide RNA (gRNA), where sequence modifications in the protospacer region can site-specifically target the Cas9 protein. Considerable work has been done to understand the nature of the base-pairing interactions between highly related chimeric RNAs and target sites, and their tolerance to mismatches, in order to predict and assess off-target effects (see, e.g., Non-Patent Document 7 (including supplementary material), incorporated herein by reference in its entirety).
[0007] The CRISPR-Cas9 gene editing system has been successfully used in a wide range of organisms and cell systems, both to induce DSB formation using wild-type Cas9 protein and to nick single DNA strands using a mutant protein called Cas9n / Cas9 D10A (see, for example, Non-Patent Document 6 and Non-Patent Document 8, each of which is incorporated herein by reference in its entirety). While DSB formation results in the creation of small insertions and deletions (indels) that can disrupt gene function, Cas9n / Cas9 D10A nickase avoids the creation of indels (resulting from repair via non-homologous end joining) while stimulating the endogenous homologous recombination machinery. Therefore, Cas9n / Cas9 D10A nickase can be used to insert DNA regions into genomes with high fidelity.
[0008] In addition to genome editing, the CRISPR system has several other applications, including regulating gene expression, genetic circuit construction, and functional genomics, among others (reviewed in Non-Patent Document 8).
[0009] Various publications are cited herein, the disclosures of which are incorporated by reference in their entireties. [Prior art documents] [Non-patent literature]
[0010] [Non-Patent Document 1] Ishino et al., Journal of Bacteriology 169(12):5429-5433(1987) [Non-patent document 2] Soret et al., “CRISPR-a widespread system that provides acquired resistance against phages in bacteria and archaea”, Nature Reviews Microbiology 6(3):181-186 (2008) [Non-patent document 3] Barrangou and Marraffini, “CRISPR-Cas systems: prokaryotes upgrade to adaptive immunity”, Cell 54(2):234-244 (2014) [Non-patent document 4] Jinek et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity”, Science 337(6096):816-821(2012) [Non-Patent Document 5] Cong et al., “Multiplex genome engineering using CRISPR / Cas systems”, Science 339(6121):819-823(2013) [Non-patent document 6] Mali et al., “RNA-guided human genome engineering via Cas9”, Science 339(6121):823-826(2013) [Non-Patent Document 7] Fu et al., “Improving CRISPR-Cas nucleases using truncated guide RNAs”, Nature Biotechnology 32(3):279-284(2014) [Non-patent document 8] Sander and Joung, “CRISPR-Cas systems for editing, regulating and targeting genomes”, Nature Biotechnology 32(4):347-355(2014) Summary of the Invention [Means for solving the problem]
[0011] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that includes a Cas9 effector protein (stiCas9) capable of generating sticky ends, and a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell, and the complex does not occur in nature.
[0012] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that is capable of generating sticky ends and includes a Cas9 effector protein (stiCas9) that includes a nuclear localization sequence (NLS), and a guide polynucleotide that forms a complex with the stiCas9 and includes a guide sequence, wherein the complex does not occur in nature.
[0013] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that includes one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, and a nucleotide sequence encoding a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell, and the complex does not occur in nature.
[0014] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system comprising: (a) one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends; and (b) a nucleotide sequence encoding a guide polynucleotide that forms a complex with stiCas9 and comprises a guide sequence, wherein the nucleotide sequences of (a) and (b) are under the control of a eukaryotic promoter, and wherein the complex does not occur in nature.
[0015] In some embodiments, the CRISPR-Cas system of the present disclosure further comprises a polynucleotide comprising a tracrRNA sequence. In some embodiments, the guide polynucleotide of the CRISPR-Cas system, the tracrRNA sequence, and the stiCas9 can form a complex, wherein the complex does not occur in nature.
[0016] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system comprising one or more vectors that form a complex with stiCas9 and include a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, and a guide polynucleotide that includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell, wherein the complex does not occur in nature.
[0017] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system comprising one or more vectors that form a complex with the stiCas9 and include a guide polynucleotide that comprises a guide sequence, and a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, where the regulatory element is a eukaryotic regulatory element, and the complex comprises one or more vectors that do not occur in nature.
[0018] In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence. In some embodiments, the non-naturally occurring vector of the present disclosure further comprises a nucleotide sequence comprising a tracrRNA sequence.
[0019] In some embodiments of the CRISPR-Cas system, the complex may cleave at a site within 10 nucleotides of the protospacer adjacent motif (PAM). In some embodiments of the CRISPR-Cas system, the complex may cleave at a site within 5 nucleotides of the protospacer adjacent motif (PAM). In some embodiments of the CRISPR-Cas system, the complex may cleave at a site within 3 nucleotides of the protospacer adjacent motif (PAM).
[0020] In some embodiments of the CRISPR-Cas system, the target sequence is 5' of a protospacer adjacent motif (PAM), and the PAM comprises a 3' G-rich motif. In embodiments of the CRISPR-Cas system, the target sequence is 5' of a protospacer adjacent motif (PAM), and the PAM sequence is NGG (where N is A, C, G, or T).
[0021] In some embodiments of the CRISPR-Cas system, the sticky end comprises a single-stranded polynucleotide overhang of 3 to 40 nucleotides. In some embodiments of the CRISPR-Cas system, the sticky end comprises a single-stranded polynucleotide overhang of 4 to 20 nucleotides. In some embodiments of the CRISPR-Cas system, the sticky end comprises a single-stranded polynucleotide overhang of 5 to 10 nucleotides.
[0022] In some embodiments of the CRISPR-Cas system, stiCas9 is derived from a bacterial species with a Type II-B CRISPR system. In some embodiments of the CRISPR-Cas system, stiCas9 comprises a domain having at least 80% identity, 85% identity, 90% identity, or 95% identity to any of SEQ ID NOs: 10-97 or 192-195. In some embodiments, stiCas9 comprises a domain that matches to the TIGR03031 protein family using an E-value cutoff of 1E-5. In some embodiments, stiCas9 comprises a domain that matches to the TIGR03031 protein family using an E-value cutoff of 1E-10.
[0023] In some embodiments of the CRISPR-Cas system, the bacterial species from which stiCas9 is derived is Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, Ruminobacter sp. RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon 274 strain F0058, Wolinella succinogenes, or the like. succinogenes, Burkholderiales bacterium YL45, Ruminobacter amylophilus, Campylobacter sp. P0111, Campylobacter sp. RM9261, Campylobacter lanienae strain RM8001, Campylobacter lanienae strain P0121, Turicimonas muris, Legionella londiniensis, Salinivibrio sharmensis, Leptospira spp. sp.) isolate FW.030, Moritella sp. isolate NORP46, Endozoicomonas sp.) S-B4-1U, Tamilnaduibacter salinus, Vibrio natriegens, Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis, or Parendozoicomonas haliclonae.
[0024] In some embodiments of the CRISPR-Cas system, the target sequence is 5' of a protospacer adjacent motif (PAM), the PAM sequence is YG (where Y is a pyrimidine), and the stiCas9 is derived from the bacterial species F. novicida.
[0025] In some embodiments of the CRISPR-Cas system, the stiCas9 comprises one or more nuclear localization signals. In some embodiments of the CRISPR-Cas system, the eukaryotic cell is an animal or human cell. In some embodiments of the CRISPR-Cas system, the eukaryotic cell is a human cell. In some embodiments of the CRISPR-Cas system, the eukaryotic cell is a plant cell.
[0026] In some embodiments of the CRISPR-Cas system, the guide sequence is linked to a direct repeat sequence.
[0027] In some embodiments, the delivery particle comprises a CRISPR-Cas system of the present disclosure, hi some embodiments, the stiCas9 and guide polynucleotide are present in a complex within the delivery particle.
[0028] In some embodiments, the guide polynucleotide further comprises a tracrRNA sequence. In some embodiments, the complex within the delivery particle further comprises a polynucleotide comprising a tracrRNA sequence.
[0029] In some embodiments, the delivery particle further comprises a lipid, a sugar, a metal, or a protein.
[0030] In some embodiments, the vesicle comprises a CRISPR-Cas system of the present disclosure.
[0031] In some embodiments, the stiCas9 and guide polynucleotide are present in a complex within a vesicle.
[0032] In some embodiments, the complex within the vesicle further comprises a polynucleotide comprising a tracrRNA sequence. In some embodiments, the vesicle is an exosome or a liposome.
[0033] In some embodiments of the CRISPR-Cas system, one or more nucleotide sequences encoding stiCas9 are codon-optimized for expression in eukaryotic cells.
[0034] In some embodiments of the CRISPR-Cas system, the nucleotides encoding the Cas9 effector protein and the guide polynucleotide are present on a single vector.
[0035] In some embodiments of the CRISPR-Cas system, the nucleotide encoding the Cas9 effector protein and the guide polynucleotide are a single nucleic acid molecule.
[0036] In some embodiments, the viral vector comprises a CRISPR-Cas system of the present disclosure. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus.
[0037] In some embodiments, the disclosure provides a eukaryotic cell comprising a Cas9 effector protein (stiCas9) capable of generating sticky ends, and a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence is capable of hybridizing to a target sequence in the eukaryotic cell, wherein the complex comprises a non-naturally occurring CRISPR-Cas system.
[0038] In some embodiments, the disclosure provides a eukaryotic cell comprising a CRISPR-Cas system comprising a Cas9 effector protein (stiCas9) capable of generating sticky ends (the Cas9 effector protein is derived from a bacterial species having a type II-B CRISPR system).
[0039] In some embodiments, the disclosure provides a method for site-specific modification of a target sequence in a eukaryotic cell, comprising: (1) introducing into the cell (a) a Cas9 effector protein (stiCas9) capable of generating sticky ends; and (b) a guide polynucleotide (the complex is not naturally occurring) that forms a complex with the stiCas9 and includes a guide sequence (the guide sequence can hybridize to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell); (2) generating sticky ends in the target sequence by the Cas9 effector protein and the guide polynucleotide; and (3) ligating (a) the sticky ends together or (b) a polynucleotide sequence of interest (SoI) to the sticky ends, thereby modifying the target sequence.
[0040] In some embodiments, the disclosure provides a method for site-specific modification of a target sequence in a eukaryotic cell, comprising: (1) introducing into the cell (a) a nucleotide sequence encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, and (b) a guide polynucleotide (the complex is not naturally occurring) that forms a complex with stiCas9 and includes a guide sequence (the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell); (2) generating sticky ends in the target sequence by the Cas9 effector protein and the guide polynucleotide; and (3) ligating (a) the sticky ends together or (b) a polynucleotide sequence of interest (SoI) to the sticky ends, thereby modifying the target sequence.
[0041] In some embodiments, the method for providing site-specific modification of a target sequence in a eukaryotic cell further comprises introducing into the cell a polynucleotide comprising a tracrRNA sequence.
[0042] In some embodiments of the method, the guide polynucleotide, the tracrRNA sequence, and the stiCas9 can form a complex, wherein the complex does not occur in nature.
[0043] In some embodiments of the method, the complex may be cleaved at a site within 10 nucleotides of the protospacer adjacent motif (PAM). In some embodiments of the method, the complex may be cleaved at a site within 5 nucleotides of the protospacer adjacent motif (PAM). In some embodiments of the method, the complex may be cleaved at a site within 3 nucleotides of the protospacer adjacent motif (PAM).
[0044] In some embodiments of the method, the target sequence is 5' to a protospacer adjacent motif (PAM), and the PAM comprises a 3' G-rich motif. In some embodiments of the method, the target sequence is 5' to the PAM, and the PAM sequence is NGG (where N is A, C, G, or T).
[0045] In some embodiments of the method, the sticky end comprises a single-stranded polynucleotide overhang of 3 to 40 nucleotides. In some embodiments of the method, the sticky end comprises a single-stranded polynucleotide overhang of 4 to 20 nucleotides. In some embodiments of the method, the sticky end comprises a single-stranded polynucleotide overhang of 5 to 10 nucleotides.
[0046] In some embodiments of the method, the stiCas9 is derived from a bacterial species that has a Type II-B CRISPR system.
[0047] In some embodiments of the method, the eukaryotic cell is an animal or human cell. In some embodiments of the method, the eukaryotic cell is a human cell. In some embodiments of the method, the eukaryotic cell is a plant cell.
[0048] In some embodiments of the method, the modification is a deletion of at least a portion of the target sequence. In some embodiments of the method, the modification is a mutation of the target sequence. In some embodiments of the method, the modification is an insertion of a sequence of interest into the target sequence.
[0049] In some embodiments, the method further comprises introducing an exonuclease to remove overhangs generated from stiCas9.
[0050] In some embodiments of the method, the exonuclease is Cas4, Artemis, or TREX4. In some embodiments of the method, the Cas4 is derived from a bacterial species that has a type II-B CRISPR system.
[0051] In some embodiments of the methods, the polynucleotides encoding the components of the complex are introduced on one or more vectors.
[0052] In some embodiments, the disclosure provides a method for introducing a sequence of interest (SoI) into a chromosome in a cell, the chromosome comprising a target sequence (TSC) comprising region 1 and region 2, and comprising: (a) a vector containing a target sequence (TSV) (the TSV contains region 2, region 1, and SoI); (b) a first Cas9-endonuclease dimer capable of generating sticky ends in the TSC (the first monomer of the first Cas9-endonuclease dimer cleaves in region 1 of the TSC, and the second monomer of the first Cas9-endonuclease dimer cleaves in region 2); and (c) a second Cas9-endonuclease dimer capable of generating sticky ends in the TSV (the first monomer of the second Cas9-endonuclease dimer cleaves in region 2 of the TSV, and the second monomer of the second Cas9-endonuclease dimer cleaves in region 1). Including introducing; The introduction of (a) the vector, (b) the first Cas9-endonuclease dimer, and (c) the second Cas9-endonuclease dimer is directed to a method that results in insertion of the SoI into the chromosome of the cell.
[0053] In some embodiments, the disclosure provides a method for introducing a sequence of interest (SoI) into a chromosome in a cell, the chromosome comprising a target sequence (TSC) comprising region 1 and region 2, and comprising: (a) a vector (the vector contains sticky ends) containing a target sequence (TSV) (the TSV contains region 2 and region 1 and an SoI); (b) a first Cas9-endonuclease dimer capable of generating sticky ends in the TSC (the first monomer of the first Cas9-endonuclease dimer cleaves in region 1 of the TSC, and the second monomer of the first Cas9-endonuclease dimer cleaves in region 2); Including introducing; The introduction of the vector in (a) and the first Cas9-endonuclease dimer in (b) is directed to a method that results in the insertion of the SoI into the chromosome of the cell.
[0054] In some embodiments, the first and second Cas9-endonuclease dimers are the same. In some embodiments, the first and second Cas9-endonuclease dimers are different.
[0055] In some embodiments, the method further comprises introducing into the cell a first guide polynucleotide that forms a complex with the first monomer of the first Cas9-endonuclease dimer and comprises a first guide sequence, where the first guide sequence hybridizes to a TSC comprising region 1 but does not hybridize to the vector.
[0056] In some embodiments, the method further comprises introducing into the cell a first guide polynucleotide that forms a complex with the first monomer of the first Cas9-endonuclease dimer and comprises a first guide sequence, wherein the first guide sequence hybridizes to a TSC and a TSV.
[0057] In some embodiments, the method further comprises introducing into the cell a second guide polynucleotide that forms a complex with a second monomer of the first Cas9-endonuclease dimer and comprises a second guide sequence, where the second guide sequence hybridizes to a TSC comprising region 2 but does not hybridize to the vector.
[0058] In some embodiments, the method further comprises introducing into the cell a second guide polynucleotide that forms a complex with a second monomer of the first Cas9-endonuclease dimer and comprises a second guide sequence, wherein the second guide sequence hybridizes to a TSC and a TSV.
[0059] In some embodiments, the method further comprises introducing into the cell a third guide polynucleotide that forms a complex with the first monomer of the second Cas9-endonuclease dimer and comprises a third guide sequence, wherein the third guide sequence hybridizes to a TSV comprising region 2 but does not hybridize to the chromosome.
[0060] In some embodiments, the method further comprises introducing into the cell a third guide polynucleotide that forms a complex with the first monomer of the second Cas9-endonuclease dimer and comprises a third guide sequence, wherein the third guide sequence hybridizes to TSC and TSV.
[0061] In some embodiments, the method further comprises introducing into the cell a fourth guide polynucleotide that forms a complex with the second monomer of the second Cas9-endonuclease dimer and comprises a fourth guide sequence, wherein the fourth guide sequence hybridizes to a TSV comprising region 1 but does not hybridize to the chromosome.
[0062] In some embodiments, the method further comprises introducing into the cell a fourth guide polynucleotide that forms a complex with the second monomer of the second Cas9-endonuclease dimer and comprises a fourth guide sequence, wherein the fourth guide sequence hybridizes to TSC and TSV.
[0063] In some embodiments, the method includes introducing into the cell a first, a second, a third, and a fourth guide polynucleotide.
[0064] In some embodiments, the method further comprises introducing into the cell a polynucleotide comprising a tracrRNA sequence.
[0065] In some embodiments, the endonuclease in the first and second monomers of the first Cas9-endonuclease dimer is a Type IIS endonuclease, hi some embodiments, the endonuclease in the first and second monomers of the second Cas9-endonuclease dimer is a Type IIS endonuclease.
[0066] In some embodiments, the endonucleases in the first and second Cas9-endonuclease dimers are type IIS endonucleases. In some embodiments, the endonucleases in the first and second Cas9-endonuclease dimers are independently selected from the group consisting of BbvI, BgcI, BfuAI, BmpI, BspMI, CspCI, FokI, MboII, MmeI, NmeAIII, and PleI. In some embodiments, the endonuclease in the first and second Cas9-endonuclease dimers is FokI. In some embodiments, the first and second Cas9-endonuclease dimers are introduced into cells as polynucleotides encoding the first and second Cas9-endonuclease dimers.
[0067] In some embodiments, the polynucleotides encoding the first and second Cas9-endonuclease dimers are present on one vector, hi some embodiments, the polynucleotides encoding the first and second Cas9-endonuclease dimers are present on two or more vectors.
[0068] In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a modified Cas9. In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a catalytically inactive Cas9. In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI. In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a Cas9 with nickase activity. In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI.
[0069] In some embodiments, the Cas9-endonuclease dimer comprises a single amino acid substitution in Cas9 relative to wild-type Cas9. In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI. In some embodiments, the single amino acid substitution is D10A or H840A. In some embodiments, the single amino acid substitution is D10A. In some embodiments, the single amino acid substitution is H840A. In some embodiments, the Cas9-endonuclease dimer comprises a double amino acid substitution relative to wild-type Cas9. In some embodiments, the double amino acid substitutions are D10A and H840A.
[0070] In some embodiments, wild-type Cas9 is effective against Streptococcus pyogenes, Staphylococcus aureus, Staphylococcus pseudintermedius, Planococcus antarcticus, Streptococcus sanguinis, Streptococcus thermophilus, Streptococcus mutans, Coribacterium glomerans, Lactobacillus farciminis, Catenibacterium mitsuokai, Lactobacillus rhamnosus, rhamnosus, Bifidobacterium bifidum, Oenococcus kitahara, Fructobacillus fructosus, Finegoldia magna, Veillonella atyipca, Solobacterium moorei, Acidaminococcus sp. D21, Eubacterium yurri, Coprococcus catus, Fusobacterium nucleatum, Filifactor allosus alocis, Peptoniphilus duerdenii, or Treponema denticola.
[0071] In some embodiments, the sticky ends comprise 5-overhangs. In some embodiments, the sticky ends comprise 3-overhangs. In some embodiments, the first, second, or both Cas9-endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 3 to 40 nucleotides. In some embodiments, the first, second, or both Cas9-endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 4 to 20 nucleotides. In some embodiments, the first, second, or both Cas9-endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 5 to 15 nucleotides.
[0072] In some embodiments of the method, upon insertion, the target sequence in the chromosome and the target sequence in the plasmid are not rearranged.
[0073] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is an animal or human cell. In some embodiments, the cell is a plant cell.
[0074] In some embodiments of the method for introducing a sequence of interest (SoI) into a chromosome in a cell, the vector (a), the first Cas9-endonuclease dimer (b), the second Cas9-endonuclease dimer (c), or a combination thereof is introduced into the cell via a delivery particle, vesicle, or viral vector. In some embodiments, the vector (a), the first Cas9-endonuclease dimer (b), the second Cas9-endonuclease dimer (c), or a combination thereof is introduced into the cell via a delivery particle. In some embodiments, the delivery particle comprises a lipid, a sugar, a metal, or a protein.
[0075] In some embodiments of the method for introducing a sequence of interest (SoI) into a chromosome in a cell, the (a) vector, the (b) first Cas9-endonuclease dimer, the (c) second Cas9-endonuclease dimer, or a combination thereof, is introduced into the cell via a vesicle. In some embodiments, the vesicle is an exosome or a liposome.
[0076] In some embodiments of the method for introducing a sequence of interest (SoI) into a chromosome in a cell, the polynucleotide capable of expressing the vector of (a), the first Cas9-endonuclease dimer of (b), the second Cas9-endonuclease dimer of (c), or a combination thereof is introduced into the cell via a viral vector. In some embodiments, the vector of (a) is a viral vector. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus.
[0077] In some embodiments, the first monomer of the first Cas9-endonuclease dimer forms a complex with the first guide polynucleotide, and the second monomer of the first Cas9-endonuclease dimer forms a complex with the second guide polynucleotide. In some embodiments, the first monomer of the second Cas9-endonuclease dimer forms a complex with the third guide polynucleotide, and the second monomer of the second Cas9-endonuclease dimer forms a complex with the fourth guide polynucleotide. In some embodiments, the first monomer of the first Cas9-endonuclease dimer forms a complex with the first guide polynucleotide sequence and the tracrRNA sequence, and the second monomer of the first Cas9-endonuclease dimer forms a complex with the second guide polynucleotide sequence and the tracrRNA sequence. In some embodiments, a first monomer of the second Cas9-endonuclease dimer forms a complex with a third guide polynucleotide sequence and a tracrRNA sequence, and a second monomer of the second Cas9-endonuclease dimer forms a complex with a fourth guide polynucleotide sequence and a tracrRNA sequence. In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a nuclear localization signal.
[0078] In some embodiments of the method for introducing a sequence of interest (SoI) into a chromosome in a cell, the cell comprises a stem cell or a stem cell line.
[0079] In some embodiments, the present disclosure provides a method of modifying one or more nucleotides in a target polynucleotide sequence in a cell, comprising: (a) A vector containing an insertion cassette (IC) (IC is in the 5'-3' direction, (i) a first region that is homologous to a portion of the target polynucleotide sequence; (ii) a second region comprising a mutation of the target polynucleotide sequence of one or more nucleotides; (iii) a first nuclease binding site; (iv) a polynucleotide sequence encoding a marker gene; (v) a second nuclease binding site; (vi) a third region comprising a mutation of the target polynucleotide sequence of one or more nucleotides; and (vii) introducing a fourth region that is homologous to a portion of the target polynucleotide sequence, wherein the first region and the fourth region are 95% to 100% identical to the target polynucleotide sequence; (b) inserting the IC into the target polynucleotide sequence via homologous recombination to generate a first modified target polynucleotide; (c) selecting cells expressing the marker gene; (d) subjecting the first modified target polynucleotide to a site-specific nuclease to generate a second modified target polynucleotide having a sticky end; and (e) subjecting the second modified target polynucleotide having sticky ends to a ligase, where the ligase ligates the sticky ends at the second region and the third region, to generate a ligated modified target nucleic acid that includes one or more modified nucleotides compared to the target polynucleotide sequence. The present invention covers a method including:
[0080] In some embodiments of the methods of modifying one or more nucleotides in a target polynucleotide sequence in a cell, the first modified target nucleic acid is isolated from the cell after (c).
[0081] In some embodiments, the site-specific nuclease is exogenous to the cell. In some embodiments, the ligase is exogenous to the cell. In some embodiments, the first modified target protein is present in the cell after (c). In some embodiments, the site-specific nuclease is introduced into the cell as a polynucleotide encoding the site-specific nuclease. In some embodiments, the ligase is introduced into the cell as a polynucleotide encoding the ligase.
[0082] In some embodiments, the site-specific nuclease is a recombinant site-specific nuclease. In some embodiments, the ligase is a recombinant ligase. In some embodiments, the site-specific nuclease is a Cas9 effector protein. In some embodiments, the Cas9 effector protein is type II-B Cas9. In some embodiments, the site-specific nuclease is a Cas9-endonuclease fusion protein. In some embodiments, the endonuclease in the Cas9-endonuclease fusion protein is a type IIS endonuclease. In some embodiments, the endonuclease in the Cas9-endonuclease fusion protein is FokI.
[0083] In some embodiments, the Cas9-endonuclease fusion protein comprises a modified Cas9. In some embodiments, the modified Cas9 comprises a catalytically inactive Cas9. In some embodiments, the catalytically inactive Cas9 is fused to a FokI endonuclease.
[0084] In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with nickase activity, and the endonuclease is FokI. In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with a D10A substitution. In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with a H840A substitution.
[0085] In some embodiments, the site-specific nuclease is a Cpf1 effector protein. In some embodiments, the site-specific nuclease is Cas9, Cpf1, or Cas9-FokI.
[0086] In some embodiments of the methods of modifying one or more nucleotides in a target polynucleotide sequence in a cell, the sticky end of the second modified target polynucleotide in (d) comprises a 5' overhang. In some embodiments, the sticky end of the second modified target polynucleotide in (d) comprises a 3' overhang. In some embodiments, the site-specific nuclease can generate a sticky end comprising a single-stranded polynucleotide of 3 to 40 nucleotides. In some embodiments, the nuclease can generate a sticky end comprising a single-stranded polynucleotide of 4 to 20 nucleotides. In some embodiments, the nuclease can generate a sticky end comprising a single-stranded polynucleotide of 5 to 15 nucleotides.
[0087] In some embodiments of the method for modifying one or more nucleotides in a target polynucleotide sequence in a cell, the target polynucleotide sequence is present in a plasmid. In some embodiments, the target polynucleotide sequence is present in a chromosome.
[0088] In some embodiments, the present disclosure is directed to an engineered guide RNA that forms a complex with a stiCas9 protein, the engineered guide RNA comprising: (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell; and (b) a tracrRNA sequence capable of binding to a Cas9 protein, wherein the tracrRNA differs from a naturally occurring tracrRNA sequence by at least 10 nucleotides and improves the nuclease efficiency of the Cas9 protein. In some embodiments, the tracrRNA sequence has at least 10 fewer nucleotides than the naturally occurring tracrRNA. In some embodiments, the tracrRNA sequence has at least 10 more nucleotides than the naturally occurring tracrRNA. In some embodiments, the guide sequence comprises at least 90% sequence identity to any one of SEQ ID NOs: 104-125 or 196-199. In some embodiments, the tracrRNA sequence comprises at least 90% sequence identity to any one of SEQ ID NOs: 148-171. In some embodiments, the guide RNA comprises at least 90% sequence identity to any one of SEQ ID NOs: 172-191.
[0089] In some embodiments, the present disclosure is directed to a CRISPR-Cas system comprising an engineered guide RNA as described herein. In some embodiments, the system does not comprise a tracrRNA sequence.
[0090] In some embodiments, the present disclosure is directed to an engineered Cas9-guide RNA complex comprising any combination of Cas9, a guide sequence, and a tracrRNA sequence found in Figure 40B. In some embodiments, the present disclosure is directed to a method of producing an engineered guide RNA that binds to a Cas9 protein, the method comprising: (a) providing a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell; (b) modifying a naturally occurring tracrRNA sequence by removing at least 10 nucleotides from the tracrRNA sequence to form a modified tracrRNA sequence; and (c) ligating the guide sequence to the modified tracrRNA sequence to generate the engineered guide RNA. In some embodiments, the present disclosure is directed to a non-naturally occurring CRISPR-Cas system that includes: (a) a Cas9 effector protein (stiCas9) that can generate sticky ends; and (b) a guide RNA that forms a complex with the stiCas9 and includes a guide sequence, where the guide sequence can hybridize to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell; where the complex does not occur in nature and does not include a tracrRNA sequence. [Brief explanation of the drawings]
[0091] [Figure 1] Schematic diagram of different repair mechanisms by Cas9. Figure 1a shows gene knockout. Figure 1b shows base editing. Figure 1c shows gene knock-in via the non-homologous end joining (NHEJ) pathway. Figure 1d shows gene knock-in via the homology-directed recombination (HDR) pathway. [Figure 2] Schematic diagram of different mechanisms of gene insertion by Cas9. Homology-directed recombination (HDR) is shown on the left, and non-homologous end joining (NHEJ) is shown on the right. [Figure 3]Schematic diagrams and results for gene insertion using different Cas9 effector proteins. Figures 3a-b show gene insertion mediated by Cas9 generating blunt ends. Figures 3c-d show gene insertion mediated by Cas9 generating overhangs (i.e., "sticky ends"). The bottom panel of Figure 3 shows the frequency of gene insertion by different Cas9 proteins in 3a-3f using Homology-Independent Targeted Insertion (HITI). [Figure 4A-B] This is described by Shmakov et al., Nature Reviews Microbiology 15:169-182 (2017). Figure 4A shows a phylogenetic tree of different types of CRISPR systems and representative bacterial species that harbor each type of CRISPR system. Figure 4B shows a close-up of type II and type V CRISPR systems, with the arrow indicating the operon containing the cas4 gene. [Figure 5A-E] Chylinski et al., Nucleic Acids Research 42(10):6091-6105 (2014). Figures 5A-D show a phylogenetic tree of type II CRISPR systems. Figure 5E shows the distinct signature genes associated with each subfamily of type II CRISPR systems. [Figure 6A-C] Figure 6A shows results obtained for DNA cleavage using Cas9 protein from Francisella novicida. Mutational signatures for genomic loci in engineered HEK293 cell lines targeted by Cas9 from Francisella novicida and Cas9 from Streptococcus pyogenes are compared. Figure 6A discloses SEQ ID NOS: 204-205 and 284, respectively, in order of appearance. Figures 6B-C are phylogenetic trees of the type II CRISPR system. Cas9 proteins selected for in vitro validation are italicized. [Figure 7]1 is a schematic representation of the ObLiGaRe method for gene insertion using zinc finger nucleases (ZFNs) as described in U.S. Pat. No. 9,567,608. [Figure 8] Schematic representation of the Cas9-PiTCH method for gene insertion described by Sakuma et al., Nature Protocols 11(1):118-133 (2016). [Figure 9] 9A-9C are schematic representations of three different Cas9-FokI fusion proteins: Figure 9a: a fusion of enzymatically inactive Cas9 (deadCas9) with FokI; Figure 9b: a fusion of Cas9 with D10A mutation (Cas9nD10A) with FokI; Figure 9c: a fusion of Cas9 with H840A mutation (Cas9nH840A) with FokI. Figures 9a-c disclose SEQ ID NO: 206. [Figure 10] Schematic representation of the different DNA digests produced by different Cas9-FokI fusion proteins in Figures 9 and 10. Figure 10 discloses SEQ ID NO: 206 as "TCCCCTCCACCCCACAGTGGGGCCACTAGGGACAGGATTGGTGACAGAAAAGCCCCATCCTTAGGCCT" and discloses the cleavage sequences as SEQ ID NOs: 285-289, respectively, in order of appearance. [Figure 11] Figure 11 discloses SEQ ID NO: 206. Schematic representation of the cleavage site generated by Cas9nD10A-FokI. [Figure 12] 12 is a schematic representation of the gene insertion method using Cas9nD10A-FokI.gRNA:guide RNA;PAM;protospacer adjacent motif. Figure 12 discloses the "genomic" sequence as SEQ ID NOs:206-208, the "vector" sequence as SEQ ID NOs:209-211, and the "knock-in" sequence as SEQ ID NO:212, all in order of appearance, respectively. [Figure 13] Figure 13 discloses SEQ ID NO: 206. Schematic representation of the cleavage site generated by Cas9nH840A-FokI. [Figure 14]14 is a schematic representation of the gene insertion method using Cas9nH840A-FokI. gRNA:guide RNA; PAM; protospacer adjacent motifs. Figure 14 discloses the "genomic" sequence as SEQ ID NOs:206 and 213-214, the "vector" sequence as SEQ ID NOs:215-217, and the "knock-in" sequence as SEQ ID NO:218, all in order of appearance, respectively. [Figure 15]
[0033] Figure 15 relates to the experiments described in Example 1. Figure 15 is a schematic representation of the gene insertion method using Cas9nD10A-FokI (Figure 15) and Cas9nH840A-FokI (Figure 15). Figures 15a-b disclose SEQ ID NO: 206. [Figure 16] This relates to the experiment described in Example 1. The target site (AAVS1 locus) is represented. "Plan A" refers to the gene insertion method using Cas9nD10A-FokI; "Plan B" refers to the gene insertion method using Cas9nH840A-FokI. Figure 16 discloses SEQ ID NO: 219. [Figure 17] This relates to the experiment described in Example 1. Representative sequences obtained from gene insertion using Cas9nD10A-FokI are shown. Figure 17 discloses SEQ ID NOs: 220-235, respectively, in order of appearance. [Figure 18] This relates to the experiments described in Example 1. Representative sequences obtained from a gene insertion method using Cas9nH840A-FokI are shown. Figure 18 discloses SEQ ID NOs: 236-258, respectively, in order of appearance. [Figure 19] Relates to the experiments described in Example 2. Figure 19 shows the design of the set of 10 guide RNAs (gRNAs) used to target the AAVS1 locus. [Figure 20] 21 relates to the experiment described in Example 2. FIG. 22 is a plasmid map of the "donor" plasmid containing the gene to be inserted into the AAVS1 locus using the gRNA of FIG. [Figure 21]
[0023] Figure 1 relates to the experiment described in Example 2.
[0024] Figure 1 is a schematic diagram of the procedure for selecting cells containing the correctly inserted gene (mCherry+ cells). [Figure 22]1 relates to the experiment described in Example 2. The results of gene insertion frequency using spacers of different lengths are shown. [Figure 23] Relates to the experiment described in Example 3. Figure 23 is a plasmid map of the "donor" plasmid containing the gene to be inserted into the SERPINA1 locus. [Figure 24] 24 is a schematic representation of the gene insertion method using deadCas9-FokI, relating to the experiment described in Example 3. FIG. 24 discloses SEQ ID NO: 206. [Figure 25] 1 is a comparison of the efficiency of various methods used for targeted gene insertion as described in Examples 2-4. [Figure 26] Pertaining to the experiments described in Example 4. Figure 26 is a schematic diagram of seamless mutagenesis. [Figure 27] Relates to the experiment described in Example 4. Schematic representation of the first step of seamless mutagenesis: recombination of a cassette containing a resistance marker into a target sequence using homology arms. [Figure 28] Relates to the experiment described in Example 4. Schematic representation of the cassette integrated into the target sequence: resistance marker flanked on both sides by nuclease binding and cleavage sites. [Figure 29] Relates to the experiment described in Example 4. Schematic representation of the second step of seamless mutagenesis: nuclease digestion and subsequent ligation at the cleavage site (shown in Figure 28) resulting in removal of the resistance marker and seamlessly generated mutations. [Figure 30]Various sequenced species, such as Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, Ruminobacter sp. RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon 274 strain F0058, and Wolinella succinogenes, were sequenced. succinogenes) (SEQ ID NOS: 10-80). [Figure 31]Various bacteria were sequenced, such as Burkholderiales bacterium, Campylobacter sp., Turicimonas muris, Salinivibrio sharmensis, Leptospira sp., Moritella sp., Endozoicomonas sp., Tamilnaduibacter salinus, Vibrio natriegens, Ruminobacter amylophilus, Vibrio sagaiensis, Arcobacter porcinus, and others. porcinus, Desulfofustis sp., and Succinatimonas sp. (SEQ ID NOs: 81-97). [Figure 32] Included are the nucleotide sequences (SEQ ID NOs: 101-103) of the guide RNA, tracrRNA, and crRNA sequences used in the experiments described in Example 8 for the Cas9 protein from MH0245_GL0161830_1. [Figure 33A-B] Figure 33A shows an exemplary 4-nucleotide 5' overhang generated by a Type II-B Cas9 protein. Figure 33A discloses SEQ ID NO: 259. Figure 33B shows an exemplary Type II-B cas operon. The cas9, cas1, cas2, and cas4 genes are represented by arrows. A CRISPR array is labeled downstream of the operon. [Figure 34A-C]This figure relates to the experiment described in Example 7. Figure 34A shows an electrophoresis gel image demonstrating the in vitro nuclease activity of the Cas9 protein from Francisella novicida (FnCas9). Figure 34B shows a Sanger sequencing plot demonstrating that FnCas9 generates sticky ends with 5' overhangs. Figure 34B discloses SEQ ID NOs: 204-205 and 284, respectively, in order of appearance. Figure 34C shows a RIMA comparison of mutation patterns between the Streptococcus pyogenes Cas9 protein (SpyCas9) and FnCas9. [Figure 35A-C]
[0071] Figure 35A shows an electrophoresis gel image demonstrating the in vitro nuclease activity of the Cas9 protein (MHCas9) from the sequenced gut metagenome MH0245. Figure 35B shows a Sanger sequencing plot demonstrating that MHCas9 generates sticky ends with 5' overhangs. Figure 35B discloses SEQ ID NOs: 260-262, respectively, in order of appearance. Figure 35C shows an electrophoresis gel image demonstrating MHCas9 activity in HEK293-REMINDEL cells validated by the Cell 1 assay. [Figure 36A-C] This is related to the experiment described in Example 8. Figure 36A shows the sequences of crRNA and tracrRNA from MHCas9. Figure 36A discloses SEQ ID NO: 263. Figure 36B shows a schematic diagram of the crRNA / tracrRNA secondary structure. Figure 36C shows a truncated phylogenetic tree with Cas9 proteins from Sulfurospirillum sp. SCADC (ssCas9), Wolinella succinogenes (WsCas9), Legionella pneumophila (LpCas9), Francisella novicida (FnCas9), and MH0245 (MHCas9). [Figure 37]
[0023] Figure 1 shows a phylogenetic tree constructed from the amino acid sequences of Cas9 proteins from various bacterial species described herein. Sequence alignment was performed using the MUSCLE algorithm, CLC Genomics Workbench v.9. [Figure 38] Phylogenetic tree constructed from the amino acid sequences of Cas9 proteins from various species of the genus Campylobacter. Sequence alignment was performed using the MUSCLE algorithm, CLC Genomics Workbench v.9. [Figure 39] Included are the nucleotide sequences of crRNAs for the various Cas9 proteins described herein (SEQ ID NOs: 104-147). [Figure 40A] Included are the nucleotide sequences of tracrRNAs for the various Cas9 proteins described herein (SEQ ID NOS: 148-171). [Figure 40B] Various combinations of Cas9 protein, crRNA(+), crRNA(-) and tracrRNA are included. [Figure 41A-T] Various sgRNAs (also referred to as "chimeric gRNAs") designed by the method described in Example 9 are illustrated, including the sequences of the sgRNAs (SEQ ID NOs: 172-191). Figure 41A also discloses a hairpin sequence as SEQ ID NO: 264. [Figure 42A-L]Optimization and trimming of sgRNAs as described in Example 9 are illustrated, as well as potential target sites for further modification. Figure 42A discloses SEQ ID NOs: 265-266, respectively, in order of appearance. Figure 42B discloses SEQ ID NOs: 267-268, respectively, in order of appearance. Figure 42C discloses SEQ ID NOs: 269 and 173, respectively, in order of appearance. Figure 42D discloses SEQ ID NOs: 270-271, respectively, in order of appearance. Figure 42E discloses SEQ ID NOs: 178 and 272, respectively, in order of appearance. Figure 42F discloses SEQ ID NOs: 179 and 273, respectively, in order of appearance. Figure 42G discloses SEQ ID NOs: 180 and 274, respectively, in order of appearance. Figure 42H discloses SEQ ID NOs: 176 and 275, respectively, in order of appearance. Figure 42I discloses SEQ ID NOs: 174 and 276, respectively, in order of appearance. Figure 42J discloses SEQ ID NOs: 191 and 277, respectively, in order of appearance. Figure 42K discloses SEQ ID NOs: 184 and 278, respectively, in order of appearance. Figure 42L discloses SEQ ID NOs: 279-280, respectively, in order of appearance. [Figure 43]
[0043] Figure 43 illustrates a bidirectional expression construct for the Type II-B CRISPR-Cas system. As shown in the inset, the top strand expresses a crRNA and spacer for a single guide RNA without tracrRNA. The bottom strand expresses a crRNA and spacer for a dual guide RNA with tracrRNA. Figure 43 discloses SEQ ID NOs: 137, 281, and 191, respectively, in order of appearance. [Figure 44] Figure 44 shows the predicted secondary structure of the single guide RNA scaffold for the Cas9 protein described herein. Figure 44 discloses SEQ ID NOs: 137, 139, 282, 122, 110, 129, 120, 124, and 104, respectively, in order of appearance. [Figure 45] Four different engineered RNAs and their respective cleavage efficiencies by MHCas9 are generally described. [Figure 46] We demonstrate the cleavage efficiency and functionality of guide RNAs of lengths 19, 20, 21, 22, and 23 by three different Cas9 systems: SpyCas9, Cl1Cas9, and MHCas9. [Figure 47]Included are the amino acid sequences of Cas9 proteins from various sequenced bacteria, such as Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis, and Parendozoicomonas haliclonae (SEQ ID NOs: 192-195). [Figure 48] Included are the nucleotide sequences of crRNAs for the various Cas9 proteins described herein (SEQ ID NOs: 196-203). [Figure 49A-B]
[0049] Referring to Example 11, Figure 49A shows an exemplary method for determining the PAM sequence of a Cas9 protein. Figure 49A discloses SEQ ID NO: 283. Figure 49B shows preferred PAM sequences for SpCas9 (top) and MHCas9 (bottom) determined by the method shown in Figure 49A. [Figure 50A-B] Refers to Example 12. Figure 50A shows a schematic of Cas9 cleavage resulting in accurate repair. Figure 50B shows a schematic of Cas9 cleavage combined with end processing by exonucleases, such as TREX2 or Artemis, resulting in imprecise repair and increased modification. [Figure 51A-B]
[00130] Figure 51A shows an overview of a method for testing the effect of adding an end-processing enzyme (FnCas4 or TREX2) to various Cas9s (SpCas9, FnCas9, Cl1Cas9, or MHCas9) using three different guide RNAs. Figure 51B shows the results for each of the Cas9 proteins using either a mock end-processing enzyme, FnCas4, or TREX2, and each of the three guide RNAs. [Figure 52A-C]Refers to Example 13. Figures 52A, 52B, and 52C show different types of mutations generated by SpCas9, Cl1Cas9, or MHCas9, respectively, when all three Cas9 proteins cleave at the same sequence. Figures 52A-C disclose SEQ ID NO: 290. [Figure 53A-B]
[00130] Figure 53A shows a schematic of the RuvC and HNH domains of a Type II-A Cas9 protein cleaving a double-stranded DNA sequence complexed with a guide RNA, generating blunt or single-nucleotide overhangs. Figure 53B shows a schematic of the RuvC and HNH domains of a Type II-B Cas9 protein cleaving a double-stranded DNA sequence complexed with a guide RNA, generating cohesive ends with 3- or 4-nucleotide overhangs. DETAILED DESCRIPTION OF THE INVENTION
[0092] CRISPR-Cas9 system is widely used in gene editing due to its ability to form targeted double-strand breaks.It is known that Cas9 protein generates blunt ends when cleaving, which provides lower specificity for the insertion and / or modification of target sequence compared with sticky ends.Cas9 protein that can generate sticky ends is also called stiCas9 and is described herein.The advantages of using stiCas9 protein for the insertion and / or modification of target sequence are described herein.
[0093] The present disclosure provides non-naturally occurring CRISPR-Cas systems; eukaryotic cells comprising CRISPR-Cas systems; methods for providing site-specific modification of a target sequence; methods for introducing a sequence of interest into a chromosome in a cell; and methods for modifying one or more nucleotides in a target polynucleotide sequence in a cell.
[0094] definition As used herein, "a" or "an" may mean one or more. As used herein in the specification and claims, and when used in conjunction with the word "comprising," the words "a" or "an" may mean one or more. As used herein, "another" or "further" may mean at least a second or more.
[0095] Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error for the method / device being employed to determine the value or the variation that exists between test subjects. Typically, the term is meant to encompass approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, or 20%, depending on the context.
[0096] Although the use of the term "or" in the claims is used to mean "and / or" unless expressly stated to refer to alternatives only or unless the alternatives are mutually exclusive, the present disclosure supports a definition that refers to alternatives only and "and / or."
[0097] As used in the specification and claims, the words "comprising" (and any form of comprising, e.g., "comprise" and "comprises"), "having" (and any form of having, e.g., "have" and "has"), "including" (and any form of including, e.g., "includes" and "include"), or "containing" (and any form of containing, e.g., "contains" and "contain") are inclusive or open-ended, and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed herein can be implemented with respect to any method, system, host cell, expression vector, and / or composition of the disclosure. Furthermore, the methods and proteins of the disclosure can be achieved using the compositions, systems, host cells, and / or vectors of the disclosure.
[0098] The use of the term "for example" and its corresponding abbreviation "eg" (whether italicized or not) means that the cited defined term is representative of and embodiments of the present disclosure and is not limited to the specific examples referenced or cited, unless expressly stated otherwise.
[0099] "Nucleic acid," "nucleic acid molecule," "nucleotide," "nucleotide sequence," "oligonucleotide," or "polynucleotide" refers to a polymeric compound comprising covalently linked nucleotides. The term "nucleic acid" includes ribonucleic acid (RNA) and deoxyribonucleic acid (DNA), both of which may be single-stranded or double-stranded. DNA includes, but is not limited to, complementary DNA (cDNA), genomic DNA, plasmid or vector DNA, and synthetic DNA. In some embodiments, the present disclosure provides a polynucleotide encoding any one of the polypeptides disclosed herein, e.g., directed to a polynucleotide encoding a Cas protein or a variant thereof.
[0100] "Gene" refers to a collection of nucleotides that encodes a polypeptide, and includes cDNA and genomic DNA nucleic acid molecules. "Gene" also refers to a nucleic acid fragment that can act as regulatory sequences preceding (5' non-coding sequences) and following (3' non-coding sequences) the coding sequence.
[0101] A nucleic acid molecule is "hybridizable" or "hybridized" to another nucleic acid molecule, e.g., cDNA, genomic DNA, or RNA, if the single-stranded form of the nucleic acid molecule can anneal to the other nucleic acid molecule under appropriate conditions of temperature and solution ionic strength. Hybridization and washing conditions are well known and are exemplified in Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 thereof (incorporated herein by reference in its entirety). The conditions of temperature and ionic strength determine the "stringency" of hybridization. Stringency conditions can be adjusted to screen moderately similar fragments, e.g., homologous sequences from distantly related organisms, against highly similar fragments, e.g., genes replicating functional enzymes from closely related organisms. For preliminary screening of homologous nucleic acids, a T of 55°C is used. m Low stringency hybridization conditions corresponding to a higher T can be used, e.g., 5xSSC, 0.1% SDS, 0.25% milk, and no formamide; or 30% formamide, 5xSSC, 0.5% SDS. Moderate stringency hybridization conditions correspond to a higher T m For example, high stringency hybridization conditions correspond to 40% formamide and 5x or 6x SCC. m For example, 50% formamide corresponds to 5x or 6x SCC. Hybridization requires that two nucleic acids contain complementary sequences, although depending on the stringency of hybridization, mismatches between bases are possible.
[0102] The term "complementary" is used to describe the relationship between nucleotide bases that are capable of hybridizing to one another. For example, with respect to DNA, adenosine is complementary to thymine and cytosine is complementary to guanine. Thus, the present disclosure also includes isolated nucleic acid fragments that are complementary to the complete sequences disclosed or used herein and nucleic acid sequences substantially similar thereto.
[0103] A DNA "coding sequence" is a double-stranded DNA sequence that is transcribed and translated into a polypeptide in a cell in vitro or in vivo when placed under the control of appropriate regulatory sequences. A "suitable regulatory sequence" refers to a nucleotide sequence located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence that influences the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include promoters, translation leader sequences, introns, polyadenylation recognition sequences, RNA processing sites, effector binding sites, and stem-loop structures. The boundaries of a coding sequence are determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxyl) terminus. Coding sequences can include, but are not limited to, prokaryotic sequences, cDNA from mRNA, genomic DNA sequences, and even synthetic DNA sequences. If a coding sequence is intended for expression in a eukaryotic cell, a polyadenylation signal and transcription termination sequence are usually located 3' to the coding sequence.
[0104] "Open reading frame", abbreviated as ORF, means a nucleic acid sequence, either DNA, cDNA or RNA, of a length that includes a translation initiation signal or start codon, e.g., ATG or AUG, and a stop codon, and that can potentially be translated into a polypeptide sequence.
[0105] The term "homologous recombination" refers to the insertion of a foreign DNA sequence into another DNA molecule, for example, the insertion of a vector into a chromosome. Preferably, the vector targets a specific chromosomal site for homologous recombination. For specific homologous recombination, the vector contains a sufficiently long region of homology with the chromosomal sequence to allow complementary binding and integration of the vector into the chromosome. A longer region of complementarity and a greater degree of sequence similarity can increase the efficiency of homologous recombination.
[0106] Polynucleotides according to the present disclosure can be propagated using methods known in the art. Once a suitable host system and growth conditions are established, recombinant expression vectors can be propagated and prepared in large quantities. Expression vectors described herein that can be used include, but are not limited to, the following vectors or their derivatives: human or animal viruses, such as vaccinia virus or adenovirus; insect viruses, such as baculovirus; yeast vectors; bacteriophage vectors (e.g., lambda), and plasmid and cosmid DNA vectors.
[0107] As used herein, the terms "promoter," "promoter sequence," or "promoter region" refer to a DNA regulatory region / sequence that can bind RNA polymerase and participate in initiating transcription of a downstream coding or non-coding sequence. In some examples of the present disclosure, the promoter sequence includes a transcription initiation site and extends upstream to include a minimum number of bases or elements used to initiate transcription at a level detectable above background. In some embodiments, the promoter sequence includes a transcription initiation site and a protein binding domain responsible for binding RNA polymerase. Eukaryotic promoters often, but not always, contain "TATA" boxes and "CAT" boxes. Various promoters, including inducible promoters, can be used to drive the various vectors of the present disclosure.
[0108] A "vector" is any means for cloning and / or transferring a nucleic acid into a host cell. A vector can be a replicon to which another DNA segment can be attached, resulting in replication of the attached segment. A "replicon" is any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of DNA replication in vivo, i.e., capable of replication under its own control. In some embodiments of the present disclosure, the vector is an episomal vector that is eliminated / disappeared from a population of cells after many cell generations, e.g., by asymmetric partitioning. The term "vector" includes both viral and non-viral means for introducing nucleic acid into cells in vitro, ex vivo, or in vivo. Numerous vectors known in the art can be used to manipulate nucleic acids, such as to incorporate response elements and promoters into genes. Possible vectors include, for example, plasmids or modified viruses, such as bacteriophages, e.g., lambda derivatives, or plasmids, e.g., PBR322 or pUC plasmid derivatives, or Bluescript vectors. For example, insertion of a DNA fragment corresponding to the response element and promoter into a suitable vector can be accomplished by ligating the appropriate DNA fragment into a selection vector which has complementary cohesive termini. Alternatively, the ends of the DNA molecules can be enzymatically modified, or any site can be produced by ligating nucleotide sequences (linkers) into the DNA termini. Such vectors can be engineered to contain a selectable marker gene that provides for selection of cells which have incorporated the marker into their genome. Such markers allow identification and / or selection of host cells that have incorporated the marker and express the protein encoded by the marker.
[0109] Viral vectors, particularly retroviral vectors, have been used in a wide range of gene delivery applications in cells and live animal subjects. Viral vectors that can be used include, but are not limited to, retrovirus, adeno-associated virus, pox, baculovirus, vaccinia, herpes simplex, Epstein-Barr, adenovirus, geminivirus, and caulimovirus vectors. Non-viral vectors include, but are not limited to, plasmids, liposomes, charged lipids (cytofectin), DNA-protein complexes, and biopolymers. In addition to the nucleic acid, the vector may also contain one or more regulatory regions and / or selectable markers useful for selecting, measuring, and monitoring the results of nucleic acid transfer (tissues of transfer, duration of expression, etc.).
[0110] Vectors can be introduced into desired host cells by well-known methods, including, but not limited to, transfection, transduction, cell fusion, and lipofection. Vectors can include various regulatory elements, such as promoters. In some embodiments, vector design can be based on constructs designed by Mali et al., "Cas9 as a versatile tool for engineering biology," Nature Methods 10:957-63 (2013). In some embodiments, the present disclosure provides expression vectors comprising any of the polynucleotides described herein, e.g., expression vectors comprising a polynucleotide encoding a Cas protein or a variant thereof. In some embodiments, the present disclosure provides expression vectors comprising a polynucleotide encoding a Cas9 protein or a variant thereof.
[0111] The term "plasmid" refers to an extrachromosomal element, often carrying genes that are not part of the cell's central metabolism, usually in the form of a circular double-stranded DNA molecule. Such elements can be autonomously replicating sequences, genome-integrating sequences, phage, or nucleotide sequences from any source, linear, circular, or supercoiled, single- or double-stranded DNA or RNA, in which multiple nucleotide sequences are combined or recombined into unique constructs capable of introducing into a cell promoter fragments and DNA sequences for selected gene products, along with appropriate 3' untranslated sequences.
[0112] As used herein, "transfection" refers to the introduction of an exogenous nucleic acid molecule, e.g., a vector, into a cell. A "transfected" cell contains an exogenous nucleic acid molecule inside the cell, and a "transformed" cell is a cell in which the exogenous nucleic acid molecule induces a phenotypic change in the cell. The transfected nucleic acid molecule can be integrated into the host cell's genomic DNA and / or maintained extrachromosomally by the cell for either a transient or long-term period. Host cells or organisms that express an exogenous nucleic acid molecule or fragment are referred to as "recombinant," "transformed," or "transgenic" organisms. In some embodiments, the present disclosure provides host cells comprising any of the expression vectors described herein, e.g., an expression vector comprising a polynucleotide encoding a Cas protein or variant thereof. In some embodiments, the present disclosure provides host cells comprising an expression vector comprising a polynucleotide encoding a Cas9 protein or variant thereof.
[0113] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein to refer to polymeric forms of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones.
[0114] As used herein, "amino acid" refers to a compound containing both a carboxyl (-COOH) and an amino (-NH2) group. "Amino acid" refers to both natural and unnatural, i.e., synthetic, amino acids. Natural amino acids with three-letter and one-letter abbreviations include alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V).
[0115] An "amino acid substitution" refers to a polypeptide or protein that contains one or more substitutions of a wild-type or naturally occurring amino acid at that amino acid residue with an amino acid that is different from the wild-type or naturally occurring amino acid. The substituted amino acid may be a synthetic or naturally occurring amino acid. In some embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. Substitution mutants can be described using an abbreviation system. For example, a substitution mutation in which the fifth amino acid residue has been substituted can be abbreviated as "X5Y," where "X" is the wild-type or naturally occurring amino acid to be replaced, "5" is the amino acid residue position within the amino acid sequence of the protein or polypeptide, and "Y" is the substitution, or non-wild-type or non-naturally occurring amino acid.
[0116] An "isolated" polypeptide, protein, peptide, or nucleic acid is a molecule that has been removed from its natural environment. It should also be understood that an "isolated" polypeptide, protein, peptide, or nucleic acid can be formulated with an excipient, e.g., a diluent or auxiliary agent, and still be considered isolated.
[0117] The term "recombinant," when used in reference to a nucleic acid molecule, peptide, polypeptide, or protein, means one that is of, or results from, a new combination of genetic material not known to exist in nature. Recombinant molecules can be produced by any of the well-known techniques available in the field of recombinant technology, including, but not limited to, polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid-phase synthesis of nucleic acid molecules, peptides, or proteins.
[0118] The term "domain," when used with reference to a polypeptide or protein, refers to a distinct functional and / or structural unit within a protein. A domain may be responsible for a specific function or interaction that contributes to the overall role of the protein. Domains can exist in a variety of biological contexts. Similar domains can be found in proteins with different functions. Alternatively, domains with low sequence identity (i.e., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5%, or less than about 1% sequence identity) may have the same function. In some embodiments, the Cas9 domain is matched to the TIGR03031 protein family using an E-value cutoff of 1E-5. In some embodiments, the Cas9 domain is matched to the TIGR03031 protein family using an E-value cutoff of 1E-10. In some embodiments, the Cas9 domain is a RuvC domain. In some embodiments, the Cas9 domain is an HNH domain.
[0119] As used herein, the term "sequence similarity" or "% similarity" refers to the degree of identity or correspondence between nucleic acid or amino acid sequences. As used herein, "sequence similarity" refers to nucleic acid sequences in which changes in one or more nucleotide bases result in the substitution of one or more amino acids but do not affect the functional properties of the protein encoded by the DNA sequence. "Sequence similarity" also refers to modifications of nucleic acids that do not substantially affect the functional properties of the resulting transcript, such as the deletion or insertion of one or more nucleotide bases. Thus, it is understood that the present disclosure encompasses more than the specified exemplary sequences. Each of the proposed modifications is well within the routine skill in the art, as is determining retention of biological activity of the encoded product.
[0120] Furthermore, those skilled in the art will recognize that similar sequences encompassed by the present disclosure are also defined by their ability to hybridize to the sequences exemplified herein under stringent conditions. Similar nucleic acid sequences of the present disclosure are nucleic acids whose DNA sequences are at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% identical to the DNA sequences of the nucleic acids disclosed herein. Similar nucleic acid sequences of the present disclosure are nucleic acids whose DNA sequences are about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 99%, at least about 99%, or about 100% identical to the DNA sequences of the nucleic acids disclosed herein.
[0121] As used herein, "sequence similarity" refers to two or more amino acid sequences in which more than about 40% of the amino acids are identical or in which more than about 60% of the amino acids are functionally identical. Functionally identical or functionally similar amino acids have chemically similar side chains. For example, amino acids can be grouped according to functional similarity as follows: Positively charged side chains: Arg, His, Lys; Negatively charged side chains: Asp, Glu; Polar uncharged side chains: Ser, Thr, Asn, Gln; Hydrophobic side chains: Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp; Others: Cys, Gly, Pro.
[0122] In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 99% identical.
[0123] In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% functionally identical. In some embodiments, similar amino acid sequences of the present disclosure have amino acids that are about 40%, at least about 40%, at least about 45%, at least about 45%, about 50%, at least about 50%, about 55%, at least about 55%, about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% identical.
[0124] In some embodiments, similar amino acid sequences of the disclosure have about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% functionally identical amino acids.
[0125] Sequence similarity is determined by sequence alignment using routine methods in the art, such as BLAST, MUSCLE, Clustal (e.g., ClustalW and ClustalX), and T-Coffee (including variants such as M-Coffee, R-Coffee, and Expresso).
[0126] The term "sequence identity" or "% identity" with respect to nucleic acid or amino acid sequences refers to the percentage of residues in the compared sequences that are identical when the sequences are aligned over a specified comparison window. In some embodiments, sequence identity is determined by aligning only a specified portion of two or more sequences. In some embodiments, sequence similarity is determined by aligning only a specified domain of two or more sequences. The comparison window can be a segment of at least 10 to over 1000 residues, at least 20 to about 1000 residues, or at least 50 to 500 residues, over which sequences can be aligned and compared. Alignment methods for determining sequence identity are well known and can be performed using publicly available databases, e.g., BLAST. "Percent identity" or "% identity," when referring to amino acid sequences, can be determined by methods known in the art. For example, in some embodiments, the "percent identity" of two amino acid sequences is determined using the algorithm of Karlin and Altschul, Proceedings of the National Academy of Sciences USA 87:2264-2268 (1990), as modified as in Karlin and Altschul, Proceedings of the National Academy of Sciences USA 90:5873-5877 (1993). Such an algorithm is incorporated into BLAST programs, e.g., the BLAST+ or NBLAST and XBLAST programs described in Altschul et al., Journal of Molecular Biology, 215:403-410 (1990). BLAST protein searches can be performed using a program, e.g., the XBLAST program, score=50, wordlength=3, etc., to obtain amino acid sequences homologous to the protein molecules of the present disclosure. If gaps exist between the two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Research 25(17):3389-3402 (1997).When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (eg, XBLAST and NBLAST) can be used.
[0127] In some embodiments, a polypeptide or nucleic acid molecule has 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, or at least 99% or 100% sequence identity to a reference polypeptide or nucleic acid molecule, respectively (or a fragment of the reference polypeptide or nucleic acid molecule). In some embodiments, a polypeptide or nucleic acid molecule has about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% sequence identity to a reference polypeptide or nucleic acid molecule, respectively (or a fragment of the reference polypeptide or nucleic acid molecule).
[0128] CRISPR-Cas system In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that includes: (a) a Cas9 effector protein capable of generating sticky ends ("sticky-end Cas9" or "stiCas9"); and (b) a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell; the complex does not exist in nature.
[0129] Generally, CRISPR or CRISPR-Cas systems are characterized by elements (also referred to as protospacers in endogenous CRISPR systems) that promote the formation of a CRISPR complex at the site of the target sequence. With respect to the formation of a CRISPR complex, a "target sequence" refers to a sequence that a guide polynucleotide is designed to target, e.g., to have complementarity, and hybridization between the target sequence and the guide polynucleotide promotes the formation of a CRISPR complex. A section of a guide polynucleotide whose complementarity to a target sequence may be important for cleavage activity is referred to herein as a guide sequence. A target sequence may comprise any polynucleotide, e.g., a DNA or RNA polynucleotide, and may be located within a target locus of interest. In some embodiments, the target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, the target sequence is located on a chromosome (TSC). In some embodiments, the target sequence is located on a vector (TSV).
[0130] The Cas proteins described herein are components of CRISPR-Cas systems that can be used, inter alia, for genome editing, gene regulation, genetic circuit construction, and functional genomics. The Cas1 and Cas2 proteins are believed to be universal to all currently identified CRISPR systems, while the Cas3, Cas9, and Cas10 proteins are believed to be specific to Type I, Type II, and Type III CRISPR systems, respectively.
[0131] Since the initial publication of the CRISPR-Cas9 system (a type II system), Cas9 variants have been identified in a range of bacterial species, and a number have been functionally characterized. See, e.g., Chylinski et al., "Classification and evolution of type II CRISPR-Cas systems," Nucleic Acids Research 42(10):6091-6105 (2014), Ran et al., "In vivo genome editing using Staphylococcus aureus Cas9," Nature 520(7546):186-91 (2015), and Esvelt et al., "Orthogonal Cas9 proteins for RNA-guided gene regulation and editing," Nature Methods 10(11):1116-1121 (2013), each of which is incorporated herein by reference in its entirety.
[0132] The present disclosure encompasses novel effector proteins for type II CRISPR-Cas systems, of which Cas9 is an exemplary effector protein. Thus, the terms "Cas9," "Cas9 protein," and "Cas9 effector protein" are used interchangeably herein to describe effector proteins that can provide sticky ends when used in a CRISPR-Cas9 system. In some embodiments, the term Cas9 refers to type II-B Cas9. In some embodiments, the term Cas9 also refers to engineered Cas9 variants, such as deadCas9-FokI, Cas9n, and Cas9n. D10A -FokI and Cas9n H840A -Refers to FokI, etc.
[0133] In some embodiments, the Cas9 effector protein is functional in prokaryotic or eukaryotic cells for in vitro, in vivo, or ex vivo applications.
[0134] The term Cas9 effector protein generally refers to an effector protein with Cas9-like function that has both RuvC and HNH nuclease domains. In some embodiments, the RuvC and HNH domains of a Cas9 effector protein each cleave one strand of a double-stranded target DNA. Thus, for example, when the RuvC and HNH domains cleave their respective strands at the same position, the result of cleavage is a double-stranded target DNA with blunt ends. When the RuvC and HNH domains cleave their respective strands at different positions (i.e., cleave at an "offset"), the result of cleavage is a double-stranded target DNA with an overhang. In some embodiments, the RuvC and HNH domains of a stiCas9 protein cleave at a 3-nucleotide offset. In embodiments, the RuvC and HNH domains of a stiCas9 protein cleave at a 4-nucleotide offset. In embodiments, the RuvC and HNH domains of a stiCas9 protein cleave at a 5-nucleotide offset. In embodiments, the RuvC and HNH domains of the stiCas9 protein cleave at an offset of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides.
[0135] In some embodiments, the term Cas9 effector protein refers to a Cas9 having a RuvC domain and an HNH domain, where the RuvC domain and the HNH domain cleave at different positions on each strand of a double-stranded target DNA. In some embodiments, the RuvC domain of the Cas9 effector protein cleaves one strand of the double-stranded target DNA (e.g., which may be referred to as the "non-target strand") at about -10, -9, -8, -7, or -6 nucleotides from the PAM, and the HNH domain of the Cas9 effector protein cleaves the other strand of the double-stranded target DNA (e.g., which may be referred to as the "target strand") at about -5, -4, -3, -2, or -1 nucleotides from the PAM.
[0136] In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -8 nucleotide from the PAM. In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -7 nucleotide from the PAM. In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -6 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -4 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -3 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -2 nucleotide from the PAM.
[0137] In some embodiments, the term Cas9 effector protein refers to Cas9, which has the TIGR03031 protein family, as identified by HMMER searches, specifically the program hmmscan (HMMER version 3.1b2). The present disclosure also relates to the identification and engineering of effector proteins associated with type II CRISPR-Cas systems. In some embodiments, the effector protein comprises a single-subunit effector module. In some embodiments, wild-type Cas9 effectors or engineered versions of Cas9 proteins are fused to one or more functional domains, such as a nuclear localization signal (NLS) and a FokI nuclease. The present disclosure encompasses computational methods and algorithms for predicting and identifying components of novel type II-B CRISPR-Cas systems.
[0138] In some embodiments, computational methods for identifying novel type II-B CRISPR-Cas loci include those described below and previously described in Shmakov et al., Nature Reviews Microbiology 15, 169-182 (2017). The presence and location of CRISPR-Cas loci in a given nucleotide sequence can be identified, for example, by using the protein sequence of one of the known Cas proteins, e.g., Cas1, as a seed in TBLASTN against the nucleotide sequence using an E-value cutoff of 0.01. Another approach to identifying the presence and location of CRISPR-Cas loci is to search for CRISPR arrays in the nucleotide sequence using a program such as CRISPRfinder or PILER-CR using default parameters. Once a CRISPR-Cas locus is identified, sequences containing up to 10 kbp upstream and downstream of the CRISPR-Cas locus can be extracted. The presence of genes in the extracted nucleotide sequence can be determined using software such as GeneMark or MetaGeneMark using default parameters. The identified genes are then translated into protein sequences and annotated to show their predicted functions using homology searches against databases of proteins with known functions (i.e., Cas1, Cas2, Cas4, Cas9, etc.), e.g., RPS-BLAST, BLAST, or HMMR.
[0139] CRISPR-Cas loci identified using the above method were investigated for the presence of both Cas9 and Cas4 proteins in the same CRISPR-Cas locus, as they are highly likely to contain type IIB Cas9. To further increase the probability of type IIB Cas9, Cas9 proteins were searched for membership in the TIGRFAM:TIGR03031 family using hmmscan.
[0140] In some embodiments, a method for identifying a novel Type II-B CRISPR-Cas locus includes identifying a Cas9 protein in the same locus as a Cas4 protein. In some embodiments, a method for identifying a novel Type II-B CRISPR-Cas locus includes translating a publicly available metagenomic gene catalog into amino acid sequences and scanning each amino acid sequence with a TIGR03031 protein family profile to identify matches above a predetermined cutoff E-value, such as 1E-5 to 1E-10.
[0141] TIGRFAM is a collection of protein families featuring curated multiple sequence alignments, hidden Markov models, and associated information designed to aid in automated protein functional identification through sequence homology. A hidden Markov model (HMM) applied to sequence alignments refers to a statistical model for a continuous column of a protein multiple sequence alignment. Typically, a protein profile HMM is developed from the curated multiple sequence alignment using position-based scoring for each amino acid, insertions, and deletions across the length of the sequence. Scores are reported in bits of information and as E-values. An E-value below a "confidence cutoff" or "confidence threshold," e.g., 0.001, is recognized as a positive "hit" or positive identification. Therefore, sequences identified using a low E-value cutoff are more likely to belong to a specific protein family. In some embodiments, the E-value cutoff is 1E-10. In some embodiments, the E-value cutoff is 1E-5. In some embodiments, the confidence cutoff E-value is at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1.
[0142] In some embodiments, identification of all predicted protein-coding genes is performed by comparing the identified genes to Cas protein-specific profiles and annotating them according to the NCBI Conserved Domain Database (CDD), a protein annotation resource consisting of a collection of fully annotated multiple sequence alignment models of ancient domains and full-length proteins. These are available as position-specific score matrices (PSSMs) for rapid identification of conserved domains in protein sequences via RPS-BLAST. CDD content includes NCBI-curated domains, which use 3D structural information to clearly define domain boundaries and provide insights into sequence / structure / function relationships, as well as domain models imported from a number of external resource databases (Pfam, SMART, COG, PRK, TIGRFAM). Protein databases are described, for example, in Finn et al., Nucleic Acids Research Database Issue 44:D279-D285(2016); Letunic et al., Nucleic Acids Research, doi:gkx922(2017); Tatusov et al., Science 278(5338):631-637(1997); and Haft et al., Nucleic Acids Research Database Issue 41:D387-D395(2013), each of which is incorporated herein in its entirety.
[0143] In some embodiments, novel type II-B CRISPR-Cas loci are identified using HMMER (or any version of HMMER, e.g., HMMER2 or HMMER3) to search for conserved domains. HMMER is a free, commonly used software package for sequence analysis, identification of homologous protein or nucleotide sequences, and sequence alignment. HMMER implements a probabilistic model called a profile hidden Markov model. HMMER can be used with profile databases, such as Pfam, SMART, COG, PRK, or TIGRFAM. HMMER can also be used with query sequences, for example, searching a protein query sequence against a database (i.e., phmmer) or using an iterative search (i.e., jackhmmer). In some embodiments, novel type II-B CRISPR-Cas loci are identified by searching for the presence of specific domains in specific protein families. In some embodiments, the TIGRFAM protein family is TIGRFAM:TIGR03031. In some embodiments, a specific domain matches the TIGR03031 protein family using an E-value cutoff of at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1. In some embodiments, a specific domain has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to any of the TIGR03031 domains identified herein. In some embodiments, the specific domain has at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to any one of SEQ ID NOs: 10-97 or 192-195.In some embodiments, the specific domain has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence identity to any one of SEQ ID NOs: 10-97 or 192-195.
[0144] In some embodiments, stiCas9 is derived from a bacterial species with a type II-B CRISPR system. In some embodiments, the type II-B CRISPR system comprises a cas4 gene. The CRISPR systems discussed herein are classified as type I, type II, and type III. All type II CRISPR systems comprise the cas1, cas2, and cas9 genes on a cas operon. Type II CRISPR systems are further categorized as type II-A, type II-B, and type II-C. In some embodiments, type II-B CRISPR systems are identified by the presence of the cas4 gene on the cas operon. The cas4 gene is not found in either type II-A or type II-C CRISPR systems.
[0145] Type II CRISPR systems can also be classified according to the sequences and / or domains of individual cas genes, such as the sequence and / or domain of cas9. Protein domains can be identified by conserved sequences or motifs and classified into families, superfamilies, and subfamilies. For example, protein domains can be classified according to PFAM or TIGRFAM. Thus, Cas proteins can be identified and classified by protein domain. For example, type II-A Cas9 proteins, such as Cas9 from Streptococcus pyogenes, belong to the TIGR01865 TIGRFAM protein family. In contrast, type II-B Cas9 proteins belong to the TIGR03031 TIGRFAM protein family.
[0146] Thus, in some embodiments, a stiCas9 of the present disclosure comprises a domain having at least 95% sequence similarity to any of SEQ ID NOs: 10-97 or 192-195. In some embodiments, a stiCas9 of the present disclosure comprises a domain having at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% sequence similarity to any of SEQ ID NOs: 10-97 or 192-195. In some embodiments, a stiCas9 of the present disclosure comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1.
[0147] In some embodiments, the Type II-B Cas9 is derived from any species that has a Type II-B CRISPR system. In some embodiments, the Type II-B Cas9 is derived from the following bacterial species: Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, Ruminobacter sp. RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon 274 strain F0058, Wolinella succinogenes, or the like. succinogenes, Burkholderiales bacterium YL45, Ruminobacter amylophilus, Campylobacter sp. P0111, Campylobacter sp. RM9261, Campylobacter lanienae strain RM8001, Campylobacter lanienae strain P0121, Turicimonas muris, Legionella londiniensis, Salinivibrio sharmensis, Leptospira spp. sp.) isolate FW.030, Moritella sp. isolate NORP46, Endozoicomonas sp.) S-B4-1U, Tamilnaduibacter salinus, Vibrio natriegens, Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis, or Parendozoicomonas haliclonae.
[0148] In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Legionella pneumophila Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Francisella novicida Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the gamma proteobacterium HTCC5015 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Parasutterella excrementihominis Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Sutterella wadsworthensis Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Sulfurospirillum sp. SCADC Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Ruminobacter sp. RM87 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Burkholderiales bacterium 1_1_47 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Bacteroidetes oral taxon 274 strain F0058 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Wolinella succinogenes Cas9 protein.In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Burkholderia bacterium YL45 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Ruminobacter amylophilus strain DSM1361 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Campylobacter sp. P0111 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Campylobacter sp. RM9261 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Campylobacter lanienae strain RM8001 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Camplylobacter lanienae strain P0121 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Turicimonas muris Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Legionella londiniensis Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Salinivibrio sharmensis Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Leptospira sp. isolate FW.030 Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Moritella sp. isolate NORP46 Cas9 protein.In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Endozoicomonas sp. S-B4-1U Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Tamilnaduibacter salinus Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Vibrio natriegens Cas9 protein. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Arcobacter skirrowii Cas9. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Francisella philomiragia Cas9. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of the Francisella hispaniensis Cas9. In some embodiments, the term Cas9 refers to a polypeptide comprising the amino acid sequence of Parendozoicomonas haliclonae Cas9. In some embodiments, the term Cas9 refers to a Cas9 polypeptide from a metagenomic sequence catalog. In some embodiments, the term Cas9 refers to a polypeptide comprising any of SEQ ID NOs: 10-97 or 192-195. See Figure 30, SEQ ID NOs: 10-80; Figure 31, SEQ ID NOs: 81-97; and Figure 47, SEQ ID NOs: 192-195.
[0149] In some embodiments, the stiCas9 protein comprises a domain having a sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of any one of SEQ ID NOs: 10-97 or 192-195. In some embodiments, the stiCas9 protein is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or about 100% identical to the amino acid sequence of any one of SEQ ID NOs: 10-97 or 192-195.
[0150] As used herein, the terms "sticky end," "staggered end," or "cohesive end" refer to a nucleic acid fragment having strands of unequal length. In contrast to "blunt ends," sticky ends are produced by staggered cuts in nucleic acids, typically DNA. A sticky or cohesive end has a protruding single strand with unpaired nucleotides, or "overhangs," such as a 3' or 5' overhang. Each overhang can anneal to another complementary overhang to form base pairs. Two complementary sticky ends can anneal together through interactions, such as hydrogen bonding. The stability of the annealed sticky ends depends on the melting temperature of the paired overhangs. Two complementary sticky ends can be joined together by chemical or enzymatic ligation, for example, using a DNA ligase.
[0151] Cas9 proteins have previously been known to generate double-stranded DNA breaks at blunt ends (see, e.g., Jinek et al., 2012). The present disclosure provides Cas9 proteins capable of generating sticky ends, also referred to herein as "stiCas9" or "sticky Cas9." DNA fragments with sticky ends offer advantages over blunt ends in further applications, such as interfragment nucleic acid insertion and overall fragment rejoining. Blunt-ended DNA sequences do not provide specificity for nucleic acid insertion; i.e., nucleic acids can be inserted at either blunt end. On the other hand, sticky ends only pair with complementary sticky ends, thus allowing transgene integration in a preferred orientation. In some embodiments, sticky ends facilitate DNA insertion via non-homologous end joining and microhomology-mediated end joining.
[0152] In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 3 to 40 nucleotides. In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 4 to 20 nucleotides. In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 5 to 15 nucleotides. In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides. In some embodiments, the sticky end generated by stiCas9 is a 5' overhang. In some embodiments, the sticky end generated by stiCas9 is a 3' overhang.
[0153] The compositions and methods described herein may include a guide polynucleotide. In some embodiments, the guide polynucleotide is an RNA molecule. An RNA molecule that binds to CRISPR-Cas components and targets them to a specific location within a target DNA is referred to herein as a "guide RNA," "gRNA," or "small molecule guide RNA," and may also be referred to herein as a "DNA-targeting RNA." A guide polynucleotide, e.g., a guide RNA, comprises at least two nucleotide segments: at least one "DNA-binding segment" and at least one "polypeptide-binding segment." A "segment" refers to a portion, section, or region of a molecule, e.g., a contiguous stretch of nucleotides of a guide polynucleotide molecule. The definition of "segment" is not limited to a specified number of total base pairs, unless otherwise specifically defined.
[0154] In some embodiments, the DNA binding segment of the guide polynucleotide hybridizes with the target sequence in eukaryotic cells, but does not hybridize with the sequence in bacterial cells.As used herein, the sequence in bacterial cells refers to the polynucleotide sequence derived from bacterial organisms, i.e., the naturally occurring bacterial polynucleotide sequence, or the sequence derived from bacteria.For example, the sequence can be bacterial chromosome or bacterial plasmid, or any other polynucleotide sequence naturally found in bacterial cells.
[0155] In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to Cas9. In some embodiments, the polypeptide binding segment of the guide polynucleotide binds to stiCas9.
[0156] In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.
[0157] The guide polynucleotide, e.g., guide RNA, can be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or is introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide, e.g., the guide RNA.
[0158] The "DNA-binding segment" (or "DNA-targeting sequence") of a guide polynucleotide, e.g., a guide RNA, comprises a nucleotide sequence that is complementary to a specific sequence within the target DNA.
[0159] A guide polynucleotide, e.g., a guide RNA, of the present disclosure can comprise a polypeptide-binding sequence / segment. The polypeptide-binding segment (or "protein-binding sequence") of a guide polynucleotide, e.g., a guide RNA, interacts with the polynucleotide-binding domain of a Cas protein of the present disclosure. Such polypeptide binding segments or sequences are known to those skilled in the art, for example, they are disclosed in U.S. Patent Application Publication Nos. 2014 / 0068797, 2014 / 0273037, 2014 / 0273226, 2014 / 0295556, 2014 / 0295557, 2014 / 0349405, 2015 / 0045546, 2015 / 0071898, 2015 / 0071899, and 2015 / 0071906, the disclosures of which are incorporated herein in their entireties.
[0160] In some embodiments of the present disclosure, stiCas9 and a guide polynucleotide can form a complex. A "complex" is a group of two or more associated nucleic acids and / or polypeptides. In some embodiments, a complex is formed when all components of the complex are present together, i.e., a self-assembly complex. In some embodiments, a complex is formed through chemical interactions, such as hydrogen bonds, between different components of the complex. In some embodiments, a guide polynucleotide forms a complex with stiCas9 through secondary structure recognition of the guide polynucleotide by stiCas9. In some embodiments, the stiCas9 protein is inactive, i.e., does not exhibit nuclease activity, until it forms a complex with a guide polynucleotide. Binding of a guide RNA induces a conformational change in stiCas9, converting stiCas9 from an inactive form to an active form, i.e., a catalytically active form. In embodiments of the present disclosure, the complex of stiCas9 and a guide polynucleotide does not occur in nature.
[0161] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that comprises a Cas9 effector protein (stiCas9) that is capable of generating sticky ends and that comprises a nuclear localization signal (NLS), and a guide polynucleotide that forms a complex with stiCas9 and that comprises a guide sequence, wherein the complex does not occur in nature.
[0162] In some embodiments, stiCas9 comprises one or more nuclear localization signals. A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for transport into the cell nucleus by nuclear transport; i.e., proteins with an NLS are transported into the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the protein surface. Exemplary nuclear localization sequences include, but are not limited to, NLSs from SV40 large T antigen, w, EGL-13, c-Myc, and TUS proteins. In some embodiments, an NLS comprises the sequence PKKKRKV (SEQ ID NO: 1). In some embodiments, an NLS comprises the sequence AVKRPAATKKAGQAKKKKLD (SEQ ID NO: 2). In some embodiments, an NLS comprises the sequence PAAKRVKLD (SEQ ID NO: 3). In some embodiments, an NLS comprises the sequence MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 4). In some embodiments, the NLS comprises the sequence KLKIKRPVK (SEQ ID NO: 5). Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK (SEQ ID NO: 6) in the yeast transcriptional repressor Matα2, and the PY-NLS.
[0163] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system that includes: (a) one or more nucleotides encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends; and (b) a nucleotide sequence encoding a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell, wherein the complex does not occur in nature.
[0164] In some embodiments, the stiCas9 protein is encoded by one or more polynucleotides. In some embodiments, the polynucleotide is DNA. In some embodiments, the polynucleotide is RNA.
[0165] In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Legionella pneumophila Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Francisella novicida Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a gamma proteobacterium HTCC5015 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Parasutterella excrementihominis Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Sutterella wasworthensis Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Sulfurospirillum sp. SCADC Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Ruminobacter sp. RM87 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Burkholderiales bacterium 1_1_47 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Bacteroidetes oral taxon 274 strain F0058 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Wolinella succinogenes Cas9 protein.In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Burkholderia bacterium YL45 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Ruminobacter amylophilus strain DSM1361 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Campylobacter sp. P0111 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Campylobacter sp. RM9261 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Campylobacter lanienae strain RM8001 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Camplylobacter lanienae strain P0121 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Turicimonas muris Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Legionella londiniensis Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Salinivibrio sharmensis Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Leptospira sp. isolate FW.030 Cas9 protein.In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Moritella sp. isolate NORP46 Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Endozoicomonas sp. S-B4-1U Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Tamilnaduibacter salinus Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Vibrio natriegens Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from the Arcobacter skirrowii Cas9 protein. In some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Francisella philomiragia Cas9 protein, hi some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Francisella hispaniensis Cas9 protein, or hi some embodiments, stiCas9 is encoded by one or more polynucleotides derived from a Parendozoicomonas haliclonae Cas9 protein.
[0166] In some embodiments, a stiCas9 of the present disclosure comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1.
[0167] In some embodiments, the guide polynucleotide of the CRISPR-Cas system is encoded by a nucleotide sequence. In some embodiments, the nucleotide sequence is DNA. In some embodiments, the guide polynucleotide is a guide RNA. In some embodiments, the guide sequence of the guide polynucleotide is a DNA targeting sequence.
[0168] In some embodiments, the nucleotide sequence encoding stiCas9 is codon-optimized. An example of a codon-optimized sequence is one optimized for expression in a eukaryote, such as a human (i.e., optimized for expression in a human), or for another eukaryote, animal, or mammal discussed herein; see, for example, the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (given knowledge in the art and this disclosure, codon-optimized encoding nucleic acid molecules, particularly for effector proteins (e.g., Cas9), are within the skill of those in the art). Other examples are possible, including codon optimization for host species other than humans, or codon optimization for specific organs. In some embodiments, the enzyme-coding sequence encoding the DNA / RNA-targeting Cas protein is codon-optimized for expression in a specific cell, such as a eukaryotic cell. Eukaryotic cells can be cells of or derived from specific organisms, such as plants or mammals, including, but not limited to, humans, or non-human eukaryotic organisms or animals or mammals discussed herein, such as mice, rats, rabbits, dogs, livestock, or non-human mammals or primates. In some embodiments, methods of altering the germline genetic identity of humans and / or methods of altering the genetic identity of animals that are likely to cause suffering without any substantial medical benefit to humans or animals, and even animals resulting from such methods, are excluded. Generally, codon optimization refers to a method of modifying a nucleic acid sequence for improved expression in a target host cell by replacing at least one codon (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) of a native sequence with a codon that is more frequently or most frequently used in the genes of the target host cell while maintaining the native amino acid sequence. Different species exhibit specific biases for certain codons for specific amino acids.Codon bias (differences in codon frequency between organisms) often correlates with the efficiency of messenger RNA (mRNA) translation, which in turn is thought to depend, inter alia, on the characteristics of the codon being translated and the availability of specific transfer RNA (tRNA) molecules. The dominance of selected tRNAs in a cell generally reflects the codons most frequently used in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon frequency tables are readily available, for example, in the "Codon Usage Database" (www.kazusa.orjp / codon / ), and these tables can be adapted in a number of ways. See Nakamura et al., "Codon usage tabulated from the international DNA sequence databases: status for the year 2000," Nucleic Acids Research 28:292 (2000). Computer algorithms are also available for codon optimization of specific sequences for expression in specific host cells. In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in a sequence encoding a DNA / RNA-targeting Cas protein correspond to the most frequently used codon for a particular amino acid. For codon frequencies in yeast, see the online Yeast Genome database (www.yeastgenome.org / community / codon_usage.shtml) or Bennetzen and Hall, "Codon selection in yeast," Journal of Biological Chemistry, 257(6):3026-31 (1982).For codon usage in plants, e.g., algae, see Campbell and Gowri, "Codon usage in higher plants, green algae, and cyanobacteria," Plant Physiology 92(1):1-11 (1990); and Murray et al., "Codon usage in plant genes," Nucleic Acids Research 17(2):477-98 (1989); or Morton, "Selection on the codon bias of chloroplast and cyanelle genes in different plant and algal lineages," Molecular Evolution 46(4):449-59 (1998). In some embodiments, one or more of SEQ ID NOs: 10-97 or 192-195 are codon-optimized.
[0169] In some embodiments, the nucleotide sequence encoding stiCas9 is codon-optimized for expression in eukaryotic cells. In some embodiments, the nucleotide sequence encoding stiCas9 is codon-optimized for expression in animal cells. In some embodiments, the nucleotide sequence encoding stiCas9 is codon-optimized for expression in human cells. The nucleotide sequence encoding stiCas9 is codon-optimized for expression in plant cells. Codon optimization is the adjustment of codons to match the expression with the host's tRNA abundance to increase the yield and efficiency of recombinant or heterologous protein expression. Codon optimization methods are routine in the art and can be performed using software programs such as Integrated DNA Technologies' Codon Optimization tool, Entelechon's Codon Usage Table analysis tool, GENEMAKER's Blue Heron software, Aptagen's Gene Forge software, DNA Builder Software, General Codon Usage Analysis software, publicly available OPTIMIZER software, and Genscript's OptimumGene algorithm.
[0170] In some embodiments, the CRISPR-Cas system of the present disclosure further comprises a tracrRNA. The "tracrRNA" or trans-activating CRISPR-RNA forms an RNA duplex with the pre-crRNA or pre-CRISPR-RNA, which is then cleaved by the RNA-specific ribonuclease RNase III to form a crRNA / tracrRNA hybrid. In some embodiments, the guide RNA comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide RNA activates the Cas9 protein.
[0171] In some embodiments of the present disclosure, stiCas9, a guide polynucleotide, and a tracrRNA can form a complex. In some embodiments, the complex of stiCas9, a guide polynucleotide, and a tracrRNA does not occur in nature.
[0172] In some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system comprising one or more vectors: (a) a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends; and (b) a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence (the guide sequence can hybridize to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell), wherein the complex does not occur in nature. Those skilled in the art will understand that a vector that includes a "guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence" also includes a vector that includes a polynucleotide sequence that can be transcribed into a guide polynucleotide. For example, a DNA vector can be transcribed to generate a guide RNA sequence.
[0173] In some embodiments, the disclosure provides a non-naturally occurring CRISPR-Cas system comprising one or more vectors that form a complex with the stiCas9 and include a guide polynucleotide that comprises a guide sequence, and a regulatory element operably linked to one or more nucleotide sequences encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, where the regulatory element is a eukaryotic regulatory element, and the complex comprises one or more vectors that do not occur in nature.
[0174] In some embodiments, the regulatory element is a promoter. In some embodiments, the regulatory element is a bacterial promoter. In some embodiments, the regulatory element is a viral promoter. In some embodiments, the regulatory element is a eukaryotic regulatory element, i.e., a eukaryotic promoter. In some embodiments, the eukaryotic regulatory element is a mammalian promoter.
[0175] "Operably linked" means that the nucleotide sequence of interest, i.e., the nucleotide encoding the Cas9 protein, is linked to a regulatory element in a manner that allows for expression of the nucleotide sequence. Thus, in some embodiments, the vector is an expression vector.
[0176] In some embodiments, the guide polynucleotide of the vector comprising the CRISPR-Cas system is encoded by a nucleotide sequence. In some embodiments, the nucleotide sequence is DNA. In some embodiments, the guide polynucleotide is a guide RNA. In some embodiments, the guide sequence of the guide polynucleotide is a DNA targeting sequence.
[0177] In some embodiments, stiCas9 and a guide polynucleotide can form a complex. In some embodiments, the complex of stiCas9 and a guide polynucleotide does not occur in nature.
[0178] In some embodiments, the vector further comprises a nucleotide sequence comprising a tracrRNA sequence. In some embodiments, the guide RNA comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide RNA activates the Cas9 protein.
[0179] In some embodiments, the CRISPR-Cas system described herein can cleave at a site within 10 nucleotides of a protospacer adjacent motif. The protospacer adjacent motif, or PAM, is a 2-6 base pair nucleotide sequence located within 1 nucleotide of the region complementary to the guide RNA. When the Cas9 protein is activated (e.g., by forming a complex with the guide polynucleotide), it searches the target DNA by binding to a sequence that matches its PAM sequence. See, e.g., Sternberg et al., "DNA interrogation by the CRISPR RNA-guided endonuclease Cas9," Nature 507(7490):62-67 (2014), which is incorporated herein by reference in its entirety. When a potential target sequence with the appropriate PAM is recognized and the guide RNA is properly paired with the target region, the nuclease domains of Cas9 (i.e., the RuvC and HNH domains) cleave the target DNA.
[0180] In some embodiments, the RuvC and HNH domains of the Cas9 protein of the present disclosure each cleave one strand of a target DNA sequence. In embodiments, the cleavage sites of the RuvC and HNH domains of the stiCas9 protein are offset, i.e., each domain cleaves at a different position on its respective strand of the target DNA, resulting in an overhang. In embodiments, the RuvC and HNH domains of the stiCas9 protein cleave at a 3-nucleotide offset. In embodiments, the RuvC and HNH domains of the stiCas9 protein cleave at a 4-nucleotide offset. In embodiments, the RuvC and HNH domains of the stiCas9 protein cleave at a 5-nucleotide offset. In embodiments, the RuvC and HNH domains of the stiCas9 protein cleave at an offset of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, or about 40 nucleotides.
[0181] In some embodiments, the RuvC and HNH domains of the Cas9 effector proteins of the present disclosure cleave at different positions on each strand of a double-stranded target DNA. In some embodiments, the RuvC domain of the Cas9 effector protein cleaves one strand of the double-stranded target DNA at about -10, -9, -8, -7, or -6 nucleotides from the PAM (e.g., which may be referred to as the "non-target strand"), and the HNH domain of the Cas9 effector protein cleaves the other strand of the double-stranded target DNA at about -5, -4, -3, -2, or -1 nucleotides from the PAM (e.g., which may be referred to as the "target strand").
[0182] In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -8 nucleotide from the PAM. In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -7 nucleotide from the PAM. In some embodiments, the RuvC domain cleaves one strand of a double-stranded target DNA at about -6 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -4 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -3 nucleotide from the PAM. In some embodiments, the HNH domain cleaves one strand of a double-stranded target DNA at about -2 nucleotide from the PAM.
[0183] In some embodiments of the present disclosure, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 10 nucleotides of a protospacer adjacent motif (PAM). In some embodiments, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 5 nucleotides of the PAM. In some embodiments, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 3 nucleotides of the PAM. In some embodiments, the PAM is located downstream (i.e., in the 3' direction) of the target sequence. In some embodiments, the PAM is located upstream (i.e., in the 5' direction) of the target sequence. In some embodiments, the PAM is located within the target sequence.
[0184] Different bacterial species recognize different PAM sequences. One method for identifying preferred PAM sequences for the Cas9 protein of the present disclosure is illustrated in Figure 49A and involves, for example, generating a plasmid library of various PAM sequences flanking a target sequence, contacting the plasmid library with the Cas9 protein, and then sequencing the plasmid library to determine which PAM sequences are "depleted" (i.e., not detected in the sequencing results). A "depleted" PAM sequence is one that is recognized and affected (i.e., cleaved) by the Cas9 protein.
[0185] For example, the PAM sequence recognized by Streptococcus pyogenes Cas9 is 5'-NGG-3' (where N is any nucleotide). Different PAMs are associated with the Cas9 proteins of Neisseria meningitidis, Treponema denticola, and Streptococcus thermophilus. The Cas9 protein of Francisella novicida has been engineered to recognize the PAM 5'-YG-3' (where Y is a pyrimidine).
[0186] In some embodiments, the PAM comprises a 3' G-rich motif. In some embodiments, the PAM sequence is NGG (N is A, C, T, U, or G). In some embodiments, the PAM sequence is NGA (N is A, C, T, U, or G). In some embodiments, the PAM sequence is YG (Y is a pyrimidine (i.e., C, T, or U)).
[0187] In some embodiments, the target sequence is 5' to the PAM, and the PAM contains a 3' G-rich motif. In some embodiments, the target sequence is 5' to the PAM, and the PAM sequence is NGG (where N is A, C, T, U, or G). In some embodiments, the target sequence is 5' to the PAM, and the PAM sequence is YG (where Y is a pyrimidine), and stiCas9 is derived from the bacterial species Francisella novicida.
[0188] In some embodiments, stiCas9 comprises one or more nuclear localization signals. A "nuclear localization signal" or "nuclear localization sequence" (NLS) is an amino acid sequence that "tags" a protein for transport into the cell nucleus by nuclear transport; i.e., proteins with an NLS are transported into the cell nucleus. Typically, an NLS comprises a positively charged Lys or Arg residue exposed on the protein surface. Exemplary nuclear localization sequences include, but are not limited to, NLSs from SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, and TUS proteins. In some embodiments, an NLS comprises the sequence PKKKRKV (SEQ ID NO: 1). In some embodiments, an NLS comprises the sequence AVKRPAATKKAGQAKKKKLD (SEQ ID NO: 2). In some embodiments, an NLS comprises the sequence PAAKRVKLD (SEQ ID NO: 3). In some embodiments, an NLS comprises the sequence MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 4). In some embodiments, the NLS comprises the sequence KLKIKRPVK (SEQ ID NO: 5). Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNP A1, the sequence KIPIK (SEQ ID NO: 6) in the yeast transcriptional repressor Matα2, and the PY-NLS.
[0189] In some embodiments, the guide polynucleotides of the present disclosure comprise a guide sequence that hybridizes to a target sequence in a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is a human, rodent, or bovine cell line or cell strain. Examples of such cells, cell lines, or cell strains include, but are not limited to, mouse myeloma (NS0) cell lines, Chinese hamster ovary (CHO) cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK (baby hamster kidney cells), VERO, SP2 / 0, YB2 / 0, YO, C127, L cells, COS (e.g., COS1 and COS7), QC1-3, HEK-293, VERO, PER.C6, HeLA, EB1, EB2, EB3, oncolytic, or hybridoma cell lines. In some embodiments, the eukaryotic cell is a CHO cell line. In some embodiments, the eukaryotic cell is a CHO cell. In some embodiments, the cell is a CHO-K1 cell, a CHO-K1 SV cell, a DG44 CHO cell, a DUXB11 CHO cell, a CHOS, a CHO GS knockout cell, a CHO FUT8 GS knockout cell, a CHOZN, or a CHO-derived cell. A CHO GS knockout cell (e.g., a GSKO cell) is, for example, a CHO-K1 SV GS knockout cell. A CHO FUT8 knockout cell is, for example, Potelligent® CHOK1 SV (Lonza Biologics, Inc.). The eukaryotic cell may be an avian cell, cell line, or cell strain, such as, for example, an EBx® cell, EB14, EB24, EB26, EB66, or EBvl3.
[0190] In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a stem cell. The stem cell can be, for example, a pluripotent stem cell, such as an embryonic stem cell (ESC), an adult stem cell, an induced pluripotent stem cell (iPSC), a tissue-specific stem cell (e.g., a hematopoietic stem cell), or a mesenchymal stem cell (MSC). In some embodiments, the human cell is a differentiated form of any of the cells described herein. In some embodiments, the eukaryotic cell is a cell derived from any primary cell in culture.
[0191] In some embodiments, the eukaryotic cells are hepatocytes, such as human hepatocytes, animal hepatocytes, or non-parenchymal cells. For example, the eukaryotic cells can be adherent metabolic test human hepatocytes, adherent induction test human hepatocytes, adherent Qualyst Transporter Certified™ human hepatocytes, suspension test human hepatocytes (e.g., 10-donor and 20-donor pooled hepatocytes), human hepatic Kupffer cells, human hepatic stellate cells, dog hepatocytes (e.g., single and pooled Beagle hepatocytes), mouse hepatocytes (e.g., CD-1 and C57BI / 6 hepatocytes), rat hepatocytes (e.g., Sprague-Dawley, Wistar Han, and Wistar hepatocytes), monkey hepatocytes (e.g., cynomolgus or rhesus monkey hepatocytes), cat hepatocytes (e.g., domestic shorthair hepatocytes), and rabbit hepatocytes (e.g., New Zealand White hepatocytes).
[0192] In some embodiments, the eukaryotic cell is a plant cell. For example, the plant cell may be from a crop plant, such as cassava, maize, sorghum, wheat, or rice. The plant cell may be from an algae, tree, or vegetable. The plant cell may be from a monocotyledonous or dicotyledonous plant, or from a crop or cereal plant, a productive plant, a fruit, or a vegetable. For example, the plant cell can be from a tree, such as a citrus tree, e.g., an orange, grapefruit, or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, e.g., an almond, walnut, or pistachio tree; a Solanaceae plant, e.g., potato, a Brassica plant, a Lactuca plant; a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, ginger, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0193] In some embodiments, the guide polynucleotide of a CRISPR-Cas system is linked to a direct repeat sequence. Direct repeats, or DR sequences, are arrays of repetitive sequences in a CRISPR locus interspersed with short stretches of non-repetitive sequences (spacers). The spacer sequence targets a protospacer adjacent motif (PAM) on the target sequence. When the non-coding portion of the CRISPR locus (i.e., the guide polynucleotide and tracrRNA) is transcribed, the transcript is cleaved into short crRNAs containing individual spacer sequences in the DR sequence, which direct the Cas9 nuclease to the PAM. In some embodiments, the DR sequence is RNA. In some embodiments, the DR sequence is encoded by a nucleic acid. In some embodiments, the DR sequence is linked to a guide polynucleotide. In some embodiments, the DR sequence is linked to a guide sequence of a guide polynucleotide. In some embodiments, the DR sequence comprises a secondary structure. In some embodiments, the DR sequence comprises a stem-loop structure. In some embodiments, the DR sequence is 10-20 nucleotides. In some embodiments, the DR sequence is at least 16 nucleotides. In some embodiments, the DR sequence is at least 16 nucleotides and comprises a single stem loop. In some embodiments, the DR sequence comprises an RNA aptamer. In some embodiments, the secondary structure or stem loop of the DR is recognized by a nuclease for cleavage. In some embodiments, the nuclease is a ribonuclease. In some embodiments, the nuclease is RNase III.
[0194] Various means for delivery of CRISPR-Cas systems are known in the art. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered via a delivery particle. The delivery particle is a biological delivery system or formulation comprising the particle. As defined herein, a "particle" is an entity having a maximum diameter of about 100 microns (μm). In some embodiments, a particle has a maximum diameter of about 10 μm. In some embodiments, a particle has a maximum diameter of about 2000 nanometers (nm). In some embodiments, a particle has a maximum diameter of about 1000 nm. In some embodiments, a particle has a maximum diameter of about 900 nm, about 800 nm, about 700 nm, about 600 nm, about 500 nm, about 400 nm, about 300 nm, about 200 nm, or about 100 nm. In some embodiments, a particle has a diameter of about 25 nm to about 200 nm. In some embodiments, a particle has a diameter of about 50 nm to 150 nm. In some embodiments, the particles have a diameter of about 75 nm to about 100 nm.
[0195] The delivery particle can be provided in any form, for example, but not limited to, a solid, semi-solid, emulsion, or colloidal particle. In some embodiments, the delivery particle is a lipid-based system, a liposome, a micelle, a microvesicle, an exosome, or a gene gun. In some embodiments, the delivery particle comprises a CRISPR-Cas system. In some embodiments, the delivery particle comprises a CRISPR-Cas system comprising stiCas9 and a guide polynucleotide. In some embodiments, the delivery particle comprises a CRISPR-Cas system comprising stiCas9 and a guide polynucleotide, wherein the stiCas9 and the guide polynucleotide are present in a complex. In some embodiments, the delivery particle comprises a CRISPR-Cas system comprising stiCas9, a guide polynucleotide, and a polynucleotide comprising tracrRNA. In some embodiments, the delivery particle comprises a CRISPR-Cas system comprising stiCas9, a guide polynucleotide, and tracrRNA.
[0196] In some embodiments, the delivery particle further comprises a lipid, a sugar, a metal, or a protein. In some embodiments, the delivery particle is a lipid envelope. Delivery of mRNA using lipid envelopes or lipid-containing delivery particles is described, for example, in Su et al., "In vitro and in vivo mRNA delivery using lipid-enveloped pH-responsive polymer nanoparticles," Molecular Pharmacology 8(3):774-784 (2011).
[0197] In some embodiments, the delivery particle is a sugar-based particle, such as GalNAc. Sugar-based particles are described in WO 2014 / 118272 and Nair et al., Journal of the American Chemical Society 136(49):16958-16961 (2014), each of which is incorporated herein by reference in its entirety.
[0198] In some embodiments, the delivery particle is a nanoparticle. The nanoparticles encompassed by the present disclosure can be provided in various forms, for example, as solid nanoparticles (e.g., metals, such as silver, gold, iron, titanium), non-metallic, lipid-based solids, polymers, nanoparticle suspensions, or combinations thereof. Metal, dielectric, and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles), can be prepared. Nanoparticles made of semiconductor materials can be labeled quantum dots if they are small enough (typically less than 10 nm) that quantization of electronic energy levels occurs. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications and can be adapted for similar purposes in the present disclosure.
[0199] The preparation of delivery particles is further described in U.S. Patent Application Publication Nos. 2011 / 0293703, 2012 / 0251560, and 2013 / 0302401; and U.S. Patent Nos. 5,543,158, 5,855,913, 5,895,309, 6,007,845, and 8,709,843, each of which is incorporated herein by reference in its entirety.
[0200] In some embodiments, the vesicle comprises the CRISPR-Cas system of the present disclosure. A "vesicle" is a small, chambered structure with fluid enclosed by a lipid bilayer. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered by the vesicle. In some embodiments, the vesicle comprises stiCas9 and a guide polynucleotide. In some embodiments, the vesicle comprises stiCas9 and a guide polynucleotide, wherein the stiCas9 and the guide polynucleotide are in a complex. In some embodiments, the vesicle comprises a CRISPR-Cas system comprising stiCas9, a guide polynucleotide, and a polynucleotide comprising a tracrRNA. In some embodiments, the vesicle comprises a CRISPR-Cas system comprising stiCas9, a guide polynucleotide, and a tracrRNA.
[0201] In some embodiments, the vesicle containing stiCas9 and the guide polynucleotide is an exosome or liposome. In some embodiments, the vesicle is an exosome. In some embodiments, exosomes are used to deliver the CRISPR-Cas system of the present disclosure. Exosomes are endogenous nanovesicles (i.e., having a diameter of about 30 to about 100 nm) that transport RNA and proteins and can deliver RNA to the brain and other target organs. Genetically engineered exosomes for delivery of exogenous biological materials into target organs are described, for example, by Alvarez-Erviti et al., Nature Biotechnology 29:341 (2011), El-Andaloussi et al., Nature Protocols 7:2112-2116 (2012), and Wahlgren et al., Nucleic Acids Research 40(17):el30 (2012), each of which is incorporated herein by reference in its entirety.
[0202] In some embodiments, the vesicle containing stiCas9 and the guide polynucleotide is a liposome. In some embodiments, liposomes are used to deliver the CRISPR-Cas system of the present disclosure. Liposomes are spherical vesicle structures with at least one lipid bilayer and can be used as vehicles for the administration of nutrients and pharmaceuticals. Liposomes are often composed of phospholipids, particularly phosphatidylcholine, but can also be composed of other lipids, such as egg phosphatidylethanolamine. Types of liposomes include, but are not limited to, multilamellar vesicles, small unilamellar vesicles, large unilamellar vesicles, and cochlear vesicles. See, e.g., Spuch and Navarro, "Liposomes for Targeted Delivery of Active Agents against Neurodegenerative Diseases (Alzheimer's Disease and Parkinson's Disease)," Journal of Drug Delivery 2011, Article ID 469679 (2011). Liposomes for delivery of biological materials, e.g., CRISPR-Cas components, are described, e.g., by Morrissey et al., Nature Biotechnology 23(8):1002-1007 (2005), Zimmerman et al., Nature Letters 441:111-114 (2006), and Li et al., Gene Therapy 19:775-780 (2012), each of which is incorporated herein by reference in its entirety.
[0203] In some embodiments, the nucleotides encoding Cas9 and the guide polynucleotide are present on a single vector. In some embodiments, the nucleotides encoding Cas9, the guide polynucleotide (or nucleotides that can be transcribed into a guide polynucleotide), and the tracrRNA are present on a single vector. In some embodiments, the nucleotides encoding Cas9, the guide polynucleotide (or nucleotides that can be transcribed into a guide polynucleotide), the tracrRNA, and the direct repeat sequence are present on a single vector. In some embodiments, the vector is an expression vector. In some embodiments, the vector is a mammalian expression vector. In some embodiments, the vector is a human expression vector. In some embodiments, the vector is a plant expression vector.
[0204] In some embodiments, the Cas9-encoding nucleotide and the guide polynucleotide are a single nucleic acid molecule. In some embodiments, the Cas9-encoding nucleotide, the guide polynucleotide, and the tracrRNA are a single nucleic acid molecule. In some embodiments, the Cas9-encoding nucleotide, the guide polynucleotide, the tracrRNA, and the direct repeat sequence are a single nucleic acid molecule. In some embodiments, the single nucleic acid molecule is an expression vector. In some embodiments, the single nucleic acid molecule is a mammalian expression vector. In some embodiments, the single nucleic acid molecule is a human expression vector. In some embodiments, the single nucleic acid molecule is a plant expression vector.
[0205] In some embodiments, a viral vector comprises the CRISPR-Cas system of the present disclosure. In some embodiments, the CRISPR-Cas system of the present disclosure is delivered by a viral vector. In some embodiments, the viral vector comprises stiCas9 and a guide polynucleotide. In some embodiments, the viral vector comprises stiCas9 and a guide polynucleotide, wherein the stiCas9 and the guide polynucleotide are present in a complex. In some embodiments, the viral vector comprises a CRISPR-Cas system comprising a polynucleotide comprising stiCas9, a guide polynucleotide, and a tracrRNA. In some embodiments, the viral vector comprises a CRISPR-Cas system comprising stiCas9, a guide polynucleotide, and a tracrRNA. In some embodiments, the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus. Examples of viral vectors are provided herein.
[0206] In some embodiments, adeno-associated viruses (AAV) and / or lentiviruses can be used as viral vectors that contain elements of the CRISPR-Cas system described herein. In some embodiments of the present disclosure, the Cas proteins are expressed intracellularly by cells transduced with the viral vector.
[0207] For many therapeutic strategies, including those contemplated by the present disclosure, Cas protein expression may only be required transiently. Consequently, in some embodiments of the present disclosure, delivery of Cas proteins into cells is achieved using non-integrating viral vectors. In other embodiments, expression of CRISPR-Cas system components is required for long periods of time, for example, when used in genetic circuits that are permanently integrated into the genome of target cells. Such applications are discussed by Agustin-Pavon, et al., "Synthetic biology and therapeutic strategies for the degenerating brain," Bioessays 36(10):979-990 (2014), which is incorporated herein by reference in its entirety.
[0208] In some embodiments, the disclosed Cas proteins and methods are used in ex vivo gene editing, e.g., CAR-T type therapy. These embodiments may involve the modification of cells from human donors. In these instances, viral vectors may be used; however, there is the additional option of directly transfecting the Cas proteins (along with in vitro transcribed guide RNA and donor DNA) into cultured cells.
[0209] In some embodiments, the disclosure provides a eukaryotic cell comprising a non-naturally occurring CRISPR-Cas system that includes (a) a Cas9 effector protein (stiCas9) capable of generating sticky ends, and (b) a guide polynucleotide that forms a complex with the stiCas9 and includes a guide sequence, where the guide sequence is capable of hybridizing to a target sequence in the eukaryotic cell. In some embodiments, the eukaryotic cell comprises a vector comprising the CRISPR-Cas system of the disclosure.
[0210] In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is an animal cell. In some embodiments, the eukaryotic cell is a human cell, such as a human stem cell. In some embodiments, the eukaryotic cell is a plant cell. Examples of various types of eukaryotic cells are provided herein.
[0211] In some embodiments, the disclosure provides a eukaryotic cell comprising a CRISPR-Cas system comprising a Cas9 effector protein (stiCas9) capable of generating sticky ends (the Cas9 effector protein is derived from a bacterial species having a Type II-B CRISPR system). In some embodiments, the eukaryotic cell comprises a stiCas9 comprising a domain that matches to the TIGR03031 protein family using an E-value cutoff of at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1. In some embodiments, the eukaryotic cell comprises a stiCas9 comprising a polypeptide sequence that has at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence similarity to any one of SEQ ID NOs: 10-97 or 192-195. In some embodiments, the eukaryotic cell comprises a stiCas9 comprising a polypeptide sequence that has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to any one of SEQ ID NOs: 10-97 or 192-195.
[0212] In some embodiments, the Cas9 protein of the present disclosure is part of a fusion protein containing one or more heterologous protein domains (e.g., about or at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more domains in addition to the Cas9 protein). The Cas9 fusion protein can include any additional protein sequences and, optionally, a linker sequence between any two domains. Examples of protein domains that can be fused to the Cas9 protein include, but are not limited to, epitope tags, reporter gene sequences, and protein domains with one or more of the following activities: methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, RNA cleavage activity, and nucleic acid binding activity. Non-limiting examples of epitope tags include histidine (His) tags, V5 tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include, but are not limited to, glutathione-5-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), autofluorescent proteins such as blue fluorescent protein (BFP), and mCherry. In some embodiments, the Cas9 protein is fused to a protein or protein fragment that binds to a DNA molecule or other cellular molecule, such as, but not limited to, maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD), GAL4 DNA binding domain, and herpes simplex virus (HSV) BP16 protein. Additional domains that can form part of fusion proteins comprising the Cas9 protein are described in US Patent Application Publication No. 20110059502, which is incorporated herein by reference in its entirety.In some embodiments, a tagged Cas9 protein is used to identify the location of the target sequence.
[0213] In some embodiments, the Cas9 protein can form a component of an inducible system. The inducible nature of the system allows for spatiotemporal control of gene editing or gene expression using forms of energy, including, but not limited to, electromagnetic radiation, acoustic energy, chemical energy, and thermal energy. Non-limiting examples of inducible systems include tetracycline-inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcriptional activation systems (e.g., FKBP, ABA), or light-inducible systems (phytochrome, LOV domain, or cryptochrome). In some embodiments, the Cas9 protein is part of a light-inducible transcriptional effector (LITE) for directing changes in transcriptional activity in a sequence-specific manner. The light component can include the Cas9 protein, a light-responsive cytochrome heterodimer (e.g., from Arabidopsis thaliana), and a transcriptional activation / repression domain. Further examples of inducible DNA-binding proteins and methods of their use are provided in International Patent Application Publication Nos. WO 2014 / 018423 and WO 2014 / 093635; U.S. Patent Nos. 8,889,418 and 8,895,308; and U.S. Patent Application Publication Nos. 2014 / 0186919, 2014 / 0242700, 2014 / 0273234, and 2014 / 0335620, each of which is incorporated herein by reference in its entirety.
[0214] Methods for site-specific modification In some embodiments, the disclosure provides a method for providing site-specific modification of a target sequence in a eukaryotic cell, comprising: (1) introducing into the cell (a) a Cas9 effector protein (stiCas9) capable of generating sticky ends; and (b) a guide polynucleotide (the complex is not naturally occurring) that forms a complex with the stiCas9 and includes a guide sequence (the guide sequence can hybridize to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell); (2) generating sticky ends in the target sequence by the Cas9 effector protein and the guide polynucleotide; and (3) ligating (a) the sticky ends together or (b) a polynucleotide sequence of interest (SoI) to the sticky ends, thereby modifying the target sequence.
[0215] "Modifications" of a target sequence encompass single nucleotide substitutions, multiple nucleotide substitutions, insertions (ie, knock-ins) and deletions (ie, knock-outs) of nucleic acids, frameshift mutations, and other nucleic acid modifications.
[0216] In some embodiments, the modification is a deletion of at least a portion of the target sequence, which can be cleaved at two different sites to generate complementary sticky ends, and the complementary sticky ends can be religated, thereby removing the portion of the sequence between the two sites.
[0217] In some embodiments, the modification is a mutation of the target sequence. Site-specific mutagenesis in eukaryotic cells is achieved by the use of site-specific nucleases that promote homologous recombination of an exogenous polynucleotide template (also called a "donor polynucleotide" or "donor vector") containing the desired mutation. In some embodiments, a sequence of interest (SoI) comprises the desired mutation.
[0218] In some embodiments, the modification is the insertion of a sequence of interest (SoI) into the target sequence. The SoI can be introduced as an exogenous polynucleotide template. In some embodiments, the exogenous polynucleotide template comprises a sticky end. In some embodiments, the exogenous polynucleotide template comprises a sticky end that is complementary to a sticky end in the target sequence.
[0219] The exogenous polynucleotide template can be of any suitable length, for example, about or at least about 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 500, or 1000 or more nucleotides in length. In some embodiments, the exogenous polynucleotide template is complementary to a portion of the polynucleotide comprising the target sequence. When optimally aligned, the exogenous polynucleotide template overlaps with one or more nucleotides of the target sequence (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides). In some embodiments, when the exogenous polynucleotide template and the polynucleotide comprising the target sequence are optimally aligned, the nearest neighbor nucleotide of the exogenous polynucleotide template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides of the target sequence.
[0220] In some embodiments, the exogenous polynucleotide is DNA, such as a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear single-stranded or double-stranded piece of DNA, an oligonucleotide, a PCR fragment, a naked nucleic acid, or a nucleic acid complexed with a delivery vehicle, such as a liposome.
[0221] In some embodiments, the exogenous polynucleotide is inserted into the target sequence using the cell's endogenous DNA repair pathway. Endogenous DNA repair pathways include the non-homologous end joining (NHEJ) pathway, the microhomology-mediated end joining (MMEJ) pathway, and the homology-directed repair (HDR) pathway. NHEJ, MMEJ, and HDR pathways repair double-stranded DNA breaks. In NHEJ, a homologous template is not required to repair the break in the DNA. NHEJ repair can be error-prone, but errors are reduced when the DNA break contains a compatible overhang. NHEJ and MMEJ are mechanistically distinct DNA repair pathways due to the different subsets of DNA repair enzymes involved in each. Unlike NHEJ, which can be as accurate as error-prone, MMEJ is always error-prone and results in both deletions and insertions at the site under repair. MMEI-associated deletions result from microhomologies (2-10 base pairs) on both sides of the double-stranded break. In contrast, HDR requires a homologous template to direct repair, but HDR repair is typically high fidelity and low error-prone. In some embodiments, the error-prone nature of NHEJ and MMEJ repair is exploited to introduce nonspecific nucleotide substitutions in the target sequence. In some embodiments, stiCas9 cleaves the target sequence in a manner that facilitates HDR repair.
[0222] During the repair process, an exogenous polynucleotide template containing an SoI can be introduced into the target sequence. In some embodiments, an exogenous polynucleotide template containing an SoI flanked by upstream and downstream sequences is introduced into a cell, where the upstream and downstream sequences share sequence similarity with either side of the site of integration in the target sequence. In some embodiments, the exogenous polynucleotide containing an SoI comprises, for example, a mutant gene. In some embodiments, the exogenous polynucleotide comprises a sequence that is endogenous or exogenous to the cell. In some embodiments, the SoI comprises a polynucleotide that encodes a protein or a non-coding sequence, such as a microRNA. In some embodiments, the SoI is operably linked to a regulatory element. In some embodiments, the SoI is a regulatory element. In some embodiments, the SoI comprises a resistance cassette, e.g., a gene that confers resistance to an antibiotic. In some embodiments, the SoI comprises a mutation of the wild-type target sequence. In some embodiments, the SoI disrupts or corrects the target sequence by creating a frameshift mutation or a nucleotide substitution. In some embodiments, the SoI comprises a marker. Introduction of a marker into the target sequence can facilitate screening for targeted integration. In some embodiments, the marker is a restriction site, a fluorescent protein, or a selectable marker. In some embodiments, the SoI is introduced as a vector containing the SoI.
[0223] The upstream and downstream sequences in the exogenous polynucleotide template are selected to promote homologous recombination between the target sequence and the exogenous polynucleotide. The upstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence upstream of the targeting site for integration (i.e., the target sequence). Similarly, the downstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence downstream of the targeting site for integration. Thus, in some embodiments, the exogenous polynucleotide template containing the SoI inserts into the target sequence by homologous recombination at the upstream and downstream sequences. In some embodiments, the upstream and downstream sequences in the exogenous polynucleotide template have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the upstream and downstream sequences of the targeted genomic sequence, respectively. In some embodiments, the upstream or downstream sequence has about 20 to 2000 base pairs, or about 50 to 1750 base pairs, or about 100 to 1500 base pairs, or about 200 to 1250 base pairs, or about 300 to 1000 base pairs, or about 400 to 750 base pairs, or about 500 to 600 base pairs. In some embodiments, the upstream or downstream sequence has about 50, about 100, about 250, about 500, about 100, about 1250, about 1500, about 1750, about 2000, about 2250, or about 2500 base pairs.
[0224] In some embodiments, the modification in target sequence is the inactivation of the expression of target sequence in cell.For example, when CRISPR complex binds to target sequence, target sequence is inactivated, so that this sequence is not transcribed, and the encoded protein is not produced, or the sequence does not function as its wild-type sequence functions.For example, protein or microRNA coding sequence can be inactivated, so that protein is not produced.
[0225] In some embodiments, a regulatory sequence can be inactivated so that it no longer functions as a regulatory sequence. Examples of regulatory sequences include promoters, transcription terminators, enhancers, and other regulatory elements described herein. Inactivated target sequences can include deletion mutations (i.e., deletion of one or more nucleotides), insertion mutations (i.e., insertion of one or more nucleotides), or nonsense mutations (i.e., substitution of a single nucleotide with another nucleotide such that a stop codon is introduced). In some embodiments, inactivation of a target sequence results in a "knockout" of the target sequence.
[0226] In some embodiments, stiCas9 and the guide polynucleotide form a complex, and the guide polynucleotide hybridizes to the target sequence to be modified. In some embodiments, stiCas9 generates sticky ends in the target sequence that hybridizes to the guide polynucleotide.
[0227] In embodiments of the methods, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 3 to 40 nucleotides. In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 4 to 20 nucleotides. In some embodiments, the sticky end generated by stiCas9 comprises a single-stranded polynucleotide overhang of 5 to 15 nucleotides. In some embodiments, the sticky end generated by stiCas9 is a 5' overhang.
[0228] In some embodiments of the method, the stiCas9 is derived from a bacterial species having a type II-B CRISPR system. The type II-B Cas9 protein discussed herein belongs to the TIGR03031 TIGRFAM protein family. Thus, in some embodiments, the stiCas9 of the present disclosure comprises a domain that matches the TIGR03031 protein family using a 1E-5 profile cutoff value. In some embodiments, the stiCas9 of the present disclosure comprises a domain that matches the TIGR03031 protein family using a 1E-10 profile cutoff value. In some embodiments, the stiCas9 of the present disclosure comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of at least 1E-10, at least 1E-9, at least 1E-8, at least 1E-7, at least 1E-6, at least 1E-5, at least 1E-4, at least 1E-3, at least 1E-2, or at least 1E-1.
[0229] In embodiments of the method, the Type II-B Cas9 protein is derived from any species that has a Type II-B CRISPR system. In some embodiments, the Type II-B Cas9 is derived from the following bacterial species: Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, Ruminobacter sp. RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon 274 strain F0058, Wolinella succinogenes, or the like. succinogenes, Burkholderiales bacterium YL45, Ruminobacter amylophilus, Campylobacter sp. P0111, Campylobacter sp. RM9261, Campylobacter lanienae strain RM8001, Campylobacter lanienae strain P0121, Turicimonas muris, Legionella londiniensis, Salinivibrio sharmensis, Leptospira spp. sp.) isolate FW.030, Moritella sp. isolate NORP46, Endozoicomonas sp.) S-B4-1U, Tamilnaduibacter salinus, Vibrio natriegens, Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis, or Parendozoicomonas haliclonae.
[0230] In some embodiments of the methods, the guide polynucleotide is a guide RNA. In some embodiments, the guide polynucleotide comprises at least two nucleotide segments: at least one "DNA-binding segment" or "guide sequence" and at least one "polypeptide-binding segment." In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to Cas9. In some embodiments, the polypeptide-binding segment of the guide polynucleotide binds to stiCas9.
[0231] In embodiments of the methods, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.
[0232] In some method embodiments, stiCas9 and the guide polynucleotide can form a complex. In some embodiments, the complex is formed when all components of the complex are present together, i.e., a self-assembly complex. In some embodiments, the complex is formed through chemical interactions, such as hydrogen bonds, between different components of the complex. In some embodiments, the guide polynucleotide forms a complex with stiCas9 through secondary structure recognition of the guide polynucleotide by stiCas9. In some embodiments, the stiCas9 protein is inactive, i.e., does not exhibit nuclease activity, until it forms a complex with the guide polynucleotide. Binding of the guide RNA induces a conformational change in stiCas9, converting stiCas9 from an inactive form to an active, i.e., catalytically active, form. In some method embodiments, the complex of stiCas9 and the guide polynucleotide does not occur in nature.
[0233] In embodiments of the methods, the sticky ends generated by stiCas9 are ligated together (i.e., chemically linked together). Ligation can be performed, for example, with a DNA ligase, such as T4 ligase or DNA ligase IV. In some embodiments, the sticky ends are ligated with an error-prone ligase that introduces one or more nucleotide substitutions. In some embodiments, a polynucleotide sequence of interest (SoI) is ligated to the sticky ends. In some embodiments, the SoI contains a mutation of interest.
[0234] In some embodiments, sticky ends are generated in SoI that are complementary to the sticky ends generated in target sequence.In some embodiments, sticky ends in SoI are generated by stiCas9.In some embodiments, SoI is ligated into sticky ends using cell's endogenous DNA repair pathway.Endogenous DNA repair pathway is described herein.
[0235] In some embodiments, the disclosure provides a method for site-specific modification of a target sequence in a eukaryotic cell, comprising: (1) introducing into the cell (a) a nucleotide sequence encoding a Cas9 effector protein (stiCas9) capable of generating sticky ends, and (b) a guide polynucleotide (the complex is not naturally occurring) that forms a complex with stiCas9 and includes a guide sequence (the guide sequence is capable of hybridizing to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell); (2) generating sticky ends in the target sequence by the Cas9 effector protein and the guide polynucleotide; and (3) ligating (a) the sticky ends together or (b) a polynucleotide sequence of interest (SoI) to the sticky ends, thereby modifying the target sequence.
[0236] In some embodiments of the methods, the stiCas9 is encoded by a nucleotide sequence. In some embodiments, the nucleotide is DNA. In some embodiments, the stiCas9 protein comprises a domain comprising a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to a nucleotide sequence of any of SEQ ID NOs: 10-97 or 192-195.
[0237] In method embodiments, the CRISPR-Cas system of the present disclosure further comprises a tracrRNA. In some embodiments, the guide RNA comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide RNA activates the Cas9 protein. In method embodiments, the stiCas9, guide polynucleotide, and tracrRNA can form a complex. In some embodiments, the complex of stiCas9, guide polynucleotide, and tracrRNA does not occur in nature.
[0238] In embodiments of the method, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 10 nucleotides of a protospacer adjacent motif (PAM). In some embodiments, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 5 nucleotides of the PAM. In some embodiments, a complex comprising stiCas9 and a guide polynucleotide may be cleaved at a site within 3 nucleotides of the PAM. In some embodiments, the PAM is downstream (i.e., in the 3' direction) of the target sequence. In some embodiments, the PAM is upstream (i.e., in the 5' direction) of the target sequence. In some embodiments, the PAM is located within the target sequence.
[0239] In method embodiments, the PAM comprises a 3' G-rich motif. In some embodiments, the PAM sequence is NGG (N is A, C, T, U, or G). In some embodiments, the PAM sequence is NGA (N is A, C, T, U, or G). In some embodiments, the PAM sequence is YG (Y is a pyrimidine (i.e., C, T, or U)). In method embodiments, the target sequence is 5' to the PAM and the PAM comprises a 3' G-rich motif. In some embodiments, the target sequence is 5' to the PAM and the PAM sequence is NGG (N is A, C, T, U, or G).
[0240] In method embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is an animal cell. In some embodiments, the eukaryotic cell is a human cell, e.g., a human stem cell. In some embodiments, the eukaryotic cell is a plant cell. Examples of various types of eukaryotic cells are provided herein. In method embodiments, stiCas9 and guide polynucleotide are introduced into the eukaryotic cell via a delivery particle. In method embodiments, stiCas9 and guide polynucleotide are introduced into the eukaryotic cell via a vesicle. In method embodiments, stiCas9 and guide polynucleotide are introduced into the eukaryotic cell via a vector. In method embodiments, stiCas9 and guide polynucleotide are introduced into the eukaryotic cell via a viral vector. In method embodiments, polynucleotides encoding components of a complex comprising stiCas9 and guide polynucleotide are introduced on one or more vectors. Examples of vectors and methods of vector delivery into cells (e.g., transfection) are provided herein.
[0241] In some embodiments, the methods of the present disclosure further comprise introducing an exonuclease into the eukaryotic cell to remove overhangs generated from stiCas9. In some embodiments, the exonuclease is a 5' to 3' exonuclease. In some embodiments, the exonuclease is a 3' to 5' exonuclease. In some embodiments, the exonuclease is added before the ligation step of the method. In some embodiments, the exonuclease is added instead of the ligation step of the method. Non-limiting examples of 5' to 3' exonucleases include lambda exonuclease, RecJ, exonuclease V, exonuclease VIII, T5 exonuclease, T7 exonuclease, Artemis, and Cas4. Non-limiting examples of 3' to 5' exonucleases include TREX1, TREX2, Werner syndrome (WRN) protein, p53, MRE11, RAD1, RAD9, APE1, and VDJP protein. In some embodiments, the exonuclease is Cas4, Artemis, or TREX2.
[0242] The introduction of Cas4, Artemis, TREX2, or other similar exonucleases allows for end-processing of sticky ends before ligation occurs, thereby reducing the chance of accurate ligation and thus increasing the efficiency of mutagenesis, and competes with endogenous DNA repair enzymes to bias repair toward one of the other repair pathways (e.g., NHEJ or MMEJ), modulating mutation patterns. For example, Cas4, Artemis, or TREX2 can increase the efficiency of mutagenesis by competing with endogenous end-processing enzymes and thus promoting error-prone repair. Cas4, Artemis, or TREX2 can also facilitate HDR repair by extending single-strand overhangs. Additional roles for Cas4, Artemis, or TREX2 may include, for example, changing mutation patterns toward more desirable indels.
[0243] Method for site-specific gene insertion (ObLiGaRe 2.0) In some embodiments, the present disclosure provides a method for introducing a sequence of interest (SoI) into a chromosome in a cell based on a derivative of the ObLiGaRe method described in U.S. Patent No. 9,567,608. ObLiGaRe (obligate ligation-gated recombination) reflects the etymological meaning of the Latin verb obligare (to directly ligate). This is broadly applicable in various cell systems and provides an additional approach for genetic engineering. While U.S. Patent No. 9,567,608 uses zinc finger nucleases to target and cleave the target sequence, the present disclosure provides the use of a first Cas9-endonuclease dimer, e.g., Cas9-FokI, and a second Cas9-endonuclease dimer. The method of site-specific gene insertion described herein is informatively referred to as the abbreviation "ObLiGaRe 2.0" to distinguish it from the ObLiGaRe method described in U.S. Patent No. 9,567,608.
[0244] In some embodiments, the disclosure provides a method for introducing a sequence of interest (SoI) into a chromosome in a cell, the chromosome comprising a target sequence (TSC) comprising region 1 and region 2, the method comprising: introducing into the cell: (a) a vector comprising a target sequence (TSV), the TSV comprising region 2 and region 1 and the SoI; (b) a first Cas9-endonuclease dimer capable of generating sticky ends in the TSC, the first monomer of the first Cas9-endonuclease dimer cleaving at region 1 of the TSC and the second dimer of the first Cas9-endonuclease dimer cleaving at region 2 of the TSC; and (c) introducing a second Cas9-endonuclease dimer capable of generating sticky ends in a TSV, wherein the first monomer of the second Cas9-endonuclease dimer cleaves in region 2 of the TSV and the second monomer of the second Cas9-endonuclease dimer cleaves in region 1), wherein introduction of the (a) vector, (b) the first Cas9-endonuclease dimer, and (c) the second Cas9-endonuclease dimer results in insertion of the SoI into the chromosome of the cell.
[0245] In some embodiments, the disclosure is directed to a method of introducing a sequence of interest (SoI) into a chromosome in a cell, the chromosome comprising a target sequence (TSC) comprising region 1 and region 2, the method comprising introducing into the cell: (a) a vector comprising a target sequence (TSV), the TSV comprising region 2 and region 1 and the SoI, the vector comprising a sticky end; and (b) a first Cas9-endonuclease dimer capable of generating sticky ends in the TSC, the first monomer of the first Cas9-endonuclease dimer cleaving at region 1 of the TSC and the second monomer of the first Cas9-endonuclease dimer cleaving at region 2; wherein introduction of the vector of (a) and the first Cas9-endonuclease dimer of (b) results in insertion of the SoI into the chromosome of the cell.
[0246] The disclosed method provides efficient and precise gene targeting without using homology in the vector (or "donor plasmid"). The disclosed method provides a strategy for site-specific gene insertion using non-homologous end joining (NHEJ) or microhomology-mediated end joining (MMEJ) pathways. The design and location of the cleavage sites (i.e., Region 1 and Region 2) in the vector is sufficient to achieve precise end joining of the vector within the genomic site (i.e., Region 1 and Region 2), i.e., the target sequence (TSC) in the chromosome of the cell.
[0247] In some embodiments, the TSV is a circular vector, i.e., a plasmid. In some embodiments, the TSV is a linearized vector or linear DNA, such as a PCR product or an annealed oligonucleotide duplex with ends complementary to the TSC after cleavage. In some embodiments, the TSV comprises sticky ends. In some embodiments, the sticky ends in the TSV are generated by a Cas9-endonuclease dimer. In some embodiments, the sticky ends in the TSV are generated before the TSV is introduced into a cell. In some embodiments, the sticky ends in the TSV are generated after the TSV is introduced into a cell.
[0248] In some embodiments, the target sequence on a chromosome (TSC) comprises Region 1 and Region 2 in a 5'-3' fashion. As used herein, sequence directionality (e.g., 5'-3') refers to the direction when reading the "coding" or "sense" strand of a double-stranded DNA sequence (typically presented as the top strand of a double-stranded DNA sequence).
[0249] Figure 12 depicts an embodiment of the present disclosure. In Figure 12, the TSC is represented by the sequence in the "genomic" frame (left side), which includes Region 1 and Region 2 (part of which overlaps with Region 1) on the "coding" strand (shown as the top strand).
[0250] As shown in the "Genome" box of Figure 12, there is a first PAM sequence upstream of Region 1 (i.e., 5' with respect to the coding strand) and on the "non-coding" or "antisense" DNA strand (shown as the bottom strand). The non-coding strand includes a region that hybridizes to the first guide polynucleotide ("gRNA1"). gRNA1 hybridizes to a sequence upstream (i.e., 5' with respect to the non-coding strand) of the first PAM sequence. This gRNA1 hybridization sequence includes a portion of Region 1 and several nucleotides outside of Region 1. As indicated by the direction of the arrow, gRNA1 hybridizes to the non-coding strand of the target sequence.
[0251] As shown in the "Genome" box of Figure 12, downstream (i.e., 3' with respect to the coding strand) and on the coding strand of Region 2 is a second PAM strand. The coding strand includes a region that hybridizes to a second guide polynucleotide ("gRNA2"). gRNA2 hybridizes to a sequence upstream (i.e., 5' with respect to the coding strand) of the second PAM sequence. This gRNA2 hybridization sequence includes a portion of Region 2 and several nucleotides outside of Region 2. As indicated by the direction of the arrow, gRNA2 hybridizes to the coding strand of the target sequence.
[0252] In some embodiments, a target sequence on a vector (TSV) comprises, in a 5'-3' fashion, region 2, immediately followed by region 1, and an SoI. Figure 12 depicts an embodiment of the present disclosure. In Figure 12, the TSV is represented by the sequence in the "Vector" box (right side) and comprises region 2 on the "coding" strand, followed by region 1 (without any overlap between the two regions).
[0253] As shown in the "Vector" box in Figure 12, there is a third PAM sequence upstream (i.e., 5' with respect to the coding strand) and on the "non-coding" side of region 2. The non-coding strand includes a region that hybridizes to a third guide polynucleotide ("gRNA3"). gRNA3 hybridizes to a sequence upstream (i.e., 5' with respect to the non-coding strand) of the third PAM sequence. This gRNA3 hybridization sequence includes a portion of region 2 and several nucleotides outside of region 2. As indicated by the direction of the arrow, gRNA3 hybridizes to the non-coding strand of the target sequence.
[0254] As shown in the "Vector" box in Figure 12, there is a fourth PAM sequence downstream (i.e., 3' with respect to the coding strand) and on the coding strand of region 1. The coding strand includes a region that hybridizes to a fourth guide polynucleotide ("gRNA4"). gRNA4 hybridizes to a sequence upstream (i.e., 5' with respect to the coding strand) of the fourth PAM sequence. This gRNA4 hybridization sequence includes a portion of region 1 and several nucleotides outside of region 1. As indicated by the direction of the arrow, gRNA4 hybridizes to the coding strand of the target sequence.
[0255] Figure 14 depicts another embodiment of the present disclosure. Figure 14 is similar to Figure 14, except that there is a gap of several nucleotides between region 1 and region 2 on the TSC, and there is a gap of several nucleotides between region 2 and region 1 on the TSV. However, the arrangement of the regions relative to each other and the orientation of the guide polynucleotides are the same in Figure 14 and Figure 12.
[0256] Thus, in some embodiments, a target sequence on a chromosome (i.e., a TSC) comprises region 1 and region 2, where a portion of region 1 overlaps with a portion of region 2. In other embodiments, a TSC comprises region 1 and region 2, where region 1 and region 2 are separated by one or more nucleotides. In some embodiments, region 1 and region 2 overlap by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides. In some embodiments, region 1 and region 2 are separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides.
[0257] In some embodiments, the target sequence on the vector (i.e., the TSV) comprises region 2 and region 1, where region 2 immediately precedes region 1 without any nucleotides between them. In other embodiments, the TSV comprises region 2 and region 1, where region 2 and region 1 are separated by one or more nucleotides. In some embodiments, region 2 and region 1 are separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nucleotides.
[0258] In some embodiments of the methods, the Cas9-endonuclease dimer generates sticky ends in the target sequence. The Cas9 proteins described herein generate site-specific degradation in nucleic acids. In some embodiments, the Cas9 protein generates site-specific double-strand breaks in DNA. The ability of Cas9 to target specific sequences in nucleic acids (i.e., site specificity) is achieved by Cas9 complexing with a guide polynucleotide, e.g., a guide RNA, that hybridizes to a defined sequence. Thus, a complex comprising Cas9 and a guide polynucleotide has at least two distinct functions: (1) specific targeting of a nucleic acid sequence, and (2) nuclease activity that generates degradation at or near the targeted nucleic acid sequence. In some embodiments, the Cas9-guide polynucleotide complex is modified so that it performs only one of the two functions. In some embodiments, Cas9 is modified to remove nuclease activity but retain the ability to complex with a guide polynucleotide such that Cas9 can still target specific nucleic acid sequences.
[0259] The wild-type Cas9 described herein is a monomeric protein containing a nucleic acid binding domain (which interacts with a guide polynucleotide) and a cleavage domain (which cleaves the target nucleic acid). In certain instances, it is advantageous to achieve higher targeting specificity using a dimeric nuclease, i.e., a nuclease that is not active until both monomers of the dimer are present at the target sequence. The binding and cleavage domains of naturally occurring nucleases (e.g., Cas9) and modular binding and cleavage domains that can be fused to create nuclease-binding specific target sites are well known to those skilled in the art. For example, the binding domain of an RNA-programmable nuclease (e.g., Cas9) or a Cas9 protein with an inactive DNA cleavage domain can be used as a binding domain (e.g., to bind to a gRNA and direct binding to the target site) to specifically bind to a desired target site, and fused or conjugated to a cleavage domain, e.g., the cleavage domain of the endonuclease FokI, to create an engineered nuclease that cleaves the target region. Cas9-FokI fusion proteins are further described in, for example, U.S. Patent Application Publication No. 2015 / 0071899 and Guilinger et al., "Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification," Nature Biotechnology 32:577-582 (2014), each of which is incorporated herein by reference in its entirety.
[0260] In some embodiments, the engineered nuclease recognizes a palindromic double-stranded target site, e.g., a double-stranded DNA target site. The target sites of many naturally occurring nucleases, such as naturally occurring DNA restriction nucleases, are well known to those skilled in the art. In some embodiments, DNA nucleases, such as EcoRI, HindIII, or BamHI, recognize palindromic double-stranded DNA target sites of 4 to 10 base pairs in length and cleave each of the two DNA strands at a specific position within the target site. In some embodiments, the endonuclease cleaves the double-stranded nucleic acid target site symmetrically, i.e., cleaves both strands at the same position, resulting in ends containing base-paired nucleotides, also referred to herein as blunt ends. In some embodiments, the endonuclease cleaves the double-stranded nucleic acid target site asymmetrically, i.e., cleaves each strand at a different position, resulting in ends containing unpaired nucleotides, i.e., sticky ends or overhangs. In some embodiments, the overhang is a 5'-overhang, i.e., the unpaired nucleotide forms the 5'-end of a DNA strand. In some embodiments, the overhang is a 3'-overhang, i.e., the unpaired nucleotide forms the 3'-end of a DNA strand. The overhang can be "attached" to (i.e., bonded to) the end of another double-stranded DNA molecule that contains a complementary unpaired nucleotide.
[0261] In some embodiments, fusion proteins are provided that include two domains: (i) an RNA-programmable nuclease (e.g., a Cas9 protein, or fragment thereof) domain fused or linked to (ii) a nuclease domain. For example, in some embodiments, the Cas9 protein (e.g., the Cas9 domain of the fusion protein) comprises a nuclease-inactivated Cas9 (e.g., a Cas9 lacking DNA cleavage activity; "dCas9") that retains RNA (gRNA)-binding activity and thus can bind to a target site complementary to the gRNA. In some embodiments, the nuclease fused to the nuclease-inactivated Cas9 domain is any nuclease that requires dimerization (e.g., the joining of two nuclease monomers) to cleave a target nucleic acid (e.g., DNA). In some embodiments, the nuclease fused to the nuclease-inactivated Cas9 is a monomer of a FokI DNA cleavage domain, thereby producing a Cas9 variant designated Cas9-FokI. The FokI DNA cleavage domain is known and, in embodiments, corresponds to amino acids 388 to 583 of FokI (NCBI Accession No. J04623). In some embodiments, the FokI DNA cleavage domain corresponds to amino acids 300 to 583, 320 to 583, 340 to 583, or 360 to 583 of FokI.(See also Wah et al., "Structure of FokI has implications for DNA cleavage," Proceedings of the National Academy of Sciences USA 95(18):10564-9 (1996); Li et al., "TAL nucleases (TALNs): hybrid proteins composed of TAL effectors and FokI DNA-cleavage domain," Nucleic Acids Research 39(1):359-72 (2011); Kim et al., "Hybrid restriction enzymes: zinc finger fusions to FokI cleavage domain," Proceedings of the National Academy of Sciences USA 93:1156-1160 (1996); each of which is incorporated herein by reference in its entirety.)
[0262] In some embodiments, dimers of Cas9-endonuclease fusion proteins, such as Cas9-FokI dimers, are provided. For example, in some embodiments, the Cas9-FokI fusion protein forms a dimer with itself to mediate cleavage of a target nucleic acid. In some embodiments, the Cas9-endonuclease fusion protein, or a dimer thereof, is associated with one or more gRNAs. In some embodiments, the dimer contains two fusion proteins, each having a Cas9 domain with gRNA-binding activity, such that the target nucleic acid is targeted using two distinct gRNA sequences that exhibit complementarity to two distinct regions of the nucleic acid target. See, e.g., Figures 10 and 11. Thus, in some embodiments, cleavage of the target nucleic acid does not occur until both fusion proteins bind to the target nucleic acid (e.g., as specified by gRNA:target nucleic acid base pairing) and the nuclease domains dimerize (e.g., FokI DNA cleavage domains; as a result of their proximity upon binding of the Cas9:gRNA domains of the fusion proteins), e.g., cleaving the target nucleic acid in the region between the bound Cas9 fusion proteins. This is illustrated by the schematics shown in Figures 10 and 11. This approach represents a significant improvement over wild-type Cas9 and other Cas9 variants, such as nickases, which do not require dimerization of the nuclease domain to cleave nucleic acids (Ran et al., "Double Nicking by RNA-Guided CRISPR Cas9 for Enhanced Genome Editing Specificity," Cell 154:1380-1389 (2013); Mali et al., "CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering," Nature Biotechnology 31:833-838 (2013)).These nickase variants can induce cleavage, or nicking, upon binding of a single nickase to a nucleic acid, which can occur at on- and off-target sites, and nicking is known to induce mutagenesis. The variants provided herein require binding of two Cas9 variants in close proximity to each other to induce target nucleic acid cleavage, thereby reducing the chance of inducing off-target cleavage. In some embodiments, the Cas9 variant fused to a nuclease domain (e.g., Cas9-Fokl) has an on-target:off-target modification ratio that is at least 2-fold, at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least 110-fold, at least 120-fold, at least 130-fold, at least 140-fold, at least 150-fold, at least 175-fold, at least 200-fold, at least 250-fold or more higher than the on-target:off-target modification ratio of wild-type Cas9 or another Cas9 variant (e.g., a nickase). In some embodiments, a Cas9 variant fused to a nuclease domain (e.g., Cas9-FokI) has an on-target:off-target modification ratio that is about 60-180-fold, about 80-160-fold, about 100-150-fold, or about 120-140-fold higher than the on-target:off-target modification ratio of wild-type Cas9 or other Cas9 variants. Methods for determining the on-target:off-target modification ratio are known. In some embodiments, the on-target:off-target modification ratio is determined by measuring the number or amount of modifications at known Cas9 off-target sites in a gene. For example, Cas9 off-target sites for the CLTA, EMX, and VEGF genes are known, and modifications at these sites can be measured and compared between a test protein and a control. Target sites and their corresponding known off-target sites are amplified from genomic DNA isolated from cells (e.g., HEK293) treated with a particular Cas9 protein or variant. Modifications are then analyzed by high-throughput sequencing.Sequences containing insertions or deletions of two or more base pairs in potential genomic off-target sites and present in significantly greater numbers (p-value <0.005, Fisher's exact test) in target gRNA-treated samples relative to control gRNA-treated samples are considered Cas9 nuclease-induced genomic modifications.
[0263] In some embodiments, the methods of the present disclosure provide a Cas9-endonuclease dimer comprising a first Cas9-endonuclease monomer and a second Cas9-endonuclease monomer. In some method embodiments, the endonuclease of the Cas9-endonuclease is a Type IIS endonuclease. In some embodiments, the endonuclease of the first monomer in the first Cas9-endonuclease dimer is a Type IIS endonuclease. In some embodiments, the endonuclease of the second monomer in the first Cas9-endonuclease dimer is a Type IIS endonuclease. In some embodiments, the endonuclease of the first monomer and the second monomer in the first Cas9-endonuclease dimer are Type IIS endonucleases. In some embodiments, the endonuclease of the first monomer in the second Cas9-endonuclease dimer is a Type IIS endonuclease. In some embodiments, the endonuclease of the second monomer in the second Cas9-endonuclease dimer is a Type IIS endonuclease. In some embodiments, the endonucleases of the first and second monomers in the second Cas9-endonuclease dimer are Type IIS endonucleases. In some embodiments, the endonucleases of the first and second Cas9-endonuclease dimers are Type IIS endonucleases.
[0264] Endonucleases, or restriction enzymes, are traditionally classified into four types based on subunit composition, cleavage site, sequence specificity, and cofactor requirements. However, amino acid sequencing has revealed enormous diversity among restriction enzymes, revealing that there are many more than four distinct types at the molecular level.
[0265] "Type IIS" endonucleases, such as FokI and AlwI, cleave outside their recognition sequences on one side. Type IIS restriction enzymes are intermediate in size, 400-650 amino acids in length, and they recognize contiguous and asymmetric sequences. They contain two distinct domains, one for DNA binding and the other for DNA cleavage. They mostly bind DNA as monomers but are thought to cleave DNA cooperatively through dimerization of the cleavage domains of adjacent enzyme molecules. For this reason, some Type IIS enzymes exhibit even greater activity toward DNA molecules containing multiple recognition sites. Non-limiting examples of Type IIS endonucleases include AcuI, AlwI, BaeI, BbsI, BbvI, BccI, BceAI, BcgI, BciVI, BcoDI, BfuAI, BmrI, BpmI, BpuEI, BsaI, BsaXI, BseRI, BsgI, BsmAI, BsmBI, BsmFI, BsmI, BspCNI, BspMI, BspQI, BsrDI, BsrI, BtgZI, BtsCI, BtsI, CspCI, EarI, EciI, FauI, FokI, HgaI, HphI, HpyAV, MboII, MlyI, MmeI, MnlI, NmeAIII, PleI, SapI, and SfaNI. In some embodiments, the endonucleases in the first and second Cas9-endonuclease dimers are independently selected from the group consisting of BbvI, BgcI, BfuAI, BmpI, BspMI, CspCI, FokI, MboII, MmeI, NmeAIII, and PleI. In some embodiments, the endonuclease in the first and second Cas9-endonuclease dimers is FokI. DNA cleavage by FokI occurs only upon dimerization of two FokI monomers. FokI cleavage of DNA generates sticky ends with four-base pair overhangs.
[0266] The endonuclease in the Cas9-endonuclease fusion protein may be an engineered Fokl nuclease, e.g., an engineered Fokl dimer. In some embodiments, the engineered Fokl dimer is an obligate heterodimer, i.e., two non-identical monomers are required to form a functional (catalytically active) dimer.
[0267] In some embodiments, the first and second Cas9-endonuclease dimers are the same. In some embodiments, the first and second Cas9-endonuclease dimers are different.
[0268] In some embodiments, the method provides that the first, second, or both Cas9-endonuclease dimers comprise a modified Cas9. In some embodiments, the modified Cas9 is a catalytically inactive Cas9 ("deadCas9"). In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a catalytically inactive Cas9. Catalytically inactive Cas9s cannot cleave DNA (i.e., the cleavage domain of Cas9 is inactivated); however, they retain the ability to target nucleic acid sequences by forming a complex with a guide polynucleotide (e.g., a guide RNA). Catalytically inactive Cas9s have been described in the art, for example, by Jinek et al. (2012) and Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression," Cell 152(5):1173-1183 (2013). In some embodiments, the catalytically inactive Cas9 comprises a double amino acid substitution relative to wild-type Cas9. In some embodiments, the Cas9-endonuclease dimer comprises a double amino acid substitution relative to wild-type Cas9. In some embodiments, the double amino acid substitution is D10A and H840A. In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI, and the Cas9 in the first, second, or both Cas9-endonuclease dimers is catalytically inactive Cas9 ("deadCas9-FokI"). In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI, and the Cas9 in the first, second, or both Cas9-endonuclease dimers comprises a D10A / H840A double amino acid substitution.
[0269] In some embodiments, the modified Cas9 is a Cas9 with nickase activity ("Cas9 nickase" or "Cas9n"). In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a Cas9 with nickase activity. A Cas9 nickase can cleave only one strand of double-stranded DNA (i.e., "nick" the DNA). Cas9 nickases are described, for example, in Cho et al., "Analysis of off-target effects of CRISPR / Cas-derived RNA-guided endonucleases and nickases," Genome Research 24:132-141 (2013), Ran et al. (Cell 2013), and Mali et al. (Nature Biotechnology 2013). In some embodiments, the Cas9 nickase comprises a single amino acid substitution relative to wild-type Cas9. In some embodiments, the Cas9-endonuclease dimer comprises a single amino acid substitution relative to wild-type Cas9. In some embodiments, the single amino acid substitution is D10A ("Cas9n (D10A) In some embodiments, the single amino acid substitution is H840A ("Cas9n (H840A) "). In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI, and the Cas9 in the first, second, or both Cas9-endonuclease dimers is a Cas9 nickase. In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI, and the Cas9 in the first, second, or both Cas9-endonuclease dimers comprises a D10A single amino acid substitution ("Cas9n"). (D10A) In some embodiments, the endonuclease in the first, second, or both Cas9-endonuclease dimers is FokI, and the Cas9 in the first, second, or both Cas9-endonuclease dimers comprises a single amino acid substitution H8410A ("Cas9n (H840A) -FokI").
[0270] In some embodiments, wild-type Cas9 is effective against Streptococcus pyogenes, Staphylococcus aureus, Staphylococcus pseudintermedius, Planococcus antarcticus, Streptococcus sanguinis, Streptococcus thermophilus, Streptococcus mutans, Coribacterium glomerans, Lactobacillus farciminis, Catenibacterium mitsuokai, Lactobacillus rhamnosus, rhamnosus, Bifidobacterium bifidum, Oenococcus kitahara, Fructobacillus fructosus, Finegoldia magna, Veillonella atyipca, Solobacterium moorei, Acidaminococcus sp. D21, Eubacterium yurri, Coprococcus catus, Fusobacterium nucleatum, Filifactor allosus alocis, Peptoniphilus duerdenii, or Treponema denticola.
[0271] In some embodiments, the sticky ends generated by the Cas9 endonuclease comprise a 5' overhang. In some embodiments, the sticky ends generated by the Cas9 endonuclease comprise a 3' overhang. In some embodiments, the first, second, or both Cas9 endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 3 to 40 nucleotides. In some embodiments, the first, second, or both Cas9 endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 4 to 30 nucleotides. In some embodiments, the first, second, or both Cas9 endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of 5 to 20 nucleotides. In some embodiments, the first, second, or both Cas9 endonuclease dimers generate sticky ends comprising single-stranded polynucleotides of about 5 nucleotides, about 10 nucleotides, about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, or about 30 nucleotides. In some embodiments, the deadCas9-FokI dimer generates a sticky end that includes a 4-nucleotide 5' overhang. (D10A) The Cas9-FokI dimer generates a sticky end containing a 27-nucleotide 5' overhang. (H840A) The -FokI dimer generates sticky ends containing 23-nucleotide 3'-overhangs.
[0272] In embodiments of the methods, the sequence of interest (SoI) is carried by a donor plasmid. The donor plasmid can be of any suitable length, for example, about or at least about 10, 15, 20, 25, 50, 75, 100, 150, 200, 250, 500, or 1000 or more nucleotides in length. In some embodiments, the donor plasmid is complementary to a portion of the chromosome containing the TSC. When optimally aligned, the donor plasmid template overlaps with one or more nucleotides of the TSC (e.g., about or at least about 1, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides). In some embodiments, when the donor plasmid and the chromosome containing the TSC are optimally aligned, the nearest nucleotide of the donor plasmid is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 100, 1500, 2000, 2500, 5000, 10000 or more nucleotides of the TSC.
[0273] In some embodiments, the SoI is DNA, such as a DNA plasmid, a bacterial artificial chromosome (BAC), a yeast artificial chromosome (YAC), a viral vector, a linear piece of DNA, a PCR fragment, a naked nucleic acid, or a nucleic acid complexed with a delivery vehicle, such as a liposome.
[0274] In some embodiments, SoI is inserted into TSC using the endogenous DNA repair pathway of cells.In some embodiments, SoI is inserted into TSC using the components of non-homologous end joining (NHEJ) repair pathway.During repair process, the donor plasmid that contains SoI can be introduced into TSC.
[0275] In some embodiments, a donor plasmid containing an SoI flanked by upstream and downstream sequences is introduced into cells, where the upstream and downstream sequences share sequence similarity with either side of the site of integration in the TSC. In some embodiments, the exogenous polynucleotide containing the SoI comprises, for example, a mutant gene. In some embodiments, the exogenous polynucleotide comprises a sequence that is endogenous or exogenous to the cell. In some embodiments, the SoI comprises a polynucleotide that encodes a protein or a non-coding sequence, such as a microRNA. In some embodiments, the SoI is operably linked to a regulatory element. In some embodiments, the SoI is a regulatory element. In some embodiments, the SoI comprises a resistance cassette, e.g., a gene that confers resistance to an antibiotic. In some embodiments, the SoI comprises a mutation of a wild-type target sequence. In some embodiments, the SoI disrupts the target sequence by creating a frameshift mutation or a nucleotide substitution. In some embodiments, the SoI comprises a marker. Introduction of a marker into the target sequence can facilitate screening for targeted integration. In some embodiments, the marker is a restriction site, a fluorescent protein, or a selectable marker. In some embodiments, the SoI is introduced as a vector containing the SoI.
[0276] The upstream and downstream sequences in the exogenous polynucleotide template are selected to promote homologous recombination between the target sequence and the exogenous polynucleotide. The upstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence upstream of the target site for integration (i.e., the target sequence). Similarly, the downstream sequence is a nucleic acid sequence that shares sequence similarity with the sequence downstream of the target site for integration. Thus, in some embodiments, an exogenous polynucleotide template comprising Sol inserts into the target sequence by homologous recombination at the upstream and downstream sequences. In some embodiments, the upstream and downstream sequences in the exogenous polynucleotide template have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with the upstream and downstream sequences of the targeted genomic sequence, respectively. In some embodiments, the upstream or downstream sequence has about 20 to 2000 base pairs, or about 50 to 1750 base pairs, or about 100 to 1500 base pairs, or about 200 to 1250 base pairs, or about 300 to 1000 base pairs, or about 400 to 750 base pairs, or about 500 to 600 base pairs. In some embodiments, the upstream or downstream sequence has about 50, about 100, about 250, about 500, about 100, about 1250, about 1500, about 1750, about 2000, about 2250, or about 2500 base pairs.
[0277] In some embodiments, upon insertion of the SoI, the target sequence in the chromosome and the target sequence in the plasmid are not reconstituted. That is, in some embodiments, the resulting sequence in the chromosome (i.e., the sequence resulting from insertion of the SoI) does not hybridize to any of the first, second, third, or fourth guide polynucleotides. Thus, in some embodiments, the resulting sequence in the chromosome containing the SoI is not susceptible to cleavage by either the first or second Cas9-endonuclease dimer or the monomers in the first or second Cas9-endonuclease dimer. As illustrated in Figures 13 and 15, the resulting "knock-in" sequence (the "predicted 5' junction") is distinct from the "genomic" and "vector" sequences, and the "knock-in" sequence does not have a sequence hybridizable to any of gRNA1, gRNA2, gRNA3, or gRNA4.
[0278] In some embodiments, the disclosed methods further include introducing into the cell a first guide polynucleotide that forms a complex with the first monomer of the first Cas9-endonuclease dimer and includes a first guide sequence (the first guide sequence hybridizes to a TSC that includes region 1 but does not hybridize to the vector). As illustrated in Figures 13 and 15, the first guide sequence (shown as "gRNA1") binds to a portion of region 1 and several nucleotides outside of region 1 on the non-coding strand of target DNA in the genome. gRNA1 does not hybridize to any other sequence in the genome or vector. In some embodiments, the first guide polynucleotide forms a complex with the first monomer of the first Cas9-endonuclease dimer through interaction with the binding domain of Cas9.
[0279] In some embodiments, the disclosed methods further include introducing into the cell a second guide polynucleotide that forms a complex with the second monomer of the first Cas9-endonuclease dimer and includes a second guide sequence (the second guide sequence hybridizes to a TSC that includes region 2 but does not hybridize to the vector). As illustrated by Figures 13 and 15, the second guide sequence (shown as "gRNA2") binds to a portion of region 2 on the coding strand of target DNA in the genome. gRNA2 does not hybridize to any other sequence in the genome or vector. In some embodiments, the second guide polynucleotide forms a complex with the second monomer of the first Cas9-endonuclease dimer through interaction with the binding domain of Cas9.
[0280] In some embodiments, the disclosed methods further include introducing into the cell a third guide polynucleotide that forms a complex with the first monomer of the second Cas9-endonuclease dimer and includes a third guide sequence (the third guide sequence hybridizes to a TSV that includes region 2 but does not hybridize to the genome). As illustrated in Figures 13 and 15, the third guide sequence (shown as "gRNA3") binds to a portion of region 2 and several nucleotides outside of region 2 on the non-coding strand of the target DNA in the vector. gRNA3 does not hybridize to any other sequence in the genome or vector. In some embodiments, the third guide polynucleotide forms a complex with the first monomer of the second Cas9-endonuclease dimer through interaction with the binding domain of Cas9.
[0281] In some embodiments, the disclosed methods further include introducing into the cell a fourth guide polynucleotide that forms a complex with the second monomer of the second Cas9-endonuclease dimer and includes a fourth guide sequence (the fourth guide sequence hybridizes to a TSC that includes region 1 but does not hybridize to the genome). As illustrated by Figures 13 and 15, the fourth guide sequence (shown as "gRNA4") binds to a portion of region 1 on the coding strand of the target DNA in the vector. gRNA4 does not hybridize to any other sequence in the genome or vector. In some embodiments, the fourth guide polynucleotide forms a complex with the second monomer of the second Cas9-endonuclease dimer through interaction with the binding domain of Cas9.
[0282] In some embodiments, the guide polynucleotide can bind to both the TSC and the TSV. Thus, in some embodiments, the method further comprises introducing into the cell a first guide polynucleotide that forms a complex with a first monomer of the first Cas9-endonuclease dimer and comprises a first guide sequence, where the first guide sequence hybridizes to the TSC and the TSV.
[0283] In some embodiments, the method further comprises introducing into the cell a second guide polynucleotide that forms a complex with a second monomer of the first Cas9-endonuclease dimer and comprises a second guide sequence, wherein the second guide sequence hybridizes to a TSC and a TSV.
[0284] In some embodiments, the method further comprises introducing into the cell a third guide polynucleotide that forms a complex with the first monomer of the second Cas9-endonuclease dimer and comprises a third guide sequence, wherein the third guide sequence hybridizes to TSC and TSV.
[0285] In some embodiments, the method further comprises introducing into the cell a fourth guide polynucleotide that forms a complex with the second monomer of the second Cas9-endonuclease dimer and comprises a fourth guide sequence, wherein the fourth guide sequence hybridizes to TSC and TSV.
[0286] In some embodiments, the first, second, third, and / or fourth guide polynucleotides are the same. In some embodiments, the first, second, third, and / or fourth guide polynucleotides are different.
[0287] In some embodiments, the disclosed methods comprise introducing into a cell first, second, third, and fourth guide polynucleotides. In some embodiments, a first monomer of a first Cas9-endonuclease dimer forms a complex with the first guide polynucleotide, and a second monomer of the first Cas9-endonuclease dimer forms a complex with the second guide polynucleotide. In some embodiments, a first monomer of a second Cas9-endonuclease dimer forms a complex with the third guide polynucleotide, and a second monomer of the second Cas9-endonuclease dimer forms a complex with the fourth guide polynucleotide.
[0288] In some embodiments, the first monomer of the first Cas9-endonuclease dimer forms a complex with the first guide polynucleotide, the second monomer of the first Cas9-endonuclease dimer forms a complex with the second guide polynucleotide, the first monomer of the second Cas9-endonuclease dimer forms a complex with the third guide polynucleotide, and the second monomer of the second Cas9-endonuclease dimer forms a complex with the fourth guide polynucleotide. In some embodiments, the first and second guide polynucleotides guide the first Cas9-endonuclease dimer to a target sequence on a chromosome of the cell, and the third and fourth guide polynucleotides guide the second Cas9-endonuclease dimer to a target sequence on a vector introduced into the cell.
[0289] In some embodiments, the disclosed methods further include introducing a tracrRNA into a cell. In some embodiments, the guide polynucleotide comprises a crRNA / tracrRNA hybrid. In some embodiments, the tracrRNA component of the guide polynucleotide activates Cas9 in the Cas9 endonuclease. In some embodiments, the Cas9 endonuclease, guide polynucleotide, and tracrRNA can form a complex. In some embodiments, the complex comprises the Cas9 endonuclease, two guide polynucleotides, and two tracrRNA sequences. In some embodiments, the complex of the Cas9 endonuclease, guide polynucleotide, and tracrRNA does not occur in nature.
[0290] In some embodiments, a first monomer of a first Cas9-endonuclease dimer forms a complex with a first guide polynucleotide sequence and a tracrRNA sequence, and a second monomer of the first Cas9-endonuclease dimer forms a complex with a second guide polynucleotide sequence and a tracrRNA sequence. In some embodiments, a first monomer of a second Cas9-endonuclease dimer forms a complex with a third guide polynucleotide sequence and a tracrRNA sequence, and a second monomer of the second Cas9-endonuclease dimer forms a complex with a fourth guide polynucleotide sequence and a tracrRNA sequence.
[0291] In some embodiments, a first monomer of a first Cas9-endonuclease dimer forms a complex with a first guide polynucleotide and tracrRNA, a second monomer of the first Cas9-endonuclease dimer forms a complex with a second guide polynucleotide and tracrRNA, a first monomer of a second Cas9-endonuclease dimer forms a complex with a third guide polynucleotide and tracrRNA, and a second monomer of the second Cas9-endonuclease dimer forms a complex with a fourth guide polynucleotide and tracrRNA. In some embodiments, the first guide polynucleotide and tracrRNA and the second guide polynucleotide and tracrRNA guide the first Cas9-endonuclease dimer to a target sequence on a chromosome of a cell, and the third guide polynucleotide and tracrRNA and the fourth guide polynucleotide and tracrRNA guide the second Cas9-endonuclease dimer to a target sequence on a vector introduced into the cell.
[0292] In some embodiments of the method, the TSV, first and / or second Cas9-endonuclease dimer are introduced into the cell as polynucleotides encoding the first and second Cas9-endonuclease dimers. In some embodiments, the polynucleotides encoding the TSV, first and / or second Cas9-endonuclease dimers are codon-optimized for expression in eukaryotic cells. In some embodiments, the polynucleotides encoding the TSV, first and / or second Cas9-endonuclease dimers are codon-optimized for expression in mammalian cells. Codon optimization methods and techniques are described herein.
[0293] In some embodiments, the TSV, the first and / or second Cas9-endonuclease dimer are introduced into a cell as a single nucleic acid molecule. In some embodiments, the polynucleotides encoding the TSV, the first and / or second Cas9-endonuclease dimer are present on a single vector. In some embodiments, the polynucleotides encoding the first and / or second Cas9-endonuclease dimer, one or more guide polynucleotides, and one or more tracrRNA sequences are present on a single vector. In some embodiments, the vector is an expression vector. In some embodiments, the vector is a eukaryotic expression vector. In some embodiments, the vector is a mammalian expression vector. In some embodiments, the vector is a human expression vector. In some embodiments, the vector is a plant expression vector.
[0294] In some embodiments, the polynucleotides encoding the TSV, the first and / or second Cas9-endonuclease dimer are present on two or more vectors. In some embodiments, the polynucleotides encoding the TSV, the first and / or second Cas9-endonuclease dimer, one or more guide polynucleotides, and one or more tracrRNA sequences are present on two or more vectors. In some embodiments, the vector is an expression vector. In some embodiments, the vector is a eukaryotic expression vector. In some embodiments, the vector is a mammalian expression vector. In some embodiments, the vector is a human expression vector. In some embodiments, the vector is a plant expression vector.
[0295] In some embodiments of the methods, the cell is a eukaryotic cell. In some embodiments, the eukaryotic cell is an animal or human cell. In some embodiments, the eukaryotic cell is a human, rodent, or bovine cell line or cell strain. Examples of such cells, cell lines, or cell strains include, but are not limited to, mouse myeloma (NS0) cell lines, Chinese hamster ovary (CHO) cell lines, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK (baby hamster kidney cells), VERO, SP2 / 0, YB2 / 0, YO, C127, L cells, COS (e.g., COS1 and COS7), QC1-3, HEK-293, VERO, PER.C6, HeLA, EB1, EB2, EB3, oncolytic, or hybridoma cell lines. In some embodiments, the eukaryotic cell is a CHO cell line. In some embodiments, the eukaryotic cell is a CHO cell. In some embodiments, the cells are CHO-K1 cells, CHO-K1 SV cells, DG44 CHO cells, DUXB 11 CHO cells, CHOS, CHO GS knockout cells, CHO FUT8 GS knockout cells, CHOZN, or CHO-derived cells. CHO GS knockout cells (e.g., GSKO cells) are, for example, CHO-K1 SV GS knockout cells. CHO FUT8 knockout cells are, for example, Potelligent® CHOK1 SV (Lonza Biologics, Inc.). The eukaryotic cells may be avian cells, cell lines, or cell strains, such as EBx® cells, EB14, EB24, EB26, EB66, or EBvl3.
[0296] In some embodiments, the eukaryotic cell is a human cell. In some embodiments, the human cell is a stem cell. The stem cell can be, for example, a pluripotent stem cell, such as an embryonic stem cell, an adult stem cell, an induced pluripotent stem cell (iPSC), a tissue-specific stem cell (e.g., a hematopoietic stem cell), or a mesenchymal stem cell (MSC). In some embodiments, the human cell is a differentiated form of any of the cells described herein. In some embodiments, the eukaryotic cell is a cell derived from any primary cell in culture. In some embodiments, the cell is a stem cell or a stem cell line.
[0297] In some embodiments, the eukaryotic cells are hepatocytes, such as human hepatocytes, animal hepatocytes, or non-parenchymal cells. For example, the eukaryotic cells can be adherent metabolic test human hepatocytes, adherent induction test human hepatocytes, adherent Qualyst Transporter Certified™ human hepatocytes, suspension test human hepatocytes (e.g., 10-donor and 20-donor pooled hepatocytes), human hepatic Kupffer cells, human hepatic stellate cells, dog hepatocytes (e.g., single and pooled Beagle hepatocytes), mouse hepatocytes (e.g., CD-1 and C57BI / 6 hepatocytes), rat hepatocytes (e.g., Sprague-Dawley, Wistar Han, and Wistar hepatocytes), monkey hepatocytes (e.g., cynomolgus or rhesus monkey hepatocytes), cat hepatocytes (e.g., domestic shorthair hepatocytes), and rabbit hepatocytes (e.g., New Zealand White hepatocytes).
[0298] In some embodiments, the eukaryotic cell is a plant cell. For example, the plant cell can be from a crop plant, such as cassava, maize, sorghum, wheat, or rice. The plant cell can be from algae, a tree, or a vegetable. The plant cell can be from a monocotyledonous or dicotyledonous plant, or from a crop or cereal plant, a productive plant, a fruit, or a vegetable. For example, the plant cell may be from a tree, such as a citrus tree, e.g., an orange, grapefruit, or lemon tree; a peach or nectarine tree; an apple or pear tree; a nut tree, e.g., an almond, walnut, or pistachio tree; a Solanaceae plant, i.e., potato; a Brassica plant, a Lactuca plant, a Spinacia plant; a Capsicum plant; cotton, tobacco, asparagus, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, ginger, lettuce, spinach, strawberry, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.
[0299] In an embodiment of the method, a first Cas9-endonuclease dimer capable of generating sticky ends in a TSC and a second Cas9-endonuclease dimer capable of generating sticky ends in a TSV are introduced into a cell via a delivery particle, vesicle, or viral vector.
[0300] In some embodiments, the TSV, first and / or second Cas9-endonuclease dimers are delivered into cells via a delivery particle. Examples of delivery particles are provided herein. In some embodiments, the delivery particle is a lipid-based system, a liposome, a micelle, a microvesicle, an exosome, or a gene gun. In some embodiments, the delivery particle comprises both monomers of the Cas9-endonuclease dimer. In some embodiments, the delivery particle comprises both monomers of both Cas9-endonuclease dimers. In some embodiments, the delivery particle comprises a Cas9-endonuclease and a guide polynucleotide. In some embodiments, the delivery particle comprises a Cas9-endonuclease and a guide polynucleotide, wherein the Cas9-endonuclease and the guide polynucleotide are in a complex. In some embodiments, the delivery particle comprises a polynucleotide encoding the Cas9-endonuclease, a polynucleotide encoding the guide polynucleotide, and a polynucleotide comprising a tracrRNA. In some embodiments, the delivery particle comprises a Cas9-endonuclease, a guide polynucleotide, and a tracrRNA. In some embodiments, the delivery particle comprises a first and / or second Cas9-endonuclease dimer, a first, second, third, and / or fourth guide polynucleotide, and a tracrRNA. In some embodiments, the delivery particle comprises a polynucleotide encoding one or more Cas9-endonucleases, a polynucleotide encoding the first, second, third, and / or fourth guide polynucleotide, and a polynucleotide encoding a tracrRNA.
[0301] In some embodiments, the delivery particle further comprises a lipid, a sugar, a metal, or a protein. In some embodiments, the delivery particle is a lipid envelope. In some embodiments, the delivery particle is a sugar-based particle, such as GalNAc. In some embodiments, the delivery particle is a nanoparticle. Examples of nanoparticles are described herein. The preparation of delivery particles is further described in U.S. Patent Application Publication Nos. 2011 / 0293703, 2012 / 0251560, and 2013 / 0302401; and U.S. Patent Nos. 5,543,158, 5,855,913, 5,895,309, 6,007,845, and 8,709,843, each of which is incorporated herein by reference in its entirety.
[0302] In some embodiments, the TSV, first and / or second Cas9-endonuclease dimers are delivered into cells via vesicles. A "vesicle" is a small, fluid-chambered structure enclosed by a lipid bilayer. Examples of vesicles are provided herein. In some embodiments, the vesicle comprises both monomers of the Cas9-endonuclease dimer. In some embodiments, the vesicle comprises both monomers of both Cas9-endonuclease dimers. In some embodiments, the vesicle comprises a Cas9-endonuclease and a guide polynucleotide. In some embodiments, the vesicle comprises a Cas9-endonuclease and a guide polynucleotide, wherein the Cas9-endonuclease and the guide polynucleotide are in a complex. In some embodiments, the vesicle comprises a polynucleotide encoding the Cas9-endonuclease, a polynucleotide encoding the guide polynucleotide, and a polynucleotide comprising a tracrRNA. In some embodiments, the vesicle comprises a Cas9 endonuclease, a guide polynucleotide, and a tracrRNA. In some embodiments, the vesicle comprises a first and / or second Cas9 endonuclease dimer, a first, second, third, and / or fourth guide polynucleotide, and a tracrRNA. In some embodiments, the vesicle comprises a polynucleotide encoding one or more Cas9 endonucleases, a polynucleotide encoding the first, second, third, and / or fourth guide polynucleotide, and a polynucleotide encoding a tracrRNA.
[0303] In some embodiments, the vesicle is an exosome or a liposome. In some embodiments, the first and / or second Cas9-endonuclease dimer is delivered into cells via an exosome. Exosomes are endogenous nanovesicles (i.e., having a diameter of about 30 to about 100 nm) that transport RNA and proteins and can deliver RNA to the brain and other target organs. Genetically engineered exosomes for delivery of exogenous biological materials into target organs are described, for example, by Alvarez-Erviti et al., Nature Biotechnology 29:341 (2011), El-Andaloussi et al., Nature Protocols 7:2112-2116 (2012), and Wahlgren et al., Nucleic Acids Research 40(17):e130 (2012), each of which is incorporated herein by reference in its entirety.
[0304] In some embodiments, the TSV, first and / or second Cas9-endonuclease dimers are delivered into cells via liposomes. Liposomes are spherical vesicular structures with at least one lipid bilayer and can be used as vehicles for the administration of nutrients and pharmaceuticals. Liposomes are often composed of phospholipids, particularly phosphatidylcholine, but can also be composed of other lipids, such as egg phosphatidylethanolamine. Types of liposomes include, but are not limited to, multilamellar vesicles, small unilamellar vesicles, large unilamellar vesicles, and cochlear vesicles. See, e.g., Spuch and Navarro, "Liposomes for Targeted Delivery of Active Agents against Neurodegenerative Diseases (Alzheimer's Disease and Parkinson's Disease)," Journal of Drug Delivery 2011, Article ID 469679 (2011). Liposomes for delivery of biological materials, e.g., CRISPR-Cas components, are described, e.g., by Morrissey et al., Nature Biotechnology 23(8):1002-1007 (2005), Zimmerman et al., Nature Letters 441:111-114 (2006), and Li et al., Gene Therapy 19:775-780 (2012), each of which is incorporated herein by reference in its entirety.
[0305] In some embodiments of the method, the TSV, first and / or second Cas9-endonuclease dimer are delivered into the cell via a viral vector. In some embodiments, the viral vector comprises both monomers of the Cas9-endonuclease dimer. In some embodiments, the viral vector comprises both monomers of both Cas9-endonuclease dimers. In some embodiments, the viral vector comprises a TSV. In some embodiments, the viral vector comprises a Cas9-endonuclease and a guide polynucleotide. In some embodiments, the viral vector comprises a Cas9-endonuclease and a guide polynucleotide, wherein the Cas9-endonuclease and the guide polynucleotide are in a complex. In some embodiments, the viral vector comprises a polynucleotide encoding the Cas9-endonuclease, a polynucleotide encoding the guide polynucleotide, and a polynucleotide comprising a tracrRNA. In some embodiments, the viral vector comprises a first and / or second Cas9-endonuclease dimer, a first, second, third, and / or fourth guide polynucleotide, and a tracrRNA. In some embodiments, the viral vector comprises a polynucleotide encoding one or more Cas9-endonucleases, a polynucleotide encoding the first, second, third, and / or fourth guide polynucleotide, and a polynucleotide encoding the tracrRNA. In some embodiments, the viral vector comprises a TSV and a polynucleotide encoding one or more Cas9-endonucleases, a polynucleotide encoding the first, second, third, and / or fourth guide polynucleotide, and a polynucleotide encoding the tracrRNA.
[0306] In some embodiments, the viral vector is an adenovirus, lentivirus, or adeno-associated virus. Examples of viral vectors are provided herein. Viral transfection with adeno-associated virus (AAV) and lentivirus vectors (administration can be local, targeted, or systemic) has been used as a delivery method for in vivo gene therapy. In an embodiment of the present disclosure, the Cas protein is expressed intracellularly by the transduced cells.
[0307] In some embodiments, the first, second, or both Cas9-endonuclease dimers comprise a nuclear localization signal. In some embodiments, the first, second, or both monomers of the first Cas9-endonuclease dimer comprise a nuclear localization signal. In some embodiments, the first, second, or both monomers of the second Cas9-endonuclease dimer comprise a nuclear localization signal. In some embodiments, the first, second, or both monomers of the first, second, or both Cas9-endonuclease dimers comprise a nuclear localization signal. Nuclear localization signals ("NLS") are described herein. Exemplary nuclear localization sequences include, but are not limited to, NLSs from SV40 large T antigen, nucleoplasmin, EGL-13, c-Myc, and TUS proteins. In some embodiments, the NLS comprises the sequence PKKKRKV (SEQ ID NO: 1). In some embodiments, the NLS comprises the sequence AVKRPAATKKAGQAKKKKLD (SEQ ID NO: 2). In some embodiments, the NLS comprises the sequence PAAKRVKLD (SEQ ID NO: 3).
[0308] In some embodiments, the NLS comprises the sequence MSRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 4). In some embodiments, the NLS comprises the sequence KLKIKRPVK (SEQ ID NO: 5). Other nuclear localization sequences include, but are not limited to, the acidic M9 domain of hnRNPA1, the sequence KIPIK (SEQ ID NO: 6) in the yeast transcriptional repressor Matα2, and PY-NLS.
[0309] Seamless mutagenesis method In some embodiments, the present disclosure provides a method for seamlessly modifying one or more nucleotides in a target polynucleotide sequence in a cell. "Seamless mutagenesis" refers to site-specific mutagenesis (i.e., substitution, deletion, or insertion of one or more nucleotides) without any other flanking changes, for example, in the presence of a selection gene used to introduce the mutation. Seamless DNA engineering for mutagenesis in protein-coding regions is advantageous because any extraneous sequences introduced during the mutagenesis step may interfere with protein expression. The present disclosure provides seamless mutagenesis using a two-step selection / counterselection strategy, which first involves inserting a selection cassette, e.g., an antibiotic resistance gene and its associated counterselection gene, at the target site. The cassette is then seamlessly replaced with the desired sequence by selection against the counterselection gene, typically involving administration of a small molecule, e.g., streptomycin or sugar. Common options for counterselection markers include sacB, rpsL, and markers that can be selected for and counterselected in the appropriate host background, such as galK, thyA, and tolC.Conventional methods for seamless mutagenesis are described, for example, in Wang et al., "Improved seamless mutagenesis by recombineering using ccdB for counterselection," Nucleic Acids Research 42(5):e37 (2014); Zhang et al., "A new logic for DNA engineering using recombination in Escherichia coli," Nature Genetics 20(2):123-128 (1998); Westenberg et al., "Counter-selection recombineering of the baculovirus genome: a strategy for seamless modification of repeat-containing BACs," Nucleic Acids Research 38:e166 (2010); Wong et al., "Efficient and seamless DNA recombineering using a thymidylate synthase A selection system in Escherichia coli," Nucleic Acids Research 33:e59 (2005), each of which is incorporated herein by reference in its entirety.
[0310] In some embodiments, the disclosure provides a method for modifying one or more nucleotides in a target polynucleotide sequence in a cell, comprising: (1) introducing into the cell a vector comprising an insertion cassette (IC), the IC comprising, in a 5'-3' direction: (a) a first region that is homologous to a portion of the target polynucleotide sequence; (b) a second region that comprises one or more nucleotide mutations in the target polynucleotide sequence; (c) a first nuclease binding site; (d) a polynucleotide sequence encoding a marker gene; (e) a second nuclease binding site; (f) a third region that comprises one or more mutations in the target polynucleotide sequence; and (g) a fourth region that is homologous to a portion of the target polynucleotide sequence, wherein the first region and the fourth region are homologous to a portion of the target polynucleotide sequence. (2) inserting an IC into the target polynucleotide sequence via homologous recombination to generate a first modified target polynucleotide; (3) selecting cells that express the marker gene; (4) subjecting the first modified target polynucleotide to a site-specific nuclease to generate a second modified target polynucleotide having sticky ends; and (5) subjecting the second modified target polynucleotide having sticky ends to a ligase (the ligase ligates the sticky ends at a second region and a third region) to create a ligated modified target nucleic acid that includes one or more modified nucleotides compared to the target polynucleotide sequence.
[0311] In some embodiments, the modification of one or more nucleotides in a target polynucleotide sequence is a nucleotide substitution, i.e., a single nucleotide substitution or multiple nucleotide substitution. The modification of one or more nucleotides in a target polynucleotide sequence can result in a change in the polypeptide sequence encoded by the polynucleotide. The modification of one or more nucleotides in a target polynucleotide sequence can also result in the inactivation of expression of a downstream polynucleotide sequence in a cell. For example, the downstream sequence is inactivated, so that the sequence is not transcribed, the encoded protein is not produced, or the sequence does not function as its wild-type sequence does. In some embodiments, the target polynucleotide sequence is a regulatory sequence. In some embodiments, a regulatory sequence can be inactivated so that it no longer functions as a regulatory sequence. Examples of regulatory sequences are described herein.
[0312] The method for modifying one or more nucleotides in a target polynucleotide sequence in a cell through seamless mutagenesis utilizes an insertion cassette. In some embodiments, the insertion cassette (IC) is present on a vector. Examples of vectors are provided herein. The IC described herein can be: (i) a first region that is homologous to a portion of the target polynucleotide sequence; (ii) a second region comprising a mutation of the target polynucleotide sequence of one or more nucleotides; (iii) a first nuclease binding site; (iv) a polynucleotide sequence encoding a marker gene; (v) a second nuclease binding site; (vi) a third region comprising a mutation of the target polynucleotide sequence of one or more nucleotides; and (vii) a fourth region that is homologous to a portion of the target polynucleotide sequence, wherein the first region and the fourth region are 95% to 100% identical to their respective portions of the target polynucleotide sequence.
[0313] An exemplary IC is shown in Figure 28. In Figure 28, the IC comprises, in a 5'-3' (with respect to the "top" or "coding" strand of double-stranded DNA) orientation, a first nuclease cleavage site, a first nuclease binding site, a resistance marker, a second nuclease binding site, and a second nuclease cleavage site. The first and second nuclease cleavage sites comprise the desired nucleotide mutation within the target polynucleotide sequence.
[0314] As shown in Figure 27, "homology arms" ("HA") are present upstream of the first nuclease cleavage site and downstream of the second nuclease cleavage site. A "homology arm" comprises a region that is homologous to a portion of a target polynucleotide sequence. In some embodiments, the first region of the IC that is homologous to a portion of the target polynucleotide sequence comprises an HA upstream of the first nuclease cleavage site. In some embodiments, the fourth region of the IC that is homologous to a portion of the target polynucleotide sequence comprises an HA downstream of the second nuclease cleavage site.
[0315] In some embodiments, the IC comprises a first region that is homologous to a portion of the target polynucleotide sequence. In some embodiments, the IC comprises a fourth region that is homologous to a portion of the target polynucleotide sequence. In some embodiments, the first and fourth regions in the IC have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to their respective portions of the target polynucleotide sequence. In some embodiments, the HA of the first and fourth regions in the IC has about 10 to 5,000 base pairs, about 20 to 2,000 base pairs, about 50 to 1,750 base pairs, about 100 to 1,500 base pairs, about 200 to 1,250 base pairs, about 300 to 1,000 base pairs, about 400 to about 750 base pairs, or about 500 to 600 base pairs. In some embodiments, the HA of the first and fourth regions in the IC has about 5, about 10, about 20, about 30, about 40, about 50, about 100, about 250, about 500, about 100, about 1250, about 1500, about 1750, about 2000, about 2250, or about 2500 base pairs.
[0316] In some embodiments, the IC comprises a second region comprising a mutation of one or more nucleotides in the target polynucleotide sequence. In some embodiments, the IC comprises a third region comprising a mutation of one or more nucleotides in the target polynucleotide sequence. As shown in Figures 28 and 29, the nuclease cleavage site comprises a mutation of one or more nucleotides in the target polynucleotide sequence. In some embodiments, the nuclease cleavage site is a cleavage site for any suitable nuclease. For example, the nuclease cleavage site can be a cleavage site for a restriction enzyme, such as HindIII, BamHI, EcoRI, BbvI, FokI, MmeI, etc. In some embodiments, the second region of the IC comprises a first nuclease cleavage site comprising the desired mutation. In some embodiments, the third region of the IC comprises a second nuclease cleavage site comprising the desired mutation. In some embodiments, the second and third regions of the IC are identical or substantially identical.
[0317] In some embodiments, the IC comprises a first and a second nuclease binding site. The nuclease binding site can be the binding site of any suitable nuclease. For example, it can be the nuclease binding site of a restriction enzyme, a zinc finger nuclease, a TALEN (transcription activator-like endonuclease), or a Cas9. For example, when the nuclease is Cas9, the guide RNA can be designed to hybridize to any sequence upstream of the PAM (i.e., 5' with respect to the associated DNA strand). Thus, in some embodiments, the nuclease binding site is located upstream of the PAM. In some embodiments, the first and second nuclease binding sites are identical or substantially identical.
[0318] In some embodiments, the IC comprises a polynucleotide encoding a marker gene. A "marker" gene is used to determine whether a nucleic acid sequence has been successfully inserted into a target sequence. The marker gene can be a selectable marker (e.g., a resistance or selection marker) or a screenable marker (e.g., a fluorescent or colorimetric marker).
[0319] Non-limiting examples of resistance / selection markers include antibiotic resistance genes (e.g., ampicillin resistance genes, kanamycin resistance genes, etc.) and other antibiotic resistance genes; auxotrophic markers (e.g., URA3, HIS3) and / or other host cell selectable markers; nucleic acids to facilitate insertion into donor nucleic acids, such as transposases and inverted repeat sequences for transposition into the Mycoplasma genome; and nucleic acids to assist replication and segregation in host cells, such as autonomously replicating sequences (ARS) or centromere sequences (CEN).
[0320] Screenable markers cause cells containing the marker gene to appear differently. Non-limiting examples of screenable markers include green fluorescent protein (GFP) and its variants (e.g., yellow fluorescent protein, red fluorescent protein, etc.); β-glucuronidase, which is used in the GUS assay to detect cells by their blue staining; and X-gal, which is used in blue / white screens well known to those skilled in the art.
[0321] The method for selecting cells expressing a marker gene varies depending on the marker used. For example, when an antibiotic resistance marker is used, selection involves growing a population of cells in a culture medium containing an antibiotic and recovering surviving cells. When a screenable marker, such as GFP, is used, selection involves recovering green cells. Cell recovery can be performed, for example, by manually picking colonies from a culture plate or by sorting using a flow cytometry device, such as fluorescence-activated cell sorting (FACS).
[0322] In an embodiment of the seamless mutagenesis method, the first step of the method involves introducing a vector containing an IC into a cell. The vector can be introduced into the cell using conventional methods in the art, such as transfection, transduction, cell fusion, and lipofection. Introduction of a vector into a cell is further described herein.
[0323] In an embodiment of the method of seamless mutagenesis, the second step of the method involves inserting the IC into the target polynucleotide sequence via homologous recombination to generate a first modified target polynucleotide. As illustrated in Figure 27, a resistance cassette is inserted into the target polynucleotide sequence via homologous recombination (indicated by crossovers on either side of the "GATC" sequence). The vectors described herein for specific homologous recombination contain sufficiently long regions of homology with chromosomal sequences (i.e., the first and fourth regions in the IC) to allow complementary binding and integration of the vector into the chromosome. As described herein, longer regions of homology and greater degrees of sequence similarity can increase the efficiency of homologous recombination.
[0324] In an embodiment of the method of seamless mutagenesis, the third step of the method comprises selecting cells that express a marker gene. The selection methods for cells that express a marker gene described herein depend on a selection marker. Selection methods and various types of marker genes are described herein.
[0325] In embodiments of the method of seamless mutagenesis, the fourth step of the method comprises subjecting the first modified target polynucleotide (i.e., the first modified target polynucleotide generated from step (2) above) to a site-specific nuclease to generate a second modified target polynucleotide having a sticky end. In some embodiments, the sticky ends are present in the second and third regions of the IC. The site-specific nuclease can be any site-specific nuclease that generates sticky ends, including, but not limited to, a restriction enzyme, a Cas9 endonuclease described herein, or a stiCas9 described herein. In some embodiments, the nuclease generates double-stranded DNA breaks that include sticky ends. In some embodiments, the site-specific nuclease is exogenous to the cell, i.e., the site-specific nuclease does not naturally occur in the cell. In some embodiments, the site-specific nuclease is introduced into the cell. In some embodiments, the site-specific nuclease is introduced into the cell as a polynucleotide encoding the site-specific nuclease. Methods for introducing polynucleotides (e.g., vectors, etc.) are described herein and include, for example, transfection, transduction, cell fusion, and lipofection. In some embodiments, the site-specific nuclease is a recombinant site-specific nuclease. As described herein, recombinant proteins refer to proteins that are not native to the cell that produces them, or proteins that have sequences that result from a new combination of genetic material not known to exist in nature, such as proteins expressed from exogenous nucleic acids introduced into a cell. In some embodiments, the recombinant site-specific nuclease is expressed from a nucleic acid that is not native to the cell.
[0326] In some embodiments, the site-specific nuclease is a Cas9 effector protein. Cas9 proteins are described herein. In some embodiments, the Cas9 effector protein is a type II-B Cas9. Type II-B Cas9 proteins are described herein and can generate sticky ends. The type II-B CRISPR systems described herein are identified, inter alia, by the presence of the cas4 gene on the cas operon, and the type II-B Cas9 protein is of the TIGR03031 TIGRFAM protein family. Thus, in some embodiments, the site-specific nuclease is of the TIGR03031 TIGRFAM protein family. In some embodiments, the site-specific nuclease comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of 1E-5. In some embodiments, the site-specific nuclease comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of 1E-10. Type II-B CRISPR systems have been shown to be effective against bacterial species such as Legionella pneumophila, Francisella novicida, gamma proteobacterium HTCC5015, Parasutterella excrementihominis, Sutterella wadsworthensis, Sulfurospirillum sp. SCADC, and Ruminobacter sp.) RM87, Burkholderiales bacterium 1_1_47, Bacteroidetes oral taxon 274 strain F0058, Wolinella succinogenes, Burkholderiales bacterium YL45, Ruminobacter amylophilus, Campylobacter sp. P0111, Campylobacter sp. RM9261, Campylobacter lanienae strain RM8001, Campylobacter lanienae strain P0121, Turicimonas muris, Legionella londiniensis, Salinivibrio sharmensis, Leptospira sp. isolate FW.030, Moritella sp. isolate NORP46, Endozoicomonas sp. S-B4-1U, Tamilnaduibacter salinus, Vibrio natriegens, Arcobacter skirrowii, Francisella philomiragia, Francisella hispaniensis hispaniensis, or Parendozoicomonas haliclonae.
[0327] In some embodiments, the site-specific nuclease is a Cas9-endonuclease fusion protein. Cas9-endonuclease proteins are described herein. In some embodiments, the Cas9-endonuclease fusion protein comprises a DNA targeting domain of Cas9 and a nuclease domain of an endonuclease. In some embodiments, the endonuclease in the Cas9-endonuclease fusion protein is a type IIS endonuclease. Examples of type IIS endonucleases are provided herein and include BbvI, BgcI, BfuAI, BmpI, BspMI, CspCI, FokI, MboII, MmeI, NmeAIII, and PleI. In some embodiments, the endonuclease in the Cas9-endonuclease fusion protein is FokI. DNA cleavage by FokI occurs only upon dimerization of two FokI monomers. FokI cleavage of DNA generates sticky ends with four base pair overhangs.
[0328] In some embodiments, the Cas9-endonuclease fusion protein comprises a modified Cas9. Modified Cas9s are described herein and include catalytically inactive Cas9 and Cas9 with nickase activity. In some embodiments, the modified Cas9 is a catalytically inactive Cas9 ("deadCas9"). Catalytically inactive Cas9s cannot cleave DNA (i.e., the cleavage domain of Cas9 is inactivated); however, they retain the ability to target nucleic acid sequences by forming a complex with a guide polynucleotide (e.g., a guide RNA). Catalytically inactive Cas9s are described herein. In some embodiments, the catalytically inactive Cas9 comprises a double amino acid substitution relative to wild-type Cas9. In some embodiments, the double amino acid substitution is D10A and H840A. In some embodiments, the Cas9-endonuclease fusion protein comprises a catalytically inactive Cas9 and the endonuclease is FokI.
[0329] In some embodiments, the modified Cas9 is a Cas9 with nickase activity ("Cas9 nickase" or "Cas9n"). A Cas9 nickase can cleave only one strand of double-stranded DNA (i.e., "nicks" the DNA). Cas9 nickases are described herein. In some embodiments, the Cas9 nickase comprises a single amino acid substitution relative to wild-type Cas9. In some embodiments, the single amino acid substitution is D10A ("Cas9n"). (D10A) In some embodiments, the single amino acid substitution is H840A ("Cas9n (H840A) "). In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with nickase activity, and the endonuclease is FokI. In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with a D10A mutation, and the endonuclease is FokI. In some embodiments, the Cas9-endonuclease fusion protein comprises a Cas9 with a H840A mutation, and the endonuclease is FokI.
[0330] In some embodiments, the site-specific nuclease is Cpf1. Cpf1 (centromeric promoter factor 1) is a single RNA-guided endonuclease found in the CRISPR / Cpf1 system that can generate sticky ends. The CRISPR / Cpf1 system is similar to the CRISPR / Cas9 system. However, there are several significant differences between Cas9 and Cpf1. Cpf1 does not utilize tracrRNA. The Cpf1 protein recognizes a different PAM sequence than Cas9. The PAM sequence of Cpf1 is a 5'T-rich motif, such as 5'-TTTN-3' (where N is A, T, C, or G). Cpf1 cleaves at a different site than Cas9. While Cas9 cleaves at a sequence adjacent to the PAM, Cpf1 cleaves at a sequence further away from the PAM. The Cp1 protein is further described in, for example, foreign patent application publications GB 1506509.7, U.S. Pat. No. 9,580,701, U.S. Pat. No. 2016 / 0208243, and Zetsche et al., "Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System," Cell 163(3):759-771 (2015), each of which is incorporated herein by reference in its entirety.
[0331] In some embodiments, the site-specific nuclease is Cas9, Cpf1, or Cas9-FokI.
[0332] In some embodiments, the sticky ends generated by the site-specific nuclease comprise a 5' overhang. In some embodiments, the sticky ends generated by the site-specific nuclease comprise a 3' overhang. In some embodiments, the site-specific nuclease generates sticky ends comprising single-stranded polynucleotides of 3 to 40 nucleotides. In some embodiments, the site-specific nuclease generates sticky ends comprising single-stranded polynucleotides of 4 to 30 nucleotides. In some embodiments, the site-specific nuclease generates sticky ends comprising single-stranded polynucleotides of 5 to 20 nucleotides. In some embodiments, the site-specific nuclease generates sticky ends comprising single-stranded polynucleotides of about 5 nucleotides, about 10 nucleotides, about 15 nucleotides, about 20 nucleotides, about 25 nucleotides, or about 30 nucleotides. In some embodiments, the deadCas9-FokI dimer generates sticky ends comprising a 4-nucleotide 5' overhang. In some embodiments, the Cas9n (D10A) In some embodiments, the Cas9-FokI dimer comprises a sticky end that includes a 27-nucleotide 5' overhang. (H840A) The -FokI dimer generates sticky ends containing 23-nucleotide 3' overhangs.
[0333] In an embodiment of the method, the fifth step of the method involves subjecting a second modified target polynucleotide having sticky ends to a ligase (the ligase ligates the sticky ends at the second and third regions) to generate a ligated modified target nucleic acid containing one or more modified nucleotides compared to the target polynucleotide sequence. A ligase is an enzyme that catalyzes the joining of two or more nucleic acid fragments by forming a chemical bond. In some embodiments, the ligase joins two or more DNA fragments together by catalyzing the formation of a phosphodiester bond. Any suitable ligase can be used, and suitable ligases can be determined by one of skill in the art. Non-limiting examples of ligases include E. coli ligase, T4 DNA ligase from bacteriophage T4, DNA ligase I, DNA ligase II, DNA ligase III, DNA ligase IV, and thermostable ligases, such as Ampligase® DNA ligase. The ligase may ligate blunt or sticky ends. In some embodiments, the ligase ligates sticky ends. In some embodiments, the ligase requires ATP to ligate the DNA fragments.
[0334] In some embodiments, the ligase is exogenous to the cell, i.e., the ligase does not naturally occur in the cell. In some embodiments, the ligase is introduced into the cell. In some embodiments, the ligase is introduced into the cell as a polynucleotide encoding the ligase. Methods for introducing polynucleotides (e.g., vectors, etc.) are described herein. In some embodiments, the ligase is a recombinant ligase, i.e., a ligase expressed from a nucleic acid that is not native to the cell.
[0335] In some embodiments, the ligated modified target nucleic acid contains one or more modified nucleotides compared to the target polynucleotide sequence, but does not contain a marker gene or any additional nucleotides upstream or downstream of the target polynucleotide sequence, i.e., the target polynucleotide sequence is seamlessly mutated.
[0336] In an embodiment of the method, the first modified target nucleic acid is isolated from the cell after the third step. Methods for isolating nucleic acids from cells are well established in the art, and include, for example, phenol / chloroform extraction, precipitation under low pH / high salt conditions, and solid-phase extraction. Commercially available nucleic acid isolation kits can be used, such as QIAGEN Miniprep Kit, Bio-Rad Quantum Prep® Miniprep Kit, and Zymo Research ZYMOPURE Plasmid Miniprep Kit.
[0337] In some embodiments of the method, the first modified target nucleic acid is present in the cell after the third step, i.e., the nucleic acid is not isolated from the cell. In some embodiments, steps (1) through (5) of the method are performed within the same cell. In some embodiments, components of the method are introduced into the cell. In some embodiments, a vector comprising the insertion cassette, a site-specific nuclease, and a ligase are introduced into the cell. Methods for introducing vectors and proteins into cells are described herein, including, for example, delivery via a delivery particle, vesicle, and / or vector, e.g., a viral vector.
[0338] In embodiments of the methods, the target polynucleotide sequence is present in a plasmid. Plasmids and examples thereof are described herein. In some embodiments, the plasmid containing the target polynucleotide sequence is a native bacterial plasmid (i.e., a plasmid that occurs naturally in a bacterial cell). In some embodiments, the plasmid containing the target polynucleotide sequence is an exogenous plasmid introduced into a cell. In some embodiments, the cell is a bacterial cell. In some embodiments, the plasmid is a genetically engineered plasmid. In some embodiments, modification of one or more nucleotides in the plasmid results in modification of cellular behavior. Modification of behavior can be expression of a modified protein, higher or lower levels of expression of one or more proteins, increased resistance or sensitivity to an antibiotic, altered response to a small molecule and / or protein, altered production of a small molecule and / or protein, etc.
[0339] In embodiments of the methods, the target polynucleotide sequence is present in a chromosome. The chromosome may be a prokaryotic or eukaryotic chromosome. In some embodiments, the chromosome is in a eukaryotic cell. In some embodiments, the chromosome is in a human cell. In some embodiments, the chromosome is in an animal cell. In some embodiments, the chromosome is in a plant cell. In some embodiments, modification of one or more nucleotides in the chromosome results in modified cellular behavior. The modified behavior can be expression of a modified protein, higher or lower levels of expression of one or more proteins, increased resistance or sensitivity to antibiotics, altered response to small molecules and / or proteins, altered production of small molecules and / or proteins, etc.
[0340] Genetically engineered guide RNA (sgRNA) In some embodiments, the disclosure provides an engineered guide RNA that forms a complex with a stiCas9 protein, the engineered guide RNA comprising: (a) a guide sequence capable of hybridizing to a target sequence in a eukaryotic cell; and (b) a tracrRNA sequence capable of binding to a Cas9 protein, wherein the tracrRNA differs from a naturally occurring tracrRNA sequence by at least 10 nucleotides and improves the nuclease efficiency of the Cas9 protein.
[0341] In some embodiments, a guide polynucleotide described herein, e.g., a guide RNA, forms a complex with a Cas9 protein, i.e., in some embodiments, the guide polynucleotide binds to Cas9. In some embodiments, the DNA-binding segment of the guide polynucleotide hybridizes to a target sequence in a eukaryotic cell, but not to a sequence in a bacterial cell.
[0342] In some embodiments, the guide polynucleotide is 10 to 150 nucleotides. In some embodiments, the guide polynucleotide is 20 to 120 nucleotides. In some embodiments, the guide polynucleotide is 30 to 100 nucleotides. In some embodiments, the guide polynucleotide is 40 to 80 nucleotides. In some embodiments, the guide polynucleotide is 50 to 60 nucleotides. In some embodiments, the guide polynucleotide is 10 to 35 nucleotides. In some embodiments, the guide polynucleotide is 15 to 30 nucleotides. In some embodiments, the guide polynucleotide is 20 to 25 nucleotides.
[0343] The guide polynucleotide can be introduced into the target cell as an isolated molecule, e.g., an RNA molecule, or is introduced into the cell using an expression vector containing DNA encoding the guide polynucleotide.
[0344] Naturally occurring CRISPR systems utilize a crRNA containing a region complementary to the target sequence and a tracrRNA that binds to the Cas9 protein and also hybridizes with the crRNA. The crRNA / tracrRNA hybrid forms an RNA secondary structure that allows the crRNA portion to bind to the target sequence and the tracrRNA portion to bind to the Cas9 protein. Non-limiting examples of RNA secondary structures include helices, stem-loops, and pseudoknots. In some embodiments, the Cas9 protein recognizes at least one stem-loop in the crRNA / tracrRNA hybrid for binding.
[0345] In genetically engineered CRISPR-Cas systems, such as the CRISPR-Cas systems of the present disclosure, it may be advantageous to utilize a single guide polynucleotide that can exhibit both complementarity with a target sequence and bind to a Cas9 protein. Thus, in some embodiments, the present disclosure provides a non-naturally occurring CRISPR-Cas system that includes a Cas9 effector protein (stiCas9) that can generate sticky ends; and a guide polynucleotide that forms a complex with stiCas9 and includes a guide sequence, where the guide sequence can hybridize to a target sequence in a eukaryotic cell but not to a sequence in a bacterial cell; the complex does not occur in nature and does not include a tracrRNA. In some embodiments, the guide polynucleotide forms at least one secondary structure. In some embodiments, the at least one secondary structure is one of a stem-loop, a helix, or a pseudoknot.
[0346] It may be advantageous to optimize the genetic engineering guide polynucleotides described herein to improve binding affinity for the Cas9 protein and / or increase targeting efficiency for the target sequence. See, e.g., Dang et al., Genome Biology 16:280 (2015); Nowak et al., Nucleic Acids Res 44(20):9555-9564 (2016); and Vejnar et al., Cold Spring Harb Protoc, doi:10.110l / pdb.top090894 (2016). In some embodiments, the genetic engineering guide polynucleotide, e.g., guide RNA, is shorter than the naturally occurring crRNA and tracrRNA combination. In some embodiments, the engineered guide RNA is at least 5 nucleotides shorter, at least 6 nucleotides shorter, at least 7 nucleotides shorter, at least 8 nucleotides shorter, at least 8 nucleotides shorter, at least 9 nucleotides shorter, at least 10 nucleotides shorter, at least 11 nucleotides shorter, at least 12 nucleotides shorter, at least 13 nucleotides shorter, at least 14 nucleotides shorter, at least 15 nucleotides shorter, at least 16 nucleotides shorter, at least 17 nucleotides shorter, at least 18 nucleotides shorter, at least 19 nucleotides shorter, at least 20 nucleotides shorter, at least 21 nucleotides shorter, at least 22 nucleotides shorter, at least 23 nucleotides shorter, at least 24 nucleotides shorter, at least 25 nucleotides shorter, at least 26 nucleotides shorter, at least 27 nucleotides shorter, at least 28 nucleotides shorter, at least 29 nucleotides shorter, or at least 30 nucleotides shorter than the naturally occurring crRNA and tracrRNA combination.
[0347] In some embodiments, the tracrRNA sequence is at least 5 nucleotides shorter, at least 6 nucleotides shorter, at least 7 nucleotides shorter, at least 8 nucleotides shorter, at least 8 nucleotides shorter, at least 9 nucleotides shorter, at least 10 nucleotides shorter, at least 11 nucleotides shorter, at least 12 nucleotides shorter, at least 13 nucleotides shorter, at least 14 nucleotides shorter, at least 15 nucleotides shorter, at least 16 nucleotides shorter, at least 17 nucleotides shorter, at least 18 nucleotides shorter, at least 19 nucleotides shorter, at least 20 nucleotides shorter, at least 21 nucleotides shorter, at least 22 nucleotides shorter, at least 23 nucleotides shorter, at least 24 nucleotides shorter, at least 25 nucleotides shorter, at least 26 nucleotides shorter, at least 27 nucleotides shorter, at least 28 nucleotides shorter, at least 29 nucleotides shorter, or at least 30 nucleotides shorter than the naturally occurring tracrRNA sequence.
[0348] In some embodiments, the genetic engineering guide polynucleotide is 5 to 40 nucleotides shorter, 6 to 40 nucleotides shorter, 7 to 40 nucleotides shorter, 8 to 40 nucleotides shorter, 9 to 40 nucleotides shorter, 10 to 40 nucleotides shorter, 11 to 40 nucleotides shorter, 12 to 40 nucleotides shorter, 13 to 40 nucleotides shorter, 14 to 40 nucleotides shorter, 15 to 40 nucleotides shorter, 16 to 40 nucleotides shorter, 17 to 40 nucleotides shorter, 18 to 40 nucleotides shorter, 19 to 40 nucleotides shorter, 20 to 40 nucleotides shorter, 21 to 40 nucleotides shorter, 22 to 40 nucleotides shorter, 23 to 40 nucleotides shorter, 24 to 40 nucleotides shorter, 25 to 40 nucleotides shorter, 26 to 40 nucleotides shorter, 27 to 40 nucleotides shorter, 28 to 40 nucleotides shorter, 29 to 40 nucleotides shorter, 30 to 40 nucleotides shorter, 31 to 40 nucleotides shorter, 32 to 40 nucleotides shorter, 33 to 40 nucleotides shorter, 34 to 40 nucleotides shorter, 35 to 40 nucleotides shorter, 36 to 40 nucleotides shorter, 37 to 40 nucleotides shorter, 38 to 40 nucleotides shorter, 39 to 40 nucleotides shorter, 40 to 40 nucleotides shorter, 41 to 40 nucleotides shorter, 42 to 40 nucleotides shorter, 43 to 40 nucleotides shorter, 44 to 40 nucleotides shorter, 45 to 40 nucleotides shorter, 46 to 40 nucleotides shorter, nucleotides to 40 nucleotides shorter, 22 nucleotides to 40 nucleotides shorter, 23 nucleotides to 40 nucleotides shorter, 24 nucleotides to 40 nucleotides shorter, 25 nucleotides to 40 nucleotides shorter, 26 nucleotides to 40 nucleotides shorter, 27 nucleotides to 40 nucleotides shorter, 28 nucleotides to 40 nucleotides shorter, 29 nucleotides to 40 nucleotides shorter, 30 nucleotides to 40 nucleotides shorter, 31 nucleotides to 40 nucleotides shorter, 32 nucleotides to 40 nucleotides shorter, 33 nucleotides to 40 nucleotides shorter, 34 nucleotides to 40 nucleotides shorter, 35 nucleotides to 40 nucleotides shorter, 36 nucleotides to 40 nucleotides shorter, 37 nucleotides to 40 nucleotides shorter, 38 nucleotides to 40 nucleotides short, or 39 nucleotides to 40 nucleotides short.
[0349] In some embodiments, the engineered tracrRNA is 5 to 40 nucleotides shorter, 6 to 40 nucleotides shorter, 7 to 40 nucleotides shorter, 8 to 40 nucleotides shorter, 9 to 40 nucleotides shorter, 10 to 40 nucleotides shorter, 11 to 40 nucleotides shorter, 12 to 40 nucleotides shorter, 13 to 40 nucleotides shorter, 14 to 40 nucleotides shorter, 15 to 40 nucleotides shorter, 16 to 40 nucleotides shorter, 17 to 40 nucleotides shorter, 18 to 40 nucleotides shorter, 19 to 40 nucleotides shorter, 20 to 40 nucleotides shorter, 21 to 40 nucleotides shorter nucleotides shorter, 22 nucleotides to 40 nucleotides shorter, 23 nucleotides to 40 nucleotides shorter, 24 nucleotides to 40 nucleotides shorter, 25 nucleotides to 40 nucleotides shorter, 26 nucleotides to 40 nucleotides shorter, 27 nucleotides to 40 nucleotides shorter, 28 nucleotides to 40 nucleotides short, 29 nucleotides to 40 nucleotides short, 30 nucleotides to 40 nucleotides short, 31 nucleotides to 40 nucleotides short, 32 nucleotides to 40 nucleotides short, 33 nucleotides to 40 nucleotides short, 34 nucleotides to 40 nucleotides short, 35 nucleotides to 40 nucleotides short, 36 nucleotides to 40 nucleotides short, 37 nucleotides to 40 nucleotides short, 38 nucleotides to 40 nucleotides short, or 39 nucleotides to 40 nucleotides short.
[0350] In some embodiments, the engineered guide polynucleotide, e.g., guide RNA, is longer than the naturally occurring crRNA and tracrRNA combination. In some embodiments, the engineered guide RNA is at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides longer than the naturally occurring crRNA and tracrRNA combination.
[0351] In some embodiments, the tracrRNA sequence is at least 5 nucleotides longer, at least 6 nucleotides longer, at least 7 nucleotides longer, at least 8 nucleotides longer, at least 9 nucleotides longer, at least 10 nucleotides longer, at least 11 nucleotides longer, at least 12 nucleotides longer, at least 13 nucleotides longer, at least 14 nucleotides longer, at least 15 nucleotides longer, at least 16 nucleotides longer, at least 17 nucleotides longer, at least 18 nucleotides longer, at least 19 nucleotides longer, at least 20 nucleotides longer, at least 21 nucleotides longer, at least 22 nucleotides longer, at least 23 nucleotides longer, at least 24 nucleotides longer, at least 25 nucleotides longer, at least 26 nucleotides longer, at least 27 nucleotides longer, at least 28 nucleotides longer, at least 29 nucleotides longer, or at least 30 nucleotides longer than the naturally occurring tracrRNA sequence.
[0352] In some embodiments, the genetic engineering guide polynucleotide is 5 to 40 nucleotides longer, 6 to 40 nucleotides longer, 7 to 40 nucleotides longer, 8 to 40 nucleotides longer, 9 to 40 nucleotides longer, 10 to 40 nucleotides longer, 11 to 40 nucleotides longer, 12 to 40 nucleotides longer, 13 to 40 nucleotides longer, 14 to 40 nucleotides longer, 15 to 40 nucleotides longer, 16 to 40 nucleotides longer, 17 to 40 nucleotides longer, 18 to 40 nucleotides longer, 19 to 40 nucleotides longer, 20 to 40 nucleotides longer, 21 to 24 nucleotides longer than the naturally occurring crRNA and tracrRNA combination. nucleotides to 40 nucleotides longer, 22 nucleotides to 40 nucleotides longer, 23 nucleotides to 40 nucleotides longer, 24 nucleotides to 40 nucleotides longer, 25 nucleotides to 40 nucleotides longer, 26 nucleotides to 40 nucleotides longer, 27 nucleotides to 40 nucleotides longer, 28 nucleotides to 40 nucleotides longer, 29 nucleotides to 40 nucleotides longer, 30 nucleotides to 40 nucleotides longer, 31 nucleotides to 40 nucleotides longer, 32 nucleotides to 40 nucleotides longer, 33 nucleotides to 40 nucleotides longer, 34 nucleotides to 40 nucleotides longer, 35 nucleotides to 40 nucleotides longer, 36 nucleotides to 40 nucleotides longer, 37 nucleotides to 40 nucleotides longer, 38 nucleotides to 40 nucleotides longer, or 39 nucleotides to 40 nucleotides longer.
[0353] In some embodiments, the engineered tracrRNA is 5 to 40 nucleotides longer, 6 to 40 nucleotides longer, 7 to 40 nucleotides longer, 8 to 40 nucleotides longer, 9 to 40 nucleotides longer, 10 to 40 nucleotides longer, 11 to 40 nucleotides longer, 12 to 40 nucleotides longer, 13 to 40 nucleotides longer, 14 to 40 nucleotides longer, 15 to 40 nucleotides longer, 16 to 40 nucleotides longer, 17 to 40 nucleotides longer, 18 to 40 nucleotides longer, 19 to 40 nucleotides longer, 20 to 40 nucleotides longer, or 21 to 40 nucleotides longer than the naturally occurring tracrRNA. nucleotide longer, 22 nucleotides to 40 nucleotides longer, 23 nucleotides to 40 nucleotides longer, 24 nucleotides to 40 nucleotides longer, 25 nucleotides to 40 nucleotides longer, 26 nucleotides to 40 nucleotides longer, 27 nucleotides to 40 nucleotides longer, 28 nucleotides to 40 nucleotides longer, 29 nucleotides to 40 nucleotides longer, 30 nucleotides to 40 nucleotides longer, 31 nucleotides to 40 nucleotides longer, 32 nucleotides to 40 nucleotides longer, 33 nucleotides to 40 nucleotides longer, 34 nucleotides to 40 nucleotides longer, 35 nucleotides to 40 nucleotides longer, 36 nucleotides to 40 nucleotides longer, 37 nucleotides to 40 nucleotides longer, 38 nucleotides to 40 nucleotides longer, or 39 nucleotides to 40 nucleotides longer.
[0354] In some embodiments, the genetic engineering guide polynucleotide differs from the naturally occurring crRNA and tracrRNA combination by at least one nucleotide, such that the binding affinity and / or targeting efficiency of the genetic engineering guide polynucleotide is higher than that of the naturally occurring crRNA / tracrRNA hybrid. In some embodiments, the genetic engineering guide polynucleotide differs from the crRNA / tracrRNA hybrid by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides. In some embodiments, the engineered tracrRNA differs from the naturally occurring tracrRNA by at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 nucleotides.
[0355] In some embodiments, modifications are made to naturally occurring tracrRNA to improve the nuclease efficiency of the Cas9 protein. In some embodiments, the modifications are present in the stem-loop of the tracrRNA. In some embodiments, the modifications are extensions of the stem-loop. In some embodiments, the modifications are shortenings of the stem-loop. In some embodiments, the modifications are substitutions of one or more nucleotides in the stem-loop. In some embodiments, the modifications are to the stem-loop shown in Figure 41.
[0356] In some embodiments, the nuclease efficiency of the Cas9 protein with the engineered guide RNA is improved by at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 100%. In some embodiments, the nuclease efficiency of the Cas9 protein with the engineered guide RNA is improved by at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, or at least about 10-fold.
[0357] The nuclease efficiency of the Cas9 protein can be measured, for example, to compare the nuclease efficiency of a Cas9 protein complexed with a naturally occurring guide RNA to that of a Cas9 protein complexed with an engineered guide RNA described herein. In some embodiments, the measurement method is a biochemical assay, such as measuring the rate of in vitro Cas9 nuclease activity on a linear or circular template. In some embodiments, the measurement method is measuring the targeting efficiency of the Cas9 protein, for example, using next-generation sequencing, a T7 endonuclease I assay, and / or a CeI assay. In some embodiments, the measurement method is an affinity test between the Cas9 protein and tracrRNA, for example, using a BIACORE system.
[0358] In some embodiments, the guide sequence comprises at least 90% sequence identity to any one of SEQ ID NOs: 104-125 or 196-199. In some embodiments, the tracrRNA sequence comprises at least 90% sequence identity to any one of SEQ ID NOs: 148-171. In some embodiments, the guide RNA comprises at least 90% sequence identity to any one of SEQ ID NOs: 172-191.
[0359] In some embodiments, the engineered guide RNA, or the crRNA portion of the guide RNA, has at least 90% sequence identity to any one of SEQ ID NOs: 104-125 or 196-199. In some embodiments, the guide RNA, or the crRNA portion of the guide RNA, has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 104-125 or 196-199.
[0360] In some embodiments, the protein-binding segment of the genetic manipulation guide polynucleotide, or the tracrRNA sequence, has at least 90% sequence identity to any one of SEQ ID NOs: 102 and 148-171. In some embodiments, the protein-binding segment of the genetic manipulation guide polynucleotide has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 102 and 148-171.
[0361] In some embodiments, the disclosure provides a genetic engineering guide polynucleotide for a Cas9 protein that has at least 90% sequence identity to any one of SEQ ID NOs: 172 to 191. In some embodiments, the genetic engineering guide polynucleotide has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 172 to 191.
[0362] The guide polynucleotides described herein can be designed using bioinformatics tools with biochemical validation. An exemplary method for designing a guide polynucleotide is as follows: (1) use protein BLAST to find related CRISPR operons; (2) search for crRNAs already annotated in the genome, or annotate CRISPRs using, for example, CRISPR-Finder; (3) determine the possible location of tracrRNA using an alignment tool, for example, CLC Genomics Workbench (QIAGEN); (4) search for TATAA boxes adjacent to regions with similarity to crRNA; (5) test the secondary structures of crRNA and all possible tracrRNAs found during the alignment and select crRNA / tracrRNA hybrids that produce the desired secondary structure; (6) trim the crRNA and tracrRNA to create short guide RNAs (sgRNAs). For example, the crRNA and tracrRNA sequences described herein can be combined to generate sgRNAs. In some embodiments, the crRNA and tracrRNA sequences are combined to generate the sgRNA as shown in Table 1.
[0363] [Table 1] [Example]
[0364] Example 1 - Targeted gene insertion at the AAVS1 locus This example demonstrates gene insertion into the AAVS1 locus using seamless mutagenesis (ObLiGaRe2.0 system) as disclosed herein.
[0365] Two Cas9n-FokI variants, Cas9n D10A and Cas9n H840A were generated as shown in Figures 12 and 14. Two donor vectors containing ObLiGaRe2.0 target sites (shown as region 2 and region 1 in the figures) upstream of the SA-2A-Puro selection cassette were generated as shown in Figures 13 and 15. The size of the donor vectors was 6 kb. The ObLiGaRe2.0 target sites were designed based on the AAVS1 locus as shown in Figure 16.
[0366] A plasmid encoding one of the Cas9n-FokI variants, four separately cloned guide RNAs (gRNAs), and the corresponding donor vector were co-transfected into HEK293 cells. Genomic insertion of the puromycin resistance cassette (the gene of interest on the donor plasmid) is shown schematically in Figure 15.
[0367] Puromycin-resistant cells were selected, and genomic DNA from the puromycin-resistant cells was collected and subjected to junction PCR. The PCR products were TOPO cloned and sequenced by Sanger sequencing to determine the accuracy of the junction.
[0368] Cas9n D10A The sequence of the 5' junction for gene insertion using -FokI is shown in Figure 17. H840AThe sequence of the 5' junction for gene insertion using -FokI is shown in Figure 18. Thus, using the ObLiGaRe2.0 system, the transgene cassette was successfully knocked into the AAVS1 locus with high accuracy with respect to the predicted junction.
[0369] Example 2 - Evaluation of the efficiency of targeted insertion without antibiotic selection and the effect of spacer length on gene insertion efficiency In this example, the effect of spacer length (offset sequences of the two gRNAs) on gene insertion efficiency was tested using an experimental setting that does not require antibiotic selection.
[0370] The AAVS1-exon 2 locus was selected as the target site. The gRNAs required for targeting 10 target sites with different spacer lengths were designed and cloned as shown in Figure 19. Therefore, 10 donor vectors containing the designed ObLiGaRe2.0 target sites and mCherry (under the control of the EF1a promoter) were generated as shown in Figure 20.
[0371] Cas9n H840A A plasmid encoding Cas9n-FokI and 2AGFP, two gRNAs, and a donor vector were co-transfected into HEK293 cells. Selection was performed as follows: cells were first sorted by FACS for GFP expression, indicating the introduction of active Cas9n-FokI. Cells were then subcultured for at least 10 passages and then sorted by FACS for mCherry expression, indicating insertion of mCherry at the target site. A schematic diagram of this is shown in Figure 21.
[0372] The results for the percentage of cells with mCherry versus spacer length (in base pairs) are shown in Figure 22. A spacer length of 17 bp showed the highest efficiency of mCherry insertion (approximately 20%). Thus, high efficiency of transgene insertion was achieved using ObLiGaRe2.0 without applying antibiotic selection.
[0373] Example 3 - Comparison of the efficiency of different gene insertion methods In this example, gene insertion using ObLiGaRe (using zinc finger nucleases) and ObLiGaRe2.0 was compared.
[0374] ObLiGaRe gene insertion was used for gene insertion into the AAVS1-int1 locus. ObLiGaRe2.0, using a Cas9n-FokI variant, was used with two or four gRNAs targeting three sites in the AAVS1-int1 and SERPINA1-intron1 loci. ObLiGaRe2.0 using deadCas9-FokI was also tested. The experimental procedure was performed as described in Example 2 (no antibiotic selection, and cell selection based on FACS counting of mCherry-positive cells). The donor plasmid for the SERPINA1 locus is shown in Figure 23. Genomic insertion of the gene of interest on the donor plasmid using deadCas9-FokI is shown in Figure 24.
[0375] The results obtained for each of the gene insertion methods tested are shown in Figure 25. Results were obtained from three independent biological replicates in one experiment. Error bars indicate SEM. Zinc finger nuclease-based ObLiGaRe ("AAVS1-int-ZFN") and Cas9n at the AAVS1-int1 locus D10A The efficiency for ObLiGaRe2.0 was comparable for AAVS1-int-C9nF-A. The variation in ObLiGaRe2.0 efficiency across different loci could be attributed to the efficiency of the gRNA. Obtaining high gene insertion efficiency is achieved by evaluating combinations of target sites and different spacer lengths.
[0376] Example 4 - Seamless mutagenesis In this example, we describe the general method for seamless mutagenesis provided in this disclosure. The desired result for seamless mutagenesis is shown in Figure 26, where mutations are made at the target site without any sequence changes in the target.
[0377] Step 1 of the method is shown in Figure 27. A resistance cassette flanked by homology arms is introduced into cells carrying the target sequence and inserts into the target region by homologous recombination. Cells containing the resistance cassette are selected.
[0378] A close-up of the resistance cassette is shown in Figure 28. Nuclease cleavage and binding sites are present on either side of the resistance cassette. A nuclease capable of generating an overhang, such as Cpfl or Cas9, cleaves at the nuclease cleavage site, generating an overhang containing the desired point mutation.
[0379] Step 2 of the method is shown in Figure 29. In vitro or in vivo ligation removes the resistance cassette using the compatible overhangs generated by the nuclease. Thus, the point mutation is inserted without leaving any "trace," i.e., any extra sequence. The protocol for nucleic acid digestion and ligation is described in Example 5.
[0380] Example 5 - Protocol for seamless mutagenesis using Cpf1 In this example, nucleic acid digestion and ligation is performed as follows:
[0381] digestion 1. In an RNase-free 0.5 mL tube, 1 μL of Cas9 10x buffer 1 μL of Cpf1 protein (10 μg / μL) 1 μL gRNA RNase-free HO up to 10 μL (this amount will be determined by the amount of DNA added in step 3) Add together. 2. Incubate at room temperature for 5 minutes. 3. Add 2-2.5 μg of plasmid DNA to be cleaved (this volume will vary depending on concentration; adjust the amount of water in step 1 accordingly). 4. Incubate at 37°C for 2 hours. 5. After digestion, perform gel electrophoresis using a 1.5% agarose gel at 150V.
[0382] Gel extraction 6. DNA of the appropriate length is excised from the gel. 7. Extract the DNA from the gel using a Gel Extraction Kit (eg, from QIAGEN). 8. Measure DNA concentration on the NANODROP.
[0383] Ligation 9. In a PCR tube, 25-30 ng of plasmid DNA (this amount varies depending on the concentration) 1 μL DTT 1 μL 10x T4 ligase buffer 1 μL T4 ligase H2O up to 10 μL Add together. 10. Incubate at 16°C for 2 hours. 11. Use 10 μl for transformation.
[0384] transformation 12. Thaw NEB10 beta cells (NEW ENGLAND BIOLABS) from a -80°C freezer by placing on ice for 10 minutes. Each vial contains 50 μL (enough for 3 transformations). Thaw SOC medium. 13. Add 10 μL of the ligation reaction to a 1.5 mL EPPENDORF tube and place on ice to chill. 14. After thawing, add 15 μL of NEB10 beta cells to the ligation reaction. 15. Leave on ice for 30 minutes. Heat a water bath to 42°C. 16. Heat shock the cells by placing them in a water bath at 42°C for 30 seconds, then place them on ice for 2 minutes. 17. Add 300 μL of SOC medium to the cells and incubate at 37° C. for 45 minutes. 18. Plate 100 μL of cells onto one-third of the plate or 300 μL onto the entire plate; plates contain the appropriate antibiotic.
[0385] Example 6 - Cas9 in vitro digestion protocol In this example, in vitro digestion of substrate DNA by Cas9 is performed as follows (for a 30 μL reaction): 1. The reaction mixture is heated at room temperature in the following order: 20 μL nuclease-free water 3 µL of 10x Cas9 nuclease reaction buffer 3 μL of 300 nM sgRNA (30 nM final concentration) 1 μL of 1 μM Cas9 nuclease (approximately 30 nM final concentration) Collect with Preincubate at 25°C for 10 minutes, then 3 μL of 30 nM substrate DNA Add the following. 2. Mix thoroughly and pulse-spin in a microcentrifuge. 3. Incubate at 37°C for 15 minutes. Add 4.1 μL of Proteinase K to each sample. Mix thoroughly and pulse spin in a microcentrifuge. 5. Incubate at room temperature for 10 minutes. 6. Move to fragment analysis.
[0386] Example 7 - Analysis of DNA repair profiles after Cas9 cleavage In this example, a type II-B Cas9 operon was identified by searching for the presence of cas4 in the operon using computational analysis. The Cas9 protein from Francisella novicida (FnCas9) was selected for production. Nuclease activity was demonstrated in an in vitro cleavage assay, as shown in Figure 34A. Sanger sequencing of the cleavage products revealed that FnCas9 generates 5' sticky ends in vitro, as shown in Figure 34B. The protein expression construct was validated in the HEK293 human cell line. RIMA was used to compare the mutation patterns in FnCas9 and the Cas9 protein from Streptococcus pyogenes (SpyCas9), as shown in Figure 34C.
[0387] Example 8 - Analysis of DNA cleavage profiles after Cas9 treatment The type II-B Cas9 variant (FnCas9) from Francisella novicida was shown to form sticky ends with low editing efficiency in mammalian cells, as described in Example 7. Other members of the type II-B Cas9 family were tested for sticky end generation. A new Cas9 variant from the sequenced gut metagenome MH0245 was identified (MHCas9). The sequences of the guide RNA, tracrRNA, and crRNA designed for MHCas9 are shown in Figure 33. In vitro assays demonstrated that MHCas9 could cleave DNA fragments, as shown in Figure 35A. Sanger sequencing revealed that MHCas9 generated 5' overhangs in vitro, as shown in Figure 35B. Furthermore, a Cell 1 assay was performed to validate that MHCas9 was also functional in the HEK293-REMINDEL human cell line, as shown in Figure 35C.
[0388] The sequence of the crRNA / tracrRNA from MHCas9 is shown in Figure 36A. A schematic diagram of the crRNA / tracrRNA showing its secondary structure is shown in Figure 36B. The truncated phylogenetic tree in Figure 36C shows the alignment of MHCas9 with other type II-B Cas9s, such as Cas9 from Sulfurospirillum sp. SCADCh (ssCas9), Wolinella succinogenes (WsCas9), and Legionella pneumophila (LpCas9), as well as FnCas9. As shown by the phylogenetic tree, FnCas9 and MHCas9 exhibit considerable diversity. However, the experimental results described in Example 7 and this Example indicate that MHCas9 and FnCas9 share the same cleavage mechanism.
[0389] Example 9 - sgRNA design In this example, we describe how to design sgRNAs: 1. Use Protein BLAST (NCBI, blast.ncbi.nlm.nih.gov / Blast.cgi?PAGE=Proteins) to find related CRISPR operons. For each species that appears in the search, select one of the RefSeq for further analysis. BLAST is run several times with various inputs and different settings. 2. Identify already annotated CRISPR RNAs (crRNAs). If not, annotate the crRNAs using CRISPR-Finder (crispr.i2bc.paris-saclay.fr / Server / ). 3. Use "Create Alignment" in CLC Genomics Workbench v.9.5 (QIGEN) to find the possible location of tracrRNA. Align both strands of crRNA to the sequence between the Cas4 and CRISPR repeats. 4. Search for TATAA boxes adjacent to regions showing similarity to crRNA. 5. Examine the secondary structures of the crRNA with all possible tracrRNAs (found in the alignment) and select those that create the desired structure. 6. Trimming the crRNA and tracrRNA to generate short guide RNAs (sgRNAs).
[0390] Figures 41A-T illustrate various sgRNAs designed by the methods described herein. Figures 42A-L illustrate optimization of sgRNAs (also called "chimeric gRNAs") by trimming and potential target sites for further modification.
[0391] Example 10 - In vitro digestion assay of modified sgRNAs Four different guide RNAs were engineered by removing various nucleotides as outlined in Figure 45 (guide-1, guide-2, guide-3, guide-4). The modified guide RNAs were then compared to the original guide RNA in an in vitro digestion assay. Figure 45 demonstrates that some modifications improved the digestion efficiency of MHCas9.
[0392] Guide RNA length was further investigated in three different Cas9 systems: SpyCas9, Cl1Cas9, and MHCas9. Guide RNAs ranging in length from 19 to 23 were prepared, and the new Cas9 variants and engineered guide RNAs were then transfected into reporter cell lines and subjected to the Surveyor™ nuclease assay (Integrated DNA Technologies, Skokie, IL). Figure 46 demonstrates the in vitro cleavage efficiency and functionality of the new Cas9 variants Cl1 and MH.
[0393] Example 11 - PAM sequence for MHCas9 We investigated PAM sequences favorable to MHCas9 using the method shown schematically in Figure 49A. We generated a library of 64 pooled plasmids covering various PAM sequence combinations and target cleavage sites. The libraries were separately digested with SpCas9 and MHCas9. Forward and reverse primers for the plasmids were used to amplify the region containing the target cleavage site and PAM, and the amplified region was then sequenced by next-generation sequencing. Plasmids containing PAM sequences favorable to either SpCas9 or MHCas9 were digested and therefore not amplified or sequenced. On the other hand, plasmids containing PAM sequences unfavorable to SpCas9 or MHCas9 were not digested and could be amplified.
[0394] The results for "depleted" PAM sequences for SpCas9 and MHCas9 are shown in Figure 49B. Compared to SpCas9, MHCas9 has a less stringent preference for the "NGG" PAM sequence.
[0395] Example 12 - Combination of Cas9 protein with exonuclease Cleavage by Type II-B Cas9 proteins was combined with an end-processing exonuclease enzyme to increase editing efficiency. A schematic diagram of the method is illustrated in Figure 50. As shown in Figure 50A, the overhangs generated by Type II-B Cas9 cleavage are accurately repaired by cells to return to the original sequence, thus limiting editing efficiency when insertion-deletion or substitution modifications are desired. In Figure 50B, after Type II-B Cas9 cleavage, the end-processing exonuclease enzyme Artemis or TREX2 is introduced, which further processes the cleaved overhangs at the Type II-B Cas9 cleavage site. Cellular repair of these processed ends results in imprecise repairs to the original sequence (i.e., an increased number of insertion-deletion or substitution modifications), thereby increasing editing efficiency.
[0396] To test the effect of combining Cas9 with exonucleases, type II-B Cas9 was tested for activity in human cell lines with or without terminal processing enzymes. Figure 51A shows a schematic overview of the experimental procedure. Plasmids encoding various type II-B Cas9 proteins (FnCas9, Cl1Cas9, MHCas9) and type II-A SpCas9 were introduced into HEK293 cells along with plasmids encoding the terminal processing enzymes FnCas4 or TREX2 and plasmids encoding three different guide RNA sequences. Genomic DNA from HEK293 cells was harvested 72 hours after transfection and analyzed by next-generation sequencing.
[0397] The results are shown in Figure 51B. Cells transfected with the control plasmid showed only background levels of modifications (due to natural variation in sequencing). FnCas9, MHCas9, and SpCas9 all showed varying amounts of genome modifications in the presence or absence of end-processing enzymes. In general, introduction of Cas9 with end-processing enzymes showed an increased number of modifications compared to without end-processing enzymes.
[0398] Example 13 - Mutational pattern analysis of Cas9 proteins Mutation pattern analysis was performed for the breaks generated by different Cas9s. HEK293 cells were transfected with SpCas9, Cl1Cas9, or MHCas9 and their respective guide RNAs. After 72 hours, the cells were lysed, and genomic DNA was extracted and subjected to next-generation amplicon sequencing. Bioinformatics tools were used to analyze the sequencing reads and quantify the relative frequency of each mutation among the detected altered reads.
[0399] The results are shown in Figure 52. Figures 52A, 52B, and 52C show the mutation patterns for the same target sequence after induction of cleavage using SpCas9, Cl1Cas9, and MHCas9, respectively. The target sequence is shown at the top of each panel. These results demonstrate that mutation patterns at the same locus after induction of cleavage using different Cas9 proteins are different, indicating distinct modes of nuclease activity for the different Cas9s.
[0400] One non-limiting hypothesis for the difference in nuclease activity may be that the RuvC and HNH nuclease domain conformations differ between Type II-A and Type II-B Cas9 proteins. As illustrated in Figure 53, Type II-A Cas9 (Panel A) exhibits identical cleavage sites for its RuvC and HNH domains (e.g., approximately 3 nucleotides upstream of the NGG PAM sequence), which result in blunt ends or single-nucleotide overhangs. On the other hand, Type II-B Cas9 (Panel B) exhibits offset cleavage sites for RuvC and HNH (e.g., approximately 7 and 3 nucleotides upstream of the NGG PAM sequence, respectively), which result in "sticky" ends, i.e., 3-4 nucleotide overhangs.
Claims
1. a) a Cas9 effector protein derived from Campylobacter lanienae strain RM8001 and capable of generating sticky ends; and b) a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence; A CRISPR-Cas system comprising:
2. a) a Cas9 effector protein that comprises a domain having at least 90% identity to SEQ ID NO: 84 and is capable of generating sticky ends; and b) a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence; A CRISPR-Cas system comprising:
3. a) a nucleotide sequence encoding a Cas9 effector protein derived from Campylobacter lanienae strain RM8001 and capable of generating sticky ends; and b) a nucleotide sequence encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence. A CRISPR-Cas system comprising:
4. a) a nucleotide sequence encoding a Cas9 effector protein that comprises a domain having at least 90% identity to SEQ ID NO: 84 and is capable of generating sticky ends; and b) a nucleotide sequence encoding a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence. A CRISPR-Cas system comprising:
5. A CRISPR-Cas system as described in claim 2 or 4, wherein the Cas9 effector protein comprises a domain having at least 95% or 100% identity to SEQ ID NO:
84.
6. A CRISPR-Cas system described in any one of claims 1 to 4, wherein the guide sequence can hybridize to a target sequence in a eukaryotic cell but does not hybridize to a sequence in a bacterial cell.
7. A CRISPR-Cas system as described in claim 3 or 4, wherein the nucleotide sequences of (a) and / or (b) are under the control of a eukaryotic promoter.
8. The CRISPR-Cas system of any one of claims 1 to 4, wherein the guide polynucleotide comprises a tracrRNA sequence.
9. The CRISPR-Cas system of any one of claims 1 to 4, further comprising a separate polynucleotide comprising a tracrRNA sequence.
10. 10. The CRISPR-Cas system of claim 9, wherein the guide polynucleotide, the tracrRNA sequence, and the Cas9 effector protein are capable of forming a complex.
11. A CRISPR-Cas system as described in claim 3 or 4, wherein the nucleotide sequences (a) and / or (b) are on one or more vectors.
12. The CRISPR-Cas system described in claim 11, wherein the nucleotide sequence encoding the Cas9 effector protein is operably linked to a regulatory element.
13. The CRISPR-Cas system of claim 12, wherein the regulatory element is a eukaryotic regulatory element.
14. 5. The system of any one of claims 1 to 4, wherein the complex is capable of cleavage at a site within 10 nucleotides of a protospacer adjacent motif (PAM).
15. The system of any one of claims 1 to 4, wherein the complex is capable of cleavage at a site within 5 nucleotides of a protospacer adjacent motif (PAM).
16. 5. The system of any one of claims 1 to 4, wherein the complex is capable of cleavage at a site within 3 nucleotides of a protospacer adjacent motif (PAM).
17. The system of claim 6 , wherein the target sequence is 5′ of a protospacer adjacent motif (PAM), the PAM comprising a 3′ G-rich motif.
18. 18. The system of claim 17, wherein the PAM sequence is NGG (where N is A, C, G, or T).
19. The system according to any one of claims 1 to 4, wherein the sticky ends comprise single-stranded polynucleotide overhangs of 3 to 40 nucleotides.
20. The system according to any one of claims 1 to 4, wherein the sticky ends comprise single-stranded polynucleotide overhangs of 4 to 20 nucleotides.
21. The system according to any one of claims 1 to 4, wherein the sticky ends comprise single-stranded polynucleotide overhangs of 5 to 15 nucleotides.
22. 5. The system of any one of claims 1 to 4, wherein the Cas9 effector protein comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of 1E-5.
23. 5. The system of any one of claims 1 to 4, wherein the Cas9 effector protein comprises a domain that matches the TIGR03031 protein family using an E-value cutoff of 1E-10.
24. The system of any one of claims 1 to 4, wherein the Cas9 effector protein comprises one or more nuclear localization signals.
25. The system of claim 6 , wherein the eukaryotic cell is an animal or human cell.
26. The system of claim 6 , wherein the eukaryotic cell is a human cell.
27. The system of claim 6 , wherein the eukaryotic cell is a plant cell.
28. The system of any one of claims 1 to 4, wherein the guide sequence is linked to a direct repeat sequence.
29. A delivery particle comprising the system of any one of claims 1 to 4.
30. 30. The delivery particle of claim 29, wherein the Cas9 effector protein and the guide polynucleotide are in a complex.
31. The delivery particle of claim 30 , wherein the complex further comprises a polynucleotide comprising a tracrRNA sequence.
32. 30. The delivery particle of claim 29, further comprising a lipid, sugar, metal, or protein.
33. A vesicle comprising a system according to any one of claims 1 to 4.
34. 34. The vesicle of claim 33, wherein the Cas9 effector protein and the guide polynucleotide are in a complex.
35. 35. The vesicle of claim 34, further comprising a polynucleotide comprising a tracrRNA sequence.
36. 34. The vesicle of claim 33, wherein the vesicle is an exosome or a liposome.
37. 5. The system of claim 3 or 4, wherein the nucleotide sequence encoding the Cas9 effector protein is codon-optimized for expression in eukaryotic cells.
38. 12. The system of claim 11, wherein the nucleotide sequence encoding the Cas9 effector protein and the guide polynucleotide are present on a single vector.
39. 12. The system of claim 11, wherein the nucleotide sequence encoding the Cas9 effector protein and the guide polynucleotide are present on a single nucleic acid molecule.
40. A viral vector comprising the system of claim 3 or 4.
41. 41. The viral vector of claim 40, wherein the viral vector is an adenovirus, a lentivirus, or an adeno-associated virus.
42. A eukaryotic cell comprising a CRISPR-Cas system described in any one of claims 1 to 4.
43. Use of a system described in any one of claims 1 to 4 to provide site-specific modification of a target sequence in a eukaryotic cell.
44. 1. A method for providing site-specific modification of a target sequence in a eukaryotic cell in vitro, comprising: a) in said cells, i. a Cas9 effector protein derived from Campylobacter lanienae strain RM8001 and capable of generating sticky ends; and ii. a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence to introduce; and b) generating sticky ends in the target sequence with the Cas9 effector protein and the guide polynucleotide; and c) i. the sticky ends together, or ii. Insert the target polynucleotide sequence (SoI) into the sticky end. Ligate Including; thereby modifying said target sequence.
45. A method for providing site-specific modification of a target sequence in a eukaryotic cell in vitro, comprising: a) in said cells, i. a Cas9 effector protein capable of generating sticky ends, comprising a domain having at least 90% identity to SEQ ID NO: 84; and ii. a guide polynucleotide that forms a complex with the Cas9 effector protein and includes a guide sequence to introduce; and b) generating sticky ends in the target sequence with the Cas9 effector protein and the guide polynucleotide; and c) i. the sticky ends together, or ii. Insert the target polynucleotide sequence (SoI) into the sticky end. Ligate Including; thereby modifying said target sequence.
46. A method described in claim 44 or 45, wherein the guide sequence can hybridize to a target sequence in a eukaryotic cell but does not hybridize to a sequence in a bacterial cell.