A novel type of CRISPR / Cas system
The Type S CRISPR/Cas system addresses the limitations of existing CRISPR/Cas systems by providing a novel approach for targeted DNA modification and bacterial inhibition without dsDNA breaks, enhancing the toolbox for nucleic acid editing and cell manipulation.
Patent Information
- Application Number
- JP2025512572
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-31
- Filing Date
- 2023-08-31
- Publication Date
- 2025-08-28
AI Technical Summary
Existing CRISPR/Cas systems lack diversity and functionality for targeted nucleic acid editing and cell manipulation, particularly in environments where dsDNA breaks are undesirable.
Development of a novel Type S CRISPR/Cas system that does not rely on signature proteins of Types I-VI systems, utilizing Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5 components, which form a ribonucleoprotein complex capable of targeting DNA without inducing dsDNA breaks and can be fused with base editors or DNA nucleases for editing and cleavage.
The Type S CRISPR/Cas system provides new tools for nucleic acid editing and cell targeting, enabling specific DNA modification and inhibition of bacterial growth without causing dsDNA breaks, and can be used in environmental, medical, and other settings.
Smart Images

Figure 2025528454000051 
Figure 2025528454000052 
Figure 2025528454000053
Abstract
Description
[Technical Field]
[0001] The technology described herein relates to novel and synthetic CRISPR / Cas systems and their components, vectors, proteins, ribonucleoprotein complexes, and methods of using the new systems and components.
[0002] We refer to this new system as the "Type S" CRISPR / Cas system. [Background technology]
[0003] Currently, types I to VI of CRISPR / Cas systems are known.
[0004] It would be desirable to provide new systems and components thereof that can increase the toolbox and utility for CRISPR / Cas targeting in general and that can provide new types of functionality for these useful systems. Summary of the Invention
[0005] The present invention fulfills these needs by demonstrating a new system, Type S CRISPR / Cas. The basic components of this system do not rely on the signature genes of any of the previously characterized Type I-VI systems. Type S does not use Cas3 (the signature protein of Type I systems), Cas9 (the signature protein of Type II systems), Cas10 (the signature protein of Type III systems), Cas12 (the signature protein of Type V systems), or Cas13 (the signature protein of Type VI systems). We have described the new system as Type S.
[0006] Furthermore, we identified that the basic components (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5) do not contain RNA-guided DNA nucleases. Nevertheless, we surprisingly found that a ribonucleoprotein complex of these components and crRNA can effectively and specifically target a predetermined site on DNA. Surprisingly, we were also able to use this system to modify cells without inducing dsDNA breaks. Surprisingly, we were also able to use this system to inhibit bacterial cell growth. We fused several basic components of the system to base editors and found that base editing could be effectively performed. We fused several basic components of the system to DNA nucleases and found that DNA cleavage could be effectively performed. Using this knowledge, we can envision many applications for our new system. For example, one or more other effector proteins or domains (e.g., nucleases, base editors, or prime editors) can be provided along with one or more components of the system to generate a synthetic system capable of RNA-directed cell and nucleic acid modification. DNA binding of Type S systems does not require the presence of all basic components (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5). Also surprisingly, when fused to a base editor or DNA nuclease, the synthetic system does not require the presence of all basic components of Type S systems (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5) to effectively perform the function of the effector domain.
[0007] Thus, the system and its components provide new tools for nucleic acid editing and cell targeting in environmental, medical and other settings. For this purpose, the following is provided:
[0008] In the first configuration
[0009] In a first aspect:- A Cas-S1 protein or a nucleic acid encoding Cas-S1. A Cas-S2 protein or a nucleic acid encoding Cas-S2. A Cas-S3 protein or a nucleic acid encoding Cas-S3. A Cas-S4 protein or a nucleic acid encoding Cas-S4. A Cas-S5 protein or a nucleic acid encoding Cas-S5. One or more nucleic acids encoding Cas-S1, S2, S3, S4 and S5 proteins. One or more nucleic acids encoding the Cas-S4 and S5 proteins. One or more nucleic acids encoding Cas-S3, S4 and S5 proteins. One or more nucleic acids encoding Cas-S2, S3, S4 and S5 proteins. One or more nucleic acids encoding Cas-S1 and S2 proteins. One or more nucleic acids encoding Cas-S1, S2 and S3 proteins. One or more nucleic acids encoding Cas-S1, S2, S3 and S4 proteins. One or more nucleic acids comprising SEQ ID NOs: 7 to 11. One or more nucleic acids comprising SEQ ID NOs: 8 to 11. One or more nucleic acids comprising SEQ ID NOs: 9 to 11. One or more nucleic acids comprising SEQ ID NOs: 10 and 11. One or more nucleic acids comprising SEQ ID NOs: 7 to 10. One or more nucleic acids comprising SEQ ID NOs: 7 to 9. One or more nucleic acids comprising SEQ ID NOs: 7 and 8. A nucleic acid comprising SEQ ID NO:7. A nucleic acid comprising SEQ ID NO:8. A nucleic acid comprising SEQ ID NO:9. A nucleic acid comprising SEQ ID NO: 10. A nucleic acid comprising SEQ ID NO: 11.
[0010] In one embodiment, each of the sequences is operably linked to a heterologous promoter (i.e., a promoter not operably linked to the sequence in nature), hi one embodiment, each of the sequences is operably linked to a promoter that is a eukaryotic promoter or a viral (e.g., phage) promoter.
[0011] In a second aspect: A protein or ribonucleoprotein complex containing one, two, three, four or all of the Cas-S1, S2, S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S1, S2, S3, S4 and S5 proteins. A protein or ribonucleoprotein complex containing Cas-S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S2, S3, S4 and S5 proteins.
[0012] In the second configuration
[0013] In a first aspect:- A. (i) a nucleic acid vector or vectors, or (ii) a nucleic acid or nucleic acids, comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising amino acids that are at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising amino acids that are at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:2; c) a third nucleotide sequence encoding a polypeptide comprising amino acids that are at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:3; d) a nucleotide sequence encoding a polypeptide comprising amino acids that are at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising amino acids that are at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:5 A nucleic acid vector or vectors, or a nucleic acid or nucleic acids, comprising:
[0014] In another embodiment of the first aspect, there is provided (i) a nucleic acid vector or plurality of nucleic acid vectors, or (ii) a nucleic acid or plurality of nucleic acids, comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising amino acids at least 94, 95, 96, 97, 98 or 99% identical to SEQ ID NO:1; and / or b) a nucleotide sequence encoding a polypeptide comprising amino acids that are at least 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% identical to SEQ ID NO:2; and / or c) a nucleotide sequence encoding a polypeptide comprising amino acids at least 98 or 99% identical to SEQ ID NO:3; and / or d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence at least 99% identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide containing amino acids that are at least 98 or 99% identical to SEQ ID NO:5 Includes.
[0015] In a second aspect:- a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% identical to SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% identical to SEQ ID NO:3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% identical to SEQ ID NO: 5 A nucleic acid vector or nucleic acid comprising an expressible nucleotide sequence selected from:
[0016] In a third aspect:- (i) the nucleic acid vector or plurality of nucleic acid vectors, or (ii) the nucleic acid or plurality of nucleic acids according to the first or second aspect, wherein said polypeptide encoded by the vector is operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the complex being operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0017] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0018] In a fourth aspect:- A kit comprising: a) one or more polypeptides described herein, or one or more nucleic acids encoding such polypeptides; and b) crRNA or one or more nucleic acids encoding the crRNA), wherein the crRNA is cognate to a PAM having the sequence 5'-AAG-3'. Includes; The kit, wherein the polypeptide is operable for use with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0019] The polypeptide can include Cas-S1, S2, S3, S4, and S5. The polypeptide can be any polypeptide, protein, fusion protein, or complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0020] In a fifth aspect:- a) one or more polypeptides described herein, or one or more nucleic acids encoding such polypeptides; and b) crRNA or one or more nucleic acids encoding the crRNA), wherein the crRNA is cognate to a PAM having the sequence 5'-AAG-3'. A ribonucleoprotein complex containing
[0021] The polypeptide can be any polypeptide, protein, fusion protein, or complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0022] In a sixth aspect:- One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs: 7 to 11, wherein the nucleotide sequence is a) operably linked to a heterologous, synthetic, eukaryotic, or non-bacterial promoter; and / or b) One or more nucleic acid vectors or one or more nucleic acids contained in a cell that is a eukaryotic cell, a non-bacterial cell, or a cell that is not a bacterial cell (e.g., an E. coli, Psurdomonas, or Klebsiella cell) that contains an endogenous nucleotide sequence comprising SEQ ID NOs: 7-11.
[0023] In the third configuration
[0024] In a fifth aspect:- A fusion protein comprising a polypeptide (Px), wherein Px is a) comprises an amino acid sequence that is at least (about) 80% identical to a sequence selected from SEQ ID NOs: 1-5; and b) A fusion protein, which is fused to a heterologous polypeptide (Py).
[0025] Py can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0026] In a second aspect:- 1. An operable protein or proteins for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the protein comprising: a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 comprising one, more, or all polypeptides selected from: The protein or proteins, wherein the complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0027] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0028] In the fourth configuration
[0029] A cell, a) a eukaryotic cell containing the vector or protein in the first, second or third configuration; b) an animal, plant, insect or fungal cell containing the vector or protein of the first, second or third component; or c) A prokaryotic cell comprising the vector or protein of the first, second or third configuration, wherein the prokaryotic cell is not a bacterial cell (e.g., an E. coli, Pseudomonas or Klebsiella cell) that comprises an endogenous nucleotide sequence encoding a polypeptide of parts a) to e) set forth in the second configuration.
[0030] In the fifth configuration
[0031] In a first aspect:- crRNA or a nucleic acid encoding the crRNA or guide RNA; and a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 a composition (optionally an in vitro composition, or the composition is contained in a medical container) comprising one, more or all of: A: The crRNA comprises a spacer that is cognate to the first protospacer, and the protospacer is f) not seen in E coli; g) is a eukaryotic protospacer; or h) is a protospacer in an animal (optionally mammalian or human), plant, or fungal cell; or B: A composition which is an E. coli protospacer lacking the endogenous nucleotide sequence encoding the polypeptides of a) to e).
[0032] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0033] In a second aspect:- crRNA or a nucleic acid encoding the crRNA; and a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 a composition (optionally an in vitro composition or wherein the composition is contained in a medical container) comprising a nucleic acid encoding one, more or all of: A: The crRNA comprises a spacer that is cognate to the first protospacer, and the protospacer is f) not seen in E coli; g) is a eukaryotic protospacer; or h) is a protospacer in an animal (optionally mammalian or human), plant, or fungal cell; or B: A composition which is an E. coli protospacer lacking the endogenous nucleotide sequence encoding the polypeptides of a) to e).
[0034] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0035] In the sixth configuration
[0036] In a first aspect:- 1. A method for modifying a nucleic acid target site in a cell, comprising: a) contacting a cell with a vector, the vector comprising one or more nucleotide sequences for producing crRNA, the crRNA comprising a spacer that is homologous to a first protospacer, the protospacer being not found in E. coli; a protospacer of a eukaryotic cell; or a protospacer of an animal (optionally human), plant, insect, or fungal cell; b) (i) a vector by which the polypeptide and crRNA are expressed in the cell, or (ii) the nucleic acid of the composition encoding the cRNA (or the crRNA or guide RNA, whereby the cRNA is expressed in the cell) and the nucleic acid of the composition encoding the polypeptide, whereby the polypeptide is expressed in the cell; and allowing the introduction of the vector into the cell; c) The crRNA guide forms a complex with the polypeptide and guides the complex to the target site.
[0037] The vector and crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The polypeptide can be any polypeptide, protein, fusion protein, or complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0038] In a second aspect:- A method of treating or preventing a disease or condition mediated by a target cell in a subject, comprising performing any of the methods described herein to modify the target cell, wherein said contacting comprises administering a vector or composition to the subject, and wherein the modification treats or prevents the disease or condition.
[0039] The vectors and compositions can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0040] In a third aspect:- A method for introducing targeted editing into a target polynucleotide, comprising: a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 with a CRISPR / Cas system comprising a protein complex comprising one, more or all of the polypeptides selected from A method in which the crRNA comprises a spacer that is cognate to a first protospacer comprised in the polynucleotide, and the crRNA hybridizes to the protospacer to guide the complex, thereby causing the complex to edit the polynucleotide.
[0041] In a fourth aspect:- 1. A method of polynucleotide targeting (optionally in a cell or in vitro), comprising: a) contacting a polynucleotide with the protein and crRNA of the second aspect in a third configuration; b) allowing the formation of a ribonucleoprotein complex comprising the protein and crRNA, wherein the complex is guided to a target site contained in the polynucleotide and modifies the polynucleotide or replication thereof.
[0042] In a fifth aspect:- A method for regulating transcription or replication of a target DNA, comprising contacting the target DNA with a complex, wherein the complex lacks a DNA nuclease and the complex binds to the target DNA, thereby regulating transcription or replication of the target DNA.
[0043] In a sixth aspect:- A method for controlling the replication of a target RNA, comprising contacting the target RNA, or DNA encoding the RNA, with a complex, wherein the complex binds to the target RNA or DNA, thereby controlling transcription of the target RNA.
[0044] In a seventh aspect:- 1. A method for editing a target DNA, comprising contacting the target DNA with a complex, wherein the complex binds to the target DNA, thereby editing the target DNA.
[0045] In an eighth aspect:- A method for cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with a complex, wherein the complex comprises a nuclease, and the dsDNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM that is identical except for one base change, whereby the nuclease cleaves the DNA in a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0046] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0047] In a ninth aspect:- 1. A method for cleaving single-stranded DNA (ssDNA), comprising contacting the DNA with a complex, wherein the complex comprises a nickase, and the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM that is identical except for one base change, whereby the nuclease cleaves the single strand of the DNA in a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0048] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0049] In a tenth aspect:- 1. A method for marking or identifying a region of DNA, comprising contacting the DNA with a complex, wherein the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within the region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, and optionally, the complex comprises a detectable label.
[0050] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0051] In an eleventh aspect:- A method for modifying transcription of a region of DNA, comprising contacting the DNA with a complex, wherein the DNA contains a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the complex upregulates or downregulates transcription of the region of the DNA or an adjacent gene.
[0052] PAM is not CGG. Optionally, -PAM can be AAN, ANG, NAG; or -PAM can be AAG with no nucleotide change or one nucleotide change.
[0053] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0054] In a twelfth aspect:- 1. A method for modifying a target dsDNA in a cell without introducing dsDNA breaks, comprising: generating a complex in the cell, wherein the complex targets a target DNA contained in the cell, the target DNA containing a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, thereby modifying the DNA without introducing breaks in the DNA.
[0055] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0056] In a thirteenth aspect:- A method for inhibiting cell growth or proliferation without introducing lethal dsDNA breaks, comprising generating a complex in a cell, wherein the complex targets a target DNA contained in the cell, the target DNA containing a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, thereby inhibiting DNA replication without introducing breaks in the DNA, thereby inhibiting cell growth or proliferation.
[0057] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0058] In a fourteenth aspect:- A method of treating or preventing a disease or condition in a human, animal, plant or fungal subject, the method comprising performing a method according to any of the fifth to thirteenth aspects, wherein cells of the subject contain said DNA or RNA, and wherein the cells mediate said disease or condition.
[0059] In the seventh configuration
[0060] A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence which is at least (about) 70 or 80% identical to SEQ ID NO:12.
[0061] In the eighth configuration
[0062] In a first aspect:- one or more nucleic acids encoding a plurality of Cas proteins and comprising at least one nucleotide sequence for generating a crRNA, wherein the Cas proteins and the RNA are capable of forming a ribonucleoprotein CRISPR / Cas complex, and the RNA is capable of guiding the complex to a protospacer contained in a target DNA, wherein the 5' end of the protospacer is adjacent to a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'; a) the complex lacks DNA nucleases and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the Cas proteins do not include DinG, Cas3, and Cas10; and c) one or more nucleic acids, wherein the complex does not contain all of the Cas proteins of a type I, II, III, V, or VI CRISPR / Cas complex.
[0063] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0064] In a second aspect:- a ribonucleoprotein CRISPR / Cas complex comprising multiple Cas proteins and a crRNA, wherein the RNA is capable of guiding the complex to a protospacer contained in a target DNA, the 5' end of the protospacer being adjacent to a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'; a) the complex lacks DNA nucleases and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the complex lacks DinG, Cas3 and Cas10; and c) A ribonucleoprotein CRISPR / Cas complex, wherein the complex does not contain all of the Cas proteins of a type I, II, III, V, or VI CRISPR / Cas complex.
[0065] Alternatively, the PAM may be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0066] In a third aspect:- A cell comprising a nucleic acid or complex described herein, wherein the DNA is contained in the chromosome of the cell.
[0067] In a fourth aspect:- A cell comprising a nucleic acid or complex described herein, wherein the DNA is contained in the chromosome of the cell.
[0068] In a fifth aspect:- 1. A method for modifying DNA, comprising: a) contacting DNA with a nucleic acid or vector described herein; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA and modifies the DNA.
[0069] The crRNA and Cas protein may be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0070] In a sixth aspect:- 1. A method for inhibiting replication of a plasmid containing DNA, comprising: a) contacting DNA with a nucleic acid or vector described herein; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA to inhibit replication of the plasmid.
[0071] The crRNA and Cas protein may be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0072] In a seventh aspect:- 1. A method for inhibiting transcription of a nucleotide sequence contained in DNA, comprising: a) contacting DNA with a nucleic acid or vector described herein; b) allowing the formation of a ribonucleoprotein complex comprising the Cas protein and the crRNA, wherein the complex is guided to a target site contained in the nucleotide sequence and inhibits its transcription.
[0073] The crRNA and Cas protein may be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0074] In an eighth aspect:- 1. A method of inhibiting the growth or proliferation of a cell (optionally a prokaryotic cell, e.g., a bacterial cell) containing DNA, comprising: a) contacting DNA with a nucleic acid or vector described herein; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA to inhibit cell growth or proliferation.
[0075] The crRNA and Cas protein may be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0076] In a ninth aspect:- A method of treating or preventing a disease or condition mediated by a target cell in a human or animal subject, comprising performing any of the methods described herein to modify the target cell, wherein said contacting comprises administering a nucleic acid to the subject, and wherein the modification treats or prevents the disease or condition.
[0077] In a tenth aspect:- 10. One or more nucleic acids of the first aspect for use in a method of treating or preventing a disease or condition mediated by a target cell in a human or animal subject, the method comprising performing any of the methods described herein to modify the target cell, said contacting comprising administering a nucleic acid to the subject, wherein the modification treats or prevents the disease or condition. [Brief explanation of the drawings]
[0078] [Figure 1]Figure 1 shows the cloning of the components of the type S CRISPR system. A plasmid (p1624) was constructed containing the pSC101 origin of replication and tetracycline resistance marker, as well as DNA sequences encoding an RNA-guided endonuclease (rge), two ORFs downstream of the rge gene, a cognate CRISPR array, and four ORFs immediately upstream of the CRISPR array. ORFs 140 and 423 were predicted to encode an error-prone DNA polymerase, while ORFs 624 and 188 were predicted to encode a helicase and an RNA endonuclease, respectively. The target plasmid, p1631, was constructed by first creating a synthetic DNA fragment containing a spacer-like sequence from the type S CRISPR / Cas system separated by the predicted PAM sequence (AAG). This sequence was then inserted into a plasmid containing the p15A origin of replication, a chloramphenicol resistance marker (CmR), and the purple amilCP gene. [Figure 2] In Figure 2, the CRISPR-Cas system inhibits the growth of colonies transformed with the p1631 targeting plasmid. The bSNP3127 strain was transformed with either p1760 (right, see Figure 3 for details), p1624 (center, see Figure 1 for details), or a no-plasmid control (left). Each of the resulting strains was transformed with a 1:1 mixture of the target (purple) and non-target (white) plasmids and streaked onto LB plates supplemented with the appropriate antibiotic for selection. No purple colonies were observed in the presence of p1760. [Figure 3] Figure 3 is a schematic diagram of the plasmid library constructed upon deletion of a series of genes in p1760 (see Figure 3 for details). Boxes indicate gene deletions. Constructs that showed plasmid inhibition are indicated with an *. [Figure 4]In Figure 4, the five genes and CRISPR array comprise a Type S CRISPR-Cas system. The bSNP3127 strain, harboring some of the plasmids shown in Figure 3, is transformed with a 1:1 mixture of the targeting p1631 (purple) plasmid and a non-targeting (colorless) plasmid. Purple colonies harboring the tested plasmids were only observed in the absence of the CRISPR array. [Figure 5] In Figure 5, the CRISPR-Cas system uses all five genes (S1–S5) and the CRISPR array for plasmid inhibition. Strains b5700 (ΔCas-S1), b5702 (ΔCas-S3), b5703 (ΔCas-S4), and b5704 (ΔCas-S5) carried plasmids with single-gene deletions of type S CRISPR-Cas genes. These strains, along with strain b4816 (negative control, without the type S CRISPR-Cas system) and strain b5408 (positive control, with the complete type S CRISPR-Cas system), were transformed with either the p1631 targeting plasmid or the p144 non-targeting plasmid. The fold reduction in transformation efficiency for each strain is shown in this graphical representation. [Figure 6] In Figure 6, controlled induction of the Type S CRISPR-Cas system prevents colony growth with the p1631 targeting plasmid. [Figure 7] Figure 7 shows a plasmid clearance assay to evaluate the targeting activity of each spacer in the Type S CRISPR array. All four tested spacers show similar targeting efficiency. [Figure 8] In Figure 8, the Type S CRISPR-Cas system does not induce lethal dsDNA breaks in the E. coli chromosome. [Figure 9] Figure 9 is a schematic diagram of the location of the protospacers used for Type S CRISPR-Cas binding assays. [Figure 10-1] In Figure 10, Type S CRISPR-Cas binds to chromosomal DNA. Growth curves of bSNP5810 derivative strains in the presence of increasing chloramphenicol concentrations. [Figure 10-2] ※continuation [Figure 10-3] ※continuation [Figure 11] Figure 11 shows a graphical representation of Type S CRISPR-Cas PAM preference in a "PAM wheel" format (see Leenay et al., Technology, 62(1), 137-147, 2016, doi: https: / / doi.org / 10.1016 / j.molcel.2016.02.031). The inner ring corresponds to the third nucleotide position 5' upstream of the protospacer, the middle ring corresponds to the second nucleotide position 5' upstream of the protospacer, and the outer ring corresponds to the first nucleotide position 5' upstream of the protospacer. In the wheel, every sequence occupies a sector, the area of which is proportional to the sequence's relative enrichment. [Figure 12] Figure 12 is a graphical representation of Type S CRISPR / Cas targeting efficiency at the most preferred PAMs. The x-axis shows the most targeted PAM of the p2259 library member according to the generated NGS data. The y-axis shows the fold reduction in sequencing reads for each PAM when comparing the NGS data generated by the b5408 strain (Cas-S targeting) with the NGS data generated by the b6259 strain (control Cas-S). [Figure 13A]Figure 13 is a graphical representation of the transformation efficiency of b5408 (wt type S) by each member of the p1935 plasmid library compared with that by the negative control p2361 plasmid (no PAM). The wild-type protospacer is indicated at the top of each graph. The complementary spacer-protospacer section of each tested protospacer is indicated by a dot. The mutated protospacer section is indicated by its corresponding single-letter nucleotide code. The tested protospacers have their unique codes indicated on the left side of the graph. (A) PAM-proximal mismatch (stretch); (B) PAM-distal mismatch (stretch); (C) PAM-proximal single mismatch; (D) PAM-proximal double mismatch; (E) PAM-proximal triple mismatch; (F) PAM-proximal quadruple mismatch. Note: The "cross" symbol indicates the presence of low (single cross) or high (double cross) numbers of uncountable "microcolonies" on the corresponding transformation plate. Microcolonies likely represent inefficient Type S CRISPR / Cas targeting that does not completely shut down plasmid replication. As a result, colonies grow at a significantly slower rate on selective media. [Figure 13B] *See (B) above [Figure 13C] *See (C) above [Figure 13D] *See (D) above [Figure 13E] *See (E) above [Figure 13F] *See (F) above [Figure 14]Figure 14 is a graphical representation of the plasmids used in Examples 4 and 5. The type S variant expressed by each plasmid is shown on the left side of the figure. The name of each plasmid is written at the bottom left of each plasmid. Type S genes are presented as boxes with solid fill, and the corresponding gene name is shown above each box. Dotted lines correspond to deleted type S genes. Expression of type S genes was driven by the pBolA promoter, which is shown as an arrow indicating the direction of transcription. Expression of the crRNA module was driven by a "leader" sequence located immediately upstream (5') of the crRNA expression module, which is presented as a box filled with diagonal stripes. The CRISPR repeats (SEQ ID NO: 6) of the crRNA expression module are indicated by triangles, and the spacer is indicated by a circle. The location of the BsmBI recognition site is indicated by an asterisk (*). The tetracycline resistance marker gene, the SC101 origin of replication, and the repA101 replication protein gene are shown as boxes filled with checks, bricks, and grids, respectively. [Figure 15] Figure 15 is a schematic diagram of the locations of protospacers used in Type S CRISPRi assays for downregulation of GFP expression from b5815 strains containing a chromosomally integrated gfp gene. Each protospacer position is indicated by a triangle, and the ID of each protospacer is indicated above the triangle. [Figure 16A] Figure 16 is a graphical representation of chromosomal CRISPRi activity in the wild-type Type S line on GFP expression in strain b5815 over a 24-hour period. Target protospacers were placed at A) the beginning of the gfp orf and the p70a promoter region (protospacers shown in parentheses: S1F, S1R, S2F, S2R, S3F, S3R), B) the end of the gfp orf and 200 bp regions upstream and downstream of gfp (protospacers shown in parentheses: S4F, S4R, S5F, S6F, S6R), and C) 1 kb upstream (protospacers shown in parentheses: U1kF, U1kR) and 2 kb upstream (protospacers shown in parentheses: U2kF, U2kR) of the gfp orf. [Figure 16B] *See B) above [Figure 16C] *See C) above [Figure 17A] Figure 17 is a graphical representation of chromosomal CRISPRi activity of the ΔcasS3 type S system on GFP expression in strain b5815 over a 24-hour period. Target protospacers were placed at A) the beginning of the gfp orf and the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R), and B) the end of the gfp orf and 200 bp regions upstream and downstream of the gfp orf (protospacers S4F, S4R, S5F, S6F, S6R). [Figure 17B] *See B) above [Figure 18A] Figure 18 is a graphical representation of chromosomal CRISPRi activity of A) ΔcasS3 type S line, B) ΔcasS1ΔcasS3 type S line, and C) ΔcasS1 type S line on GFP expression in strain b5815 over 24 hours. Target protospacers were placed at the start of the gfp orf and the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R). Plasmids used to transform strain b5815 are indicated by (:). [Figure 18B] *See B) above [Figure 18C] *See C) above [Figure 19A] Figure 19 is a graphical representation of CRISPRi activity on the plasmids of A) ΔcasS3 Type S line, B) ΔcasS4 Type S line, C) ΔcasS4 Type S line, and D) ΔcasS1ΔcasS3 Type S line for GFP expression in strain b6386 over 24 hours. Target protospacers were placed at the start of the gfp orf and the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R). [Figure 19B] *See B) above [Figure 19C] *See C) above [Figure 19D] *See D) above. [Figure 20]Figure 20 is a schematic diagram of the plasmid combinations used for A) CasS1, B) CasS1(ΔCasS3), C) CasS3, and D) CasS4 base editing assays. An asterisk (*) indicates that the spacer is directed to one of the following: non-targeting control, S1F, S1R, S2F, S2R, S3F, or S3R. [Figure 21A] Figure 21 shows a heatmap visualization of representative sequencing results from a type S CRISPR / Cas base editing assay. PmCDA1 cytidine deaminase is fused to the C-terminus of A) the Cas-S1 subunit of a wild-type type S CRISPR / Cas system, B) the Cas-S1 subunit of a ΔCas-S3 type S CRISPR / Cas system, C) the Cas-S3 subunit of a wild-type type S CRISPR / Cas system, and D) the Cas-S2 subunit of a wild-type type S CRISPR / Cas system. Above each box, i) the code for the protospacer tested and ii) the nucleotide sequence of the protospacer (highlighted with a boxed arrow further indicating the orientation of the protospacer) are shown, along with the sequence of the immediately adjacent genomic region where the base-editing modification was detected. Below each box, i) the code of the unique Sanger sequencing result (each derived from a separate colony) and ii) all C to T modifications or G to A modifications detected in the sequencing results are shown. The color intensity of each heatmap cell represents the relative C to T (or G to A) abundance at each position according to the ratio of the heights of the corresponding C and T peaks from the Sanger sequencing results. For example, a black box represents a colony with 100% C to T mutation genotype for the corresponding position, whereas an off-white box corresponds to a mixed C to T mutation genotype of 95%–5% (a reliable detection limit for Sanger sequencing) for the corresponding position. [Figure 21B] *See B) above [Figure 21C-1] *See C) above [Figure 21C-2] *See C) above (continued) [Figure 21C-3] *See C) above (continued) [Figure 21C-4] *See C) above (continued) [Figure 21D] *See D) above. [Figure 22] Figure 22 is a schematic diagram of the plasmid combinations used in the TevCasS plasmid targeting assay. [Figure 23] Figure 23 is a graphical representation of results from a TevCasS plasmid targeting assay. The x-axis corresponds to the reduction in transformation efficiency of the MG1655 strain expressing TevCasS1 and targeting the target sequence on the TevCasS targeting plasmid relative to the transformation efficiency of the MG1655 strain expressing TevCasS1 and not targeting the target sequence on the TevCasS targeting plasmid. The y-axis corresponds to the size of the I-TevI spacer sequence in the target sequence of the TevCasS targeting plasmid. [Figure 24] Figure 24 is a graphical representation of results from a TevCasS chromosomal targeting assay. The x-axis corresponds to the reduction in transformation efficiency of the MG1655 strain expressing TevCasS1 and targeting the chromosomal TevCasS target sequence relative to the transformation efficiency of the MG1655 strain expressing TevCasS1 and not targeting the chromosomal TevCasS target sequence. The y-axis corresponds to the size of the I-TevI spacer sequence in the target sequence of the chromosomal TevCasS target sequence. [Figure 25] Figure 25 is a graphical representation of TevCasS target cleavage site selection in a "chronaplot" format. The sequence of the TevCasS target cleavage site is 5'CN1N2N3G-3'. The inner circle corresponds to the N1 nucleotide, the middle circle corresponds to the N2 nucleotide, and the outer circle corresponds to N3. In the wheel, every sequence occupies a sector, the area of which is proportional to the relative enrichment of the sequence. DETAILED DESCRIPTION OF THE INVENTION
[0079] The technology described herein relates to novel and synthetic CRISPR / Cas systems and their components, vectors, proteins, ribonucleoprotein complexes, and methods of using the new systems and components.
[0080] We call our new system "CRISPR-S," which we refer to as a "Type S CRISPR / Cas" system.
[0081] The basic components of this system do not rely on the signature genes of any of the previously characterized types I-III, V, or VI systems. Type S does not use Cas3 (the signature protein of type I systems), Cas9 (the signature protein of type II systems), Cas10 (the signature protein of type III systems), Cas12 (the signature protein of type V systems), or Cas13 (the signature protein of type VI systems). The basic components do not include the IscB, IsrB, or IshB proteins (IscB and IsrB share a common evolutionary history with Cas9). We refer to the new system as Type S. Thus, in one embodiment, the vectors herein do not encode Cas3, 9, 10, 12, and / or 13. The nucleic acids herein may not encode Cas3, 9, 10, 12, and / or 13. The complexes herein may not include Cas3, 9, 10, 12, and / or 13. The methods herein may not use Cas3, 9, 10, 12, and / or 13. In one embodiment, the vectors or nucleic acids herein do not encode TnpB. In one embodiment, the complexes described herein do not include TnpB. In one embodiment, the vectors or nucleic acids herein do not encode an HD-nuclease domain. In one embodiment, the complexes herein do not include an HD-nuclease domain. In one example, the vectors or nucleic acids herein encode Cas8. In one example, the complexes herein include Cas8.
[0082] The complex may comprise a protein for PAM specification, a protein for interacting with Cas8, a protein for processing the pre-crRNA, and a scaffolding protein, the complex comprising Cas-S1 to Cas-S5 proteins.The complex may comprise a protein for PAM specification, a protein for interacting with Cas8, and a scaffolding protein, the complex comprising Cas-S1, Cas-S2, Cas-S4, and Cas-S5 proteins.
[0083] In one example, the vector or nucleic acid comprises a CRISPR array for generating crRNA. The array may comprise one or more repeat sequences and one or more spacer sequences. The spacer sequence may be complementary to the first protospacer sequence (as defined elsewhere herein). In some embodiments, the spacer sequence of the array may be complementary to the second, third, fourth, etc. protospacer in the target sequence, e.g., when multiple editing or modification is desired. For example, the repeat sequence is SEQ ID NO: 6 or SEQ ID NO: 70, or a sequence identical except for 1, 2, 3, 4, or 5 changes, particularly 1 or 2 changes, e.g., 1 change.
[0084] The vectors or nucleic acids herein may comprise a Type S array, i.e., an array encoding a crRNA operable with at least Cas-S1, 2, 3, 4, and 5 to target the first protospacer (or, e.g., the first and second protospacers) of a nucleic acid or target sequence. The vectors or nucleic acids herein may comprise a Type S array, i.e., an array encoding a crRNA operable with at least Cas-S1, 2, 4, and 5 to target the first protospacer (or, e.g., the first and second protospacers) of a nucleic acid or target sequence. The vectors or nucleic acids herein may comprise a Type S array, i.e., an array encoding a crRNA operable with at least Cas-S2, 3, 4, and 5 to target the first protospacer (or, e.g., the first and second protospacers) of a nucleic acid or target sequence. The vector or nucleic acid herein may comprise a Type S array, i.e., an array encoding a crRNA operable with at least Cas-S2, 4 and 5 to target the first protospacer (or, e.g., the first and second protospacers) of a nucleic acid or target sequence.
[0085] The protospacer is downstream of the 5'-AAG-3'PAM in the nucleic acid or target sequence. - PAM can be AAN, ANG, NAG; or The -PAM can be either no nucleotide change or one AAG.
[0086] PAM is not CGG.
[0087] Alternatively, the PAM can be AHN, KAG, AGG, GAC, or GTG.
[0088] Alternatively, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3', such as 5'-AAC-3' and 5'-ATG-3'. In particular, the PAM is 5'-AAC-3'.
[0089] Importantly, the inventors identified that the components of the Type S system (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5) do not contain RNA-guided DNA nucleases. Nevertheless, the inventors surprisingly found that ribonucleoprotein complexes of these components and crRNA can effectively and specifically target predetermined sites on DNA. Surprisingly, this system could also be used to modify cells without inducing dsDNA breaks. See Example 1.3.4 below. Using this knowledge, the inventors can envision many applications of the new system. For example, one or more proteins or domains with effector activity (e.g., nuclease (see Example 5 hereinbelow), base editing (see Example 4 hereinbelow), or prime editing activity) can be provided along with one or more components of the system to yield a synthetic system capable of RNA- or DNA-directed modification of cells and nucleic acids.
[0090] Thus, the type S system and its components provide new tools for nucleic acid editing and cell targeting in environmental, medical and other settings.
[0091] In a first aspect, the following aspects of the novel Type S system are provided: A Cas-S1 protein or a nucleic acid encoding Cas-S1. A Cas-S2 protein or a nucleic acid encoding Cas-S2. A Cas-S3 protein or a nucleic acid encoding Cas-S3. A Cas-S4 protein or a nucleic acid encoding Cas-S4. A Cas-S5 protein or a nucleic acid encoding Cas-S5. One or more nucleic acids encoding Cas-S1, S2, S3, S4 and S5 proteins. One or more nucleic acids encoding the Cas-S4 and S5 proteins. One or more nucleic acids encoding Cas-S3, S4 and S5 proteins. One or more nucleic acids encoding Cas-S2, S3, S4 and S5 proteins. One or more nucleic acids encoding Cas-S1, S2, S4 and S5 proteins. One or more nucleic acids encoding Cas-S2, S3, S4 and S5 proteins. One or more nucleic acids encoding Cas-S2, S4 and S5 proteins. 1. A vector or nucleic acid encoding Cas-S1, S2, S3, S4 and S5 proteins, comprising SEQ ID NO:12 or a DNA sequence that is at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identical to SEQ ID NO:12. A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence that is at least (about) 80% (or at least about 80%, 85%, 90%, 91, 92, 93, 94, 95, 96, 97, 98 or at least about 99%) identical to SEQ ID NO:12.
[0092] In one embodiment, each of the sequences is operably linked to a heterologous promoter (i.e., a promoter not operably linked to the sequence in nature). In one embodiment, each of the sequences is operably linked to a promoter that is a eukaryotic promoter or a viral (e.g., phage) promoter. The promoter can be a constitutive promoter. The promoter can be an inducible promoter. The promoter can be as described in any construct, concept, aspect, example, embodiment, option, or other feature described herein.
[0093] Optionally, neither the nucleic acid nor the vector encodes a Cas nuclease. The nucleic acid may encode a non-Cas effector protein or domain, such as a nuclease that is not a Cas nuclease.
[0094] In a second aspect: A protein or ribonucleoprotein complex containing one, two, three, four or all of the Cas-S1, S2, S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S1, S2, S3, S4 and S5 proteins. A protein or ribonucleoprotein complex containing Cas-S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S2, S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S1, S2, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S2, S3, S4 and S5 proteins. Cas-protein or ribonucleoprotein complex containing S2, S4 and S5 proteins.
[0095] Optionally, the complex does not include a Cas nuclease. The complex may include a non-Cas effector protein or domain, such as a nuclease that is not a Cas nuclease.
[0096] Type S proteins are optionally capable of RNA-directed modification of DNA sequences without causing breaks in the DNA. Cas-S1 protein: Cas-S1 is a Cas protein comprising amino acids at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identical to SEQ ID NO: 1. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S1 is Cas-S1.1. Cas-S1.1 is a Cas protein comprising the amino acids of SEQ ID NO: 1.
[0097] Cas-S2 protein: Cas-S2 is a Cas protein comprising amino acids at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identical to SEQ ID NO:2. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S2 is Cas-S2.1. Cas-S2.1 is a Cas protein comprising the amino acids of SEQ ID NO:2.
[0098] Cas-S3 protein: Cas-S3 is a Cas protein comprising amino acids at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identical to SEQ ID NO:3. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, the Cas-S3 is Cas-S3.1. Cas-S3.1 is a Cas protein comprising the amino acids of SEQ ID NO:3.
[0099] Cas-S4 Protein: Cas-S4 is a Cas protein comprising amino acids at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identical to SEQ ID NO:4. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S4 is Cas-S4.1. Cas-S4.1 is a Cas protein comprising the amino acids of SEQ ID NO:4.
[0100] Cas-S5 Protein: Cas-S5 is a Cas protein comprising amino acids at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% identical to SEQ ID NO:5. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, the Cas-S5 is Cas-S5.1. Cas-S5.1 is a Cas protein comprising the amino acids of SEQ ID NO:5.
[0101] For example, any identity percentage herein is at least (about) 70%. For example, any identity percentage herein is at least (about) 80%. For example, any identity percentage herein is at least (about) 90%. For example, any identity percentage herein is at least 95%. For example, any identity percentage herein is at least about 96%. For example, any identity percentage herein is at least about 97%. For example, any identity percentage herein is at least about 98%. For example, any identity percentage herein is at least about 99%.
[0102] The percent identity of amino acid sequences is determined using the blastP algorithm with the following parameters:- Default parameters are adjusted for short input sequences, with the expectation threshold set to 0.05 and the seed sequence length for starting the alignment set to 6. Regions of low compositional complexity are masked. The scoring matrix used is "BLOSUM62", with scoring costs for creating and extending a gap of 11 and 1, respectively. A conditional composition score matrix adjustment is used to compensate for the amino acid composition of the sequences being compared.
[0103] The percent identity of nucleotide sequences is determined using the blastn algorithm with the following parameters:- Default parameters are automatically adjusted for short input sequences, with the expectation threshold set to 0.05 and the seed sequence length for starting alignment set to 28. Regions of low compositional complexity are masked. The query sequence is masked while generating the seed sequence used to scan the database, but not for extension. Matches are scored as +1 and mismatches as -2.
[0104] For example, an amino acid sequence described herein is identical to a reference SEQ ID NO: except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, for example 1 or 2, for example 1 amino acid change. For example, a nucleic acid sequence described herein is identical to a reference SEQ ID NO: except for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotide changes. For example, an amino acid sequence of a polypeptide is identical to SEQ ID NO: 1 except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, for example 1 or 2, for example 1 amino acid change. For example, the amino acid sequence of the polypeptide is identical to SEQ ID NO: 2 except for 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid changes, in particular 1 to 5, such as 1 to 3, for example 1 or 2, for example 1 amino acid change. For example, the amino acid sequence of the polypeptide is identical to SEQ ID NO: 3 except for 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid changes, in particular 1 to 5, for example 1 to 3, for example 1 or 2, for example 1 amino acid change. For example, the amino acid sequence of the polypeptide is identical to SEQ ID NO: 4 except for 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid changes, in particular 1 to 5, for example 1 to 3, for example 1 or 2, for example 1 amino acid change. For example, the amino acid sequence of the polypeptide is identical to SEQ ID NO: 5 except for 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid changes, in particular 1 to 5, such as 1 to 3, for example 1 or 2, for example 1 amino acid change.
[0105] For example, an amino acid sequence described herein is identical to a reference SEQ ID NO: 1 except that the total number of amino acid changes is (about) 5, 10, 15, 20, 25, or 30% or less of the number of amino acids in the reference sequence. For example, a nucleotide sequence described herein is identical to a reference SEQ ID NO: 1 except that the total number of nucleotide changes is (about) 5, 10, 15, 20, 25, or 30% or less of the number of nucleotides in the reference sequence. Thus, in one example, the amino acid sequence of the polypeptide is SEQ ID NO: 1, or an amino acid sequence that is identical to SEQ ID NO: 1 except that the total number of amino acid changes is (about) 5, 10, 15, 20, 25, or 30% or less of the number of amino acids in SEQ ID NO: 1, particularly (about) 20% or less, for example (about) 10% or less of the number of amino acids in SEQ ID NO: 1. Thus, in one example, the amino acid sequence of the polypeptide is SEQ ID NO: 2, or an amino acid sequence identical to SEQ ID NO: 2 except that the total number of amino acid changes is not more than (about) 5, 10, 15, 20, 25, or 30%, particularly not more than (about) 20%, for example not more than (about) 10% of the number of amino acids in SEQ ID NO: 2. Thus, in one example, the amino acid sequence of the polypeptide is SEQ ID NO: 3, or an amino acid sequence identical to SEQ ID NO: 3 except that the total number of amino acid changes is not more than (about) 5, 10, 15, 20, 25, or 30%, particularly not more than (about) 20%, for example not more than (about) 10% of the number of amino acids in SEQ ID NO: 3. Thus, in one example, the amino acid sequence of the polypeptide is SEQ ID NO: 4, or an amino acid sequence identical to SEQ ID NO: 4 except that the total number of amino acid changes is not more than (about) 5, 10, 15, 20, 25, or 30%, particularly not more than (about) 20%, for example not more than (about) 10% of the number of amino acids in SEQ ID NO: 4. Thus, in one example, the amino acid sequence of the polypeptide is SEQ ID NO: 5, or an amino acid sequence that is identical to SEQ ID NO: 5 except that the total number of amino acid changes is not more than (about) 5, 10, 15, 20, 25 or 30% of the number of amino acids in SEQ ID NO: 5, particularly not more than (about) 20%, for example not more than (about) 10% of the number of amino acids in SEQ ID NO: 5.
[0106] In one embodiment, the vector herein encodes Cas-S1, 2, 3, 4 and 5.
[0107] In one embodiment, the vector herein encodes Cas-S4 and 5.
[0108] In one embodiment, the vector herein encodes Cas-S3, 4 and 5.
[0109] In one embodiment, the vector herein encodes Cas-S2, 3, 4 and 5.
[0110] In one embodiment, the vector herein encodes Cas-S1 and 2.
[0111] In one embodiment, the vector herein encodes Cas-S1, 2 and 3.
[0112] In one embodiment, the vector herein encodes Cas-S1, 2, 3 and 4.
[0113] In one embodiment, the vector herein encodes Cas-S1, 2, 4 and 5.
[0114] In one embodiment, the vectors herein encode Cas-S2, 4 and 5. In one embodiment, the vectors herein encode Cas-S1.1, 2.1, 3.1, 4.1 and 5.1.
[0115] In one embodiment, the vector herein encodes Cas-S4.1 and 5.1.
[0116] In one embodiment, the vector herein encodes Cas-S3.1, 4.1 and 5.1.
[0117] In one embodiment, the vector herein encodes Cas-S2.1, 3.1, 4.1 and 5.1.
[0118] In one embodiment, the vector herein encodes Cas-S1.1 and 2.1.
[0119] In one embodiment, the vector herein encodes Cas-S1.1, 2.1 and 3.1.
[0120] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, 3.2 and 4.1.
[0121] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, 4.1 and 5.1.
[0122] In one embodiment, the vector herein encodes Cas-S2.1, 4.1 and 5.1.
[0123] In one embodiment, the complex herein comprises Cas-S1, 2, 3, 4 and 5.
[0124] In one embodiment, the complex herein comprises Cas-S4 and 5.
[0125] In one embodiment, the complex herein comprises Cas-S3, 4 and 5.
[0126] In one embodiment, the complex herein comprises Cas-S2, 3, 4 and 5.
[0127] In one embodiment, the complex herein comprises Cas-S1 and 2.
[0128] In one embodiment, the complex herein comprises Cas-S1, 2 and 3.
[0129] In one embodiment, the complex herein comprises Cas-S1, 2, 3 and 4.
[0130] In one embodiment, the complex herein comprises Cas-S1, 2, 4 and 5.
[0131] In one embodiment, the complex herein comprises Cas-S2, 4 and 5.
[0132] In one embodiment, the complex herein comprises Cas-S1.1, 2.1, 3.1, 4.1 and 5.1.
[0133] In one embodiment, the complex herein comprises Cas-S4.1 and 5.1.
[0134] In one embodiment, the complex herein comprises Cas-S3.1, 4.1 and 5.1.
[0135] In one embodiment, the complex herein comprises Cas-S2.1, 3.1, 4.1 and 5.1.
[0136] In one embodiment, the complex herein comprises Cas-S1.1 and 2.1.
[0137] In one embodiment, the complex herein comprises Cas-S1.1, 2.1 and 3.1.
[0138] In one embodiment, the complex herein comprises Cas-S1.1, 2.1, 3.1 and 4.1.
[0139] In one embodiment, the complex herein comprises Cas-S1.1, 2.1, 4.1 and 5.1.
[0140] In one embodiment, the complex herein comprises Cas-S2.1, 4.1 and 5.1.
[0141] In one embodiment, the fusion protein herein comprises Cas-S1.1, 2.1, 3.1, 4.1 or 5.1.
[0142] In one embodiment, the fusion protein herein comprises Cas-S1.1, 2.1, 3.1, 4.1 and 5.1.
[0143] In one embodiment, the fusion protein herein comprises Cas-S4.1 and 5.1.
[0144] In one embodiment, the fusion protein herein comprises Cas-S3.1, 4.1 and 5.1.
[0145] In one embodiment, the fusion protein herein comprises Cas-S2.1, 3.1, 4.1 and 5.1.
[0146] In one embodiment, the fusion protein herein comprises Cas-S1.1 and 2.1.
[0147] In one embodiment, the fusion protein herein comprises Cas-S1.1, 2.1 and 3.1.
[0148] In one embodiment, the fusion protein herein comprises Cas-S1.1, 2.1, 3.1 and 4.1.
[0149] In one embodiment, the fusion protein herein comprises Cas-S1.1, 2.1, 4.1 and 5.1.
[0150] In one embodiment, the fusion protein herein comprises Cas-S2.1, 4.1 and 5.1.
[0151] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1, 4.1, or 5.1.
[0152] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1, 4.1 and 5.1.
[0153] In one embodiment, the methods herein use Cas-S4.1 and 5.1.
[0154] In one embodiment, the methods herein use Cas-S3.1, 4.1 and 5.1.
[0155] In one embodiment, the methods herein use Cas-S2.1, 3.1, 4.1 and 5.1.
[0156] In one embodiment, the methods herein use Cas-S1.1 and 2.1.
[0157] In one embodiment, the methods herein use Cas-S1.1, 2.1 and 3.1.
[0158] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1 and 4.1.
[0159] In one embodiment, the methods herein use Cas-S1.1, 2.1, 4.1 and 5.1.
[0160] In one embodiment, the methods herein use Cas-S2.1, 4.1 and 5.1.
[0161] In one aspect, one or more vectors and nucleic acids comprising one or more nucleotide sequences selected from SEQ ID NOs: 7-11 are provided.
[0162] The following is provided:- One or more nucleic acids comprising SEQ ID NOs: 7 to 11. One or more nucleic acids comprising SEQ ID NOs: 8 to 11. One or more nucleic acids comprising SEQ ID NOs: 9 to 11. One or more nucleic acids comprising SEQ ID NOs: 10 and 11. One or more nucleic acids comprising SEQ ID NOs: 7, 8, 10 and 11. One or more nucleic acids comprising SEQ ID NOs: 8, 9, 10 and 11. One or more nucleic acids comprising SEQ ID NOs: 8, 10 and 11. A nucleic acid comprising SEQ ID NO:7. A nucleic acid comprising SEQ ID NO:8. A nucleic acid comprising SEQ ID NO:9. A nucleic acid comprising SEQ ID NO: 10. A nucleic acid comprising SEQ ID NO: 11. A nucleic acid comprising SEQ ID NOs: 7 to 11. A nucleic acid comprising SEQ ID NOs: 7, 8, 10 and 11. A nucleic acid comprising SEQ ID NOs: 8, 9, 10 and 11. A nucleic acid comprising SEQ ID NOs: 8, 10 and 11. One or more nucleic acid vectors comprising SEQ ID NOs: 7-11. One or more nucleic acid vectors comprising SEQ ID NOs: 8-11. One or more nucleic acid vectors comprising SEQ ID NOs: 9-11. One or more nucleic acid vectors comprising SEQ ID NOs: 10 and 11. One or more nucleic acid vectors comprising SEQ ID NOs: 7, 8, 10 and 11. One or more nucleic acid vectors comprising SEQ ID NOs: 8, 9, 10 and 11. One or more nucleic acid vectors comprising SEQ ID NOs: 8, 10 and 11. A nucleic acid vector comprising SEQ ID NO:7. A nucleic acid vector comprising SEQ ID NO:8. A nucleic acid vector comprising SEQ ID NO:9. A nucleic acid vector comprising SEQ ID NO:10. A nucleic acid vector comprising SEQ ID NO:11. A nucleic acid vector comprising SEQ ID NOs: 7 to 11. A nucleic acid vector comprising SEQ ID NOs: 7, 8, 10 and 11. A nucleic acid vector comprising SEQ ID NOs: 8, 9, 10 and 11. A nucleic acid vector comprising SEQ ID NOs: 8, 10 and 11. A vector comprising SEQ ID NO:12. A nucleic acid comprising SEQ ID NO: 12.
[0163] Optionally, the described nucleotide sequence or at least one of the described nucleotide sequences comprises: a) is heterologous to the nucleotide sequence; b) It is not an E. coli or Klebsiella promoter; c) Is it a eukaryotic promoter? d) is an animal promoter (optionally a mammalian or human promoter); e) is a plant promoter; f) a fungal promoter (optionally, a yeast promoter); g) is an insect promoter; h) a viral promoter (optionally a viral, AAV or lentiviral promoter); or i) Synthetic promoters The promoter is operably linked to a promoter which is
[0164] Any of the promoters described herein can be used with the vectors and nucleic acids described herein.
[0165] In one embodiment, the vector or nucleic acid is contained in a cell that is not an E. coli or Klebsiella (e.g., not in K. pneumoniae) cell. In one embodiment, the vector or nucleic acid is contained in a cell that is not a Pseudomonas cell. In one embodiment, the vector or nucleic acid is contained in a cell that is not an E. coli or Klebsiella cell that includes an endogenous nucleotide sequence including SEQ ID NOs: 7-11. In one embodiment, the vector or nucleic acid is contained in a cell (e.g., a bacterial cell) that does not include an endogenous nucleotide sequence including SEQ ID NOs: 7-11. In one embodiment, the vector or nucleic acid is contained in a eukaryotic cell. In one embodiment, the vector or nucleic acid is contained in a human cell. In one embodiment, the vector or nucleic acid is contained in an animal cell. In one embodiment, the vector or nucleic acid is contained in a plant cell. In one embodiment, the vector or nucleic acid is contained in a fungal cell.
[0166] One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs: 7 to 11, wherein the nucleotide sequence is a) operably linked to a heterologous, synthetic, eukaryotic, or non-bacterial promoter; and / or b) Also provided are one or more nucleic acid vectors or one or more nucleic acids contained in a cell that is a eukaryotic cell, a non-bacterial cell, or a cell that is not a bacterial cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) that contains an endogenous nucleotide sequence comprising SEQ ID NOs: 7-11.
[0167] One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs: 7 to 11, wherein the nucleotide sequence is contained in a cell (e.g., a bacterial cell) that does not contain an endogenous nucleotide sequence comprising SEQ ID NOs: 7 to 11.
[0168] Optionally, any vector or nucleic acid or nucleotide sequence herein is codon-optimized for use in human, animal (e.g., mammal, rodent, mouse or rat), plant or fungal cells.Optionally, any vector or nucleic acid or nucleotide sequence herein is codon-optimized for use in eukaryotic cells.Vector and nucleic acid sequence for use in eukaryotic cells may contain a nuclear localization sequence (NLS).NLS facilitates the entry of proteins and polypeptides described herein into eukaryotic cells.
[0169] Optionally, any vector or nucleic acid or nucleotide sequence herein is codon-optimized for use in prokaryotic cells, such as bacterial or archaeal cells. For example, the bacterial cell is not an E. coli cell. For example, the bacterial cell is not a Klebsiella (e.g., K. pneumoniae) cell. For example, the bacterial cell is not a Pseudomonas cell. In one example, the cell comprises a DNA or polynucleotide that is modified or targeted by the methods described herein.
[0170] In one example, the cell herein is a human, animal (e.g., mammal, rodent, mouse, or rat), plant, or fungal cell. In one example, the cell herein is a Homo sapiens, Drosophila melanogaster, Mus musculus, Rattus norvegicus, Caenorhabditis elegans, or Arabidopsis thaliana cell. In one example, the cell herein is a prokaryotic cell, such as a bacterial or archaeal cell. For example, the bacterial cell is not an E. coli cell. For example, the bacterial cell is not a Klebsiella cell. For example, the bacterial cell is not a Pseudomonas cell.
[0171] In one example, the animals or mammals herein are vertebrates.
[0172] Optionally, the crRNA disclosed herein (see, e.g., Concept 10 below) is 15-100 nucleotides in length and includes a sequence of at least 10 consecutive nucleotides (e.g., a spacer sequence) that is complementary to the target sequence. For example, the crRNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleotides in length. For example, the crRNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 consecutive nucleotides (e.g., a spacer sequence) that is complementary to the target sequence.
[0173] In one embodiment, the crRNA is as described elsewhere herein. The crRNA may comprise two repeat sequences. The crRNA may comprise at least one repeat sequence having the nucleotide sequence of SEQ ID NO: 6. The crRNA may comprise two repeat sequences (e.g., having the nucleotide sequence of SEQ ID NO: 6) and one spacer sequence that hybridizes to the first protospacer sequence in the target sequence. The crRNA may comprise at least one repeat sequence having the nucleotide sequence of SEQ ID NO: 70. The crRNA may comprise two repeat sequences (e.g., having the nucleotide sequence of SEQ ID NO: 70) and one spacer sequence that hybridizes to the first protospacer sequence in the target sequence.
[0174] The spacer sequence can be about 25 to 39 nucleotides in length. The spacer sequence can be about 29 to 35 (e.g., about 30 to 34, or about 31 to 33) nucleotides in length. The spacer sequence can be about 32 nucleotides in length. The spacer sequence can be 32 nucleotides in length. The spacer can be (about) 70% (e.g., (about) 80%, or (about) 90%, or (about) 95%) complementary to the protospacer sequence in the target sequence. The spacer can be 100% complementary to the protospacer sequence in the target sequence. The spacer and / or protospacer can be as described in Concept 10 herein.
[0175] The target sequence may be RNA. The target sequence may be single-stranded DNA.
[0176] In certain embodiments, the target sequence can be a sequence of dsDNA. The target sequence can be a sequence in a mammalian, for example, human, genome. Optionally, the target sequence comprises a sequence associated with a disease or disorder in a human or animal subject.
[0177] Optionally, the target sequence contains a point mutation associated with a disease or disorder, and the complex edits the point mutation in the target sequence. Alternatively, the target sequence is contained in a gene whose expression is erroneous (e.g., present or overexpressed) in the disease or disorder, and the complex introduces a point mutation (or introduces one or more mutations within a region of the target sequence) that disrupts or prevents expression of the gene. The point mutation(s) can be located about 10 to about 20 nucleotides upstream (5') of the 5'-AAG-3' PAM in the target sequence. The point mutation(s) can be located between about 10 and about 400 nucleotides, or between about 10 and 350 nucleotides, or between about 10 and 300 nucleotides, or between about 10 and 250 nucleotides, or between about 10 and 200 nucleotides, or between about 10 and 150 nucleotides upstream (5') or downstream (3') (e.g., downstream) of the PAM in the target sequence. The point mutation(s) can be located within about 30-120 nucleotides (e.g., about 50-100 nucleotides) upstream (5') or downstream (3') (e.g., downstream) of the PAM in the target sequence. - PAM can be AAN, ANG, NAG; or The -PAM can be either no nucleotide change or one AAG.
[0178] Alternatively, the PAM can be AHN, KAG, AGG, GAC, or GTG.
[0179] Alternatively, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3', such as 5'-AAC-3' and 5'-ATG-3'. In particular, the PAM is 5'-AAC-3'.
[0180] For example, the mutation is a C to T point mutation (or one or more mutations are C to T point mutations), and the complex deaminates the targeted C point mutation, resulting in a sequence not associated with the disease or disorder. For example, the targeted C point mutation (or one or more mutations are targeted C point mutations) is present in a DNA strand that is not complementary to the crRNA. For example, the C to T point mutation(s) is / are introduced by a Cas-S1, Cas-S3, or Cas-S4 protein, particularly a cytidine deaminase (e.g., PmCDA1 cytidine deaminase, described elsewhere herein) fused (e.g., to the C-terminus) of Cas-S1 or Cas-S3. The fusion protein may include all of Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5. Alternatively, a fusion protein containing a cytidine deaminase fused to a Cas-S1 protein may lack the Cas-S3 protein. The fusion protein may further comprise a UGI protein.
[0181] Thus, there is provided a fusion protein complex comprising a fusion protein in complex with a further protein, the fusion protein comprising a protein of the formula ABCDE (5' to 3' direction), wherein: A is a CasS protein selected from Cas-S1, Cas-S3, or Cas-S4 (particularly Cas-S1); B is optionally present, and if present, comprises a linker (e.g., any of the linkers described herein, particularly, for example, an XTEN linker having the amino acid sequence of SEQ ID NO: 17); C is a base editor, e.g., a cytidine deaminase (e.g., a PmCDA1 cytidine deaminase described elsewhere herein, e.g., having the amino acid sequence of SEQ ID NO: 19); D is optionally present, and if present, comprises a linker (e.g., any of the linkers described herein, particularly the 10 amino acid linker of SEQ ID NO: 23); E is a uracil DNA glycosylase inhibitor (UGI) protein (described in more detail elsewhere herein, e.g., having the amino acid sequence of SEQ ID NO: 21); and Additional proteins in a complex with the fusion protein optionally include Cas-S2 and Cas-S5 proteins and Cas-S1 and / or Cas-S4 proteins, such that the fusion protein complex comprises at least Cas-S1, Cas-S2, Cas-S4 and Cas-S5 proteins.
[0182] In one embodiment, the fusion protein complex may further comprise a Cas-S3 protein. Also provided is a first vector expressing a fusion protein of the formula ABCDE, and a second vector expressing an additional protein of the fusion protein complex.
[0183] In one embodiment, the fusion protein complex may further comprise a crRNA as described elsewhere herein.
[0184] In another example, the mutation is an A-to-G point mutation, and the complex deaminates the targeted A point mutation, resulting in a sequence that is not associated with the disease or disorder. For example, the targeted A point mutation is present on the DNA strand that is not complementary to the crRNA.
[0185] In one example, an isolated composition, vector, nucleic acid, or protein described herein is provided. Isolated can mean, for example, that the composition, vector, nucleic acid, or protein is provided to exclude E. coli cells.
[0186] Optionally, the target sequence is contained in a gene whose expression is erroneous (e.g., present or overexpressed) in a disease or disorder, and the complex introduces a double-strand (or single-strand) break in the target sequence, thereby disrupting or preventing expression of the gene. Optionally, the target sequence is contained in a pathogenic bacterium that causes a disease or disorder, and the complex introduces a double-strand (or single-strand) break in the bacterial chromosome, thereby killing (e.g., selectively killing) the bacterium. Optionally, the target sequence is contained in an antibiotic resistance gene contained in a pathogenic bacterium that causes a disease or disorder, and the complex introduces a double-strand (or single-strand) break in the target sequence, thereby disrupting or preventing expression of the antibiotic gene, resulting in resensitization of the bacterium to the antibiotic. The double-strand (or single-strand) break can be located upstream (5') of the PAM sequence in the target sequence (e.g., about 29 to about 40 nucleotides (e.g., 30 to 32 nucleotides) upstream (5') of the PAM sequence). The PAM sequence can be any described herein, for example, selected from AHN, KAG, AGG, GAC, and GTG (see also concept 39 herein). A double-stranded break can be introduced by a fusion protein described elsewhere herein in which Py is a nuclease (e.g., any nuclease described herein). A single-stranded break in a chromosome (or in double-stranded DNA) can be introduced by a fusion protein described elsewhere herein in which Py is a nickase (e.g., any nickase described herein).
[0187] For example, double-strand breaks are introduced by nucleases. For example, single-strand breaks are introduced by nickases. For example, double-strand breaks are introduced by Cas-S1 or Cas-S4 proteins, particularly I-Tev nucleases (e.g., I-TevI nucleases described elsewhere herein) fused (e.g., to the N-terminus) to Cas-S1. The fusion protein complex may include all of Cas-S1, Cas-S2, Cas-S4, and Cas-S5, and may optionally further include Cas-S3.
[0188] Thus, there is provided a fusion protein complex comprising a fusion protein in complex with a further protein, the fusion protein comprising a protein of the formula ABC (in the 5' to 3' direction), wherein: A is an I-Tev nuclease (e.g., any of the I-TevI proteins described herein, e.g., having the amino acid sequence of SEQ ID NO: 45); B is optionally present, and if present, comprises a linker (e.g., any of the linkers described herein, particularly, e.g., an XTEN linker having the amino acid sequence of SEQ ID NO: 17); C is a CasS protein selected from Cas-S1 and Cas-S4; and Additional proteins in a complex with the fusion protein optionally include Cas-S2 and Cas-S5 proteins and a Cas-S1 or Cas-S4 protein, such that the fusion protein complex includes at least Cas-S1, Cas-S2, Cas-S4 and Cas-S5 proteins.
[0189] In one embodiment, the fusion protein complex may further comprise a Cas-S3 protein. Also provided is a first vector expressing a fusion protein of formula ABC, and a second vector expressing an additional protein of the fusion protein complex.
[0190] In one embodiment, the fusion protein complex may further comprise a crRNA as described elsewhere herein.
[0191] In one embodiment, a fusion protein complex comprising an I-Tev nuclease recognizes and cleaves a target sequence that includes, in the 5' to 3' direction: a) an I-TevI cleavage site nucleotide sequence (as described elsewhere herein, in particular, the cleavage site nucleotide sequence is 5'-CNNNG-3', e.g., the I-TevI cleavage site nucleotide sequence has the nucleotide sequence of SEQ ID NO: 42); b) an I-TevI spacer nucleotide sequence (as described elsewhere herein, but in particular, the I-TevI spacer nucleotide sequence is about 31 nucleotides in length, e.g., the I-TevI spacer nucleotide sequence is at least (about) 80% (e.g., (about) 90%) identical to the nucleotide sequence of SEQ ID NO: 43 (e.g., 100% identical to the nucleotide sequence of SEQ ID NO: 43); c) a PAM sequence (as described elsewhere herein, but in particular, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', and 5'-ACA-3', e.g., 5'-AAC-3' and 5'-ATG-3', e.g., the PAM is 5'-AAC-3'); and d) a protospacer sequence (described elsewhere herein, but particularly when the protospacer is a spacer sequence, is about 32 nucleotides in length, and optionally sequences a) through d) are each immediately adjacent to one another.
[0192] In one embodiment, a fusion protein complex containing an I-Tev nuclease introduces a double-stranded break within the cleavage site nucleotide sequence. For example, if the cleavage site nucleotide sequence comprises the sequence of SEQ ID NO: 42, the fusion protein introduces a cleavage after the second C on the forward strand of the target sequence and after the first T on the reverse strand of the target sequence (because the cleavage site nucleotide sequence on the reverse strand contains the sequence 5'-CGTTG-3'). The I-TevI nuclease nicks the forward strand immediately after the second C and immediately after the first T on the reverse strand. The double nicking of the cleavage site results in staggered double-stranded DNA breaks, with each site of cleavage having a single-stranded dinucleotide overhang: 5'-AC-3' for the forward strand and 5'-GT-3' for the reverse strand.
[0193] In one example, the vector, complex, nucleic acid, or protein is not contained in a cell, and the cell contains endogenous nucleotide sequences encoding Cas-S1, S2, S3, S4, and S5 proteins. In one example, the vector, complex, nucleic acid, or protein is not contained in a cell of a species selected from the species in Table 3. For example, the vector, complex, nucleic acid, or protein is not contained in an E. coli cell. For example, the vector, complex, nucleic acid, or protein is not contained in a Ktedonobacter, e.g., Ktedonobacter racemifer, cell. For example, the vector, complex, nucleic acid, or protein is not contained in an Allochromatium, e.g., Allochromatium warmingii, cell. For example, the vector, complex, nucleic acid, or protein is not contained in an Ignatius, e.g., Ignatius tetrasporus, cell.
[0194] In one example, the vector, complex, nucleic acid, or protein is not contained in the cell, and the cell comprises an endogenous nucleotide sequence encoding a polypeptide comprising SEQ ID NO:1, a polypeptide comprising SEQ ID NO:2, a polypeptide comprising SEQ ID NO:3, a polypeptide comprising SEQ ID NO:4, and a polypeptide comprising SEQ ID NO:5.
[0195] concept: The following numbered concepts relating to new systems, their components and uses are also provided: Any concept may be combined with any configuration, aspect, example, embodiment, option, or other feature disclosed herein.
[0196] Concept 1A. A nucleic acid vector or vectors comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO: 3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO: 5 A nucleic acid vector or a plurality of nucleic acid vectors comprising:
[0197] Concept 1B. A nucleic acid vector or vectors comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 94% (e.g., 95%) identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 98% (e.g., 99%) identical to SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 99% (e.g., 100%) identical to SEQ ID NO:3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 98% (e.g., 99%) identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 88% (e.g., 89%) identical to SEQ ID NO: 5 A nucleic acid vector or a plurality of nucleic acid vectors comprising:
[0198] Concept 1C. A nucleic acid vector or vectors comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 98% (e.g., 99%) identical to SEQ ID NO:2; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 98% (e.g., 99%) identical to SEQ ID NO:4; and c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 88% (e.g., 89%) identical to SEQ ID NO:5 A nucleic acid vector or a plurality of nucleic acid vectors comprising:
[0199] Concept 1C-1. The nucleic acid vector or vectors of Concept 1C, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence at least 94% (e.g., 95%) identical to SEQ ID NO:1.
[0200] Concept 1C-2. The nucleic acid vector or vectors of Concept 1C or 1C-1, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence at least 99% (e.g., 100%) identical to SEQ ID NO:3.
[0201] Concept 1D. A nucleic acid vector or vectors comprising an expressible nucleotide sequence, the sequence being: a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least about 80% (e.g., about 90%) identical to SEQ ID NO:2; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least about 80% (e.g., about 90%) identical to SEQ ID NO:4; and c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least about 80% (e.g., about 90%) identical to SEQ ID NO:5 A nucleic acid vector or a plurality of nucleic acid vectors comprising:
[0202] Concept 1D-1. A nucleic acid vector or vectors according to Concept 1D, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence at least about 80% (e.g., about 90%) identical to SEQ ID NO:1.
[0203] Concept 1D-2. A nucleic acid vector or vectors according to Concept 1D or 1D-1, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence at least about 80% (e.g., about 90%) identical to SEQ ID NO:3.
[0204] Instead of a vector in these concepts, one or more nucleic acids containing the described components are provided instead. The disclosure herein regarding vectors should also be read mutatis mutandis as applying to nucleic acids. For example, the nucleic acid can be contained in a chromosome of a cell, e.g., a prokaryotic cell. For example, the nucleic acid can be contained in a plasmid. For example, the modification can be a modification of the chromosome or plasmid, particularly a chromosome. Thus, the crRNA can contain a spacer capable of hybridizing to a protospacer contained in a chromosome or plasmid, particularly a chromosome. Thus, the crRNA can contain a spacer that is homologous to a protospacer contained in a chromosome or plasmid, particularly a chromosome.
[0205] Concept 2. At least one of the nucleotide sequences is a) is heterologous to at least one nucleotide sequence; b) It is not an E. coli or Klebsiella promoter; c) Is it a eukaryotic promoter? d) is an animal promoter (optionally a mammalian or human promoter); e) is a plant promoter; f) a fungal promoter (optionally, a yeast promoter); g) is an insect promoter; h) a viral promoter (optionally a viral, AAV or lentiviral promoter); or i) Synthetic promoters 2. The vector of claim 1, wherein the vector is operably linked to a promoter which is
[0206] In any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein relating to promoters, the promoter may be a constitutive promoter. In any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein relating to promoters, the promoter may be an inducible promoter. Suitable promoters are well known to those skilled in the art. Inducible promoters may be particularly useful when it is desired to express type S proteins and complexes only under certain defined parameters. In other situations where it is desired to constantly produce the type S proteins and complexes described herein, constitutive promoters may be used.
[0207] In any configuration, concept, aspect, example, embodiment, option or other feature herein relating to a promoter, the promoter may be a BolA promoter, e.g., a pBolA promoter comprising the nucleotide sequence of SEQ ID NO: 14. In any configuration, concept, aspect, example, embodiment, option or other feature herein relating to a promoter, the promoter may be a p70a promoter, e.g., a p70a promoter comprising the nucleotide sequence of SEQ ID NO: 15. In any configuration, concept, aspect, example, embodiment, option or other feature herein relating to a promoter, the promoter may be an arabinose-inducible promoter, e.g., a pBAD promoter comprising the nucleotide sequence of SEQ ID NO: 24.
[0208] Concept 3. The vector of Concept 1 or Concept 2, wherein at least one of the nucleotide sequences is operably linked to a constitutive promoter.
[0209] In one embodiment, all of the nucleotide sequences are operably linked to a constitutive promoter. In one embodiment, all of the nucleotide sequences are contained in a single operon under the control of a single constitutive promoter.
[0210] Concept 4. The vector of Concept 1 or Concept 2, wherein at least one of the nucleotide sequences is operably linked to an inducible promoter.
[0211] In one embodiment, all of the nucleotide sequences are operably linked to an inducible promoter. In one embodiment, all of the nucleotide sequences are contained in a single operon under the control of a single inducible promoter.
[0212] Concept 5. The vector of any one of Concepts 1-4, lacking a nucleotide sequence encoding the polypeptide of Concept 1A or 1B part c).
[0213] Concept 6. The vector of any one of Concepts 1-5, lacking a nucleotide sequence encoding the polypeptide of Concept 1A or 1B part a).
[0214] Concept 7A. A nucleic acid vector comprising an expressible nucleotide sequence, a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 90% identical to SEQ ID NO:2; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 90% identical to SEQ ID NO:4; and c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (approximately) 90% identical to SEQ ID NO: 5 A nucleic acid vector comprising:
[0215] Concept 7B. a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO: 3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 80% (e.g., about 90%) identical to SEQ ID NO: 5 A nucleic acid vector comprising an expressible nucleotide sequence selected from:
[0216] Concept 8A. The vector of Concept 7A, further comprising a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 90% identical to SEQ ID NO:1.
[0217] Concept 8B. The vector of Concept 7A or 7B, further comprising a nucleotide sequence encoding a Cas protein domain operable with the polypeptides of parts a), b) and c) to bind to a nucleotide sequence selected from double-stranded DNA, single-stranded DNA and RNA.
[0218] Such Cas proteins can be selected from Cas proteins of naturally occurring type I, II, III, IV, V, or VI CRISPR / Cas complexes that are well known to those of skill in the art. The methods provided herein can be used to identify proteins that are operable with Cas-S2, Cas-S4, and Cas-S5 proteins to provide additional activities.
[0219] Concept 9. The vector of Concept 7A or Concept 8A or 8B, further comprising a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least (about) 90% identical to SEQ ID NO:3.
[0220] Concept 10. A method for generating a crRNA comprising one or more nucleotide sequences for generating a crRNA, the crRNA comprising a spacer that is homologous to a first protospacer in a target sequence, and optionally, the protospacer comprises: a) It is not found in E. coli; b) is a eukaryotic protospacer; or c) A vector according to any of the preceding concepts, which is a protospacer of an animal (optionally human), plant, insect, or fungal cell.
[0221] The protospacer may not be found in bacteria that contain endogenous nucleotide sequences encoding the polypeptides a) through e) described in Concept 1.
[0222] The crRNA optionally includes a spacer having a sequence of at least 10 contiguous nucleotides complementary to the protospacer. The crRNA optionally includes a spacer having a sequence of at least 15 contiguous nucleotides complementary to the protospacer. The crRNA optionally includes a spacer having a sequence of at least 20 contiguous nucleotides complementary to the protospacer. The crRNA optionally includes a spacer having a sequence of at least 25 contiguous nucleotides complementary to the protospacer. The crRNA optionally includes a spacer having a sequence of at least 28 contiguous nucleotides complementary to the protospacer.
[0223] In any configuration, concept, aspect, example, embodiment, option, or other feature herein relating to a protospacer and a spacer, the spacer and protospacer can be about 10 to 40 nucleotides in length. For example, the spacer and protospacer can be about 25 to 40 nucleotides, e.g., about 25 to 38 nucleotides, e.g., about 28 to 35 nucleotides. The spacer and protospacer can be about 25 to 39 nucleotides in length. The spacer and protospacer can be about 29 to 35 (e.g., about 30 to 34, or about 31 to 33) nucleotides in length. For example, the spacer and protospacer can be about 32 nucleotides in length. For example, the spacer and protospacer can be about 28 nucleotides in length. For example, the spacer and protospacer can be 32 nucleotides in length. For example, the spacer and protospacer can be 28 nucleotides in length. The spacer and / or protospacer can have any of the features described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0224] In any configuration, concept, aspect, example, embodiment, option, or other feature herein relating to a protospacer and a spacer, the spacer may be (about) 90% complementary to the protospacer. The spacer may be (about) 70% complementary to the protospacer. The spacer may be (about) 80% complementary to the protospacer. The spacer may be (about) 95% complementary to the protospacer. The spacer may be 100% complementary to the protospacer over its entire length.
[0225] In one embodiment, when the protospacer is about 32 (e.g., 32) nucleotides in length (or longer), nucleotides 1-28 of the spacer 5' to the PAM sequence are complementary to the protospacer sequence.
[0226] In any configuration, concept, aspect, example, embodiment, option, or other feature herein relating to a crRNA, the crRNA includes at least one repeat sequence. For example, the crRNA includes two repeat sequences. The repeat sequence can be any sequence capable of forming a hairpin loop and recognized by the type S system. The repeat sequence can be about 20 to 25 nucleotides in length (e.g., 22 or 23 nucleotides). The repeat sequence can be a nucleotide sequence that is at least about 80%, e.g., about 90% (e.g., at least about 95, 96, 97, or 98%) identical to the nucleotide sequence of SEQ ID NO:6. The repeat sequence can be a nucleotide sequence that is at least about 80%, e.g., about 90% (e.g., at least about 95, 96, 97, or 98%) identical to the nucleotide sequence of SEQ ID NO:70. The repeat sequence can be the nucleotide sequence of SEQ ID NO:70.
[0227] In any of the nucleic acids or vectors described herein, the crRNA can be encoded by a CRISPR array (e.g., a CRISPR array described elsewhere herein).
[0228] According to any configuration, concept, aspect, example, embodiment, option or other feature herein, the plant may be a monocotyledonous or dicotyledonous plant. Optionally, the plant is selected from the group consisting of corn, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oat, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower and tomato.
[0229] As known to those skilled in the art, "cognate to" or "cognate with" or "complementary to" refer to components that can operate together. For example, crRNA is known to operate with PAM to guide Cas to a protospacer in a target nucleic acid. In this sense, crRNA is cognate to PAM and cognate to the protospacer.
[0230] Concept 11.a) Nuclear localization sequence (NLS); b) a phage packaging sequence (optionally a pac or cos site); c) Plasmid origin of replication; d) Plasmid transfer origin; e) bacterial plasmid backbone; f) eukaryotic plasmid backbone; g) structural protein genes of a human virus (optionally, AAV or lentivirus); h) viral (optionally AAV or lentiviral) rep and / or cap sequences; i) a sequence encoding a selection marker or selectable marker; j) a eukaryotic promoter; or k) the nucleotide sequence of a human, animal, plant or fungal gene A vector according to any of the preceding concepts, comprising:
[0231] In any configuration, concept, aspect, example, embodiment, option, or other feature, the vector, protein, or complex herein further comprises an NLS when the target sequence is in a eukaryotic cell. NLSs are known to those skilled in the art and include, but are not limited to, monopartite or bipartite NLSs. The NLS sequence may be derived from SV40 large T antigen (see Kalderon et al., Cell, 39(3), 499-509, 1984, doi:https: / / doi.org / 10.1016 / 0092-8674(84)90457-4, the entire contents of which are incorporated herein by reference), nucleoplasmin, importin α, c-myc, EGL-13, TUS protein, hnRNP A1, or yeast transcriptional repressor Matα2.
[0232] Any of the vectors described herein can be delivered to bacterial cells via a phage particle. Any of the vectors described herein can be delivered to bacterial cells via a phagemid particle packaged within a phage particle. Any of the vectors described herein can be delivered to bacterial cells via a conjugative plasmid. Concept 12.a) Plasmid Vectors (Optionally, Conjugative Plasmids)
[0233] b) a transposon vector (optionally a conjugative transposon); c) a viral vector (optionally a phage, AAV or lentiviral vector); d) a phagemid (optionally a packaged phagemid); or e) Nanoparticles (Optionally, Lipid Nanoparticles) 2. The vector of any of the preceding concepts, wherein
[0234] Concept 13A. One or more polypeptides of any of the preceding concepts.
[0235] Concept 13B. One or more polypeptides expressed from the vector of any of the preceding concepts.
[0236] Concept 13C. One or more polypeptides of any preceding claim expressed from a vector of any preceding concept.
[0237] Concept 14A. A fusion protein comprising a polypeptide (Px), wherein Px is a) comprises an amino acid sequence that is at least (about) 80% identical (e.g., (about) 90% identical) to a sequence selected from SEQ ID NOs: 1-5; and b) A fusion protein, which is fused to a heterologous polypeptide (Py).
[0238] The 14th concept also provides:- A fusion protein comprising Cas-S1 fused to a heterologous polypeptide (Py).
[0239] A fusion protein comprising Cas-S2 fused to a heterologous polypeptide (Py).
[0240] A fusion protein comprising Cas-S3 fused to a heterologous polypeptide (Py).
[0241] A fusion protein comprising Cas-S4 fused to a heterologous polypeptide (Py).
[0242] A fusion protein comprising Cas-S5 fused to a heterologous polypeptide (Py).
[0243] Concept 14B. A fusion protein complex comprising: (i) a) comprises an amino acid sequence that is at least (about) 80% identical (e.g., (about) 90% identical) to a sequence selected from SEQ ID NOs: 1, 3, and 4; and b) fused to a heterologous polypeptide (Py); fusion protein polypeptide (Px); (ii) a polypeptide comprising an amino acid sequence at least (about) 80% identical (e.g., (about) 90% identical) to SEQ ID NO: 2; (iii) a polypeptide comprising an amino acid sequence at least (about) 80% identical (e.g., (about) 90% identical) to SEQ ID NO: 5; and (iv) a crRNA comprising a spacer sequence that is homologous to the first protospacer in the target sequence; and (v) optionally, a polypeptide that is at least (about) 80% identical (e.g., (about) 90% identical) to a sequence selected from SEQ ID NOs: 1, 3, and 4, and that is not based on the amino acid sequence of a polypeptide described in part (i) a); and (vi) Optionally, a fusion protein complex comprising a polypeptide that is at least (about) 80% identical (e.g., (about) 90% identical) to a sequence selected from SEQ ID NOs: 1, 3, and 4, and that is not based on the amino acid sequence of a polypeptide described in part (i)a) and, if present, is not based on the amino acid sequence of a polypeptide described in part (v).
[0244] Concept 14C. A fusion protein complex comprising: (i) a) comprises a Cas-S1, Cas-S3, or Cas-S4 protein; and b) fused to a heterologous polypeptide (Py) fusion protein polypeptide (Px); (ii) Cas-S2 protein; (iii) Cas-S5 protein; (iv) a crRNA comprising a spacer sequence that is homologous to the first protospacer in the target sequence; and (v) optionally, a protein selected from Cas-S1, Cas-S3, and Cas-S4 proteins, and which is of a different class of CasS proteins than the proteins described in part (i)a); and (vi) optionally, a protein selected from Cas-S1, Cas-S3 and Cas-S4 proteins, which is of a different class of CasS proteins than the proteins described in part (i)a) and, if present, is of a different class of CasS proteins than the proteins described in part (v); A fusion protein complex comprising:
[0245] In one example, the fusion protein includes two or more of Cas-S1 to S5, for example, Cas-S1 and S2, or S4 and S5.
[0246] In any configuration, concept, aspect, example, embodiment, option, or other feature of the fusion protein, Py may be fused (directly or indirectly) to the N-terminus of the CasS protein. In any configuration, concept, aspect, example, embodiment, option, or other feature of the fusion protein, Py may be fused (directly or indirectly) to the C-terminus of the CasS protein.
[0247] The fusion can be direct or indirect (e.g., via a peptide linker, e.g., (G4S) n via a linker, where n=1, 2, 3, 4, 5, 6, 7, 8, 9, or 10). The peptide linker can be about 10-20 amino acids in length (e.g., about 16 amino acids in length). The peptide linker can be about 8-12 amino acids in length (e.g., about 10 amino acids in length). The peptide linker can have the amino acid sequence of SEQ ID NO: 17. The peptide linker can have the amino acid sequence of SEQ ID NO: 23. The linker can be as described elsewhere herein (e.g., in concepts 27 or 29 herein).
[0248] Here, "heterologous" refers to a polypeptide not found in nature fused to Px. In one example, Py comprises an I-Tev nuclease or a mutH protein.
[0249] In any configuration, concept, aspect, example, embodiment, option, or other feature, Py is an I-Tev nuclease, such as an I-TevI nuclease described elsewhere herein. In one example, the I-TevI nuclease is a protein comprising the amino acid sequence of SEQ ID NO: 45. In one embodiment, the I-TevI nuclease is encoded by a nucleic acid sequence that encodes an amino acid sequence comprising the amino acids of SEQ ID NO: 45. In one embodiment, the I-TevI nuclease is encoded by the nucleic acid sequence of SEQ ID NO: 46. In one example, the I-TevI nuclease is a protein having an amino acid sequence at least about 80% identical (e.g., about 85%, 90%, 95%, 96%, 97%, or 98% identical) to the amino acids of SEQ ID NO: 45, capable of cleaving an I-TevI cleavage site nucleotide sequence having the nucleotide sequence of SEQ ID NO: 42, and capable of recognizing an I-TevI spacer nucleotide sequence having the nucleotide sequence of SEQ ID NO: 43.
[0250] Generally, when one component is heterologous to another, the components are not found in nature or in the same cell.
[0251] For example, the selected sequence is SEQ ID NO: 1. For example, the selected sequence is SEQ ID NO: 2. For example, the selected sequence is SEQ ID NO: 3. For example, the selected sequence is SEQ ID NO: 4. For example, the selected sequence is SEQ ID NO: 5.
[0252] Py can include an effector domain, such as a domain that has nuclease activity, nickase activity, recombinase activity, reverse transcriptase activity, helicase activity, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcription activation activity, or transcription repression activity.
[0253] Py can be naturally occurring. Py can be a synthetic sequence. Py can be a synthetically mutated derivative of a naturally occurring sequence.
[0254] Optionally, the effector domain is a nucleic acid editing domain. For example, the nucleic acid editing domain includes a deaminase domain. The deaminase domain can be a cytidine deaminase domain. The cytidine deaminase domain can be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. The cytidine deaminase domain can be at least (about) 80%, at least (about) 85%, at least (about) 90%, at least (about) 92%, at least (about) 95%, at least (about) 96%, at least (about) 97%, at least (about) 98%, at least (about) 99%, or at least (about) 99.5% identical to the cytidine deaminase domain of any one of SEQ ID NOS: 350-389 disclosed in U.S. Application No. 16 / 976,047 or WO 2019 / 168953, the sequences of which are expressly incorporated herein by reference.
[0255] Optionally, Py comprises a uracil glycosylase inhibitor (UGI) domain. The UGI domain may comprise the amino acid sequence of SEQ ID NO: 500, as disclosed in U.S. Application No. 16 / 976,047 or WO 2019 / 168953, the sequences of which are expressly incorporated herein by reference. The UGI domain may have the amino acid sequence of SEQ ID NO: 21. The UGI domain may have the nucleotide sequence of SEQ ID NO: 20.
[0256] Optionally, Py comprises the amino acid sequence of SEQ ID NO: 534. Optionally, Py can comprise the amino acid sequence of SEQ ID NO: 538. These sequences are disclosed in U.S. Application No. 16 / 976,047 or WO 2019 / 168953, the sequences of which are expressly incorporated herein by reference.
[0257] Optionally, Py comprises a fusion protein and further comprises a second UGI domain. Optionally, Py comprises the amino acid sequence of SEQ ID NO: 540. Optionally, Py comprises the amino acid sequence of SEQ ID NO: 541. These sequences are disclosed in U.S. Application No. 16 / 976,047 or WO 2019 / 168953, the sequences of which are expressly incorporated herein by reference.
[0258] Optionally, the deaminase domain is an adenosine deaminase domain. The fusion protein may include a second adenosine deaminase domain. For example, the first adenosine deaminase domain and the second adenosine deaminase domain include an ecTadA domain or a variant thereof. The first adenosine deaminase domain and the second adenosine deaminase domain may include the amino acid sequence of any one of SEQ ID NOs: 400-458. The first adenosine deaminase may include the amino acid sequence of SEQ ID NO: 400. The second adenosine deaminase may include the amino acid sequence of SEQ ID NO: 458. Optionally, Py includes the amino acid sequence of SEQ ID NO: 535 or 539. These sequences are disclosed in U.S. Application No. 16 / 976,047 or International Publication No. WO 2019 / 168953, the sequences of which are expressly incorporated herein by reference.
[0259] Base editors that can be used with the Cas-S system are well known in the art, including systems in which the CasS protein is fused to a cytosine or adenosine deaminase domain and directed to a target sequence to make the desired modification.
[0260] In one embodiment, the base editor is selected from the following: 1) Cytosine base editor (CBE) that converts C:G to T:A (Komar et al., Nature, 533:420-4, 2016, incorporated herein by reference in its entirety) 2) Adenine base editor (ABE) that converts A:T to G:C (Gaudelli et al., Nature, 551(7681), 464-471, 2017, incorporated herein by reference in its entirety) 3) Cytosine Guanine Base Editor (CGBE), which converts C:G to G:C (Chen et al., Biorxiv, 2020; Kurt et al., Nature Biotechnology, 2020, both of which are incorporated by reference in their entirety) 4) Cytosine Adenine Base Editor (CABE), which converts C:G to A:T (Zhao et al., Nature Biotechnology, 2020, incorporated herein by reference in its entirety) 5) Adenine Cytosine Base Editor (ACBE) that converts A:T to C:G (International Publication No. WO 2020 / 181180, incorporated herein by reference in its entirety) 6) Adenine thymine base editor (ATBE) that converts A:T to T:A (International Publication No. WO 2020 / 181202, incorporated herein by reference in its entirety)
[0261] 7) Thymine adenine base editors (TABEs) that convert T:A to A:T (WO 2020 / 181193 (or U.S. Patent Application Publication No. 2022 / 0170013); WO 2020 / 181178; WO 2020 / 181195, each of which is incorporated by reference in its entirety).
[0262] Base editors differ in the base-modifying enzymes they use. CBEs rely on ssDNA cytidine deaminases, including APOBEC1, rAPOBEC1, APOBEC1 mutants or evolved forms (evoAPOBEC1), and APOBEC homologs (APOBEC3A (eA3A), Anc689), cytidine deaminase 1 (CDA-1), evoCDA-1, FERNY, and evoFERNY.
[0263] ABE relies on the deoxyadenosine deaminase activity of the tandem fusion TadA-TadA*, which is an evolved version of TadA, an E. coli / itRNA adenosine deaminase enzyme that can convert adenosine to inosine on ssDNA. TadA* includes TadA-8a-e and TadA-7.10.
[0264] Examples of DNA base editor proteins include, but are not limited to, BE1, BE2, BE3, BE4, BE4-GAM, HF-BE3, Sniper-BE3, Target-AID, Target-AID-NG, ABE, EE-BE3, YE1-BE3, YE2-BE3, YEE-BE3, BE-PLUS, SaBE3, SaBE4, SaBE4-GAM, Sa(KKH)-BE3, VQR-BE3, VRER-BE3, EQR-BE3, xBE3, Cas12a-BE, Ea3A-BE3, A3A-BE3, TAM, CRISPR-X, A BE7.9, ABE7.10, ABE7.10*, xABE, ABESa, VQR-ABE, VRER-ABE, Sa(KKH)-ABE, ABE8e, SpRY-ABE, SpRYCBE, SpG-CBE4, SpG-ABE, SpRY-CBE4, SpCas9-NG- Includes ABE, SpCas9-NG-CBE4, enAsBE1.1, enAsBE1.2, enAsBE1.3, enAsBE1.4, AsBE1.1, AsBE1.4, CRISPR-Abest, CRISPR-Cbest, eA3A-BE3 and AncBE4.
[0265] In one embodiment, Py is a cytidine deaminase. The cytidine deaminase can be PmCDA1 cytidine deaminase. The cytidine deaminase can be from a sea lamprey. In one embodiment, the PmCDA1 cytidine deaminase comprises the amino acid sequence of SEQ ID NO: 19. In one embodiment, the PmCDA1 cytidine deaminase is encoded by a nucleic acid sequence that encodes an amino acid sequence comprising the amino acids of SEQ ID NO: 19. In one embodiment, the PmCDA1 cytidine deaminase is encoded by the nucleic acid sequence of SEQ ID NO: 18. In one example, the PmCDA1 cytidine deaminase is a protein having an amino acid sequence at least about 80% identical (e.g., about 85%, 90%, 95%, 96%, 97%, or 98% identical) to the amino acids of SEQ ID NO: 19 and capable of mutating C to T in a target nucleic acid.
[0266] Concept 15. A fusion protein complex comprising at least two additional polypeptides, two additional polypeptides each having an amino acid sequence that is at least 90% identical to a sequence selected from the sequences of SEQ ID NO:2, SEQ ID NO:4, and SEQ ID NO:5; and 15. The fusion protein of claim 14, wherein the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:2, SEQ ID NO:4 and SEQ ID NO:5.
[0267] Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 2, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 4 and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 4, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2 and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 5, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2 and SEQ ID NO: 4.
[0268] Concept 16. The fusion protein complex of Concept 15, wherein the protein complex comprises at least three additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5.
[0269] Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:2, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:3, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:2, SEQ ID NO:4, and SEQ ID NO:5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:5, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:2, SEQ ID NO:3, and SEQ ID NO:4.
[0270] Concept 17A. The fusion protein complex of Concept 15, wherein the protein complex comprises at least three additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:4, and SEQ ID NO:5.
[0271] Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 1, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2, SEQ ID NO: 4, and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 2, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 1, SEQ ID NO: 4, and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 5, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 1, SEQ ID NO: 2, and SEQ ID NO: 4.
[0272] Concept 17B. The fusion protein complex of Concept 15, wherein the protein complex comprises at least four additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5.
[0273] Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 1, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 2, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 4, and SEQ ID NO: 5. Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO: 4, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3 ... Px may comprise an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:5, and the further polypeptide comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4.
[0274] Concept 18. Px comprises an amino acid sequence selected from amino acid sequences that are at least 90% identical to a sequence selected from the sequences of SEQ ID NOs: 1, 3 and 4; the fusion protein is in the form of a fusion protein complex comprising at least two additional polypeptides, two further polypeptides each having an amino acid sequence at least 90% identical to a sequence selected from the sequences of SEQ ID NO: 2 and SEQ ID NO: 5; (i) optionally, a polypeptide that is at least 90% identical to a sequence selected from SEQ ID NOs: 1, 3, and 4 and is not based on the amino acid sequence of a polypeptide described in Part I; and (ii) optionally, a polypeptide that is at least 90% identical to a sequence selected from SEQ ID NOs: 1, 3, and 4, and that is not based on the amino acid sequence of a polypeptide described in Part I, and if present, is not based on the amino acid sequence of a polypeptide described in part (i). 15. The fusion protein of concept 14.
[0275] Concept 19. The fusion protein or fusion protein complex of any one of concepts 14-18, wherein Py is fused to the C-terminus of the amino acid sequence of part I.
[0276] Alternatively, Py is fused to the N-terminus of the part I amino acid sequence.
[0277] Concept 20. The fusion protein or fusion protein complex of any one of Concepts 14-19, wherein Py is optionally an enzyme selected from a deaminase (e.g., cytosine or adenine deaminase), a base editor, a prime editor, a helicase, a reverse transcriptase, a methyltransferase, a methylase, an acetylase, an acetyltransferase, a transcriptional activator or inactivator, a translational activator or inactivator, or a nuclease (e.g., a DNA nuclease, an RNA nuclease, a nickase, or a dead nuclease, e.g., dCas).
[0278] In any configuration, concept, aspect, example, embodiment, option, or other feature, Py is a nuclease, such as any of the nucleases disclosed herein, particularly an I-TevI nuclease. Py can be a DNA nuclease. Py can be an RNA nuclease. Py can be a nickase or a dead nuclease. In one embodiment, Py is a base editor, such as any of the base editors disclosed herein, particularly a PmCDA1 cytidine deaminase. In one embodiment, Py is a prime editor.
[0279] Concept 21. The fusion protein or fusion protein complex of Concept 20, wherein Py is a nuclease that is an I-TevI nuclease.
[0280] In any configuration, concept, aspect, example, embodiment, option, or other feature, in Part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence (e.g., any CasS protein described herein). In one embodiment, in Part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence that is at least (about) 80% (e.g., (about) 90%) identical to the sequence of SEQ ID NO: 1. The I-TevI nuclease can be as described elsewhere herein (e.g., as in Concept 14).
[0281] Concept 22. The fusion protein or fusion protein complex of Concept 21, wherein the I-TevI nuclease does not comprise an intact DNA binding domain.
[0282] I-TevI contains 245 amino acids and consists of an N-terminal catalytic domain and a C-terminal DNA-binding domain connected by a long, flexible linker. The crystal structure of the DNA-binding domain of I-TevI (comprising residues 130-245) complexed with the 20-bp primary binding region of its DNA target reveals the presence of an extended segment containing a zinc finger (comprising residues 151-167) that contacts the main chain with DNA from the minor groove, a minor-groove-binding α-helix (comprising residues 183-194), and a helix-turn helix (comprising residues 204-245). The N-terminal catalytic domain was shown to comprise residues 1-92. Biochemical data indicate that zinc fingers do not contribute to DNA-binding affinity or enzyme specificity, but rather have a novel function, acting as distance determinants that control the relative positions of the catalytic and DNA-binding domains; see, e.g., Roey et al., 2005, DOI:10.1007 / 0-387-27421-9_7, which is incorporated herein by reference in its entirety.
[0283] Kleinstiver et al., G3 Genes|Genomes|Genetics, 4(6), 1155-1165, 2014, doi: https: / / doi.org / 10.1534 / g3.114.011445 (incorporated herein by reference in its entirety) showed that various truncations of the N-terminal catalytic domain retained activity when fused to TALEN constructs. The largest construct was residues 1-206 of I-TevI. The smallest unit that was functional when fused to a TALEN was residues 1-162.
[0284] Thus, in one embodiment, I-TevI does not contain a complete helix-turn-helix domain. I-TevI may comprise the amino acid sequence of SEQ ID NO: 45. I-TevI may not contain a helix-turn-helix domain. I-TevI may comprise amino acids 1-195-203 of SEQ ID NO: 45. I-TevI may not contain a minor groove-binding α-helix. I-TevI may comprise amino acids 1-167 of SEQ ID NO: 45. I-TevI may not contain a complete zinc finger domain. I-TevI may comprise amino acids 1-162 of SEQ ID NO: 45.
[0285] Concept 23. The fusion protein or fusion protein complex of Concept 21 or 22, wherein the I-TevI nuclease comprises an N-terminal catalytic domain.
[0286] Thus, I-TevI may comprise amino acids 1 to 92 of SEQ ID NO: 45. I-TevI may comprise the amino acid sequence of SEQ ID NO:45.
[0287] Concept 24. The fusion protein or fusion protein complex of any one of Concepts 21 to 23, wherein the I-TevI nuclease is a bacteriophage I-TevI nuclease.
[0288] In any configuration, concept, aspect, example, embodiment, option, or other feature, the I-TevI nuclease can be T4 bacteriophage I-TevI. The I-TevI nuclease can be derived from T4 bacteriophage. The I-TevI nuclease can be encoded by an intron of bacteriophage T4. The I-TevI nuclease can include amino acids 1-206 of naturally occurring I-TevI nuclease. The I-TevI nuclease can include the amino acid sequence of SEQ ID NO:45. The I-TevI nuclease can be encoded by the nucleotide sequence of SEQ ID NO:44.
[0289] Concept 25. The fusion protein or fusion protein complex of claim 20, wherein Py comprises a base editor selected from an adenine base editor (ABE, e.g., adenosine deaminase), a cytosine base editor (CBE, e.g., cytidine deaminase or apolipoprotein B mRNA editing complex (APOBEC) family deaminase), a cytosine guanine base editor (CGBE), a cytosine adenine base editor (CABE), an adenine cytosine base editor (ACBE), an adenine thymine base editor (ATBE), a thymine adenine base editor (TABE), and a uracil DNA glycosylase inhibitor (UGI) protein.
[0290] In any configuration, concept, aspect, example, embodiment, option, or other feature relating to a fusion protein or fusion protein complex, Py is a CBE. In one embodiment, Py is a PmCDA1 cytidine deaminase (described elsewhere herein, e.g., in Concept 14). In one embodiment, Py is a PmCDA1 cytidine deaminase in combination with a UGI protein. In one embodiment, Py is a CBE linked to a UGI protein via a linker (described elsewhere herein). In one embodiment, Py is a PmCDA1 cytidine deaminase linked to a UGI protein via a linker (described elsewhere herein). In one embodiment, Py is a CBE linked to a UGI protein via a linker (described elsewhere herein), and the CBE is linked to a CasS protein described herein via a linker (described elsewhere herein). In one embodiment, Py is a PmCDA1 cytidine deaminase linked to a UGI protein via a linker (described elsewhere herein), and PmCDA1 is linked to a CasS protein described herein via a linker (described elsewhere herein). In one embodiment, Py is a CBE linked to a UGI protein via a linker (described elsewhere herein), and the UGI protein is linked to a CasS protein described herein via a linker (described elsewhere herein). In one embodiment, Py is a PmCDA1 cytidine deaminase linked to a UGI protein via a linker (described elsewhere herein), and the UGI protein is linked to a CasS protein described herein via a linker (described elsewhere herein). In one embodiment, the CasS protein is selected from a CasS1 protein, a CasS3 protein, and a CasS4 protein.
[0291] Concept 26. The fusion protein or fusion protein complex of concept 25, wherein the base editor is a sea lamprey base editor.
[0292] In any configuration, concept, aspect, example, embodiment, option or other feature, the base editor is a PmCDA1 cytidine deaminase having an amino acid sequence encoded by SEQ ID NO: 19. In one embodiment, the base editor is derived from sea lamprey.
[0293] Concept 27. The fusion protein or fusion protein complex of Concept 25 or Concept 26, wherein the base editor is fused to the C-terminus of the amino acid sequence of Part I.
[0294] In any configuration, concept, aspect, example, embodiment, option, or other feature, the base editor (e.g., PmCDA1 cytidine deaminase) is fused to the amino acid sequence of Part I via a peptide linker. The peptide linker can be as described in Concept 14 or Concept 29. In any configuration, concept, aspect, example, embodiment, option, or other feature, the base editor (e.g., PmCDA1 cytidine deaminase) is fused to the CasS protein via a peptide linker. The linker can be (approximately) 10-20 amino acids in length. The linker can be (approximately) 10-22, (approximately) 11-21, (approximately) 12-20, (approximately) 13-19, (approximately) 14-18, or (approximately) 15-17 amino acids in length. The linker can be (approximately) 16 amino acids in length. The linker can be a linker having the amino acid sequence of SEQ ID NO: 17. The linker can be any linker described herein.
[0295] Concept 28. The fusion protein or fusion protein complex of Concept 27, wherein the UGI protein is fused to a base editor.
[0296] In any configuration, concept, aspect, example, embodiment, option, or other feature, the UGI protein has the amino acid sequence of SEQ ID NO: 21. The UGI protein can be as described elsewhere herein (e.g., in Concept 14). The UGI protein can be fused directly to the base editor. The UGI protein can be fused to the base editor via a linker, for example, a peptide linker described elsewhere herein.
[0297] Concept 29. The fusion protein or fusion protein complex of Concept 28, wherein the UGI protein is fused to the base editor via a peptide linker.
[0298] In any configuration, concept, aspect, example, embodiment, option, or other feature, the UGI protein is fused to the base editor via a peptide linker that is (approximately) 8-12 amino acids in length. The peptide linker may be (approximately) 9-11 amino acids in length. The linker may be (approximately) 10 amino acids in length. The linker may have the amino acid sequence of SEQ ID NO: 23. The linker may be fused to the C-terminus of the base editor, and the UGI protein is fused to the C-terminus of the linker. The linker may be fused to the N-terminus of the base editor, and the UGI protein is fused to the N-terminus of the linker. The linker may be as described elsewhere herein (e.g., in Concept 14 or 27).
[0299] Concept 30. The fusion protein or fusion protein complex of any one of Concepts 14-29, combined with a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence.
[0300] The crRNA, spacer, and / or protospacer sequences can have any of the features described elsewhere herein (e.g., in Concept 10 or the first aspect of the detailed description). In any configuration, concept, aspect, example, embodiment, alternative, or feature, the protospacer is not found in a bacterium comprising an endogenous nucleotide sequence encoding a polypeptide of a) through e) described in Concept 1A or 1B. In any configuration, concept, aspect, example, embodiment, alternative, or feature, the protospacer is a protospacer from a eukaryotic cell. In any configuration, concept, aspect, example, embodiment, alternative, or feature, the protospacer is heterologous to the source of Py. For example, if Py is a naturally occurring polypeptide, the protospacer is derived from a different species (or genus) than Py. In any configuration, concept, aspect, example, embodiment, alternative, or feature, the protospacer is a protospacer from an animal (optionally human), plant, insect, or fungal cell.
[0301] In one embodiment, the fusion protein complex forms a ribonucleoprotein complex with a crRNA that includes a spacer that is homologous to the first protospacer in the target sequence. In one embodiment, the fusion protein forms a ribonucleoprotein complex with a crRNA that includes a spacer that is homologous to the first protospacer in the target sequence.
[0302] Concept 31. One or more nucleic acid vectors comprising one or more nucleotide sequences encoding the fusion protein or fusion protein complex of any one of Concepts 14-30.
[0303] The vector may further encode a crRNA as described elsewhere herein (e.g., in the first aspect, Concept 10, or Concept 30 of the detailed description). In one embodiment, first and second vectors are provided that encode a fusion protein complex as described elsewhere herein. In one embodiment, the first vector encodes the fusion protein Px, and the second vector encodes an additional protein of the fusion protein complex.
[0304] The vector may encode second, third, fourth, etc. crRNAs that contain spacers cognate to the second, third, fourth, etc. protospacers in the target sequence. Expression of multiple crRNAs can be useful for multiplex editing, e.g., to target multiple genes in a cell or organism.
[0305] Concept 32. The vector of Concept 31, wherein the fusion protein is encoded by a single first nucleotide sequence and at least two additional polypeptides are encoded by second nucleotide sequences.
[0306] In one embodiment, the fusion protein described herein is encoded by a single first nucleotide sequence contained in a first vector, and at least two additional CasS polypeptides of the complex are encoded by second nucleotide sequences contained in a second vector. In one embodiment, all of the additional CasS polypeptides of the fusion protein complex described herein are encoded by the second nucleotide sequence. In one embodiment, all of the additional CasS polypeptides of the complex described herein are encoded by the second nucleotide sequence contained in the second vector.
[0307] Concept 33. A fusion protein complex is a) a first nucleotide sequence encoding Py fused to an amino acid sequence that is at least (about) 80% (e.g., (about) 90%) identical to the sequence of SEQ ID NO:1, SEQ ID NO:3, or SEQ ID NO:4; and b) a second nucleotide sequence encoding at least two additional polypeptides encoding amino acid sequences that are at least (about) 80% (e.g., (about) 90%) identical to the sequences of SEQ ID NOs: 1, 2, 3, 4 and / or 5; 32. The vector of claim 31, wherein the vector is encoded by
[0308] In any configuration, concept, aspect, example, embodiment, option or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence at least (about) 80% (e.g., (about) 90%) identical to each of the sequences set forth in SEQ ID NOs: 1, 2, 3, 4, and 5. In any configuration, concept, aspect, example, embodiment, option or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence at least (about) 80% (e.g., (about) 90%) identical to each of the sequences set forth in SEQ ID NOs: 2, 4, and 5. In any configuration, concept, aspect, example, embodiment, option or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence at least (about) 80% (e.g., (about) 90%) identical to each of the sequences set forth in SEQ ID NOs: 2, 3, 4, and 5. In any configuration, concept, aspect, example, embodiment, option or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence at least (about) 80% (e.g., (about) 90%) identical to each of the sequences set forth in SEQ ID NOs: 1, 2, 4, and 5.
[0309] Concept 34A. The vector of Concept 32 or Concept 33, wherein the crRNA is encoded by a first nucleotide sequence.
[0310] Concept 34B. The vector of Concept 32 or Concept 33, wherein the crRNA is encoded by a second nucleotide sequence.
[0311] When multiple crRNAs are encoded to target multiple protospacers, they can all be encoded by a first nucleotide sequence. When multiple crRNAs are encoded to target multiple protospacers, they can all be encoded by a second nucleotide sequence.
[0312] Concept 35A. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is heterologous to at least one of the nucleotide sequences.
[0313] Concept 35B. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is a eukaryotic promoter.
[0314] The eukaryotic promoter can be any of those described herein.
[0315] Concept 35C. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is an animal promoter.
[0316] In one embodiment, the promoter is a mammalian or human promoter. The animal, mammalian or human promoter can be any of those described herein.
[0317] Concept 35D. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is a plant promoter.
[0318] The plant promoter can be any of those described herein.
[0319] Concept 35E. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is a fungal promoter.
[0320] In one embodiment, the promoter is a yeast promoter. The fungal or yeast promoter can be any of those described herein.
[0321] Concept 35F. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is an insect promoter.
[0322] The insect promoter can be any of those described herein.
[0323] Concept 35G. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is a viral promoter.
[0324] In one embodiment, the promoter is a viral, AAV, or lentiviral promoter. The viral, AAV, or lentiviral promoter can be any of those described herein.
[0325] Concept 35H. The vector of any one of Concepts 31-34, wherein at least one of the nucleotide sequences is operably linked to a promoter that is a synthetic promoter.
[0326] The synthetic promoter can be any of those described herein.
[0327] Concept 36. The vector of any one of Concepts 31-35, wherein at least one of the nucleotide sequences is operably linked to a constitutive promoter.
[0328] Concept 37. The vector of any one of Concepts 31-35, wherein at least one of the nucleotide sequences is operably linked to an inducible promoter.
[0329] Concept 38A. A vector, protein, fusion protein or fusion protein complex according to any construct, concept, aspect, example, embodiment, option or other feature comprising a crRNA, wherein the crRNA comprises two repeat sequences (e.g., two repeat sequences).
[0330] The crRNA and other features of the repeat sequence may have other features described elsewhere herein (e.g., in the first aspect of the detailed description, or in Concept 10).
[0331] Concept 38B. A vector, protein, fusion protein or fusion protein complex according to any construct, concept, aspect, example, embodiment, option or other feature comprising a crRNA, wherein the crRNA comprises at least one repeat sequence comprising the nucleotide sequence of SEQ ID NO:6.
[0332] Concept 38C. A vector, protein, fusion protein or fusion protein complex according to any of the configurations, concepts, aspects, examples, embodiments, options or other features comprising crRNA, wherein the crRNA comprises two repeat sequences comprising the nucleotide sequence of SEQ ID NO:6.
[0333] Concept 38D. A vector, protein, fusion protein or fusion protein complex according to any construct, concept, aspect, example, embodiment, option or other feature comprising a crRNA, wherein the crRNA comprises at least one repeat sequence comprising the nucleotide sequence of SEQ ID NO: 70.
[0334] Concept 38E. A vector, protein, fusion protein or fusion protein complex according to any of the configurations, concepts, aspects, examples, embodiments, options or other features comprising a crRNA, wherein the crRNA comprises two repeat sequences comprising the nucleotide sequence of SEQ ID NO: 70.
[0335] Concept 39. The vector, protein, fusion protein or fusion protein complex according to any configuration, concept, aspect, example, embodiment, option or other feature, comprising a crRNA comprising a spacer that is cognate to a first protospacer in a target sequence, wherein the protospacer is immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3'.
[0336] In any construct, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer cognate to a first protospacer in a target sequence, the protospacer may be immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer cognate to the first protospacer in the target sequence, the protospacer may be immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3'. In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer cognate to the first protospacer in the target sequence, the protospacer may be immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence selected from 5'-AAC-3' and 5'-ATG-3'. In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence, the protospacer may be immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence that is 5'-AAC-3'. In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence, the protospacer may be immediately adjacent to a protospacer adjacent motif (PAM) sequence having any of the features described elsewhere herein.
[0337] Concept 40. A vector, protein, fusion protein or fusion protein complex according to any configuration, concept, aspect, example, embodiment, option or other feature, comprising a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence, wherein the spacer sequence is (approximately) 25-39 nucleotides in length.
[0338] The spacer sequence may be (approximately) 28 to 32 nucleotides in length. The spacer sequence may be approximately 32 nucleotides in length. The spacer sequence may be 32 nucleotides in length.
[0339] In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence, the spacer sequence can be 32 nucleotides in length. In any configuration, concept, aspect, example, embodiment, option or other feature comprising a crRNA comprising a spacer that is cognate to the first protospacer in the target sequence, the spacer sequence can have any of the features described elsewhere herein (e.g., in the first aspect, Concept 10 and Concept 30 of the Detailed Description).
[0340] Concept 41. The vector, protein, fusion protein or fusion protein complex according to any construct, concept (e.g., as in Concept 40), aspect, example, embodiment, option or other feature, comprising a crRNA comprising a spacer that is homologous to the first protospacer in the target sequence, wherein nucleotides 1-28 of the spacer are identical to the complement of nucleotides 1-28 of the protospacer sequence that is immediately 5' to the PAM sequence in the target sequence.
[0341] The protospacer may have any of the features described elsewhere herein (eg, in the first aspect, Concept 10 and Concept 30 of the detailed description).
[0342] Side concept 42. A vector, protein, fusion protein or fusion protein complex according to any construct, concept, aspect, example, embodiment, option or other feature, comprising a crRNA comprising a spacer that is homologous to a first protospacer in the target sequence, wherein the spacer is identical to the complement of the protospacer in the target sequence.
[0343] The spacer can be identical over its entire length to the complement of the protospacer in the target sequence. The spacer can be (about) 80% (e.g., (about) 90%) identical over its entire length to the complement of the protospacer in the target sequence. The spacer can be identical to the complement of the protospacer in the target sequence over the first 1-28 nucleotides immediately 5' to the PAM sequence in the target sequence, and (about) 80% (e.g., (about) 90%) identical over the remaining nucleotides in the protospacer.
[0344] Concept 43. The vector, fusion protein, or fusion protein complex according to any configuration, concept, aspect, example, embodiment, option, or other feature, comprising a fusion protein wherein Py is a nuclease that is an I-TevI nuclease, wherein the I-TevI nuclease recognizes an I-TevI cleavage site nucleotide sequence that is 5' or 3' (e.g., 5') to a protospacer in a target sequence.
[0345] The I-TevI cleavage site nucleotide sequence can be as described elsewhere herein. In one embodiment, the I-TevI cleavage site nucleotide sequence is 5' to the protospacer in the target sequence. The I-TevI cleavage site nucleotide sequence can be (approximately) 3 to 7 nucleotides in length, for example (approximately) 4 to 6 nucleotides in length. The I-TevI cleavage site nucleotide sequence can be about 5 nucleotides in length. The I-TevI cleavage site nucleotide sequence can include the motif 5'-CNNNG-3'. The I-TevI cleavage site nucleotide sequence may be selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', 5'-CAGCG-3', 5'-CACTG-3', 5'-CGACG-3', 5'-CGATG-3', 5'-CTCCG-3', 5'-CTACG-3', 5'-CTGTG-3', 5'-CTGCG-3', 5'-CCGTG-3', 5'-CCTAG-3', 5'-CCATG-3', 5'-CAAAG-3', 5'-CAGAG-3', 5'-CATTG-3' and 5'-CATCG-3'. The I-TevI cleavage site nucleotide sequence may be selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', and 5'-CAGCG-3', or selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', and 5'-CTGAG-3'. The I-TevI cleavage site nucleotide sequence may be selected from 5'-CCACG-3', 5'-CACCG-3', and 5'-CATAG-3'. The I-TevI cleavage site nucleotide sequence may be selected from 5'-CCACG-3' and 5'-CACCG-3'. The I-TevI cleavage site nucleotide sequence may have the nucleotide sequence of SEQ ID NO: 42.
[0346] Concept 44. The vector, fusion protein, or fusion protein complex of any construct, concept, aspect, example, embodiment, alternative, or other feature, comprising a fusion protein wherein Py is a nuclease that is an I-TevI nuclease, wherein the I-TevI nuclease recognizes an I-TevI spacer nucleotide sequence that is 5' or 3' (e.g., 5') to a protospacer in a target sequence; The I-TevI spacer nucleotide sequence of the vector, fusion protein, or fusion protein complex can be as described elsewhere herein. In one embodiment, the I-TevI spacer nucleotide sequence is 5' to the protospacer in the target sequence. The I-TevI spacer nucleotide sequence can be (approximately) 10 to 35 nucleotides in length, e.g., 15 to 35 or 20 to 35 nucleotides in length. The I-TevI spacer nucleotide sequence can be (approximately) 29 to 33 nucleotides in length. The I-TevI spacer nucleotide sequence can be (approximately) 30 to 32 nucleotides in length. The I-TevI spacer nucleotide sequence can be (approximately) 31 nucleotides in length. The I-TevI spacer nucleotide sequence can be at least (approximately) 80% (e.g., (approximately) 90%) identical to the nucleotide sequence of SEQ ID NO:43. The I-TevI spacer nucleotide sequence can be 100% identical to the nucleotide sequence of SEQ ID NO:43.
[0347] Concept 45. The vector, fusion protein, or fusion protein complex according to any construct, concept, aspect, example, embodiment, option, or other feature, comprising a fusion protein wherein Py is a nuclease that is an I-TevI nuclease, and wherein the target sequence is a) I-TevI cleavage site nucleotide sequence; b) I-TevI spacer nucleotide sequence; c) a PAM sequence; and d) Protospacer sequence A vector, a fusion protein or a fusion protein complex comprising:
[0348] In one embodiment, sequences a) through d) are each immediately adjacent to one another. The I-TevI cleavage site nucleotide sequence is as described elsewhere herein (e.g., as described in Concept 43). The I-TevI spacer nucleotide sequence is as described elsewhere herein (e.g., as described in Concept 44). The PAM sequence is as described elsewhere herein (e.g., as described in Concept 39). The protospacer sequence is as described elsewhere herein (e.g., as in the first aspect of the detailed description, Concept 10, and Concept 30).
[0349] Concept 46A. A cell comprising one or more vectors, fusion proteins, or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option, or other feature described elsewhere herein.
[0350] Concept 46B. A eukaryotic cell comprising one or more vectors, fusion proteins, or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option, or other feature described elsewhere herein.
[0351] Concept 46C. An animal, plant, insect or fungal cell comprising one or more vectors, fusion proteins or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option or other feature described elsewhere herein.
[0352] Concept 46D. A prokaryotic cell comprising one or more vectors, fusion proteins, or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option, or other feature described elsewhere herein.
[0353] Concept 46E. A prokaryotic cell comprising one or more vectors, fusion proteins, or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option, or other feature described elsewhere herein, wherein the prokaryotic cell is not a bacterial cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) comprising an endogenous nucleotide sequence encoding a polypeptide of a) through e) described in Concept 1A.
[0354] In one embodiment, the cell is a vertebrate, mammalian, or human cell.
[0355] Concept 47A. A composition comprising one or more vectors, fusion proteins, or fusion protein complexes according to any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein.
[0356] The compositions can be as described elsewhere herein. In one embodiment, the compositions described herein can be in vitro compositions. In one embodiment, the compositions described herein can be contained in a medical container.
[0357] Also provided are pharmaceutical compositions comprising one or more vectors, fusion proteins, or fusion protein complexes according to any of the constructs, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein, and further comprising a diluent, excipient, or carrier, for use in any of the methods of medical treatment described herein.
[0358] Also provided is one or more vectors, fusion proteins or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option or other feature described elsewhere herein for use in therapy.
[0359] Also provided is one or more vectors, fusion proteins or fusion protein complexes according to any construct, concept, aspect, example, embodiment, option or other feature described elsewhere herein for use in treating any disease described herein by performing any of the methods described herein.
[0360] The pharmaceutical composition may be contained within a medical device (e.g., an ampoule, a syringe, or an inhaler). The pharmaceutical composition may be formulated into a tincture, a capsule, or a sustained-release formulation. The pharmaceutical composition may be an oral tablet contained within a blister pack.
[0361] In one embodiment, a pharmaceutical composition comprising one or more vectors, fusion proteins, or fusion protein complexes described herein is formulated for oral or rectal administration. In one embodiment, the pharmaceutical composition is formulated for oral administration. In one embodiment, the pharmaceutical composition is formulated as a capsule or coated tablet.
[0362] Formulations comprising one or more vectors, fusion proteins, or fusion protein complexes described herein may be lyophilized before encapsulation. In one embodiment, pharmaceutical compositions comprising one or more vectors, fusion proteins, or fusion protein complexes described herein are lyophilized formulations. In one embodiment, pharmaceutical compositions comprising one or more vectors, fusion proteins, or fusion protein complexes described herein are encapsulated formulations that are released into the lower intestine of a subject.
[0363] Concept 47B. crRNA or a nucleic acid encoding the crRNA; and a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 a composition (optionally an in vitro composition, or the composition is contained in a medical container) comprising one, more or all of: A: The crRNA comprises a spacer that is cognate to the first protospacer, and the protospacer is f) not found in E. coli; g) is a eukaryotic protospacer; or h) is a protospacer of an animal (optionally mammalian or human), plant or fungal cell; or B: A composition which is an E. coli protospacer lacking an endogenous nucleotide sequence encoding a polypeptide as described in parts a) to e) above.
[0364] In one example, the composition includes (a) to (e). In one example, the composition includes (a) to (d). In one example, the composition includes (a) to (c). In one example, the composition includes (a) to (b). In one example, the composition includes (d) to (e). In one example, the composition includes (c) to (e). In one example, the composition includes (b) to (e).
[0365] In one example, the composition includes (a). In one example, the composition includes (b). In one example, the composition includes (c). In one example, the composition includes (d). In one example, the composition includes (e).
[0366] The crRNA is operable within a cell to guide the polypeptides, proteins, fusion proteins, and complexes described herein to a target sequence contained in a nucleic acid within the cell. Complexes are provided that include one, more, or all of polypeptides (a) through (e) and a crRNA, where the crRNA is capable of guiding the complex to a protospacer contained in the target DNA. For example, the complex includes (a) through (e).
[0367] Concept 47C. crRNA or a nucleic acid encoding the crRNA; and a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 a composition (optionally an in vitro composition or wherein the composition is contained in a medical container) comprising a nucleic acid encoding one, more or all of: A: The crRNA comprises a spacer that is cognate to the first protospacer, and the protospacer is f) not found in E. coli; g) is a eukaryotic protospacer; or h) is a protospacer of an animal (optionally mammalian or human), plant or fungal cell; or B: A composition which is an E. coli protospacer lacking an endogenous nucleotide sequence encoding a polypeptide as described in parts a) to e) above.
[0368] In one example, the composition comprises a vector encoding at least two, three, or four of the polypeptides. In one example, (a)-(e) are encoded by the same vector. In one example, the composition comprises a vector encoding all of (a)-(e).
[0369] In one example, the composition nucleic acid encodes (a) to (e). In one example, the composition nucleic acid encodes (a) to (d). In one example, the composition nucleic acid encodes (a) to (c). In one example, the composition nucleic acid encodes (a) to (b). In one example, the composition nucleic acid encodes (d) to (e). In one example, the composition nucleic acid encodes (c) to (e). In one example, the composition nucleic acid encodes (b) to (e).
[0370] The crRNA is operable within a cell to guide a polypeptide to a target sequence contained in a nucleic acid within the cell. A complex is provided that includes one, more, or all of the polypeptides (a) through (e) and a crRNA, where the crRNA is capable of guiding the complex to a protospacer contained in the target DNA. For example, the complex includes (a) through (e).
[0371] Also provided are methods. Any method herein can be an in vitro method. Optionally, the method is carried out on cells contained in a subject, such as a human, an animal (e.g., a mammal, such as a rodent, mouse or rat), a plant, an insect, or a fungus (e.g., yeast).
[0372] As used herein, modifying (e.g., modifying a target sequence of DNA, polynucleotide, or cell) can be modifying the cutting, editing, blocking, marking, or labeling of the target sequence.
[0373] Concept 48A. A method for modifying a nucleic acid target site in a cell, comprising: (I) contacting a cell with one or more vectors, proteins, fusion proteins, fusion protein complexes, or compositions described in any construct, concept, aspect, example, embodiment, option, or other feature, comprising a crRNA that includes a spacer that is cognate to the first protospacer in the target sequence; (II) (i) a vector protein, fusion protein, or fusion protein complex, by which the polypeptide and crRNA are expressed in the cell; or (ii) the cRNA (or nucleic acid encoding the crRNA, whereby the cRNA is expressed in the cell) of the composition and the nucleic acid encoding the polypeptide or protein of the composition, whereby the polypeptide or protein is expressed in the cell and allowing the introduction of the vector into the cell; (III) A method in which the crRNA forms a complex with a polypeptide or protein and guides the complex to the target sequence.
[0374] When the method involves the introduction of a vector expressing a CasS protein or the introduction of a CasS protein, the modifying can be the downregulation of a gene containing the target sequence. When the method involves the introduction of a vector expressing a CasS protein or the introduction of a CasS protein, the modifying can be the downregulation of a gene adjacent to (e.g., within 2 kb of) the target sequence.
[0375] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be cleaving the target sequence (e.g., where Py comprises a nuclease or nickase).
[0376] Where the method involves introducing a vector expressing a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, modifying can be the introduction of one or more mutations in the target sequence (e.g., where Py comprises a nuclease, nickase, prime editor, or base editor).
[0377] Where the method includes introducing a vector that expresses a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be performing base editing in the target sequence (e.g., where Py is the base editor).
[0378] Where the method involves introducing a vector expressing a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be to perform prime editing at the target sequence (e.g., where Py is the prime editor).
[0379] Where the method involves introducing a vector that expresses a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be deaminating one or more nucleotides in the target sequence (e.g., where Py is a deaminase, such as a cytosine or adenine deaminase).
[0380] Where the method involves introducing a vector that expresses a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be methylating one or more nucleotides in the target sequence (e.g., where Py is a methyltransferase).
[0381] Where the method involves introducing a vector that expresses a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be methylating one or more nucleotides in the target sequence (e.g., where Py is a methylase).
[0382] Where the method involves introducing a vector that expresses a fusion protein or fusion protein complex, or introducing a fusion protein or fusion protein complex, the modifying can be acetylating one or more nucleotides in the target sequence (e.g., where Py is an acetylase).
[0383] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be subjecting the DNA of the target sequence to acetyltransferase activity (e.g., where Py is an acetyltransferase).
[0384] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be reverse transcribing the target sequence, or reverse transcribing RNA within the cell to generate DNA that is inserted into the target sequence (e.g., where Py is a reverse transcriptase).
[0385] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be activating transcription of a gene in the cell, for example, a gene containing or adjacent to the target sequence (e.g., where Py is a transcriptional activator).
[0386] Where the method involves the introduction of a vector that expresses a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be to inactivate transcription of a gene in the cell, for example, a gene that includes or is adjacent to the target sequence (e.g., where Py is a transcriptional inactivator).
[0387] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be activating translation of a gene in the cell, for example, a gene containing or adjacent to the target sequence (e.g., where Py is a translational activator).
[0388] Where the method involves the introduction of a vector expressing a fusion protein or fusion protein complex, or the introduction of a fusion protein or fusion protein complex, the modifying can be to inactivate translation of a gene in the cell, for example, a gene containing or adjacent to the target sequence (e.g., where Py is a translational inactivator).
[0389] Concept 48B. A method for modifying a nucleic acid target sequence in a cell, comprising: a) contacting a cell with a vector according to Concept 10 or any one of Concepts 11-12 when dependent on Concept 10, or a composition according to Concept 47; b) (i) a vector by which the polypeptide and crRNA are expressed in a cell, or (ii) the cRNA (or nucleic acid encoding the crRNA, whereby the cRNA is expressed in the cell) of the composition and the nucleic acid encoding the polypeptide of the composition, whereby the polypeptide is expressed in the cell and allowing the introduction of the vector into the cell; c) A method in which the crRNA forms a complex with a polypeptide and guides the complex to the target sequence.
[0390] It is worth noting that the present inventors have found that the orientation of spacer does not affect the efficiency of Type S CRISPR-Cas system to down-regulate the transcription of target gene, but this is not the case for other CRISPR-Cas systems that must target either the coding region or non-coding region of gene to effectively down-regulate its transcription.Thus, the complex modifies (for example, cuts or edits) the coding target sequence or non-coding target sequence, and optionally (i) the coding strand but not the non-coding strand of DNA is modified, or (ii) the non-coding strand but not the coding strand of DNA is modified.Modifying herein can be usefully (i) the coding strand but not the non-coding strand of DNA, or (ii) the non-coding strand but not the coding strand of DNA.
[0391] Concept 48C. A complex is guided to a target sequence and modifies the target sequence, and optionally the complex a) a nuclease that cleaves a target sequence; b) a deaminase that deaminates a target sequence (e.g., a cytosine or adenine deaminase); c) a base editor that edits the bases of a target sequence; d) a prime editor that primes and edits the target sequence; e) reverse transcriptase, which reverse transcribes the target sequence or reverse transcribes RNA within the cell to produce DNA that is inserted into the target sequence; f) methyltransferases that methylate DNA in cells; g) Methylase, which methylates DNA in cells; h) acetylase, which acetylates DNA within cells; i) acetyltransferases that subject intracellular DNA to acetyltransferase activity; j) a transcriptional activator that activates transcription of a gene in a cell, e.g., a gene that contains or is adjacent to a target sequence; and k) a transcriptional inactivator that inactivates transcription of a gene in a cell, e.g., a gene that contains or is adjacent to a target sequence. comprising a component selected from: Optionally, the component is encoded by a vector or nucleic acid of the composition.
[0392] In any configuration, concept, aspect, example, embodiment, option or other feature relating to modifying a target DNA or RNA, the complex further comprises a base editor (e.g., as described elsewhere herein).
[0393] In any configuration, concept, aspect, example, embodiment, option, or other feature relating to modifying a target DNA or RNA, the complex further comprises a prime editor (e.g., as described elsewhere herein).
[0394] In any configuration, concept, aspect, example, embodiment, option or other feature relating to modifying target DNA or RNA, the complex further comprises a nuclease (e.g., as described elsewhere herein).
[0395] Concept 49A. The method of Concept 48C, wherein the cell is a eukaryotic cell, an animal cell, a plant cell, an insect cell, or a fungal cell.
[0396] Concept 49B. The method of Concept 48, wherein the cell is a prokaryotic cell.
[0397] In one embodiment, the cell is a prokaryotic cell, and the cell is not a prokaryotic cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) that contains an endogenous nucleotide sequence encoding a)-e) of the polypeptides described in Concept 1.
[0398] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the animal herein is a non-human mammal or a vertebrate. Optionally, the animal cell herein is a cell of a non-human mammal or a vertebrate. Optionally, the cell is a plant cell, a bacterial cell, a fungal cell, a mammalian cell, an insect cell or an archaeal cell. The cell herein can be ex vivo or in vivo.
[0399] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the cell herein is a pluripotent cell, such as a pluripotent stem cell. Optionally, the cell herein is a stem cell, e.g., an embryonic stem cell. Optionally, the cell herein is a totipotent cell. In one embodiment of these options, the cell is a human cell. Alternatively, the cell is not a human cell and is not contained within a human or human embryo.
[0400] Optionally, any method herein is not one of the following: (a) The process for cloning humans; (b) processes for modifying the genetic identity of the human germ line; (c) the use of human embryos; (d) A process for altering the genetic identity of an animal that is likely to cause disease in humans or animals without any substantial medical benefit to them.
[0401] Optionally, any method herein is not a method of treatment of the human or animal body. Optionally, any method herein is not a surgical method.
[0402] Concept 50. The method of Concept 48 or 49, wherein the target sequence is contained in a chromosome or episome within the cell.
[0403] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is contained in a chromosome. In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is contained in a plasmid. In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is not contained in a plasmid.
[0404] Concept 51. The method of Concept 50, wherein the method inhibits replication of a chromosome or episome in a cell.
[0405] In one example, inhibition is at least (about) 20, 30, 40, 50, 60, 70, 80, 90, or 95% compared to replication in identical cells not exposed to the vector or composition. In one example, inhibition is 100%.
[0406] Concept 52. The method of any one of Concepts 48 to 51, wherein the target sequence is within (approximately) 2 kb upstream or downstream of the gene of interest.
[0407] In any configuration, concept, aspect, example, embodiment, option, or other feature relating to a target sequence, the target sequence is located upstream (5') or downstream (3') of the gene of interest within (approximately) 2 kb. The target sequence can be located upstream (5') of the gene of interest. The target sequence can be located downstream (3') of the gene of interest. The target sequence can be located within (approximately) 1.75 kb, (approximately) 1.5 kb, (approximately) 1.25 kb, (approximately) 1 kb, (approximately) 0.75 kb, (approximately) 0.5 kb, or (approximately) 0.25 kb upstream (5') or downstream (3') of the gene of interest. The target sequence can be located within the gene of interest, for example, within the promoter, exon, and / or intron of the gene of interest. In one embodiment, the target sequence is located within the promoter of the gene of interest.
[0408] Concept 53. The method of Concept 52, wherein the target sequence of the gene of interest is at least 2 kb upstream and downstream of any promoter, exon and intron of any other gene.
[0409] As shown in Example 4 herein, the CasS3 protein can have a wide range of effects, with binding effects exhibited within approximately 2 kb on either side (i.e., 5' and 3') of the protospacer sequence in the target sequence. Thus, in applications where the CasS system is used for CRISPRi (i.e., any of the methods described herein for upregulating or downregulating a gene of interest), it may be beneficial to avoid off-target effects if the protospacer in the target sequence is at least 2 kb upstream (5') and downstream (3') of any other gene (i.e., promoter, exon, or intron).
[0410] Concept 54. The method of any one of Concepts 48-53, wherein said modifying is a substitution, deletion or insertion of one or more nucleotides.
[0411] In one embodiment, the modifying is a substitution of one or more nucleotides in the target sequence. In one embodiment, when Py is a CBE, the modifying is a substitution of one or more C to T in the nucleotides of the target sequence. In one embodiment, the CBE is a PmCDA1 cytidine deaminase, e.g., as described elsewhere herein.
[0412] Concept 55. A method of treating or preventing a disease or condition mediated by a target cell in a subject, comprising performing a method described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 48-54) to modify the target cell, wherein said contacting comprises administering a vector, protein (including a fusion protein or fusion protein complex), or composition to the subject, wherein the modification treats or prevents the disease or condition.
[0413] The target cells can be, for example, pathogenic bacterial cells. They can be antibiotic-resistant bacterial cells, and the modification re-sensitizes the bacterial cells to the antibiotic.
[0414] Concept 56. The method of concept 55, wherein the subject is a human, an animal, or a plant.
[0415] Concept 57. The method of Concept 56, wherein the target cell is a cell of a subject that contains a nucleic acid defect and wherein the modification corrects the defect.
[0416] For example, certain diseases are known to be caused by single nucleotide polymorphisms (SNPs) (e.g., progeria syndrome and condylar dysplasia are caused by the missense SNP c.1580G>T SNP in the LMNA gene, which encodes the lamin A / C protein, and an SNP in the F5 gene causes factor V Leiden thrombophilia). Correction of these SNPs by targeted substitution can be used to ameliorate the disease. In 2012, there were 124 genetic diseases demonstrated to be caused by retrotransposon insertions, including cystic fibrosis (Alu), hemophilia A (L1), and X-linked dystonia-parkinsonism (SVA). See Hancks et al., Curr. Opin. Genet. Dev., 22, 191-203, 2012, doi:10.1016 / j.gde.2012.02.006, the entire contents of which are incorporated herein by reference. Certain targeted editing techniques can be used to correct these types of insertions and ameliorate the associated diseases.
[0417] Concept 58A. The method of Concept 56 or 57, wherein the modification adds new functionality to the cell.
[0418] In one embodiment, the modification upregulates or downregulates expression of the gene in the cell. In one embodiment, the modification downregulates expression of the gene in the cell. In one embodiment, the modification inhibits expression of the gene in the cell. In one embodiment, the modification blocks expression of the gene in the cell.
[0419] In another embodiment, the modification is the insertion of a new gene or a genetic pathway (e.g., an operon) that can be used to produce a substance that is beneficial to the cell. For example, the new gene may produce a beneficial metabolite or therapeutic drug. In another embodiment, the modification is the insertion of a new gene or a genetic pathway (e.g., an operon) that can be used to remove a substance that is harmful to the cell. For example, the new gene may metabolize a harmful toxin.
[0420] Concept 58B. The method of Concept 56 or 57, wherein the modification adds a new nucleotide sequence to the genome of the cell for expression of a protein encoded by the new sequence.
[0421] For example, the protein is a heterologous protein. For example, the protein is a therapeutic protein, such as an antibody, an antibody fragment (e.g., an antibody single variable domain), a TCR (T cell receptor) binding site, a TCR variable domain, a hormone, an incretin, a growth factor, an anti-cancer drug, a neurotransmitter, or an enzyme.
[0422] Concept 58C. The method of Concept 56 or 57, wherein the modification modulates expression of a nucleotide sequence contained in a chromosome or episome of the cell.
[0423] In one embodiment, the modification modulates expression of a nucleotide sequence contained in the plasmid. In one embodiment, the modification modulates expression of a nucleotide sequence not contained in the plasmid. For example, expression is upregulated. For example, expression is downregulated.
[0424] Concept 59. A vector, protein (including fusion proteins and fusion protein complexes), cell, or composition as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 1-58) for use in a method for treating or preventing a disease or condition in a human or animal, wherein the method is as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 55-58).
[0425] Any vector herein (including any construct, concept, aspect, example, embodiment, option or other feature herein) may be: a) Plasmid Vector (Optionally, a Conjugative Plasmid) b) a transposon vector (optionally a conjugative transposon); c) a viral vector (optionally a phage, AAV or lentiviral vector); d) a phagemid (optionally a packaged phagemid); or e) Nanoparticles (Optionally, Lipid Nanoparticles) Any vector that can be.
[0426] Concept 60. A method for introducing targeted editing into a target polynucleotide, comprising contacting the target polynucleotide with a fusion protein or fusion protein complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein, wherein the protein or fusion protein complex is combined with a crRNA comprising a spacer that is cognate to a first protospacer in the target sequence (e.g., as described in the first aspect of the Detailed Description, Concept 10, or Concept 30, or any one of Concepts 38-45 when dependent on Concept 30), A method in which a first protospacer in a target sequence is included in a polynucleotide, and a crRNA hybridizes to the protospacer to guide a protein (e.g., a ribonucleoprotein complex), thereby causing the protein (e.g., a ribonucleoprotein complex) to edit the polynucleotide.
[0427] The method can be performed in a parent cell containing the polynucleotide. Progeny cells or organisms derived from the edited parent cell are further provided, wherein the progeny cells or organisms retain the edit in their genome. The method can include one or more steps of culturing the edited cell to generate the progeny cell or a plurality of such progeny cells.
[0428] Concept 61A. The method of Concept 60, wherein bases or nucleic acid sequences in a polynucleotide are inserted, deleted, or substituted.
[0429] Concept 61B. The vector of Concept 10, or Concept 11 or 12 when dependent on Concept 10, the method of any one of Concepts 48-54, or the method of Concept 60 or 61, wherein the crRNA is cognate to a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0430] Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39).
[0431] Concept 62A. A container containing a plurality of proteins operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, wherein the proteins include any of the proteins described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., as described in Concept 13); A container, wherein the complex is operable with any of the protospacer adjacent motifs (PAMs) (e.g., having one of the sequences described in Concept 39) described in any configuration, concept, aspect, example, embodiment, option or other feature herein, wherein the container is not a cell and the protein is mixed with an in vitro buffer.
[0432] Concept 62B. A container containing a plurality of proteins operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, wherein the proteins comprise any of the fusion proteins or fusion protein complexes described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., as described in any one of Concepts 14-29); A container, wherein the complex is operable with any of the protospacer adjacent motifs (PAMs) (e.g., having one of the sequences described in Concept 39) described in any configuration, concept, aspect, example, embodiment, option or other feature herein, wherein the container is not a cell and the protein is mixed with an in vitro buffer.
[0433] Concept 62C. A container containing an operable protein or proteins for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the proteins comprising: a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 and A vessel, wherein the complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3', the vessel is not a cell, and the protein is mixed with an in vitro buffer.
[0434] Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39).
[0435] - 1. An operable protein or proteins for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the protein comprising: a) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 1; b) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 2; c) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 3; d) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 4; and e) a polypeptide comprising an amino acid sequence at least (about) 80% identical to SEQ ID NO: 5 and Also provided is a protein or proteins, wherein the complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0436] Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39).
[0437] Concept 63. A nucleic acid vector or plurality of nucleic acid vectors described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., described in any one of Concepts 1-12, or any one of Concepts 31-45), wherein the polypeptide or protein (including fusion proteins and fusion protein complexes) encoded by the vector is operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, and wherein the complex is operable with any of the protospacer adjacent motifs (PAMs) (e.g., having one of the sequences described in Concept 39) described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0438] In one embodiment, the protospacer adjacent motif (PAM) has the sequence 5'-AAG-3'.
[0439] Concept 64. A method for targeting a polynucleotide, comprising: a) contacting a polynucleotide with a protein (including fusion proteins and fusion protein complexes) described in any configuration, concept, aspect, example, embodiment, option or other feature herein, wherein the protein is combined with a crRNA that includes a spacer that is cognate to the first protospacer in the target sequence; b) allowing the formation of a ribonucleoprotein complex comprising the protein and crRNA, which is guided by a target sequence contained in the polynucleotide and modifies the polynucleotide or a copy thereof; A method comprising:
[0440] Targeting can be intracellular. Targeting can be in vitro.
[0441] Concept 65. The method of Concept 64, wherein the polynucleotide is contained in a chromosome or an episome.
[0442] In one embodiment, the polynucleotide is comprised in a chromosome. In one embodiment, the polynucleotide is comprised in a plasmid. In one embodiment, the polynucleotide is not comprised in a plasmid.
[0443] Concept 66A. The method of Concept 65, wherein the replication of the polynucleotide is inhibited.
[0444] Concept 66B. The method of Concept 65 for editing a polynucleotide.
[0445] Editing can insert bases or nucleic acid sequences in the polynucleotide. Editing can delete bases or nucleic acid sequences in the polynucleotide. Editing can replace bases or nucleic acid sequences in the polynucleotide.
[0446] Concept 67A. A kit comprising: a) one or more proteins described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., described in Concept 13); and b) comprises a crRNA (or one or more nucleic acids encoding the crRNA) that is homologous to any of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., a PAM having one of the sequences described in Concept 39); The kit, wherein the polypeptide is operable for use with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0447] Concept 67B. A kit comprising: a) one or more fusion proteins or fusion protein complexes described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., described in any one of Concepts 14-29), or one or more nucleic acids encoding such polypeptides; and b) a crRNA, or one or more nucleic acids encoding the crRNA), wherein the crRNA is homologous to any of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., a PAM having one of the sequences described in Concept 39); The kit, wherein the polypeptide is operable for use with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0448] Concept 67C. A kit comprising: a) one or more vectors described in any construct, concept, aspect, example, embodiment, option, or other feature herein (e.g., described in any one of Concepts 1-12, or any one of Concepts 31-45), or one or more nucleic acids encoding such polypeptides; and b) a crRNA, or one or more nucleic acids encoding the crRNA), wherein the crRNA is homologous to any of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., a PAM having one of the sequences described in Concept 39); The kit, wherein the polypeptide is operable for use with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0449] Concept 67D. A kit comprising: a) one or more polypeptides described in Concept 7, or one or more nucleic acids encoding such polypeptides; and b) crRNA (or one or more nucleic acids encoding the crRNA), wherein the crRNA is cognate to a PAM having the sequence 5'-AAG-3' Includes; The kit, wherein the polypeptide is operable for use with a crRNA or a guide RNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0450] The complexes described herein are provided. In one example, the complexes further comprise one or more Cascade Cas proteins, such as one, more, or all of type I CasA-E proteins. The complexes may further comprise Cas3 and optionally one or more homologous type I Cascade proteins. The complexes may further comprise Cas3, 9, 10, 12, or 13. Thus, the functions of a type S system or component can be provided together with the functions of one or more proteins of a different type of CRISPR / Cas system. The complexes may comprise a helicase fused to a nuclease. The nuclease may be, for example, Cas3, 9, 10, 12, or 13. For example, the nuclease is Cas3. For example, the nuclease is Cas9.
[0451] In one example, the complex further comprises a TevI nuclease, e.g., an I-TevI nuclease domain. The I-TevI can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., in Concept 14, or Concepts 21-24, or Concepts 43-45).
[0452] In one example, the complex further comprises a MutH protein.
[0453] In one example, the complex is an isolated complex. The complex can be in vitro. The complex can be contained in a medical container, such as an IV bag, a medical vial, or a medical injection device.
[0454] Concept 68A. A ribonucleoprotein complex comprising: a) one or more proteins described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., described in Concept 13); and b) a crRNA that is homologous to any of the (PAM) (e.g., having one of the sequences described in Concept 39) described in any configuration, concept, aspect, example, embodiment, option or other feature herein; A ribonucleoprotein complex comprising:
[0455] Concept 68B. A ribonucleoprotein complex comprising: a) one or more fusion proteins or fusion protein complexes described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., described in any one of Concepts 14-29); and b) a crRNA that is homologous to any of the (PAM) (e.g., having one of the sequences described in Concept 39) described in any configuration, concept, aspect, example, embodiment, option or other feature herein; A ribonucleoprotein complex comprising:
[0456] Concept 68C. A ribonucleoprotein complex comprising: a) one or more polypeptides described in Concept 5, or one or more nucleic acids encoding such polypeptides; and b) crRNA or one or more nucleic acids encoding the crRNA), wherein the crRNA is cognate to a PAM having the sequence 5'-AAG-3'. A ribonucleoprotein complex containing
[0457] Concept 68D. a) Cas nuclease; b) a DNA or RNA nuclease; or c) nuclease 68C. The complex of claim 68C, wherein
[0458] Concept 69. A method for regulating transcription or replication of a target DNA, comprising contacting the target DNA with any of the complexes described herein (e.g., the complex described in Concept 68), wherein the complex lacks a DNA nuclease, and wherein the complex binds to the target DNA, thereby regulating transcription or replication of the target DNA.
[0459] Concept 70A. A method for controlling replication of a target RNA, comprising: A method comprising contacting a target RNA, or DNA encoding the RNA, with any of the complexes described herein (e.g., with a complex described in Concept 68), wherein the complex binds to the target RNA or DNA, thereby controlling transcription of the target RNA.
[0460] Concept 70B. A method for regulating transcription of a target RNA, comprising: A method comprising contacting a target RNA, or DNA encoding the RNA, with any of the complexes described herein (e.g., with a complex described in Concept 68), wherein the complex binds to the target RNA or DNA, thereby controlling transcription of the target RNA.
[0461] Concept 71. The method of Concept 69 or 70, wherein the DNA is contained in a chromosome or episome (optionally, a plasmid).
[0462] In one embodiment, the DNA is contained in a chromosome. In one embodiment, the DNA is contained in a plasmid. In one embodiment, the DNA is not contained in a plasmid.
[0463] Concept 72. A method of editing target DNA, comprising contacting the target DNA with any of the complexes described herein (e.g., with a complex described in Concept 68), wherein the complex binds to the target DNA, thereby editing the target DNA.
[0464] Concept 73A. A method for cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with any of the complexes described herein (e.g., with the complex described in Concept 68), wherein the complex comprises a fusion protein in which Py comprises a nuclease, and the dsDNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM) sequences described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the nuclease cleaves the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0465] Concept 73B. A method for cleaving RNA, comprising contacting said RNA with any of the complexes described herein (e.g., with a complex described in Concept 68), wherein the complex comprises a fusion protein in which Py comprises an RNA nuclease, and wherein the RNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM) sequences described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the nuclease cleaves the RNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0466] Concept 73C. A method for cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with a complex of Concept 68C or 68D, wherein the complex comprises a nuclease, and the dsDNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the nuclease cleaves the DNA in a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0467] or, - PAM can be AAN, ANG, NAG; or The -PAM can be either no nucleotide change or one AAG.
[0468] The nuclease can be any nuclease described herein (e.g., an I-TevI nuclease described elsewhere herein). The nuclease can be a DNA nuclease. The nuclease can be an RNA nuclease. The nuclease can be a TALEN. The nuclease can be a zinc finger protein. The nuclease can be a Cas nuclease from any naturally occurring system. The nuclease can be a synthetic Cas nuclease.
[0469] Concept 74A. A method for cleaving single-stranded DNA (ssDNA), comprising contacting the ssDNA with any of the complexes described herein (e.g., with the complex described in Concept 68), wherein the complex comprises a fusion protein in which Py comprises a nickase, and the ssDNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM) sequences described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the nickase cleaves the ssDNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0470] Concept 74B. A method for cleaving a single strand in double-stranded DNA (dsDNA), comprising contacting the dsDNA with any of the complexes described herein (e.g., with the complex described in Concept 68), wherein the complex comprises a fusion protein in which Py comprises a nickase, and the dsDNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM) sequences described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the nickase cleaves the dsDNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0471] Concept 74C. A method for cleaving single-stranded DNA (ssDNA), comprising contacting the DNA with a complex of Concept 68C or 68D, wherein the complex comprises a nickase, and the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the nickase cleaves the single strand of the DNA in a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
[0472] The nickase can be any nuclease described herein. The nickase can be a DNA nickase. The nickase can be selected from Nt.BstNBI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nb.BbvCI, Nt.BbvCI, and Nb.BsmI (available from New England BioLabs). The nickase can be a synthetic Cas nickase (e.g., nCas9 or Cas9D10A).
[0473] Concept 75A. A method of marking or identifying a region of DNA, comprising contacting said DNA with any of the complexes described herein (e.g., with the complex of Concept 68), wherein the DNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM)s described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, and optionally, the complex comprises a fusion protein in which Py comprises a detectable label.
[0474] Concept 75B. A method for marking or identifying a region of DNA, comprising contacting said DNA with a complex of Concept 68C or 68D, wherein the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within the region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, and optionally, the complex comprises a detectable label.
[0475] Concept 76A. A method of modifying transcription of a region of DNA, comprising contacting said DNA with any of the complexes described herein (e.g., with a complex described in Concept 68), wherein the DNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM) sequences described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the complex upregulates or downregulates transcription of the DNA or a region of an adjacent gene.
[0476] Concept 76B. A method of altering transcription of a region of DNA, comprising contacting said DNA with a complex of Concept 68C or 68D, wherein the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the complex upregulates or downregulates transcription of a region of the DNA or an adjacent gene.
[0477] In one embodiment, the complex downregulates transcription. Downregulation can be at least about 50% compared to transcription in the absence of the complex. In one embodiment, downregulation is at least about 60%, 70%, 80%, or 90%. In another embodiment, downregulation is at least about 95%, 96%, 97%, 98%, or 99%. In another embodiment, downregulation is 100%, i.e., transcription of the region of DNA is completely blocked.
[0478] In some embodiments, downregulation may be desirable when a gene is overexpressed and causes a negative phenotype, but complete removal of the protein would be harmful. Downregulation of a target gene can be useful for creating and studying knockout phenotypes where producing a complete knockout would be lethal to the cell or organism.
[0479] In some embodiments, the transcription blocking of essential genes in pathogenic bacteria can be used to kill pathogenic bacteria without affecting other beneficial bacteria.Similarly, pathogenic bacteria can be killed by blocking the transcription of replication origin.The blocking of non-essential genes in cells or organisms is useful for knockout and phenotypic behavior research.
[0480] Concept 77A. A method of modifying a target dsDNA in a cell without introducing a dsDNA break, comprising generating in a cell with any of the complexes described herein (e.g., with the complex described in Concept 68), wherein the complex targets a target dsDNA contained in the cell, and the target dsDNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM)s described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39), whereby the complex binds to the dsDNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the dsDNA is modified without introducing a break in the dsDNA.
[0481] Concept 77B. A method for modifying a target dsDNA in a cell without introducing dsDNA breaks, comprising producing in a cell a complex of Concept 68C or 68D, wherein the complex targets a target DNA contained in the cell, and the target DNA contains a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM that is identical except for one base change, whereby the complex binds to the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the DNA is modified without introducing breaks in the DNA.
[0482] Concept 78. The method of Concept 77, wherein DNA is edited, e.g., the editing inserts, deletes, or substitutes bases or nucleic acid sequences in the DNA.
[0483] The modifying can be the introduction of one or more substitutions when the complex is a complex described in Concept 68B and Py is a base editor or prime editor (e.g., any of the base editors or prime editors described elsewhere herein). The modifying can be the introduction of one or more insertions and / or deletions (INDELS) when the complex is a complex described in Concept 68B and Py is a nuclease or nickase and the cell is a cell capable of repairing dsDNA or ssDNA breaks via the NHEJ machinery.
[0484] Concept 79A. A method for inhibiting cell growth or proliferation without introducing lethal dsDNA breaks, comprising producing in a cell any of the complexes described herein (e.g., the complex described in Concept 68), wherein the complex targets a target DNA contained in the cell, and the target DNA comprises a protospacer sequence flanked on its 5' side by any one of the (PAM)s described in any configuration, concept, aspect, example, or embodiment, whereby the complex binds to the DNA within a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer, thereby inhibiting DNA replication without introducing lethal dsDNA breaks in the DNA, thereby inhibiting cell growth or proliferation.
[0485] Concept 79B. A method of inhibiting cell growth or proliferation without introducing lethal dsDNA breaks, comprising producing in a cell a complex of Concept 68C or 68D, wherein the complex targets a target DNA contained in the cell, the target DNA comprising a protospacer sequence flanked on its 5' side by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, thereby inhibiting DNA replication without introducing breaks in the DNA, thereby inhibiting cell growth or proliferation.
[0486] When the complex is the complex described in Concept 68B, Py is a nickase that introduces single-strand breaks in DNA, and the cell is a cell capable of repairing dsDNA breaks or ssDNA breaks via the NHEJ mechanism, cell growth or proliferation can be inhibited by introducing one or more insertions and / or deletions (INDELS). When the complex is the complex described in Concept 68A, cell growth or proliferation can be inhibited by downregulating or blocking transcription of essential genes in the cell. Essential genes are well known to those skilled in the art. When the complex is the complex described in Concept 68A, cell growth or proliferation can be inhibited by downregulating or blocking transcription of chromosomal replication origins in the cell.
[0487] Concept 80. A method of treating or preventing a disease or condition in a human, animal, plant, or fungal subject, comprising performing a method according to any one of concepts 69-79, wherein cells of the subject contain said DNA or RNA, and wherein the cells mediate the disease or condition.
[0488] Concept 81A. The method of Concept 81, wherein the complex modifies a coding or non-coding target sequence.
[0489] In one embodiment, the coding strand of DNA is modified, but not the non-coding strand. In another embodiment, the non-coding strand of DNA is modified, but not the coding strand.
[0490] Concept 81B. The method of any one of Concepts 69-80, wherein the complex modifies (e.g., cuts, edits, blocks, marks, or labels) a coding or non-coding target sequence, and optionally (i) the coding strand of DNA but not the non-coding strand is modified, or (ii) the non-coding strand of DNA but not the coding strand is modified.
[0491] Concept 82. The method of Concept 81, wherein modifying cuts, edits, blocks, marks, or labels the target sequence.
[0492] When the complex is a complex described in Concept 68A, modifying can be blocking. When the complex is a complex described in Concept 68A, modifying can be downregulating transcription of the coding target sequence. When the complex is a complex described in Concept 68B and Py is a nuclease or nickase, modifying can be cleaving the coding target sequence. When the complex is a complex described in Concept 68B and Py is a base editor or prime editor, modifying can be editing the coding target sequence. The base editor can introduce one or more substitutions into the coding target sequence.
[0493] Concept 83A. One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs: 7-11, wherein the nucleotide sequence is operably linked to a heterologous promoter, a synthetic promoter, a eukaryotic promoter, or a non-bacterial promoter.
[0494] Concept 83B. One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs:7-11, wherein the nucleotide sequence is contained in a cell that is not a eukaryotic cell, a non-bacterial cell, or a bacterial cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) that comprises an endogenous nucleotide sequence comprising SEQ ID NOs:7-11.
[0495] Concept 83C. One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs:7-11, wherein the nucleotide sequence is operably linked to a heterologous promoter, a synthetic promoter, a eukaryotic promoter, or a non-bacterial promoter, and wherein the nucleotide sequence is contained in a cell that is not a eukaryotic cell, a non-bacterial cell, or a bacterial cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) that comprises an endogenous nucleotide sequence comprising SEQ ID NOs:7-11.
[0496] Concept 84. A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence that is at least (about) 80% identical to SEQ ID NO:12.
[0497] Concept 85. One or more nucleic acids encoding a plurality of Cas proteins and comprising at least one nucleotide sequence for generating a crRNA, wherein the Cas proteins and RNA are capable of forming a ribonucleoprotein CRISPR / Cas complex, and the RNA is capable of guiding the complex to a protospacer comprised in a target DNA, wherein the 5' end of the protospacer is adjacent to any one of the protospacer adjacent motifs (PAMs) described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39); a) the complex lacks DNA nucleases and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the Cas proteins do not include DinG, Cas3, and Cas10; and c) one or more nucleic acids, wherein the complex does not contain all of the Cas proteins of a type I, II, III, IV, V, or VI CRISPR / Cas complex.
[0498] Concept 86. The nucleic acid of Concept 85, wherein the nucleic acid is contained in a vector, or wherein the nucleic acid is contained in one or more vectors.
[0499] Concept 87. The nucleic acid of Concept 85 or 86, wherein the DNA is contained in a chromosome.
[0500] Concept 88. A ribonucleoprotein CRISPR / Cas complex comprising multiple Cas proteins and a crRNA, wherein the RNA is capable of guiding the complex to a protospacer comprised in a target DNA, wherein the 5' end of the protospacer is adjacent to any one of the protospacer adjacent motifs (PAMs) described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences described in Concept 39); a) the complex lacks DNA nucleases and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the complex lacks Cas3 and Cas10; and c) A ribonucleoprotein CRISPR / Cas complex, wherein the complex does not contain all of the Cas proteins of a type I, II, III, IV, V or VI CRISPR / Cas complex.
[0501] Concept 89. The nucleic acid of any one of Concepts 85-87 or the complex of Concept 88, wherein the crRNA is not a type I, II, III, V or VI crRNA, respectively.
[0502] Concept 90. The nucleic acid or complex of any one of concepts 85-89, wherein the complex is capable of modifying (i) the coding strand of DNA but not the non-coding strand, or (ii) the non-coding strand of DNA but not the coding strand.
[0503] Concept 91. The nucleic acid or complex of any one of concepts 85-90, wherein the complex modifies (e.g., cuts, edits, blocks, marks, or labels) a coding or non-coding target sequence of DNA.
[0504] Concept 92. The nucleic acid or complex of any of Concepts 85-91, wherein the plurality of Cas proteins comprises Cas-S1, S2, S3, S4 and S5.
[0505] For example, the plurality of Cas proteins includes Cas-S1.1, S2.1, S3.1, S4.1 and S5.1.
[0506] Concept 93. The nucleic acid or complex of any one of Concepts 85-92, wherein the complex further comprises an effector protein domain for modifying DNA, and optionally the effector domain comprises nuclease activity, nickase activity, recombinase activity, reverse transcriptase, helicase, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity, or transcriptional repression activity.
[0507] In one example, the complex is a) a nuclease that cleaves a target sequence; b) a deaminase that deaminates a target sequence (e.g., a cytosine or adenine deaminase); c) a base editor that edits the base at the target site; d) a prime editor that prime-edits the target site; e) reverse transcriptase, which reverse transcribes the target site or reverse transcribes RNA within the cell to generate DNA that is inserted into the target site; f) methyltransferases that methylate DNA in cells; g) Methylase, which methylates DNA in cells; h) acetylase, which acetylates DNA within cells; i) acetyltransferases that subject intracellular DNA to acetyltransferase activity; j) a transcriptional activator that activates transcription of a gene in a cell, e.g., a gene that contains or is adjacent to a target site; and k) a transcriptional inactivator that inactivates transcription of a gene in a cell, e.g., a gene that contains or is adjacent to a target site. The present invention includes components selected from the following:
[0508] Concept 94. A cell comprising the nucleic acid or complex of any one of Concepts 85-93, wherein the DNA is comprised in the chromosome of the cell.
[0509] Concept 95. A cell comprising the nucleic acid or complex of any one of Concepts 85-93, wherein the DNA is contained in a plasmid in the cell.
[0510] Concept 96. The cell of Concept 95, wherein the complex is capable of inhibiting cell growth or proliferation.
[0511] Concept 97. A method for modifying DNA, comprising: a) contacting DNA with a nucleic acid as defined in any one of concepts 85-87 and 89-93; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA and modifies the DNA.
[0512] Concept 98. A method for inhibiting replication of a plasmid containing DNA, comprising: a) contacting DNA with a nucleic acid as defined in any one of concepts 85-87 and 89-93; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA to inhibit replication of the plasmid.
[0513] Concept 99. A method for inhibiting transcription of a nucleotide sequence contained in DNA, comprising: a) contacting DNA with a nucleic acid as defined in any one of concepts 85-87 and 89-93; b) allowing the formation of a ribonucleoprotein complex comprising the Cas protein and the crRNA, wherein the complex is guided to a target site contained in the nucleotide sequence and inhibits its transcription.
[0514] Concept 100. A method of inhibiting the growth or proliferation of a cell (optionally a prokaryotic cell, e.g., a bacterial cell) containing DNA, comprising: a) contacting DNA with a nucleic acid as defined in any one of concepts 85-87 and 89-93; b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site in the DNA to inhibit cell growth or proliferation.
[0515] Concept 101. A method of treating or preventing a disease or condition mediated by a target cell in a human or animal subject, comprising performing a method according to any one of concepts 97-100 to modify the target cell, said contacting comprising administering a nucleic acid to the subject, wherein the modification treats or prevents the disease or condition.
[0516] Concept 102. The method of Concept 101, wherein the target cell is a cell of a subject that contains a nucleic acid defect and the modification corrects the defect.
[0517] Concept 103. Modification a) adding new functionality to the cell, and optionally the modification upregulates or downregulates expression of a gene in the cell or adds a new nucleotide sequence to the genome of the cell for expression of a protein encoded by the new nucleotide sequence; or b) modulate the expression of a nucleotide sequence contained in the chromosome or episome (optionally, a plasmid) of a cell; or c) The method of any one of concepts 101 and 102, wherein the growth or proliferation of the cells is inhibited.
[0518] The cell can be a bacterial cell of any genus or species in Table 3. The method can modify a cell of any genus or species in Table 3. The DNA or polynucleotide can be contained in a cell of any genus or species in Table 3. The method can kill or inhibit the growth or proliferation of a cell of any genus or species in Table 3. The cell can be a cell of a genus or species different from one of the genera or species in Table 3, respectively, e.g., the cell is not an E. coli cell, or the cell is not a Klebsiella cell. The cell can be an archaea, such as a methanogen. The method can kill or inhibit the growth or proliferation of an archaea. The method can edit the genome of a bacterial or archaea cell. Any prokaryotic cell (e.g., a bacterial or archaea cell) or fungal cell herein can be contained in a microbiota, such as a microbiota of a human, animal, plant, or environment (such as soil or waterway). For example, the microbiota is contained in the digestive tract (GI tract) of a human or animal. For example, the microbiota is contained in the urinary system of a human or animal, e.g., the bladder, kidney, or urethra. For example, the microbiota is contained in the blood of a human or animal. For example, the microbiota is contained in the eye, nose, or ear of a human or animal. For example, the microbiota is contained in the hair of a human or animal. For example, the microbiota is contained in the skin of a human or animal. For example, the microbiota is contained in the leaves, roots, stems, or seeds of a plant.
[0519] The disease or condition herein may be selected from any of the following: [Table A-1] [Table A-2]
[0520] Neurodegenerative or CNS Diseases or Conditions for Treatment or Prevention by the Method In one example, the neurodegenerative or CNS disease or condition is selected from the group consisting of Alzheimer's disease, geriatric psychosis, Down's syndrome, Parkinson's disease, Creutzfeldt-Jakob disease, diabetic neuropathy, Parkinsonism, Huntington's disease, Machado-Joseph disease, amyotrophic lateral sclerosis, diabetic neuropathy, and Creutzfeldt-Jakob disease. For example, the disease is Alzheimer's disease. For example, the disease is Parkinsonism.
[0521] In one example, any of the methods described herein is practiced on a human or animal subject to treat a CNS or neurodegenerative disease or condition, and the method causes downregulation of the subject's Treg cells, thereby promoting the infiltration of systemic monocyte-derived macrophages and / or Treg cells across the choroid plexus into the subject's brain, thereby treating, preventing, or reducing the progression of the disease or condition (e.g., Alzheimer's disease). In one embodiment, the method causes an increase in IFN-gamma in the subject's CNS system (e.g., in the brain and / or CSF). In one example, the method restores nerve fibers and / or reduces the progression of nerve fiber damage. In one example, the method restores nerve myelin and / or reduces the progression of nerve myelin damage. In one example, any of the methods described herein treat or prevent a disease or condition disclosed in WO 2015 / 136541 and / or the methods can be used in conjunction with any of the methods disclosed in WO 2015 / 136541 (the disclosure of which is incorporated herein by reference in its entirety to provide a disclosure of such methods, diseases, conditions and potential therapeutic agents, e.g., agents such as immune checkpoint inhibitors, e.g., anti-PD-1, anti-PD-L1, anti-TIM3 or other antibodies disclosed therein, that can be administered to a subject to treat and / or prevent CNS and neurodegenerative diseases and conditions). Cancers to be treated or prevented by this method
[0522] Cancers that can be treated include non-vascularized or substantially non-vascularized tumors, and vascularized tumors.Cancers can include non-solid tumors (e.g., hematological tumors, such as leukemia and lymphoma) or solid tumors.The types of cancers that can be treated by the methods, proteins (including fusion proteins and fusion protein complexes), vectors, complexes, and compositions described herein include, but are not limited to, carcinomas, blastomas, and sarcomas, as well as certain leukemias or lymphoid malignancies, benign and malignant tumors, and malignant tumors, such as sarcomas, carcinomas, and melanomas.Adult tumors / cancers and pediatric tumors / cancers are also included.
[0523] Hematological cancer is cancer of blood or bone marrow.Examples of hematological (or hematopoietic) cancer include acute leukemia (for example, acute lymphocytic leukemia, acute myeloid leukemia, acute myelogenous leukemia and myeloblastic, promyelocytic, myelomonocytic, monocytic and erythroleukemia), chronic leukemia (for example, chronic myelogenous (granulocytic) leukemia, chronic myelogenous leukemia and chronic lymphocytic leukemia), polycythemia vera, lymphoma, Hodgkin's disease, non-Hodgkin's lymphoma (indolent and aggressive), multiple myeloma, Waldenstrom's macroglobulinemia, heavy chain disease, myodysplastic syndrome, hairy cell leukemia and myelodysplasia.
[0524] A solid tumor is an abnormal mass of tissue that usually does not contain cysts or liquid areas. Solid tumors can be benign or malignant. Various types of solid tumors are named for the type of cells that form them (such as sarcoma, carcinoma, and lymphoma). Examples of solid tumors, such as sarcoma and carcinoma, include fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, and other sarcomas, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, lymphoid malignancies, pancreatic cancer, breast cancer, lung cancer, ovarian cancer, prostate cancer, hepatocellular carcinoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, medullary thyroid carcinoma, papillary thyroid carcinoma, pheochromocytoma, sebaceous gland carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchial carcinoma, renal cell carcinoma, hepatic carcinoma, thyroid ... Cancer, cholangiocarcinoma, choriocarcinoma, Wilms' tumor, cervical cancer, testicular tumor, seminoma, bladder cancer, melanoma, and CNS tumors (e.g., gliomas (such as brainstem gliomas and mixed gliomas), glioblastomas (also known as glioblastoma multiforme), astrocytomas, CNS lymphomas, germinomas, meningiomas, schwannomas, craniopharyngiomas, ependymomas, pinealomas, hemangioblastomas, acoustic neuromas, oligodendroglioma, meningiomas, neuroblastomas, retinoblastomas, and brain metastases).
[0525] Autoimmune diseases for treatment or prevention by the method [Table B-1] [Table B-2] [Table B-3]
[0526] Inflammatory diseases for treatment or prevention by the method [Table C]
[0527] It will be understood that the specific embodiments described herein are shown by way of illustration and not as limitations of the invention. The principal features of this invention can be employed in various embodiments without departing from the scope of the invention. Those skilled in the art will recognize, or be able to ascertain using no more than routine research, numerous equivalents to the specific procedures described herein. Such equivalents are considered to be within the scope of the present invention and are encompassed by the claims. All publications and patent applications mentioned herein are indicative of the level of skill of those skilled in the art to which this invention pertains. All publications, patent applications, issued patents, and other documents mentioned herein are incorporated by reference as if each individual publication, patent application, issued patent, or other document was specifically and individually indicated to be incorporated by reference in its entirety. Reference is made to the publications mentioned herein and equivalent publications from the United States Patent and Trademark Office (USPTO) or WIPO, the disclosures of which are incorporated by reference herein to provide a disclosure that may be used in the present invention and / or to provide one or more features (e.g., vectors) that may be included in one or more claims herein.
[0528] The use of the terms "a" or "an" in the claims and / or specification when used in conjunction with the term "comprising" can mean "one," but is also consistent with the meanings of "one or more," "at least one," and "one or more than one." The use of the term "or" in the claims is used to mean "and / or" unless expressly indicated to refer to alternatives only, or unless the alternatives are mutually exclusive, although the present disclosure supports a definition that refers to alternatives only and "and / or." Throughout this application, the term "about" is used to indicate that a value includes the inherent variation of error of the device, method being employed to determine the value, or the variation that exists among study subjects.
[0529] As used in this specification and claims, the terms "comprising" (and any form of "comprise" such as "comprise" and "comprises"), "having" (and any form of "having" such as "have" and "has"), "including" (and any form of "includes" and "include"), or "containing" (and any form of "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0530] As used herein, "or combinations thereof" or similar terms refer to all permutations and combinations of the listed items preceding the term. For example, "A, B, C, or combinations thereof" includes at least one of A, B, C, AB, AC, BC, or ABC, and is intended to also include BA, CA, CB, CBA, BCA, ACB, BAC, or CAB if order is important in a particular situation. Continuing with this example, combinations containing one or more repeats of an item or term, such as BB, AAA, MB, BBC, AAABCCCC, CBBAAA, CABABB, etc., are expressly included. Those skilled in the art will understand that there is typically no limitation on the number of items or terms in any combination unless it is clear from the context.
[0531] Any part of this disclosure may be read in combination with any other part of this disclosure, unless otherwise clear from the context.
[0532] All of the compositions and / or methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of the present invention have been described with reference to preferred embodiments, it will be apparent to those skilled in the art that variations can be made in the compositions and / or methods and in the steps or sequence of steps of the methods described herein without departing from the concept, spirit and scope of the invention. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.
[0533] The invention is described in more detail in the following non-limiting examples. [Example]
[0534] [Example 1] 1.1 Abstract This example describes the in-silico identification and in vivo (in E. coli) functional characterization of a previously uncharacterized CRISPR-Cas system whose targeting activity is independent of the predicted RNA-guided Cas endonuclease. We demonstrate that the system i) successfully targets plasmid DNA by an unknown mechanism, preventing plasmid-induced antibiotic resistance, and ii) efficiently binds to chromosomal DNA without inducing cell death. We refer to this new system as the "Type S CRISPR / Cas system."
[0535] 1.1.1 Purpose Characterize a CRISPR system that differs from currently known functionally characterized systems.
[0536] 1.2 Materials and Methods 1.2.1 Plasmid and strain construction The plasmids and strains used in this study are listed in the Appendix (see section 1.5 below). Plasmids were constructed by InFusion HD™ cloning (Takara) from PCR-generated DNA fragments. Insertion of the target sequence into the chromosome of the E. coli C-1 strain was performed using lambda red-mediated recombination. The resulting strain is b5451(ΔlacZ).
[0537] 1.2.2 Transformation assay for sequence-specific plasmid inhibition Cells were first transformed with the p1624 plasmid (shown in Figure 1) and the plasmid shown in Figure 3, which contain components of a novel CRISPR / Cas system (referred to herein as the Type S CRISPR / Cas system). To evaluate the functionality of the plasmids carrying the Type S system components, E. coli cells were transformed with these plasmids separately and subsequently with a 1:1 mixture of two other plasmids, p1631 and p892. The targeting plasmid p1631 carries the system's CRISPR array and all of the native target sequence (5 protospacers) of the amilCP gene, conferring a purple color to colonies carrying the plasmid. Plasmid p892 served as a non-targeting control and contained the same origin of replication (p15A) and antibiotic marker gene (CmR) as the targeting plasmid, but lacked the target sequence and amilCP gene.
[0538] Cells that did not contain any components of the type S system were also transformed with the p1631 / p892 plasmid mixture as a control to assess the ratio of the two plasmids in the mixture.
[0539] Transformed cells were plated on LB agar plates containing the appropriate antibiotic to select for the presence of the plasmid. Plates were incubated overnight at 37°C, and the proportion of purple and white colonies carrying the target and non-target plasmids, respectively, was determined.
[0540] 1.2.3 Plasmid inhibition using controlled expression of components of the Type S system Strain b5456, carrying plasmid p1826, identical to p1793 (Fig. 3 ), but containing the type S system under the control of an arabinose-inducible promoter (instead of pBolA), was transformed by electroporation with a 1:1 mixture of two plasmids, p1631 and p892. Cells were allowed to recover for 1.5 h and plated on LB agar containing the appropriate antibiotic with or without 0.2% arabinose.
[0541] 1.2.4 Growth inhibition assay A 1:1 mixture of E. coli C-1 and b5451 competent cells was transformed with plasmid p1826, which contains the necessary components of the Type S CRISPR-Cas system, and allowed to recover for 1.5 hours before plating on LB agar containing 20 μg / mL tetracycline to select for transformants carrying the plasmid. The plates also contained X-gal (20 ng / mL) and IPTG (1 mM) to differentiate E. coli C-1 (lacZ+, blue) and b5451 (lacZ-, white) colonies.
[0542] 1.2.5 Plasmid clearance assay The b5408 strain was transformed separately with plasmids p1935, p1936, p1937, p1938, p1631, and p144 (see Table 2 and Figure 7). Serial dilutions of the recovered transformation solution were plated onto LB agar plates supplemented with 10 μg / mL tetracycline and 20 μg / mL chloramphenicol. The plates were incubated overnight at 37°C, and the resulting colony-forming units were counted to calculate the effect of the type S system on the transformation efficiency of the b5408 strain.
[0543] 1.2.6 Chromosome binding assay The b5872 strain was transformed with the p1985, p1993, p1994, p1995, and p1996 plasmids (see Table 2). The transformation mixture was plated on LB agar plates supplemented with 10 ng / mL tetracycline and incubated overnight at 37°C. Three randomly selected single colonies per transformation were used to inoculate LB cultures supplemented with 10 ng / mL tetracycline. After 20 hours of incubation at 37°C with rigorous shaking at 220 rpm, the cultures reached late stationary phase. 1 μL from each culture was added to 99 μL of LB medium supplemented with 10 ng / mL tetracycline and either 0 μg / mL, 20 μg / mL, 25 μg / mL, 30 μg / mL, or 35 μg / mL chloramphenicol. The cultures were transferred to wells of a microtiter plate and incubated in a plate reader for 24 hours at 37°C with rigorous shaking (900 rpm). The optical density of the cultures was monitored throughout the incubation period with readings at 600 nm every 10 minutes.
[0544] 1.3 Results 1.3.1 Type S CRISPR / Cas system We cloned into a low-copy plasmid DNA encoding an RNA-guided endonuclease (rge), two open reading frames (ORFs) downstream of the rge gene, a cognate CRISPR array, and four ORFs immediately upstream (5') of the CRISPR array (Figure 1). ORFs 140 and 423 were predicted to encode an error-prone DNA polymerase, while ORFs 624 and 188 were predicted to encode a helicase and an RNA endonuclease, respectively. The resulting p1624 plasmid contained the inserted genes under the transcriptional control of a constitutive promoter. The CRISPR array, meanwhile, contained five spacers. The corresponding protospacers were preceded by a 5'-AAG-3' protospacer-adjacent motif (PAM) of the type S CRISPR-Cas system. We constructed a target plasmid (p1631) with four predicted protospacers (corresponding to spacers #1, #2, #4, and #5) flanked at their 5' ends by 5'-AAG-3'PAM (Figure 1).
[0545] We then tested the in vivo functionality of the system with p1624. The p1624 and p1631 plasmids were used for the sequence-specific plasmid inhibition assay described in Section 1.2.2 of the Materials and Methods section. Briefly, cells harboring the p1624 plasmid were transformed with a 1:1 mixture of the p1631 targeting plasmid and a control (non-targeting) plasmid. p1631 expresses a purple protein, facilitating screening. Both the targeting and non-targeting plasmids were transformed with similar efficiency, demonstrating that the current setup of the system does not induce plasmid targeting. Therefore, it was hypothesized that this system requires additional subunits for activity. We extended our work by cloning six additional ORFs (ORFs 83, 135, 136, 162, 145, and 286; see Figure 3) into p1624 upstream (5') of ORF 188. All additional ORFs were transcribed in the same direction as the previously cloned ORFs and CRISPR array. We used the resulting p1760 plasmid instead of p1624 to repeat the plasmid inhibition assay. p1760 inhibited the growth of colonies transformed with the p1631 targeting plasmid, but allowed the growth of colonies transformed with non-targeting plasmids, as all colonies appeared white (Figure 2). The results indicated that p1760 contained an active CRISPR-Cas system.
[0546] 1.3.2 Deletion analysis to identify minimal functional systems We aimed to identify the minimum number of elements required for sequence-specific plasmid inhibition / targeting by the CRISPR-Cas system. We created a library of plasmids carrying a series of gene deletions in p1760 (Figure 3) and repeated the plasmid inhibition assay (Materials and Methods, section 1.2.3).
[0547] Eleven of the 12 novel constructs we tested showed target plasmid inhibition (Figure 4). No plasmid inhibition was observed in the absence of a CRISPR array, as demonstrated by assays using p1899 (bottom row, right panel of Figure 4). Plasmid inhibition was observed in assays using p1793 (bottom row, third panel of Figure 4), which contains an operon of only five genes (encoding Cass, designated Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5).
[0548] Each of these genes was separately deleted from p1793, resulting in the construction of plasmids p2009 (ΔCas-S1), p2011 (ΔCas-S3), p2012 (ΔCas-S4), and p2013 (ΔCas-S5) (see Table 2). The plasmids were transformed into E. coli NEB10b cells, generating strains b5700 (ΔCas-S1), b5702 (ΔCas-S3), b5703 (ΔCas-S4), and b5704 (ΔCas-S5), respectively (see Table 1). Each of these strains, along with the negative and positive targeting control strains b4816 and b5408, was transformed with the targeting plasmid p1631 and plated on agar plates with the appropriate antibiotic selection. The transformation efficiency of b5408 was reduced by approximately three orders of magnitude, whereas that of all other strains was not (see Figure 5). These results demonstrate that the Type S CRISPR-Cas system requires all five subunits for plasmid inhibition. It is noteworthy that there is no predicted RNA-guided DNA endonuclease gene among these five genes, indicating that the mechanism of action of the CRISPR-Cas system likely does not depend on DNA cleavage.
[0549] To further confirm the functionality of the Type S CRISPR-Cas system, we cloned a DNA fragment containing the minimal system and a CRISPR array downstream of an arabinose-inducible promoter. The resulting plasmid (p1826) could be maintained in the same cells with the p1631 target plasmid in the absence of arabinose (Figure 6, left panel). In the presence of arabinose, the CRISPR-Cas system was expressed, resulting in the inhibition of p1631 and, surprisingly, the inhibition of bacterial growth (Figure 6, right panel).
[0550] 1.3.3 Spacers in Type S CRISPR arrays show similar targeting efficiency We evaluated the targeting efficiency of each spacer in the CRISPR array used in these experiments in the Type S CRISPR-Cas system. The array contained five spacers, and the p1631 targeting plasmid contained four protospacers from these five spacers. Using p1631 as a basis, we constructed four additional targeting plasmids, p1935, p1936, p1937, and p1938, each containing only one protospacer (see Table 2 and Figure 7, left panel). We separately transformed all five targeting plasmids and the non-targeting p144 plasmid (see Table 2) into the b5408 strain and plated serial dilutions of the transformation mixture on appropriate selective plates (Figure 7, right panel). The next day, colony-forming units were counted for analysis; all four spacers tested demonstrated similar targeting efficiencies, reducing the transformation efficiency of the corresponding targeting plasmid by three orders of magnitude compared to the non-targeting plasmid. Furthermore, the combined targeting efficiency of all four spacers was not significantly higher than the efficiency of each spacer separately.
[0551] 1.3.4 Type S CRISPR-Cas systems bind to chromosomal DNA We determined that the Type S CRISPR-Cas system does not contain a predicted RNA-guided DNA endonuclease. Therefore, we sought to determine whether the system's subunits contained a novel dsDNA cleavage domain capable of efficiently introducing lethal dsDNA breaks into the E. coli chromosome. We constructed strain b5451 by transferring four previously tested protospacers from p1631 to the E. coli C-1 chromosome and inactivating the lacZ gene. The resulting strain was distinguishable from the original non-target E. coli C-1 strain on X-gal plates, as the latter strain metabolized X-gal and converted it into a blue pigment. Both the constitutive expression system and the arabinose-inducible Type S CRISPR-Cas expression system were tested as described in Section 1.2.6 of the Materials and Methods section. Briefly, a 1:1 mixture of E. coli C-1 and b5451 competent cells was transformed with the p1826 plasmid, which contains the type S CRISPR-Cas system under the control of an arabinose-inducible promoter. The ratio of white to blue colonies remained 1:1 under appropriate antibiotic selection conditions, regardless of induction conditions (Figure 8). This result further supported the possible absence of a dsDNA endonuclease domain in the system.
[0552] We then wondered whether the type S CRISPR-Cas system has a DNA-binding function or whether its mode of action is at the RNA level. We constructed the bSNP5810 strain by introducing the crm gene into the intergenic region of E. coli C-1. We then constructed one negative control strain, designated p1985, and four crm-targeting plasmids, designated p1993, p1994, p1995, and p1996 (see Table 2). To this end, we replaced the spacer in the CRISPR array of p1793 with either a non-targeting spacer or a single spacer targeting four different positions in the promoter or coding region of the crm gene (Figure 9). We transformed the bSNP5810 strain with each of p1993, p1994, p1995, and p1996 and cultured the cells in medium supplemented with different chloramphenicol concentrations. In the absence of chloramphenicol, the growth rates of all strains carrying the CRM targeting spacer were similar to that of the control strain (Figure 10). Nevertheless, increasing the chloramphenicol concentration proportionally reduced the growth rates of all target strains (Figure 10, graphs 2-4). The growth rate of the control strain remained unaffected by the presence of chloramphenicol in the medium. These results demonstrate that the Type S CRISPR-Cas system binds steadily to the chromosome and acts as a surprisingly efficient transcription inhibitor.It is noteworthy that the orientation of the crm-targeting spacer does not affect the binding efficiency of the type S CRISPR-Cas system and the subsequent downregulation of crm expression, which is not the case for other previously known CRISPR-Cas systems that must target either the coding or non-coding regions of a gene to efficiently downregulate its transcription (Qi et al, Cell, 152(5), 1173-1183, 2013, doi: https: / / doi.org / 10.1016 / j.cell.2013.02.022; Rath et al., NAR, 43(1), 237-246, 2015, doi: https: / / doi.org / 10.1093 / nar / gku1257 and Luo et al. al., NAR, 43(1), 674-681, 2015, doi:https: / / doi.org / 10.1093 / nar / gku971).
[0553] 1.4 Conclusion We characterized a nuclease-deficient CRISPR-Cas system (which we refer to as type S CRISPR / Cas) and determined its minimal functional elements. Our studies revealed that type S CRISPR / Cas contains five Cas proteins (which we refer to as Cas-S1 through Cas-S5) and a cognate CRISPR array. None of the Cas proteins were shown to have DNA nuclease activity. Surprisingly, this system inhibits the growth of bacterial cells containing at least one target sequence flanked by an AAG PAM. Furthermore, this system can be programmed to efficiently bind to chromosomal targets and act as a transcription inhibitor without inducing lethal dsDNA breaks. Using these findings, we identified and characterized a novel, functional CRISPR / Cas system that can be used in a minimal form to inhibit plasmids, reduce bacterial cell growth, and inhibit transcription from the chromosome. We explore the utility of individual Cas proteins in the system, such as by creating fusions of type S Cas with nuclease domains to effect targeted DNA cleavage. Other fusion partners include base editors, prime editors, deaminases and reverse transcriptases.
[0554] 1.5 Supplementary Notes [Table 1]
[0555] [Table 2-1] [Table 2-2]
[0556] Example 2: Identification and characterization of PAMs in type S CRISPR / Cas systems 2.1 Background SNIPR Biome identified and characterized the Type S CRISPR / Cas system described in Example 1.
[0557] To identify potential PAM sequences for the Type S CRISPR / Cas system described in Example 1, we conducted an in-house in silico analysis of public genome databases. This analysis was based on spacers from wild-type Type S CRISPR arrays. A small number of potential protospacers were discovered that contained a 5'-AAG-3' sequence motif at the 5' end but no common sequence motif at the 3' end. It was hypothesized that this 5'-AAG-3' motif at the 5' end of the protospacer could function as an effective PAM for Type S systems. This hypothesis was confirmed in plasmid targeting / killing assays (described in Examples 1.3.1-1.3.3). Nevertheless, because the PAM sequences of most CRISPR-Cas systems are promiscuous, we decided to further experimentally investigate various sequences recognized by Type S CRISPR / Cas as effective PAMs.
[0558] 2.2 Experimental design The Type S CRISPR / Cas-based PAM identification process was based on the following plasmids: 1. PAM plasmid library (p2259); each member of this library contains the protospacer sequence present in the target plasmid p1935 (see Example 1.2.5) and a mixture of candidate 5 nt long PAMs at the 5' end of the protospacer.
[0559] 2. Targeting type S plasmid (p1793); the plasmid expresses a spacer complementary to the protospacer of the type S CRISPR / Cas system (Cas-S1 to Cas-S5) and the PAM plasmid library.
[0560] 3. Non-targeting (control) Type S plasmid (p1985): The plasmid expresses a spacer that is not complementary to the protospacers of the Type S CRISPR / Cas system (Cas-S1 through Cas-S5) and the PAM plasmid library. Furthermore, the spacer is not complementary to any part of the E. coli genome.
[0561] The steps of the PAM specific assay are as follows: 1. Construction of the PAM plasmid library designated as p2259 (see section below).
[0562] 2. Construction of targeted E. coli NEB Turbo strain (b5408) by transforming NEB Turbo strain b3127 (NEB) with p1793 (type S targeting) plasmid.
[0563] 3. Construction of the non-targeted E. coli NEB Turbo strain (b6259) by transforming the NEB Turbo strain b3127 (NEB) with the p1985 (control type S) plasmid.
[0564] 4. Transformation of b5408 (type S targeting) and b6259 (control type S) with the p2259 plasmid library and cultivation under appropriate antibiotic selection conditions.
[0565] 5. Isolation of p2259 plasmid libraries from both cultures and preparation of NGS fragments containing the PAM-protospacer region.
[0566] 6. Analysis of NGS and sequencing results.
[0567] 2.2.1 Step 1: Construction of the PAM library The PAM library plasmid p2259 was designed to contain a 5 nt long variable PAM region at the 5' end of the protospacer. 5 It should consist of (1024) different PAMs.
[0568] Construction of the PAM library plasmid p2259 was based on the previously described target plasmid p1935 (see Example 1.2.5). The p1935 plasmid was used as a template for PCR amplification using appropriately designed primers. CloneAMP PCR premix (Takara) was used in the PCR reaction. The forward primer was designed to introduce a 5-nt degenerate sequence at the 5' end of the protospacer present on the p1935 plasmid. The gel-purified PCR product was used for In-Fusion cloning using 5x In-Fusion HD Enzyme Premix (Takara). Five microliters from the In-Fusion cloning reaction was transformed into 50 μL of NEB10b electrocompetent cells (NEB) by electroporation. The cells were then allowed to recover in SOC medium (final volume 500 μL) at 37°C for 1 hour with shaking at 200 rpm. The harvested cells were transferred to 500 mL of LB liquid medium supplemented with 20 μg / mL chloramphenicol. From the 500 mL culture, 100 μL was removed and plated onto an LB agar plate supplemented with 20 μg / mL chloramphenicol. The culture and plate were incubated overnight at 37°C with rigorous shaking (200 rpm). The next day, 536 colonies were counted on the plate, and then calculated: 2.68 million successfully transformed cells were added to the 500 mL culture. Equation 1 was used for this calculation.
[0569]
number
[0570] This number of transformed cells in the culture corresponds to 2,617× coverage of the PAM library members, according to Equation 2.
[0571]
number
[0572] Glycerol stocks were made from the 500 mL culture, and the constructed p2259 plasmid library was then isolated from the remaining culture using a Qiagen Plasmid Midi Kit.
[0573] 2.2.2 Step 2: Transformation of b5408 (CRIPSR-S targeting) and b6259 (control type S) strains with the p2259 plasmid library Some members of the p2259 plasmid library are expected to contain a PAM recognized by the type S CRISPR / Cas system. These PAM-positive plasmids will be successfully targeted, and their replication will be blocked by the type S system. Therefore, b5408 (type S-targeted) cells transformed with PAM-positive members of the p2259 plasmid library should not survive in medium supplemented with chloramphenicol, because the p2259 plasmid library contains a chloramphenicol resistance marker gene. On the other hand, b5408 (type S-targeted) cells transformed with PAM-negative members of the p2259 plasmid library should survive in medium supplemented with chloramphenicol. In contrast, none of the p2259 plasmids in the library are targeted by the type S CRISPR / Cas system expressed in the b6259 (control type S) strain and survive in chloramphenicol-supplemented medium.
[0574] The p2259 plasmid library was electroporated into electrocompetent cells of b5408 (type S targeting) and b6259 (control type S), respectively. Electrocompetent cells from both strains were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Five transformation reactions per strain were performed using 40 ng of the p2259 plasmid library per reaction. Cells were allowed to recover in 500 μL of SOC medium at 37°C for 1 hour with shaking at 200 rpm. The harvests from each strain were combined, and each of the last two harvests (one per strain) was used to inoculate a 500 mL LB liquid medium culture supplemented with 20 μg / mL chloramphenicol and 10 μg / mL tetracycline. The cultures were incubated overnight at 37°C with rigorous shaking at 200 rpm. The next day, glycerol stocks from each culture were prepared and the p2259 plasmid library from each strain was isolated using a Qiagen Plasmid Midi Kit.
[0575] Plasmid isolation from a culture of b5408 (type S targeting) transformed with p2259 followed by NGS indicates which PAM sequences are not recognized by the Type-S CRISPR / Cas system and which PAM sequences are subsequently required by the system. Plasmid isolation from a culture of b6259 (control type S) transformed with p2259 followed by NGS indicates the abundance of each plasmid member in the p2259 library. This information is used to normalize the NGS results from the b5408 strain (type S targeting). 2.2.3 Step 3: Identifying PAMs by NGS
[0576] p2259 isolated from b6259 (control type S) should contain all members of the p2259 PAM library, providing quantitative information about the abundance of each PAM plasmid in the library.
[0577] The p2259 plasmids isolated from b5408 (type S targeting) and b6259 (control type S) were subjected to next-generation sequencing (NGS). The first-round NGS PCR amplification step for both p2259 preparations was performed using appropriately designed primers in 50 μL of KAPA HiFi Hotstart ReadyMix reaction (Roche) (25 cycles: 30 seconds of denaturation, 20 seconds of annealing, and 15 seconds of extension). The reaction products were purified with AMPure XP reagent (Beckman Coulter). Subsequently, Illumina index and flow cell binding sequences were added to the purified products using a second NGS PCR amplification step. The first-round NGS PCR product from strain b5408 (type S targeting) was subjected to a second NGS PCR amplification step using appropriately designed primers. The first-round NGS PCR product from strain b6259 (control CRISPR-S) was subjected to a second NGS PCR amplification step using appropriately designed primers. Both amplifications were in 50 μL KAPA HiFi Hotstart ReadyMix reactions (Roche) (10 cycles, 20 ng template DNA, 30 s denaturation, 20 s annealing, 15 s extension).
[0578] Both reaction products were purified with AMPure XP reagent (Beckman Coulter). The purified products were sequenced using an Illumina MiSeq® in a single-end 150-cycle sequence.
[0579] 2.3 Results NGS data was used to compare the relative abundance of each member of the p2259 library in strains b5408 (type S targeted) and b6259 (control). Data analysis revealed a clear preference of the type S CRISPR / Cas system for the following PAM sequences (Figures 11 and 12): 5'-AHN-3' 5'-KAG-3' 5'-AGG-3' 5'-GAC-3' 5'-GTG-3'
[0580] [Example 3] Mismatch tolerance of protospacers in type S CRISPR / Cas systems 3.1 Background CRISPR / Cas systems exhibit different spacer-protospacer complementarity requirements for DNA / RNA binding and targeting. For example, the Escherichia coli type IE CRISPR / Cas system tolerates spacer-protospacer mismatches outside the "seed" region without compromising its immune system function (see Semenova et al., PNAS, 108(25), 10098-10103, 2011, doi:https: / / doi.org / 10.1073 / pnas.110414410), and Streptococcus pyogenes Cas9 requires 15 nt of spacer-protospacer complementarity for DNA cleavage in vitro (see Jinek et al., Science, 337(6096), 816-821, 2012, doi:10.1126 / science.12258), but tolerates multiple (up to 15 nucleotides) spacer-protospacer mismatches for DNA binding in bacteria (Ciu et al., Nature Comm.,9(1912),2018,doi:https: / / doi.org / 10.1038 / s41467-018-04209-5).
[0581] Type S CRISPR / Cas systems do not contain a predicted dsDNA nuclease domain. It has previously been shown that the presence of the Cas-S3 helicase subunit is required for the plasmid clearance activity of Type S CRISPR / Cas systems (see Example 1.2.5). Spacer-protospacer mismatches are expected to adversely affect the DNA target binding stability of the CRISPR-Cas complex of the prehelicase Type S CRISPR / Cas system (Type S CRISPR-Cas complexes that do not contain a helicase subunit but contain the following subunits: Cas-S1, -S2, -S4, and -S5). Reduced DNA binding efficiency of Type S CRISPR / Cas systems is expected to result in reduced plasmid clearance. Therefore, the assay described in this example was designed to use plasmid curing / transformation efficiency as a measurable output of the effect of spacer-protospacer mismatches on the CRISPR-Cas targeting efficiency of Type S CRISPR / Cas systems.
[0582] 3.2 Experimental design Plasmid p1935 was previously designed (see Example 1.2.5) and constructed to contain a 32-bp protospacer and a PAM sequence (5'-AAG-3') targeted by the Type S CRISPR / Cas system. In this example, various variants of the p1935 plasmid were designed and constructed containing various combinations of mutations in the protospacer region. These mutations partially disrupt the complementarity of the protospacer to the targeting spacer. The accompanying transformation efficiency graph (Figure 13) shows these mutations in detail.
[0583] Using appropriately designed primers, various protospacer mutations were introduced into the protospacer sequence of the p1935 plasmid by Q5® Site-Directed Mutagenesis (NEB). The resulting PCR products were circularized using the KLD Enzyme Mix (NEB) reaction. The KLD reaction mixture was transformed into NEB-10β electrocompetent cells (NEB). The cells were then incubated in SOC medium (final volume 500 μL) at 37°C for 1 hour and allowed to recover by shaking at 200 rpm. The recovered cells were plated on LB agar plates supplemented with 20 μg / mL chloramphenicol for plasmid selection. The plates were incubated overnight at 37°C with a single colony from each plate into a 10 mL LB liquid medium culture supplemented with 20 μg / mL chloramphenicol for plasmid selection. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm). Each culture was sampled for preparation of glycerol stocks and stored at −80° C. The remaining culture was used for plasmid isolation using the Qiagen Plasmid Mini Kit.
[0584] b5408 is a NEB Turbo E. coli strain containing the plasmid p1793 (which contains the wild-type type S CRISPR / Cas system; Cas-S1 through Cas-S5 contain type S CRISPR arrays). 50 ng of each member of the p1935 plasmid library was electroporated into electrocompetent b5408 (wt type S) cells. Electrocompetent cells from both strains were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Cells were allowed to recover in 500 μL of SOC medium at 37 °C for 1 h with shaking at 200 rpm. The harvest from each strain was used to prepare serial dilutions, which were then spot-plated onto LB agar plates supplemented with the appropriate antibiotic, 20 μg / mL chloramphenicol, and 10 μg / mL tetracycline, for plasmid selection. Transformations were performed in biological triplicate. Colonies formed from each transformation were counted, and the reduced transformation efficiency of the b5408 (wt Type S) strain compared to its transformation efficiency when transformed with a negative (non-targeting) control plasmid demonstrated how the corresponding spacer-protospacer mismatch affects the targeting efficiency of the Type S CRISPR / Cas system (see Figure 13).
[0585] Control: The original plasmid p1935 is a positive control plasmid that always produced very few transformants (only escape mutants).
[0586] Plasmid p2361 is a negative control plasmid. It is identical to p1935 except for the 5'-AAG-3' PAM sequence upstream (5' side) of the protospacer, which is replaced with a 5'-CGG-3' sequence that is not predicted to be a PAM for type S CRISPR / Cas (see Example 2). Therefore, the transformation efficiency of p2361 (-ve control) should be the highest among all library members mentioned in this example.
[0587] 3.3 Results FIG. 13 shows an analysis of the transformation efficiency for each of the mismatches.
[0588] Careful examination of the graph led to the following conclusions regarding how spacer-protospacer mismatches affect the plasmid clearance efficiency of type S CRISPR / Cas systems: 1.5 consecutive PAM distal mismatches are well tolerated. 2. Eight consecutive PAM distal mismatches completely abolish targeting. 3. A single mismatch at the first or second PAM-proximal position reduces targeting by three orders of magnitude. 4. A single mismatch at the 3rd to 10th PAM-proximal positions reduces targeting by 0 to 1 order of magnitude. 5. Double mismatches at the first to third PAM-proximal positions reduce targeting by 3-4 orders of magnitude. 6. Double mismatches at PAM-proximal positions 4-10 reduce targeting by 0-1 order of magnitude. 7. Triple mismatches at positions 1-5 proximal to the PAM reduce targeting by 3-4 orders of magnitude. 8. Triple mismatches at PAM-proximal positions 4-9 reduce targeting by 0-1 order of magnitude. 9. A triple mismatch at PAM-proximal positions 8-10 reduces targeting by four orders of magnitude. 10. Quadruple mismatches at PAM-proximal positions 1-8 reduce targeting by 3-4 orders of magnitude. 11. A quadruple mismatch at PAM-proximal positions 6-9 reduces targeting by an order of magnitude. 12. A quadruple mismatch at PAM-proximal positions 7-10 reduces targeting by 5 orders of magnitude.
[0589] [Example 4] Type S CRISPR / Cas system for CRISPRi applications 4.1 Aiming In this study, we evaluated i) the efficiency of the type S CRISPR / Cas system to bind to chromosomal and plasmid targets in E. coli , acting as a CRISPRi tool, and ii) the subunits essential for DNA binding.
[0590] 4.2 Introduction We previously demonstrated that the Type S CRISPR / Cas system, while lacking a DNA nuclease domain, drives crRNA-dependent impairment of plasmid replication (see Examples 1.3.2 and 1.3.3). The crRNA-guided DNA-binding activity of the Type S CRISPR / Cas system can be utilized for the development of CRISPR interference (CRISPRi) tools in E. coli. We anticipate that it can also be used in other cell types. In this example, the CRISPRi potential of the Type S CRISPR / Cas system was evaluated by targeted downregulation of transcription and subsequent translation of the reporter green fluorescent protein (gfp) gene. The target gfp gene was either chromosomally integrated or plasmid-derived. Furthermore, this example revealed which Type S CRISPR / Cas system subunits are essential for the DNA-binding activity of the Type S CRISPR / Cas system.
[0591] 4.3 Methods and Results Five variants of type S CRISPR / Cas systems were tested for their CRISPRi potential; i) Wild type (wt)-S CRISPR / Cas system ii) Cas-S1 deletion mutant (ΔCas-S1) type S CRISPR / Cas system iii) Cas-S3 deletion mutant (ΔCas-S3) Type S CRISPR / Cas system iv) Cas-S4 deletion mutant (ΔCas-S4) Type S CRISPR / Cas system v) Cas-S1 / Cas-S3 double deletion mutant (ΔCas-S1ΔCas-S3) type S CRISPR / Cas system.
[0592] All variants were placed under the transcriptional control of the pBolA promoter (SEQ ID NO: 14) and cloned together with a non-targeting crRNA expression module to generate plasmids p1985 (wt), p2062 (ΔCas-S1), p2064 (ΔCas-S3), p2065 (ΔCas-S4), and p2160 (ΔCas-S1ΔCas-S3), respectively (Figure 14). The 5' and 3' ends of the crRNA module spacer contained recognition sites for the BsmBI-v2 (NEB) restriction enzyme, which facilitated cloning of the targeting spacer (Figure 14). All plasmids contained the repA101 replication protein gene, the SC101 low-copy replication origin, and the tetR tetracycline resistance marker gene (Figure 14).
[0593] 4.4 Chromosome CRISPRi For the CRISPRi assay of the chromosome type S CRISPR / Cas system, strains MG1655 and b5815 (gfp+ve strain) were used. b5815 is an E. coli MG1655 derivative constructed to contain the gfp gene under the control of the p70a promoter (SEQ ID NO: 15) and positioned between the ybcB and ybcC genes (Figure 15). Sixteen protospacers within or near the gfp gene of the b5815 (gfp+ve) strain were selected and targeted for the chromosome CRISPRi assay (Figure 15, Table 4, SEQ ID NOs: 25 to 41). Eight of the protospacers were located on the gfp coding strand (i.e., protospacers U2kF, U1kF, S1F, S2F, S3F, S4F, S5F, S6F) and eight were located on the non-coding strand (i.e., protospacers U2kR, U1kR, S1R, S2R, S3R, S4R, S5R, S6R).
[0594] 4.4.1 Method Construction of plasmids expressing different type S CRISPR / Cas system variants and targeting the above-mentioned protospacers was performed by digesting the p1985 (wt), p2062 (ΔCas-S1), p2064 (ΔCas-S3), p2065 (ΔCas-S4), and p2160 (ΔCas-S1ΔCas-S3) plasmids with BsmBi-v2 (NEB) restriction enzyme and then ligating pre-annealed complementary ssDNA oligos encoding the corresponding spacers into the digested plasmids using T4 DNA ligase (NEB). Five µL of each ligation reaction was transformed into 50 µL of NEB10β electrocompetent cells (NEB) by electroporation. Cells were then allowed to recover in SOC medium (final volume 500 µL) at 37 °C for 1 h with shaking at 200 rpm. 100 μL of each harvest was removed and plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol. Plates were incubated overnight at 37°C.
[0595] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 10 μg / mL tetracycline. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0596] The next day, glycerol stocks were prepared for each culture, and the remaining culture was used for plasmid isolation using the Qiagen Plasmid Mini Kit. Each plasmid was sequence-verified by next-generation sequencing. Each plasmid was transformed into electrocompetent b5815 (gfp+ve strain) cells. Electrocompetent cells were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were performed using 40 ng of each plasmid per reaction. Cells were allowed to recover in 500 μL of SOC medium at 37°C for 1 hour with shaking at 200 rpm. 100 μL from each harvest was plated onto LB agar plates supplemented with 10 μg / mL tetracycline. Plates were incubated overnight at 37°C.
[0597] The next day, three single colonies were randomly selected from each transformation and used to inoculate 5 mL cultures of M9 minimal medium supplemented with glycerol (0.4% w / v) and 10 μg / mL tetracycline. The cultures were incubated overnight at 37°C with rigorous shaking (200 rpm).
[0598] The next day, 2 μL from each culture was transferred to 198 μL of M9 medium supplemented with glycerol (0.4% w / v) and 10 μg / mL tetracycline. The resulting 200 μL cultures were transferred to separate wells of a Thermo Scientific™ Nunc microWell 96-well optical-bottom plate with a polymer base. The plates were sealed with Breathe-Easy® sealing membranes (Sigma-Aldrich) and loaded onto a Synergy H1 microplate reader (Biotek) for incubation at 37°C with continuous shaking. The optical density (absorbance at 600 nm) and green fluorescence emission (excitation at 485 nm, emission at 516 nm) of the cultures were measured at 10-minute intervals for 24 hours. Fluorescence emission data from each culture and each time point were normalized to the corresponding OD600 nm data. Normalized data from cultures belonging to biological triplicates were averaged, and the corresponding standard deviation was calculated. The final data was plotted in the graph shown in Figure 16.
[0599] 4.4.2 Results The type S CRISPR / Cas system downregulated GFP expression from b5815 (a gfp +ve strain) to the non-fluorescent control strain (b230) levels when targeting the protospacers at the beginning of the gfp orf and the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R) (Figure 16A) or the end of the gfp orf and the 200 bp region upstream and downstream of gfp (protospacers S4F, S4R, S5F, S6F, S6R) (Figure 16B), demonstrating highly efficient and tight CRISPRi performance. The only exception was targeting the S5R protospacer, which downregulated GFP expression by approximately 50%. GFP expression was also downregulated when the type S CRISPR / Cas system targeted protospacers located 1 kb upstream of the gfp orf (protospacers U1kF, U1kR), whereas GFP expression was completely unaffected for protospacers 2 kb upstream of the gfp orf (U2kF, U2kR) (Figure 16C).
[0600] The helicase-deficient ΔCas-S3 type S CRISPR / Cas system extensively or completely downregulated GFP expression from b5815 when targeting the promoter p70a region or the forward strand of the gfp orf (protospacer S1F, S1R, S2F, S2R, S3F), but had no effect on GFP expression when targeting the reverse strand of the gfp orf (protospacer S3R) (Figure 17A). Here, we demonstrated that Cas-S3 cannot unwind DNA and subsequently blocks transcription of genes located 2 kb away from the target protospacer.
[0601] To confirm that the function of the Type S CRISPR / Cas system CRISPRi is Cas-S3 dependent when targeting the extreme end or outer protospacers of a gene's ORF, targeting assays for the protospacers S4F, S4R, S5F, S5R, S6F, and S6R were repeated using the ΔCas-S3 Type S CRISPR / Cas system (Figure 17B). Indeed, GFP expression in b5815 (a gfp+ve strain) was unaffected for selected targets in the absence of Cas-S3. These results indicate that Cas-S3 is not essential for the formation of the RNP complex of the Type S CRISPR / Cas system and its binding to DNA targets.
[0602] The ΔCas-S4 Type S CRISPR / Cas system did not affect GFP expression from b5815 (a gfp+ve strain), indicating that the subunit is essential for the formation of the RNP complex of the Type S CRISPR / Cas system and / or its binding to the DNA target (Figure 18A).
[0603] The ΔCas-S1 Type S CRISPR / Cas system and the ΔCas-S1, ΔCas-S3 Type S CRISPR / Cas systems reduced GFP expression from b5815 (a GFP+ve strain) in a similar manner, but only when the reverse strand of the p70a promoter was targeted (fully or partially) (Figures 18B, 18C). These results indicate that Cas-S1, although important for its stability and / or the efficiency of its DNA-binding activity, is not essential for the formation of the RNP complex of the Type S CRISPR / Cas system and its binding to the DNA target. Furthermore, these results indicate that Cas-S1 is essential for loading of Cas-S3 into the complex.
[0604] 4.5 CRISPRi plasmids The b6386 strain was used for CRISPRi assays of the plasmid Type S CRISPR / Cas system. b6386 is an E. coli MG1655 strain transformed with the p2370 plasmid, which contains the gfp gene (under the control of the p70a promoter, SEQ ID NO: 15), the p15a origin of replication, and the cmR chloramphenicol resistance marker gene. Only four of the five Type S variants tested for chromosomal CRISPRi in Example 4.4 were also tested for plasmid CRISPRi. The wild-type Type S CRISPR / Cas system containing Cas-S1-Cas-S5 was not tested due to its demonstrated activity in inhibiting plasmid replication (see Examples 1.3.2 and 1.3.3). Six protospacers within the p70a promoter region and the gfp gene in b6386 were selected for the on-plasmid CRISPRi assay, which was the same as the chromosomal CRISPRi assay; three of the protospacers were placed on the gfp coding strand (i.e., protospacers S1F, S2F, S3F) and three were placed on the non-coding strand (i.e., protospacers S1R, S2R, S3R) (see Figure 15).
[0605] 4.5.1 Method Construction of various type S expression plasmids targeting these protospacers was carried out as previously described in Example 4.4.1.
[0606] Plasmid transformation was performed as described in Example 4.4.1 with the following two modifications: 1.100 μL from each harvest was plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol instead of 10 μg / mL tetracycline.
[0607] 2.2) The inoculum culture was supplemented with 20 μg / mL chloramphenicol instead of 10 μg / mL tetracycline.
[0608] 3. The final data was plotted on the graph shown in FIG.
[0609] 4.5.2 Results ΔCas-S3 Type S completely or efficiently (though not completely) downregulated GFP expression from b6386 (a gfp +ve strain) only when it targeted the protospacer within the promoter p70a region (S1F and S1R), regardless of the orientation of the target strand (Figure 19A). None of the other tested variants downregulated GFP expression (Figure 19B, Figure 19C, Figure 19D). This is in stark contrast to the results obtained by chromosomal CRISPRi experiments.
[0610] Example 5: Type S CRISPR / Cas system in base editing applications 5.1 Introduction Currently developed and reported CRISPR / Cas base editing tools rely on single-subunit effector modules from different class 2 CRISPR / Cas systems. Class 1 CRISPR / Cas systems have multi-subunit effector modules, complicating the development of base editing tools. As a result, the base editing capabilities of class 1 base editors remain unexplored, for example, in terms of editing specificity, efficiency, and window.
[0611] In this example, we extensively investigated the potential of the Type S CRISPR / Cas system to transform it into a flexible base editing platform. We designed and constructed a collection of base editors by fusing a 16-amino acid XTEN linker (SEQ ID NOS: 16 and 17, see Schellenberger et al., Nat. Biotechnol., 27(12), 1186-1190, 2009, doi:10.1038 / nbt.1588) or PmCDA1 cytidine deaminase from sea lamprey (SEQ ID NOS: 18 and 19, see Nishida et al., Science, 353(6305), aaf8729, 2016, doi:10.1126 / science.aaf8729) to either the N- or C-terminus of the Cas-S1, Cas-S3, or Cas-S4 subunit of the Type S CRISPR / Cas system. The Uracil DNA glycosylase inhibitor (UGI) protein (SEQ ID NOS: 20 and 21, see Mol et al., Cell, 82(5), 701-8, 1995, doi:10.1016 / 0092-8674(95)90467-0) was fused to the C-terminus of each chimera using a 10 amino acid linker (SEQ ID NOS: 22 and 23). These base editors were then tested for their ability to generate targeted chromosomal C to T modifications.
[0612] 5.2 Experimental design 5.2.1 Overview of screening strains and plasmids: The gfp gene was inserted into the chromosome of E. coli MG1655 under the transcriptional control of the p70a promoter (SEQ ID NO: 15). The resulting b5815 (gfp+ve) strain was used as the test strain for all base editing assays. Six protospacer targets for the base editing assays were selected to have a 5'-AAG-3' PAM at their 5' ends and to be located within a 249-bp-long region consisting of the last 52 bp of the p70a promoter and the first 197 bp of the gfp sequence. Three of the protospacers were on the template (with respect to transcription) strand, and three were on the non-template strand (Figure 15).
[0613] Two categories of plasmids were designed and constructed for base editing assays: The first category contained "base editor" plasmids responsible for expression of the pm_cda1 gene fused to either the N- or C-terminus of either the Cas-S1, Cas-S3, or Cas-S4 genes (see left side of Figure 20). Expression of the chimeras was set under the transcriptional control of an arabinose-inducible promoter (p BAD , SEQ ID NO: 24, see Guzman et al., J. Bacteriol., 177(14), 4121-4130, 1995, doi:10.1128 / jb.177.14.4121-4130.1995). The backbone of the plasmid contains a chloramphenicol resistance marker gene (crm) for selection against the antibiotic chloramphenicol. R ) and the p15a origin of replication.
[0614] All PCR reactions were performed using Q5 Hi Fi 2X Master Mix (NEB) with appropriately designed primers and PCR templates to construct the plasmids. PCR fragments were used in NEBuilder® HiFi DNA Assembly Master Mix reactions. Five microliters from each assembly reaction was transformed into 50 μL of NEB10β electrocompetent cells (NEB) by electroporation. Cells were then allowed to recover in SOC medium (final volume 500 μL) at 37°C for 1 hour with shaking at 200 rpm. 100 μL of each harvest was plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol. Plates were incubated overnight at 37°C.
[0615] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 20 μg / mL chloramphenicol. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0616] The next day, glycerol stocks were prepared for each culture, and the remaining culture was used for plasmid isolation using a Qiagen Plasmid Mini Kit. Each plasmid was sequence-verified by next-generation sequencing.
[0617] The second category contained i) crRNA expression modules (one each for non-targeting control, S1F, S1R, S2F, S2R, S3F, and S3R) and ii) "guide" plasmids responsible for expression of different versions of the Type S Cas operon, in which either the Cas-S1, Cas-S3, or Cas-S4 genes were deleted (see Figure 20, right-hand plasmid). Expression of the crRNA was set under the control of the original leader (promoter) sequence, as found in the CRISPR locus of the Type S CRISPR / Cas system. Expression of the operon was controlled by the E. coli genome (p bolA The plasmid backbone contained a tetracycline resistance marker gene (tet R ) and the pSC101 origin of replication.
[0618] The construction method for the “guide” plasmid was the same as in the first category, except that (i) 10 μg / mL tetracycline was used for plating onto LB agar instead of 20 μg / mL chloramphenicol, and (ii) 10 μg / mL tetracycline was used for inoculating LB liquid medium instead of 20 μg / mL chloramphenicol.
[0619] Subsequently, the p2062 (non-targeting control, ΔCas-S1), p2064 (non-targeting control, ΔCas-S3), p2065 (non-targeting control, ΔCas-S4), and p2160 (non-targeting control, ΔCas-S1, ΔCas-S3) plasmids were restriction digested with BsmBI-v2 (NEB), and the digested products were gel-purified using a Zymoclean Gel DNA Recovery Kit (Zymoresearch). Appropriately designed oligos were annealed and ligated to the digested plasmids using T4 DNA ligase (NEB). Five microliters from each ligation reaction was transformed into 50 μL of NEB10β electrocompetent cells (NEB) by electroporation. The cells were then allowed to recover in SOC medium (final volume 500 μL) at 37 °C for 1 h with shaking at 200 rpm. 100 μL of each harvest was removed and plated onto LB agar plates supplemented with 10 μg / mL tetracycline. Plates were incubated overnight at 37°C.
[0620] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 10 μg / mL tetracycline. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0621] The next day, glycerol stocks were prepared for each culture, and the remaining culture was used for plasmid isolation using a Qiagen Plasmid Mini Kit. Each plasmid was sequence-verified by next-generation sequencing.
[0622] 5.2.2 Base editing assays Each of the six "base editor" plasmids carrying the PmCDA1 cytidine deaminase fused to either the N- or C-terminus of each of the wt type S subunits, Cas-S1, Cas-S3, and Cas-S4 (see the left side of Figure 20), was transformed into electrocompetent b5815 cells. Electrocompetent cells were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were performed using 40 ng of each plasmid per reaction. Cells were allowed to recover in 500 μL of SOC medium at 37°C for 1 hour with shaking at 200 rpm. 100 μL from each harvest was plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol. Plates were incubated overnight at 37°C.
[0623] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 20 μg / mL chloramphenicol. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0624] The next day, electrocompetent cells were prepared for each of the six resulting b5815 strains using the base-editing plasmids according to the protocol described at (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Each of the six strains was then electrotransformed with the corresponding guide plasmids, as shown on the right side of Figure 20. These plasmids were responsible for the expression of i) the type S subunit, which is not expressed by the paired base-editor plasmid, and ii) the targeting and non-targeting crRNAs. Transformation reactions were performed using 40 ng of each plasmid per reaction. Cells were allowed to recover in 500 μL of SOC medium at 37°C for 1 hour with shaking at 200 rpm. 100 μL of each harvest was plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol, 10 μg / mL tetracycline, and 0.2% v / w arabinose. Plates were incubated overnight at 37°C.
[0625] The next day, single colonies were randomly selected and subjected to colony PCR using appropriately designed genome-specific primers. The PCR products were sequenced by Sanger sequencing, and the results were analyzed for the detection of C to T or G to A modifications when targeting and modifying the reverse strand.
[0626] 5.3 Results Representative results from a base editing assay are shown in Figure 21. A detailed analysis of the Sanger sequencing data from all base editing assays is provided below:
[0627] 5.3.1 PmCDA1 linked to the N-terminus of ●Cas-S1 (WT type S): No base editing Cas-S3 (WT type S): A single C to T mutation in only one of the targeted protospacers Cas-S4 (WT type S): A single C to T mutation in only one target protospacer (same term as Cas-S3-N, CDA)
[0628] 5.3.2 PmCDA1 linked to the C-terminus of ●Cas-S1: Protospacer S1F: No base editing Protospacer S1R: No base editing Protospacer S2F: Base editing at multiple positions with varying levels of efficiency (105nt editing window). Not only C to T conversions, but also a few G to A conversions. Protospacer S2R: Four positions (29nt editing window) efficiently base-edited with varying levels of efficiency per position. Not only a G to A conversion, but also one C to T conversion. Protospacer S3F: Eight positions (35nt editing window) were base-edited with varying levels of efficiency per position. Not only G to T conversions, but also one G to A conversion. Protospacer S3R: A few positions (outside the protospacer region) were base-edited with low efficiency. Streaking improved the efficiency and extended the editing window (96 nt). A mixture of C to T and G to A conversions was detected.
[0629] Cas-S1 (in the absence of Cas-S3: type S_ΔCas-S3): Protospacer S1F: No base editing Protospacer S1R: A single G to A alteration within the protospacer Protospacer S2F: Multiple positions were modified from C to T with varying levels of efficiency within the protospacer region and 7 nt downstream of the protospacer region. Two additional C to T conversions were detected within a 102 nt editing window upstream of the 5' end of the protospacer. Protospacer S2R: Four positions (48 nt editing window) efficiently base-edited with varying levels of efficiency per position. Only one weak G-to-A modification at three PAM-proximal positions, and three efficient C-to-T modifications downstream of the protospacer. Protospacer S3F: Multiple very weak modifications Protospacer S3R: Only a few very weak modifications (a mixture of C to T and G to A conversions) and one very efficient C to T modification 40 nt downstream of the 3' end of the protospacer were detected.
[0630] ●Cas-S3: Protospacer S1F: No base editing Protospacer S1R: A major single position was base-edited in the middle of the protospacer, with several additional positions edited inefficiently. Streaking extended the base-editing window (50 nt), but not only G-to-A conversions but also a few C-to-T conversions were detected. Protospacer S2F: No base editing Protospacer S2R: Multiple positions were inefficiently edited within a wide editing window (up to 208 nt). Primarily, C to T conversions were detected, as well as G to A conversions. Streaking improved the efficiency, as it resulted in three clean conversions (both G to A and C to T) within a 178 nt window. Protospacer S3F: Multiple base edits were performed (388 nt editing window). Conversions within the protospacer region were the cleanest. Only two positions were found with G to A conversions. Protospacer S3R: No base editing was initially detected, but streaking resulted in multiple G to A and C to T conversions within the 388 nt editing window.
[0631] ●Cas-S4: All targeted protospacers were edited at multiple positions within the protospacer region, with the exception of S2R, which was also edited downstream of the protospacer region, and S3R, which had a single G-to-A mutation at the second PAM-distal position of the protospacer. Protospacers S1F and S2R had a mixture of C to T and G to A modifications, whereas the rest of the protospacers had either C to T or G to A modifications depending on the orientation of the protospacer.
[0632] 5.4 Conclusion This example generated the first reported Class I CRISPR / Cas platform for efficient base editing. Seven of the developed base editors utilized the PmCDA1 cytidine deaminase. Six base editors utilized the wild-type Type S CRISPR / Cas system, and one base editor utilized the ΔCas-S3 Type S CRISPR / Cas system.
[0633] Fusion of PmCDA1 to the N-terminus of Cas-S1 did not result in any detectable modification.
[0634] Fusion of PmCDA1 to the N-terminus of Cas-S4 and Cas-S3 resulted in only a single detectable mixed (wild-type / mutant) modification for one of the spacers tested. These systems can be used when highly specific point mutations are required. However, the limited number of available spacers limits the applicability of these systems.
[0635] Fusion of PmCDA1 to the C-terminus of Cas-S1, Cas-S4, and Cas-S3 resulted in three base editing systems with very different base editing results (editing window, editing efficiency, and editing location) even on the same protospacer. In general, Cas-S1- and Cas-S3-based base editors have very wide editing windows, which can be very beneficial for generating amino acid libraries or introducing stop codons for a single protein. In the absence of Cas-S3, the editing window of the Cas-S1-based base editor was reduced, and a shift in editing location was recorded. The Cas-S4-based base editor had a narrower base editing window and was the only base editor that introduced modifications to all targeted protospacers.
[0636] Example 6: Functionalization of Type S CRISPR / Cas system for dsDNA cleavage 6.1 Purpose Type S is a nuclease-deficient CRISPR-Cas system that lacks a bioinformatically predicted or experimentally determined (see Examples 1 and 4 herein) nuclease domain in any of its subunits. The goal of this study was to expand the repertoire of genome editing reagents by fusing the nuclease domain of the I-TevI homing endonuclease (SEQ ID NOs: 44 and 45) to the N-terminus of the CasS1 and CasS4 subunits of a helicase-deficient Type S system (ΔCasS3, see Figure 22) to create a synthetic crRNA-guided DNA nuclease. The developed chimeras were tested on DNA targets derived from chromosomes and plasmids, establishing a set of rules for efficient dsDNA cleavage.
[0637] 6.2 Introduction I-TevI is a homing endonuclease consisting of an N-terminal nuclease domain (1-92 amino acids (aa)), a zinc finger-containing spacer domain (92-169 aa), and a DNA binding domain (169-245 aa) (see Dgell et al., PNAS, 98(14), 7898-7903, 2001, doi: https: / / doi.org / 10.1073 / pnas.141222498 and Kleinstiver et al., PNAS, 109(21), 8061-8066, 2012, doi: https: / / doi.org / 10.1073 / pnas.111798410). I-TevI binds to its DNA target as a monomer and introduces sequence-specific staggered dsDNA cleavage via the GIY-YIG motif present in its N-terminal nuclease domain. The naturally occurring I-TevI cleavage site sequence is 5'-CAACG-3' (forward strand) / 5'-CGTTG-3' (reverse strand). Nevertheless, I-TevI has been reported to recognize the more flexible 5'-CNNNG-3' (forward strand) / 5'-CNNNG-3' (reverse strand) I-TevI cleavage site (see both Edgell et al., 2001 and Kleinstiver et al., 2012, supra). We hypothesized that it would be possible to create synthetic TevCasSx (x = 1 or 4) crRNA-guided DNA nucleases by fusing the N-terminus of I-TevI and the zinc finger / linker domain to the N-terminus of CasS1 or CasS4.
[0638] In Example 4, CRISPRi experiments demonstrated that the helicase-deficient ΔCasS3 type S system retained strong DNA binding activity. Furthermore, in the same example, it was demonstrated that the helicase-expressing ΔCasS1 type S system exhibited the same CRISPRi / DNA binding efficiency as the helicase-deficient ΔCasS3 type S system, but the ΔCasS1ΔCasS3 type S system exhibited reduced CRISPRi / DNA binding activity. These results support a model in which CasS1 is involved in the recruitment of CasS3 to the type S complex. Therefore, we decided to use the ΔCasS3 type S helicase-deficient system to construct the TevCasSx synthetic nuclease, aiming to avoid interference with the nuclease activity of the TevCasSx chimera by CasS3 recruitment or helicase activity.
[0639] 6.3 Experimental Design 6.3.1 TevCasS Plasmid DNA Cleavage Assay TevCasS Plasmids - For the DNA cleavage assay, three categories of plasmids were designed, constructed, and used: i) TevCasS targeting plasmids, ii) TevCasSx expression plasmids, and iii) Type S guide plasmids.
[0640] Twenty-three TevCasS target plasmids (designated p2497-p2510, p2575-p2578, and p2655-p2659, see Figure 22, right-hand plasmids) were designed to contain various TevCasS target sites (SEQ ID NOS: 46-68, respectively, see Table 4) separately cloned into a low-copy plasmid backbone consisting of a low-copy pBBR1 replicon and a kanamycin resistance marker gene. The first TevCasS targeting plasmid, designated p2497 (Table 4), was designed to encompass a 71-nt-long target site with the following structure from 5' to 3': i) a 5-nt-long I-TevI 5'-CAACG-3' (SEQ ID NO: 42) cleavage site, followed by ii) a 31-nt-long native I-TevI DNA spacer sequence 5'-CTCAGTAGATGTTTTCTTGGGTCTACCGTTA-3' (SEQ ID NO: 46) that interacts with the I-TevI zinc finger domain, followed by iii) a 3-nt-long type S PAM sequence 5'-AAG-3' (SEQ ID NO: 13), followed by iv) a 32-nt-long protospacer sequence 5'-TCGTGTTGTCCACGGTTACCCGCTGGCTGGAA-3' (SEQ ID NO: 69) that is complementary to the sequence of the first spacer in the native CRISPR array of the type S CRISPR-Cas system.The target sites of the additional TevCasS target plasmids (Table 4, SEQ ID NOS: 46-68) had the same I-TevI cleavage site, PAM, and protospacer portion as the target on the p2497 plasmid, but their I-TevI DNA spacer portions were different; Various numbers of nucleotides were removed from or added to the 3' end of the DNA spacer sequence to construct target plasmids with I-TevI DNA spacer sequences of 5 nt (SEQ ID NO:59), 7 nt (SEQ ID NO:58), 9 nt (SEQ ID NO:57), 11 nt (SEQ ID NO:56), 13 nt (SEQ ID NO:55), 15 nt (SEQ ID NO:54), 17 nt (SEQ ID NO:53), 19 nt (SEQ ID NO:52), 21 nt (SEQ ID NO:51), 23 nt (SEQ ID NO:50), 25 nt (SEQ ID NO:49), 27 nt (SEQ ID NO:48), 28 nt (SEQ ID NO:64), 29 nt (SEQ ID NO:47), 30 nt (SEQ ID NO:65), 31 nt (SEQ ID NO:46), 32 nt (SEQ ID NO:66), 33 nt (SEQ ID NO:60), 34 nt (SEQ ID NO:67), 35 nt (SEQ ID NO:61), 36 nt (SEQ ID NO:68), 37 nt (SEQ ID NO:62), or 39 nt (SEQ ID NO:63) in length.
[0641] For the construction of I-TevI target plasmids, all PCR reactions were performed using Q5 Hi Fi 2x MasterMix (NEB) and appropriately designed primers and PCR templates. PCR fragments were used in KLD Enzyme Mix (NEB) reactions. Five microliters from each reaction was transformed into 50 μL of NEB10β electrocompetent cells (NEB) by electroporation. Cells were then allowed to recover in SOC medium (final volume 500 μL) at 37°C for 1 hour with shaking at 200 rpm. 100 μL of each recovery was plated onto LB agar plates supplemented with 50 μg / mL kanamycin. Plates were incubated overnight at 37°C.
[0642] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 50 μg / mL kanamycin. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0643] The next day, glycerol stocks were prepared for each culture, and the remaining culture was used for plasmid isolation using a Qiagen Plasmid Mini Kit. Each plasmid was sequence-verified by next-generation sequencing.
[0644] Plasmids expressing TevCasS1 (p2461) and TevCasS4 (p2462) (see Figure 22, middle plasmid) were designed such that the first 618 bp of the I-tevI gene (SEQ ID NO: 44), which expresses the N-terminal 206 amino acids of the I-TevI nuclease (SEQ ID NO: 45), was linked to the N-terminus of the casS1 and casS4 genes using XTEN linker expression sequences (SEQ ID NOs: 16 and 17). Expression of the chimeras was set under the transcriptional control of an arabinose-inducible promoter (p BAD , SEQ ID NO: 24) (see Guzman et al., J. Bacteriol., 177(14), 4121-4130, 1995, doi:10.1128 / jb.177.14.4121-4130.1995). The backbone of the constructed plasmid contained a chloramphenicol resistance marker gene (crm) for selection against the antibiotic chloramphenicol. R ) and the p15a origin of replication.
[0645] The construction method for the expression plasmid was the same as for the target plasmid, except that (i) the NEBuilder® HiFi DNA Assembly Master Mix was used instead of the KLD Enzyme Mix (NEB), (ii) 20 μg / mL chloramphenicol was used instead of 50 μg / mL kanamycin for plating on LB agar, and (iii) 20 μg / mL chloramphenicol was used instead of 50 μg / mL kanamycin for inoculating LB liquid medium.
[0646] Four type S guide plasmids were designed (see Figure 22, left-hand plasmid): two non-targeting (negative) control and two targeting plasmids. The guide plasmids carried i) a crRNA expression module (either a non-targeting control crRNA or a crRNA targeting the protospacer within the TevCasS target sequence) and ii) expression of either the ΔCasS1ΔCasS3 or ΔCasS1ΔCasS4 version of the type S Cas operon (Table 2). Expression of the crRNA was set under the control of the original leader (promoter) sequence, as found in the CRISPR locus of the wild-type type S CRISPR / Cas system. Expression of the operon was set under the transcriptional control of the constitutive promoter of the bolA gene (SEQ ID NO: 14) present in the E. coli genome (pbolA). The guide plasmids share the same backbone containing a tetracycline resistance marker gene (tetR) for selection on the antibiotic tetracycline, and the low-copy-number pSC101 origin of replication.
[0647] The construction method for the guide plasmid was the same as for the expression plasmid, except that (i) 10 μg / mL tetracycline was used for plating onto LB agar instead of 20 μg / mL chloramphenicol, and (ii) 10 μg / mL tetracycline was used for inoculating LB liquid medium instead of 20 μg / mL chloramphenicol.
[0648] For TevCasS plasmid-DNA cleavage assays, the following plasmid combinations were transformed into E. coli strain MG1655 (see also Figure 22, left side) by electroporation: 1. p2461 (expressing TevCasS1) and p2160 (non-targeting control guide) 2. p2461 (expressing TevCasS1) and p2511 (TevCasS targeting guide) 3. p2462 (expressing TevCasS4) and p2460 (non-targeting control guide) 4. p2462 (expressing TevCasS4) and p2512 (TevCasS targeting guide)
[0649] Electrocompetent MG1655 cells were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were performed using 40 ng of each plasmid per reaction. Cells were then left to recover in SOC medium (final volume 500 μL) at 37°C for 1 hour with shaking at 200 rpm. 100 μL of each harvest was plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol and 10 μg / mL tetracycline. Plates were incubated overnight at 37°C. The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 20 μg / mL chloramphenicol and 10 μg / mL tetracycline. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm). The next day, glycerol stocks were prepared for each culture, and the cultures were used for the preparation of electrocompetent cell plasmids according to the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Each of the four strains was transformed with the target plasmid (Figure 22, right) and a p2027 control plasmid identical to the target plasmid but lacking the TevCasS target sequence. The harvest from each strain was used to prepare serial dilutions, which were then spot-plated onto LB agar plates supplemented with the appropriate antibiotics: 20 μg / mL chloramphenicol, 10 μg / mL tetracycline, and 50 μg / mL kanamycin for plasmid selection. Transformations were performed in biological triplicate. Successful introduction of a dsDNA break into the target plasmid would result in elimination of the plasmid and subsequent growth inhibition in the presence of kanamycin.
[0650] Colonies formed from each transformation were counted. The reduction in transformation efficiency of the two strains expressing targeted crRNA compared to the two strains expressing non-targeted crRNA was calculated in two steps: 1) normalizing the transformation efficiency of each strain for each targeted plasmid by dividing it by the corresponding transformation efficiency of the control plasmid p2027, and 2) dividing the normalized transformation efficiency of the strain with targeted crRNA by the corresponding normalized transformation efficiency of the strain with non-targeted crRNA.
[0651] 6.3.2 TevCasS chromosomal-DNA cleavage assay TevCasS target sequences with I-TevI spacers of 29 nt (SEQ ID NO: 47), 30 nt (SEQ ID NO: 65), 31 nt (SEQ ID NO: 46), 32 nt (SEQ ID NO: 66), and 33 nt (SEQ ID NO: 60) were inserted into the chromosome of E. coli MG1655 to generate five TevCasS chromosomal target strains, each containing a different target. All constructed strains were transformed with the p2461 (TevCasS1 expression) plasmid. Electrocompetent cells were prepared using the protocol described (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were performed using 40 ng of each plasmid per reaction. Cells were then allowed to recover in SOC medium (final volume 500 μL) at 37°C for 1 hour with shaking at 200 rpm. 100 μL of each harvest was removed and plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol. Plates were incubated overnight at 37°C.
[0652] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid culture supplemented with 20 μg / mL chloramphenicol. The culture was incubated overnight at 37°C with rigorous shaking (200 rpm).
[0653] The next day, glycerol stocks were prepared for each culture, which were then ...
Claims
1. A nucleic acid vector or vectors comprising an expressible nucleotide sequence, said sequence comprising: a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:
5. A nucleic acid vector or a plurality of nucleic acid vectors comprising:
2. At least one of the nucleotide sequences is i. heterologous to at least one nucleotide sequence; ii. is not an E. coli, Pseudomonas or Klebsiella promoter; iii. It is a eukaryotic promoter; iv. an animal promoter (optionally a mammalian or human promoter); v. Is a plant promoter; vi. a fungal promoter (optionally a yeast promoter); vii. It is an insect promoter; viii. a viral promoter (optionally a viral, AAV, or lentiviral promoter); or ix. Synthetic promoters The vector of claim 1 , wherein the vector is operably linked to a promoter which is
3. 3. The vector of claim 1, wherein at least one of the nucleotide sequences is operably linked to a constitutive promoter.
4. 3. The vector of claim 1, wherein at least one of the nucleotide sequences is operably linked to an inducible promoter.
5. A vector according to any one of claims 1 to 4, lacking a nucleotide sequence encoding a polypeptide according to part c) of claim 1.
6. A vector according to any one of claims 1 to 5, lacking a nucleotide sequence encoding a polypeptide according to part a) of claim 1.
7. A nucleic acid vector comprising an expressible nucleotide sequence, a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:2; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:4; and c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:
5. A nucleic acid vector comprising:
8. The vector of claim 7, further comprising a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:
1.
9. 9. The vector of claim 7 or claim 8, further comprising a nucleotide sequence encoding a polypeptide comprising an amino acid sequence that is at least 90% identical to SEQ ID NO:
3.
10. one or more nucleotide sequences for generating a crRNA, wherein the crRNA comprises a spacer that is cognate to a first protospacer in a target sequence, and optionally, the protospacer comprises: (a) not found in E. coli, Pseudomonas and / or Klebsiella; (b) is not found in a bacterium containing an endogenous nucleotide sequence encoding a polypeptide of any one of a) to e) of claim 1; (c) is a eukaryotic protospacer; or (d) A vector according to any one of the preceding claims, which is a protospacer for animal (optionally human), plant, insect or fungal cells.
11. (i) Nuclear localization sequence (NLS); (ii) a phage packaging sequence (optionally a pac or cos site); (iii) a plasmid origin of replication; (iv) a plasmid transfer origin; (v) bacterial plasmid backbone; (vi) a eukaryotic plasmid backbone; (vii) a structural protein gene of a human virus (optionally, an AAV or a lentivirus); (viii) viral (optionally AAV or lentiviral) rep and / or cap sequences; (ix) a sequence encoding a selection or selectable marker; (x) a eukaryotic promoter; or (xi) the nucleotide sequence of a human, animal, plant or fungal gene; 10. The vector of any one of the preceding claims, comprising:
12. Each vector is i) a plasmid vector (optionally a conjugative plasmid) ii) a transposon vector (optionally a conjugative transposon); iii) a viral vector (optionally a phage, AAV or lentiviral vector); iv) a phagemid (optionally a packaged phagemid); or v) Nanoparticles (optionally lipid nanoparticles) 10. The vector of any one of the preceding claims, wherein
13. 10. One or more polypeptides according to any one of the preceding claims expressed from a vector according to any one of the preceding claims.
14. A fusion protein comprising a polypeptide (Px), wherein Px is I. comprises an amino acid sequence that is at least 90% identical to a sequence selected from SEQ ID NOs: 1-5; and II. Fused to a heterologous polypeptide (Py); Fusion proteins.
15. in the form of a fusion protein complex comprising at least two additional polypeptides, The two additional polypeptides each have an amino acid sequence that is at least 90% identical to a sequence selected from the sequences of SEQ ID NO: 2, SEQ ID NO: 4, and SEQ ID NO: 5; and The fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2, SEQ ID NO: 4 and SEQ ID NO:
5. The fusion protein of claim 14.
16. The protein complex comprises at least three additional polypeptides, and the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4 and SEQ ID NO:
5. The fusion protein complex of claim 15.
17. The protein complex comprises at least three additional polypeptides, and the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 4 and SEQ ID NO: 5; or the protein complex comprises at least four additional polypeptides, and the fusion protein complex comprises an amino acid sequence that is at least 90% identical to each of the sequences of SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, and SEQ ID NO:5; The fusion protein complex of claim 15.
18. Px comprises an amino acid sequence selected from amino acid sequences that are at least 90% identical to a sequence selected from the sequences of SEQ ID NOs: 1, 3 and 4; the fusion protein is in the form of a fusion protein complex comprising at least two additional polypeptides, each of the two further polypeptides has an amino acid sequence that is at least 90% identical to a sequence selected from the sequences of SEQ ID NO: 2 and SEQ ID NO: 5; (i) optionally, a polypeptide that is at least 90% identical to a sequence selected from SEQ ID NOs: 1, 3, and 4, and that is not based on the amino acid sequence of a polypeptide described in Part I; and (ii) optionally, a polypeptide that is at least 90% identical to a sequence selected from SEQ ID NOs: 1, 3, and 4, and that is not based on the amino acid sequence of a polypeptide described in part I, and if present, is not based on the amino acid sequence of a polypeptide described in part (i). The fusion protein of claim 14.
19. The fusion protein or fusion protein complex of any one of claims 14 to 18, wherein Py is fused to the C-terminus of the amino acid sequence of part I.
20. 20. The fusion protein or fusion protein complex of any one of claims 14 to 19, wherein Py is an enzyme selected from a deaminase (e.g., cytosine or adenine deaminase), a base editor, a prime editor, a helicase, a reverse transcriptase, a methyltransferase, a methylase, an acetylase, an acetyltransferase, a transcriptional activator or inactivator, a translational activator or inactivator, or a nuclease (e.g., a DNA nuclease, an RNA nuclease, a nickase, or a dead nuclease, such as dCas9); and optionally Py is selected from a base editor, a prime editor, and a nuclease.
21. 21. The fusion protein or fusion protein complex of claim 20, wherein Py is a nuclease that is an I-TevI nuclease, and optionally, in part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:
1.
22. 22. The fusion protein or fusion protein complex of claim 21, wherein the I-TevI nuclease lacks a complete DNA-binding domain.
23. 23. The fusion protein or fusion protein complex of claim 21 or claim 22, wherein the I-TevI nuclease comprises an N-terminal catalytic domain.
24. 24. The fusion protein or fusion protein complex of any one of claims 21 to 23, wherein the I-TevI nuclease is a bacteriophage I-TevI nuclease, in particular a T4 bacteriophage I-TevI, and optionally the I-TevI nuclease comprises amino acids 1 to 206 of a naturally occurring I-TevI nuclease, for example the I-TevI nuclease comprises the amino acid sequence of SEQ ID NO:
45.
25. 21. The fusion protein or fusion protein complex of claim 20, wherein Py comprises a base editor (optionally in combination with a UGI protein) selected from an adenine base editor (ABE, e.g., adenosine deaminase), a cytosine base editor (CBE, e.g., cytidine deaminase or apolipoprotein B mRNA editing complex (APOBEC) family deaminase), a cytosine guanine base editor (CGBE), a cytosine adenine base editor (CABE), an adenine cytosine base editor (ACBE), an adenine thymine base editor (ATBE), a thymine adenine base editor (TABE), and a uracil DNA glycosylase inhibitor (UGI) protein, in particular a CBE, e.g., PmCDA1 cytidine deaminase.
26. 26. The fusion protein or fusion protein complex of Claim 25, wherein the base editor is a sea lamprey base editor, e.g., the base editor is a PmCDA1 cytidine deaminase having an amino acid sequence encoded by SEQ ID NO:
19.
27. 27. The fusion protein or fusion protein complex of Claim 25 or Claim 26, wherein the base editor is fused to the C-terminus of the amino acid sequence of part I, and optionally the base editor (e.g., PmCDA1 cytidine deaminase) is fused to the amino acid sequence via a peptide linker, e.g., a linker of 10 to 20 amino acids (e.g., about 16 amino acids), such as a linker having the amino acid sequence of SEQ ID NO:
17.
28. 28. The fusion protein or fusion protein complex of Claim 27, wherein a UGI protein is fused to the base editor, optionally wherein the UGI protein has the amino acid sequence of SEQ ID NO:
21.
29. 29. The fusion protein or fusion protein complex of Claim 28, wherein the UGI protein is fused to the base editor via a peptide linker, e.g., a linker of about 8 to 12 amino acids (e.g., about 10 amino acids), e.g., a linker having the amino acid sequence of SEQ ID NO: 23, optionally wherein the linker is fused to the C-terminus of the base editor and the UGI protein is fused to the C-terminus of the linker.
30. is combined with (e.g., forms a ribonucleoprotein complex with) a crRNA that includes a spacer that is homologous to a first protospacer in the target sequence, and optionally, the protospacer is (a) is not found in a bacterium containing an endogenous nucleotide sequence encoding a polypeptide of any one of a) to e) of claim 1; (b) is a eukaryotic protospacer; (c) heterologous to the source of Py (e.g., where Py is a naturally occurring polypeptide); or (d) a protospacer of an animal (optionally human), plant, insect, or fungal cell; A fusion protein or fusion protein complex according to any one of claims 14 to 29.
31. One or more nucleic acid vectors, in particular two vectors, comprising one or more nucleotide sequences encoding a fusion protein or fusion protein complex according to any one of claims 14 to 30, and optionally a crRNA according to claim 30.
32. 32. The vector of claim 31 , wherein the fusion protein is encoded by a single first nucleotide sequence (e.g., contained in a first vector), and at least two additional polypeptides are encoded by second nucleotide sequences (e.g., contained in a second vector), and optionally all of the additional polypeptides are encoded by second nucleotide sequences (e.g., contained in a second vector).
33. the fusion protein complex a) a first nucleotide sequence encoding Py fused to an amino acid sequence that is at least 90% identical to the sequence of SEQ ID NO:1, SEQ ID NO:3, or SEQ ID NO:4; and b) a second nucleotide sequence encoding said at least two further polypeptides encoding amino acid sequences that are at least 90% identical to the sequences of SEQ ID NOs: 1, 2, 3, 4 and / or 5; It is coded by Optionally, the fusion protein complex comprises an amino acid sequence at least 90% identical to each of the sequences of SEQ ID NOs: 1, 2, 3, 4 and 5.
34. 34. The vector of claim 32 or 33, wherein the crRNA is encoded by the first nucleotide sequence or the second nucleotide sequence.
35. At least one of the nucleotide sequences is i. heterologous to at least one nucleotide sequence; ii. It is a eukaryotic promoter; iii. An animal promoter (optionally a mammalian or human promoter); iv. is a plant promoter; v. a fungal promoter (optionally a yeast promoter); vi. is an insect promoter; vii. a viral promoter (optionally a viral, AAV or lentiviral promoter); or viii. Synthetic promoters The vector according to any one of claims 31 to 34, which is operably linked to a promoter which is
36. The vector of any one of claims 31 to 35, wherein at least one of the nucleotide sequences is operably linked to a constitutive promoter.
37. The vector of any one of claims 31 to 35, wherein at least one of the nucleotide sequences is operably linked to an inducible promoter.
38. The vector of claim 10, or of claim 11 or 12 when dependent on claim 10, wherein the crRNA comprises two repeat sequences and / or the crRNA comprises a repeat sequence comprising the nucleotide sequence of SEQ ID NO: 6 or SEQ ID NO: 70; or the vector of claim 31, or of any one of claims 32 to 37 when dependent on claim 31; or the fusion protein of claim 30; or the protein of claim 13, when citing claim 10, or citing claim 11 or 12 when dependent on claim 10.
39. The protospacer is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3', e.g., 5 and 5'-TAG-3', e.g., 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', and 5'-TAG-3', e.g., 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', and 5' 10. The vector of claim 10, or of claims 11, 12 or 38 when dependent on claim 10, wherein the target sequence is immediately adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence selected from 5'-AAC-3', for example 5'-AAC-3' and 5'-ATG-3', and in particular wherein the PAM is 5'-AAC-3'; or the vector of claim 31, or of any one of claims 32 to 38 when dependent on claim 31; or the fusion protein of claim 30, or of claim 38 when dependent on claim 30; or the protein of claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10), or of claim 38 when dependent on claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10).
40. 10. The vector of claim 10, or of claims 11, 12, 38 or 39 when dependent on claim 10, wherein the spacer sequence is 25 to 39 nucleotides in length (for example 28 to 32 nucleotides in length) or about 32 nucleotides in length (for example 32 nucleotides in length); or the vector of claim 31, or of any one of claims 32 to 39 when dependent on claim 31; or the fusion protein of claim 30, or of claim 38 when dependent on claim 30 or of claim 39; or the protein of claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10), or of claim 38 or claim 39 when dependent on claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10).
41. 41. The vector, fusion protein, or protein of claim 40, wherein base pairs 1-28 of the spacer are identical to the complement of nucleotides 1-28 of the protospacer sequence immediately 5' to the PAM sequence in the target sequence.
42. 10. A vector according to claim 10, or according to claims 11, 12, 38 to 41 when dependent on claim 10, wherein the spacer is identical to the complement of the protospacer in the target sequence; or a vector according to claim 31, or according to any one of claims 32 to 41 when dependent on claim 31; or a fusion protein according to claim 30, or according to any one of claims 38 to 41 when dependent on claim 30; or a protein according to claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10), or according to any one of claims 38 to 41 when dependent on claim 13 (when relying on claim 10, or when relying on claim 11 or 12 when dependent on claim 10).
43. Py is a nuclease that is an I-TevI nuclease, wherein the I-TevI nuclease recognizes an I-TevI cleavage site nucleotide sequence that is 5' or 3' (e.g., 5') to the protospacer in the target sequence, and optionally the I-TevI cleavage site nucleotide sequence is about 3 to 7 nucleotides in length (e.g., about 5 nucleotides in length), e.g., the I-TevI cleavage site nucleotide sequence is 5'-CNNNG-3', e.g., the I-TevI cleavage site nucleotide sequence is , 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', 5'-CAGCG-3', 5'-CACTG-3', 5'-CGA CG-3',5'-CGATG-3',5'-CTCCG-3',5'-CTACG-3',5'-CTGTG-3',5'-C TGCG-3',5'-CCGTG-3',5'-CCTAG-3',5'-CCATG-3',5'-CAAAG-3',5' and 5'-CAGAG-3', 5'-CATTG-3', and 5'-CATCG-3' (e.g., selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', and 5'-CAGCG-3', or selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', and 5'-CTGAG-3', or selected from 5'- 31 , or the vector of any one of claims 32 to 42 when dependent on claim 31 , wherein the I-TevI cleavage site nucleotide sequence is selected from 5′-CCACG-3′, 5′-CACCG-3′ and 5′-CATAG-3′, or selected from 5′-CCACG-3′ and 5′-CACCG-3′), in particular the I-TevI cleavage site nucleotide sequence has the nucleotide sequence of SEQ ID NO: 42; or the fusion protein of claim 30, or the fusion protein of any one of claims 38 to 42 when dependent on claim 30.
44. 31. The vector of claim 31, or any one of claims 32 to 43 when dependent on claim 31, wherein Py is a nuclease that is an I-TevI nuclease, and wherein the I-TevI nuclease recognizes an I-TevI spacer nucleotide sequence that is 5' or 3' (e.g., 5') to the protospacer in the target sequence, and optionally the I-TevI spacer nucleotide sequence is about 10 to 35 nucleotides in length, or such as about 29 to 33 nucleotides in length (e.g., about 31 nucleotides in length), e.g., the I-TevI spacer nucleotide sequence is at least about 90% identical to the nucleotide sequence of SEQ ID NO:43 (e.g., 100% identical to the nucleotide sequence of SEQ ID NO:43).
45. Py is a nuclease that is I-TevI, and the target sequence is, in the 5' to 3' direction: a) an I-TevI cleavage site nucleotide sequence (optionally as described in claim 43); b) an I-TevI spacer nucleotide sequence (optionally as described in claim 44); c) a PAM sequence (optionally as set forth in claim 39); and d) a protospacer sequence, optionally the protospacer sequence is about 32 base pairs in length (e.g., 32 base pairs in length) or is as described in claim 30; Including, and optionally, the vector of claim 31, or of any one of claims 32 to 44 when dependent on claim 31; or the fusion protein of claim 30, or of any one of claims 38 to 44 when dependent on claim 30, wherein sequences a) to d) are respectively immediately adjacent to each other.
46. A cell, A. A eukaryotic cell comprising a vector or protein according to any one of the preceding claims; B. An animal, plant, insect or fungal cell containing a vector or protein according to any one of the preceding claims; or C. A prokaryotic cell comprising a vector or protein according to any one of the preceding claims, and optionally a cell which is not a bacterial cell (e.g., an E. coli, Pseudomonas or Klebsiella cell) comprising an endogenous nucleotide sequence encoding a polypeptide of claim 1 a) to e).
47. A composition comprising a vector or protein according to any one of claims 1 to 45, or a cell according to claim 46 (optionally an in vitro composition, or said composition is contained in a medical container).
48. 1. A method for modifying a nucleic acid target sequence in a cell, comprising: I) contacting a cell with a vector according to claim 10 or according to any one of claims 11, 12, 38 to 45 when dependent on claim 10; or a vector according to claim 31 or according to any one of claims 32 to 45 when dependent on claim 31; or a fusion protein according to claim 30 or according to any one of claims 38 to 45 when dependent on claim 30; or a protein according to claim 13 (when relying on claim 10 or when relying on claim 11 or 12 when dependent on claim 10) or according to any one of claims 38 to 45 when dependent on claim 13 (when relying on claim 10 or when relying on claim 11 or 12 when dependent on claim 10); or with a composition according to claim 44; II) (i) the vector or protein, whereby the polypeptide or protein and the crRNA are expressed in the cell, or (ii) the cRNA of the composition (or a nucleic acid encoding the crRNA or guide RNA, whereby the cRNA or guide RNA is expressed in the cell) and the nucleic acid of the composition encoding the polypeptide or protein, whereby the polypeptide or protein is expressed in the cell. and allowing the introduction of the vector into said cell; III) The crRNA forms a complex with the polypeptide or protein and guides the complex to the target sequence.
49. 49. The method of claim 48, wherein the cell is a eukaryotic cell, an animal cell, a plant cell, an insect cell, or a fungal cell, or a prokaryotic cell, optionally wherein the cell is not a prokaryotic cell (e.g., an E. coli, Pseudomonas, or Klebsiella prokaryotic cell that comprises an endogenous nucleotide sequence encoding a polypeptide of claim 1 a) to e).
50. 50. The method of claim 48 or 49, wherein the target sequence is contained in a chromosome or episome (optionally a plasmid) within the cell, in particular contained in the chromosome of the cell.
51. 51. The method of claim 50, wherein the replication of a chromosome or episome in the cell is inhibited.
52. 52. The method of any one of claims 48 to 51, wherein the target sequence is within about 2 kb upstream or downstream of a gene of interest, such as within about 1 kb upstream or downstream of a gene of interest, such as within a promoter, exon and / or intron of a gene of interest, such as within the promoter of said gene of interest.
53. 53. The method of claim 52, wherein the target sequence of the gene of interest is at least 2 kb upstream and downstream of any promoter, exon and intron of any other gene.
54. 54. The method of any one of claims 48 to 53, wherein the modifying is a substitution, deletion or insertion of one or more nucleotides, for example a substitution (for example a C to T substitution when Py is PmCDA1 cytidine deaminase).
55. 55. A method of treating or preventing a disease or condition in a subject that is mediated by target cells in the subject, comprising performing the method of any one of claims 48 to 54 to modify said target cells, wherein said contacting comprises administering said vector, protein or composition to said subject, and wherein said modification treats or prevents said disease or condition.
56. 56. The method of claim 55, wherein the subject is a human, animal, or plant.
57. 57. The method of claim 56, wherein the target cell is a cell of the subject that contains a nucleic acid defect and wherein the modification corrects the defect.
58. The modification is a) adding new functionality to said cell, optionally wherein said modification upregulates or downregulates (in particular downregulates, e.g., blocks) expression of a gene in said cell, or adds a new nucleotide sequence to the genome of the cell for expression of a protein encoded by the new sequence; or b) modulating the expression of a nucleotide sequence contained in a chromosome or episome (optionally a plasmid) of said cell; 58. The method of claim 56 or 57.
59. 59. A vector, protein, cell or composition according to any one of claims 1 to 47 for use in a method of treating or preventing a human or animal disease or condition according to any one of claims 55 to 58.
60. 1. A method for introducing targeted editing into a target polynucleotide, comprising contacting the target polynucleotide with a fusion protein according to claim 30, or any one of claims 38 to 45 when dependent on claim 30, The method, wherein the first protospacer in the target sequence is included in the polynucleotide, and the crRNA hybridizes to the protospacer to guide the protein (e.g., ribonucleoprotein complex), thereby causing the protein (e.g., ribonucleoprotein complex) to edit the polynucleotide.
61. 61. The method of claim 60, wherein bases or nucleic acid sequences in the polynucleotide are inserted, deleted, or substituted.
62. 1. A container containing a plurality of proteins operable for use with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the proteins comprising: a) any of the proteins of claim 13; or b) comprising any of the fusion proteins according to any one of claims 14 to 29; 40. A container, wherein the complex is operable with a protospacer adjacent motif (PAM) having one of the sequences described in claim 39, the container is not a cell, and the protein is mixed with an in vitro buffer.
63. 46. The nucleic acid vector or vectors of any one of claims 1 to 12 or any one of claims 31 to 45, wherein the polypeptide or protein encoded by the vector is operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the complex being operable with a protospacer adjacent motif (PAM) having one of the sequences of claim 39.
64. 1. A method for targeting a polynucleotide (optionally in a cell or in vitro), comprising: a) contacting the polynucleotide with a protein and crRNA according to claim 13 (when citing claim 10 or when citing claim 11 or 12 when dependent on claim 10), or according to any one of claims 38 to 45 when dependent on claim 13 (when citing claim 10 or when citing claim 11 or 12 when dependent on claim 10); or according to claim 30, or according to any one of claims 38 to 45 when dependent on claim 30; b) allowing the formation of a ribonucleoprotein complex comprising said protein and a crRNA or guide RNA, said complex being guided by a target sequence contained in said polynucleotide and modifying said polynucleotide or a copy thereof; A method comprising:
65. 65. The method of claim 64, wherein the polynucleotide is contained in a chromosome or episome (optionally, a plasmid).
66. a) inhibiting replication of said polynucleotide; or 66. The method of claim 65, wherein b) said polynucleotide is edited, e.g., said editing inserts, deletes, or substitutes bases or nucleic acid sequences in said polynucleotide.
67. A kit comprising: a) one or more proteins according to claim 13; or a fusion protein according to any one of claims 14 to 29; or one or more vectors according to any one of claims 1 to 12 or any one of claims 31 to 37 when dependent on any one of claims 1 to 12; and b) a crRNA or one or more nucleic acids encoding a crRNA, wherein the crRNA is homologous to a PAM having one of the sequences set forth in claim 39; Including, The kit, wherein the polypeptide is operable for use with crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
68. A ribonucleoprotein complex comprising: a) one or more proteins according to claim 13; or a fusion protein according to any one of claims 14 to 29; and b) a crRNA homologous to a PAM having one of the sequences according to claim 39; A ribonucleoprotein complex comprising:
69. 69. A method for regulating transcription or replication of a target DNA, comprising contacting the target DNA with the complex of claim 68, wherein the complex lacks DNA nucleases and wherein the complex binds to the target DNA, thereby regulating the transcription or replication of the target DNA.
70. 70. A method for regulating transcription of a target RNA, comprising contacting the target RNA, or DNA encoding the RNA, with the complex of claim 69, wherein the complex binds to the target RNA or DNA, thereby regulating the transcription of the target RNA.
71. 71. The method of claim 69 or 70, wherein the DNA is contained in a chromosome or episome (optionally, a plasmid).
72. 69. A method of editing a target DNA, comprising contacting the target DNA with the complex of claim 68, wherein the complex binds to the target DNA, thereby editing the target DNA.
73. 69. A method for cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with the complex of claim 68, wherein the complex comprises a fusion protein in which Py comprises a nuclease, and the dsDNA comprises a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby the nuclease cleaves the DNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
74. 69. A method for cleaving single-stranded DNA (ssDNA), comprising contacting the ssDNA with the complex of claim 68, wherein the complex comprises a fusion protein in which Py comprises a nickase, and the ssDNA comprises a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby the nickase cleaves a single strand of the ssDNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer.
75. 69. A method of marking or identifying a region of DNA, comprising contacting the DNA with the complex of claim 68, wherein the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby the complex binds to the DNA within the region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, and optionally the complex comprises a fusion protein in which Py comprises a detectable label.
76. 69. A method of modifying transcription of a region of DNA, comprising contacting the DNA with the complex of claim 68, wherein the DNA comprises a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby the complex binds to the DNA within the region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, whereby the complex upregulates or downregulates, in particular downregulates, transcription of the region of DNA or an adjacent gene.
77. 69. A method of modifying a target dsDNA in a cell without introducing dsDNA breaks, comprising producing in the cell a complex of claim 68, wherein the complex targets a target dsDNA contained in the cell, the target dsDNA comprising a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby the complex binds to the dsDNA within a region defined by complementary binding of the spacer sequence of the crRNA to the protospacer, thereby modifying the dsDNA without introducing dsDNA breaks.
78. 78. The method of Claim 77, wherein the DNA is edited, e.g., the editing inserts, deletes, or substitutes bases or nucleic acid sequences in the DNA.
79. 69. A method of inhibiting cell growth or proliferation without introducing lethal dsDNA breaks, comprising producing in said cell a complex of claim 68, wherein said complex targets a target DNA contained in said cell, said target DNA comprising a protospacer sequence flanked on its 5' side by a PAM having one of the sequences of claim 39, whereby said complex binds to said DNA within a region defined by complementary binding of the spacer sequence of said crRNA to said protospacer, thereby inhibiting replication of said DNA without introducing lethal dsDNA breaks in said DNA, thereby inhibiting said growth or proliferation of said cell.
80. 80. A method of treating or preventing a disease or condition in a human, animal, plant or fungal subject, comprising performing a method according to any one of claims 69 to 79, wherein cells of said subject contain said DNA or RNA and said cells mediate said disease or condition.
81. 81. The method of any one of claims 69 to 80, wherein the complex modifies a coding or non-coding target sequence, optionally (i) the coding strand of DNA is modified rather than the non-coding strand, or (ii) the non-coding strand of DNA is modified rather than the coding strand.
82. 82. The method of claim 81, wherein said modifying cuts, edits, blocks, marks or labels said target sequence.
83. One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NOs: 7-11, wherein the nucleotide sequence is a) operably linked to a heterologous, synthetic, eukaryotic or non-bacterial promoter; and / or b) one or more nucleic acid vectors or one or more nucleic acids contained in a cell that is a eukaryotic cell, a non-bacterial cell, or a cell that is not a bacterial cell (e.g., an E. coli, Pseudomonas, or Klebsiella cell) that contains an endogenous nucleotide sequence comprising SEQ ID NOs: 7-11.
84. A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence which is at least 80% identical to SEQ ID NO:
12.
85. 40. A ribonucleoprotein CRISPR / Cas complex comprising a plurality of Cas proteins and a crRNA or guide RNA, wherein the RNA is capable of guiding the complex to a protospacer comprised in a target DNA, the 5' end of the protospacer being adjacent to a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3' or cognate to a PAM having one of the sequences of claim 39; a) the complex is devoid of DNA nucleases and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the complex lacks DinG, Cas3 and Cas10; and c) A ribonucleoprotein CRISPR / Cas complex, wherein said complex does not comprise all of the Cas proteins of a type I, II, III, IV, V or VI CRISPR / Cas complex.