Novel CRISPR / Cas system
By developing an S-type CRISPR/Cas system that does not rely on known characteristic genes, using the new Cas-S protein and crRNA complex, the problem of functional limitations of the existing CRISPR/Cas system has been solved, and efficient DNA targeting and modification functions have been achieved.
Patent Information
- Application Number
- CN202380076370.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-31
- Filing Date
- 2023-08-31
- Publication Date
- 2025-06-13
AI Technical Summary
The existing CRISPR/Cas systems rely on specific characteristic genes and proteins when targeting DNA, limiting the diversity and functional expansion of the system.
A new CRISPR/Cas system, called the S-type CRISPR/Cas system, was developed, which does not rely on the characteristic genes in the known type I-VI system, uses new Cas-S1 to Cas-S5 proteins and is able to effectively target DNA with the ribonucleoprotein complex of crRNA.
The S-type CRISPR/Cas system can effectively and specifically target predefined sites on DNA, realize the functions of cell modification and DNA cleavage, and does not induce dsDNA breakage, providing new methods for nucleic acid editing and cell targeting.
Smart Images

Figure BDA0005382061370000741 
Figure BDA0005382061370000751 
Figure BDA0005382061370000771
Abstract
Description
Technical Field
[0001] The technology described herein relates to novel and synthetic CRISPR / Cas systems and their components, vectors, proteins, ribonucleoprotein complexes, and methods of using the novel systems and components.
[0002] We named the novel system the "Type S" CRISPR / Cas system. Background Art
[0003] Currently, Types I-VI CRISPR / Cas systems are known.
[0004] It would be desirable to provide novel systems and their components that can increase the toolbox and utilities commonly used for CRISPR / Cas targeting and can bring new types of functions to these useful systems. Summary of the Invention
[0006] The present invention meets these needs by providing the novel Type S CRISPR / Cas system. The basic components of this system do not rely on the characteristic genes of any of the previously elucidated Types I-VI systems. Type S does not use Cas3 (the characteristic protein of Type I systems), Cas9 (the characteristic protein of Type II systems), Cas10 (the characteristic protein of Type III systems), Cas12 (the characteristic protein of Type V systems), or Cas13 (the characteristic protein of Type VI systems). We describe the novel system as Type S.
[0007] In addition, we have identified that the basic components (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5) do not contain RNA-guided DNA nucleases. Nevertheless, we were surprised to find that the ribonucleoprotein complex of these components with crRNA can effectively and specifically target predetermined sites on DNA. Using this system, it is surprisingly also possible to modify cells without inducing dsDNA breaks. We can also surprisingly use the said system to inhibit the growth of bacterial cells. We fused several basic components of the said system with base editors and found that it can perform base editing in an effective manner. We also fused several basic components of the said system with DNA nucleases and found that it can perform DNA cleavage in an effective manner. Armed with this knowledge, we can envision many applications of the new system. For example, it is possible to provide one or more other effector proteins or domains (such as nucleases, base editors, or prime editors) together with one or more components of the said system to produce a synthetic system that can modify cells and nucleic acids in an RNA-guided manner. The DNA binding of the S-type system does not require the presence of all basic components (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5). When fused with a base editor or DNA nuclease, again surprisingly, the synthetic system does not require the presence of all basic components (Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5) of the S-type system to effectively perform the function of the effector domain.
[0008] Accordingly, the said system and its components provide new means for nucleic acid editing and cell targeting in environmental, medical, and other contexts. `
[0009] For this purpose, there is provided:-
[0010] In the first configuration
[0011] In the first aspect: -
[0012] The Cas-S1 protein or a nucleic acid encoding Cas-S1.
[0013] The Cas-S2 protein or a nucleic acid encoding Cas-S2.
[0014] The Cas-S3 protein or a nucleic acid encoding Cas-S3.
[0015] The Cas-S4 protein or a nucleic acid encoding Cas-S4.
[0016] The Cas-S5 protein or a nucleic acid encoding Cas-S5.
[0017] One or more nucleic acids encoding Cas-S1, S2, S3, S4, and S5 proteins.
[0018] One or more nucleic acids encoding Cas-S4 and S5 proteins.
[0019] One or more nucleic acids encoding Cas-S3, S4, and S5 proteins.
[0020] One or more nucleic acids encoding Cas-S2, S3, S4, and S5 proteins.
[0021] One or more nucleic acids encoding Cas-S1 and S2 proteins.
[0022] One or more nucleic acids encoding Cas-S1, S2, and S3 proteins.
[0023] One or more nucleic acids encoding Cas-S1, S2, S3, and S4 proteins.
[0024] One or more nucleic acids comprising SEQ ID NO:7-11.
[0025] One or more nucleic acids comprising SEQ ID NO:8-11.
[0026] One or more nucleic acids comprising SEQ ID NO:9-11.
[0027] One or more nucleic acids comprising SEQ ID NO:10 and 11.
[0028] One or more nucleic acids comprising SEQ ID NO:7-10.
[0029] One or more nucleic acids comprising SEQ ID NO:7-9.
[0030] One or more nucleic acids comprising SEQ ID NO:7 and 8.
[0031] A nucleic acid comprising SEQ ID NO:7.
[0032] A nucleic acid comprising SEQ ID NO:8.
[0033] A nucleic acid comprising SEQ ID NO:9.
[0034] A nucleic acid comprising SEQ ID NO:10.
[0035] A nucleic acid comprising SEQ ID NO:11.
[0036] In one embodiment, each of the sequences is operably linked to a heterologous promoter (i.e., a promoter that is not naturally operably linked to the sequence). In one embodiment, each of the sequences is operably linked to a promoter that is a eukaryotic promoter or a viral (e.g., phage) promoter.
[0037] In the second aspect: -
[0038] A protein or ribonucleoprotein complex comprising 1, 2, 3, 4, or all of the Cas-S1, S2, S3, S4, and S5 proteins.
[0039] A protein or ribonucleoprotein complex comprising the Cas-S1, S2, S3, S4, and S5 proteins.
[0040] A protein or ribonucleoprotein complex comprising the Cas-S4 and S5 proteins.
[0041] A protein or ribonucleoprotein complex comprising the Cas-S3, S4, and S5 proteins.
[0042] A protein or ribonucleoprotein complex comprising the Cas-S2, S3, S4, and S5 proteins.
[0043] In the second configuration
[0044] In the first aspect: -
[0045] (i) A nucleic acid vector or nucleic acid vectors or (ii) a nucleic acid or nucleic acids that comprise an expressible nucleotide sequence, wherein the sequence comprises
[0046] a) A nucleotide sequence encoding a polypeptide that comprises amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:1;
[0047] b) A nucleotide sequence encoding a polypeptide that comprises amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:2;
[0048] c) A third nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:3;
[0049] d) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:4; and
[0050] e) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:5.
[0051] In another embodiment of the first aspect, there is provided (i) a nucleic acid vector or nucleic acid vectors or (ii) a nucleic acid or nucleic acids which comprise an expressible nucleotide sequence, wherein said sequence comprises
[0052] a) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least 94, 95, 96, 97, 98 or 99% identity to SEQ ID NO:1; and / or
[0053] b) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98 or 99% identity to SEQ ID NO:2; and / or
[0054] c) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least 98 or 99% identity to SEQ ID NO:3; and / or
[0055] d) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least 99% identity to SEQ ID NO:4; and
[0056] e) A nucleotide sequence encoding a polypeptide, said polypeptide comprising amino acids having at least 98 or 99% identity to SEQ ID NO:5.
[0057] In the second aspect: -
[0058] A nucleic acid vector or nucleic acid comprising an expressible nucleotide sequence selected from
[0059] a) a nucleotide sequence encoding a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO:1;
[0060] b) a nucleotide sequence encoding a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO:2;
[0061] c) a nucleotide sequence encoding a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO:3;
[0062] d) a nucleotide sequence encoding a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO:4; and
[0063] e) a nucleotide sequence encoding a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO:5.
[0064] In the third aspect: -
[0065] (i) a nucleic acid vector or nucleic acid vectors or (ii) a nucleic acid or nucleic acids according to the first or second aspect, wherein the polypeptide encoded by the vector is operable with a crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide, wherein the complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0066] The crRNA can be as described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein. The PAM can alternatively be any of the PAMs described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein.
[0067] In the fourth aspect: -
[0068] A kit comprising
[0069] a) one or more polypeptides described herein or one or more nucleic acids encoding such polypeptides; and
[0070] b) a crRNA or one or more nucleic acids encoding a crRNA), wherein the crRNA is homologous to a PAM having the sequence 5'-AAG-3';
[0071] Wherein the polypeptide can be operably used with a crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide.
[0072] The polypeptide can comprise Cas-S1, S2, S3, S4, and S5. The polypeptide can be any polypeptide, protein, fusion protein, or complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The PAM can alternatively be any of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0073] In the fifth aspect: -
[0074] A ribonucleoprotein complex comprising
[0075] a) one or more polypeptides described herein or one or more nucleic acids encoding such polypeptides; and
[0076] b) a crRNA or one or more nucleic acids encoding a crRNA), wherein the crRNA is homologous to a PAM having the sequence 5'-AAG-3'.
[0077] The polypeptide can be any polypeptide, protein, fusion protein, or complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein. The PAM can alternatively be any of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0078] In the sixth aspect: -
[0079] One or more nucleic acid vectors or one or more nucleic acids comprising at least one nucleotide sequence selected from SEQ ID NO: 7-11, wherein the nucleotide sequence
[0080] a) is operably linked to a heterologous promoter, synthetic promoter, eukaryotic promoter, or non-bacterial promoter; and / or
[0081] b) is comprised in a cell that is a eukaryotic cell, non-bacterial cell, or bacterial cell that does not contain an endogenous nucleotide sequence (e.g., an E. coli, Psurdomonas, or Klebsiella cell) that contains SEQ ID NO: 7-11.
[0082] In the third configuration
[0083] In the first aspect: -
[0084] A fusion protein comprising a polypeptide (Px), wherein said Px
[0085] a) comprises an amino acid sequence having at least (about) 80% identity to a sequence selected from SEQ ID No: 1-5; and
[0086] b) is fused to a heterologous polypeptide (Py).
[0087] Py can be as described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein.
[0088] In the second aspect: -
[0089] A protein or proteins operable with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, said protein comprising one, more or all of the polypeptides selected from
[0090] a) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO: 1;
[0091] b) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO: 2;
[0092] c) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO: 3;
[0093] d) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO: 4; and
[0094] e) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO: 5;.
[0095] Wherein said complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0096] The crRNA can be as described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein. The PAM can alternatively be any of the PAMs described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein.
[0097] In the fourth configuration
[0098] A cell, wherein said cell is
[0099] a) A eukaryotic cell comprising a vector or protein in a first, second, or third configuration;
[0100] b) An animal, plant, insect, or fungal cell comprising a vector or protein in a first, second, or third configuration; or
[0101] c) A prokaryotic cell comprising a vector or protein in a first, second, or third configuration, wherein said cell is not a bacterial cell (e.g., Escherichia coli, Pseudomonas, or Klebsiella cells), and which comprises an endogenous nucleotide sequence encoding a polypeptide of parts a) to e) as described in the second configuration.
[0102] In the fifth configuration
[0103] In the first aspect: -
[0104] A composition (optionally an in vitro composition or wherein the composition is contained in a medical container), which comprises a crRNA or a nucleic acid encoding a crRNA or a guide RNA; and one, more than one, or all of the following
[0105] a) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:1;
[0106] b) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:2;
[0107] c) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:3;
[0108] d) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:4; and
[0109] e) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:5.
[0110] Wherein
[0111] A: The crRNA comprises a spacer homologous to a first protospacer, wherein said protospacer
[0112] f) Is not found in Escherichia coli;
[0113] g) Is a eukaryotic cell protospacer; or
[0114] h) An animal (optionally mammalian or human), plant, fungal cell protospacer;
[0115] Or
[0116] B: The protospacer of Escherichia coli, which lacks the endogenous nucleotide sequence encoding the polypeptides of a) to e).
[0117] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0118] In the second aspect: -
[0119] A composition (optionally an in vitro composition or a composition contained in a medical container), which comprises a crRNA or a nucleic acid encoding a crRNA; and a nucleic acid encoding one, more or all of the following
[0120] a) A polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 1;
[0121] b) A polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 2;
[0122] c) A polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 3;
[0123] d) A polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 4; and
[0124] e) A polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 5.
[0125] Wherein
[0126] A: The crRNA comprises a spacer homologous to a first protospacer, wherein the protospacer
[0127] f) is not found in Escherichia coli;
[0128] g) is a eukaryotic cell protospacer; or
[0129] h) a protospacer of an animal (optionally a mammal or a human), a plant, a fungal cell;
[0130] Or
[0131] B: The protospacer of Escherichia coli, which lacks the endogenous nucleotide sequence encoding the polypeptides of a) to e).
[0132] The crRNA can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0133] In the sixth configuration
[0134] In the first aspect: -
[0135] A method for modifying a nucleic acid target site in a cell, the method comprising
[0136] a) contacting the cell with a vector, wherein the vector comprises one or more nucleotide sequences for generating a crRNA, wherein the crRNA comprises a spacer homologous to a first protospacer, wherein the protospacer:
[0137] is not found in Escherichia coli; is a eukaryotic cell protospacer; or is an animal (optionally human), plant, insect, fungal cell protospacer;
[0138] b) allowing the introduction into the cell of
[0139] (i) the vector, whereby a polypeptide and the crRNA are expressed in the cell, or
[0140] (ii) the cRNA of the composition (or a nucleic acid encoding the crRNA or guide RNA, whereby the cRNA is expressed in the cell) and the nucleic acid of the composition encoding the polypeptide, whereby the polypeptide is expressed in the cell;
[0141] c) wherein the crRNA guide forms a complex with the polypeptide to direct the complex to the target site.
[0142] The vector and the crRNA can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein. The polypeptide can be any polypeptide, protein, fusion protein or complex described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0143] In the second aspect: -
[0144] A method for treating or preventing a disease or condition mediated by target cells in a subject, the method comprising performing any method described herein to modify the target cells, wherein the contacting comprises administering the vector or composition to the subject, and wherein the modification treats or prevents the disease or condition.
[0145] The vector and the composition can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0146] In the third aspect: -
[0147] A method for introducing targeted editing in a target polynucleotide, the method comprising contacting the target polynucleotide with a CRISPR / Cas system comprising a crRNA and a protein complex, the protein complex comprising one, more than one or all of the polypeptides selected from the following
[0148] a) a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 1;
[0149] b) a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 2;
[0150] c) a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 3;
[0151] d) a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 4; and
[0152] e) a polypeptide comprising an amino acid having at least (about) 80% identity with SEQ ID NO: 5.
[0153] Wherein the crRNA comprises a spacer homologous to a first protospacer comprised in a polynucleotide, wherein the crRNA hybridizes with the protospacer to direct the complex, whereby the complex edits the polynucleotide.
[0154] In the fourth aspect: -
[0155] A method of targeting a polynucleotide (optionally in a cell or in vitro), the method comprising
[0156] a) contacting the polynucleotide with a protein and a crRNA of the second aspect of the third configuration;
[0157] b) allowing the formation of a ribonucleoprotein complex comprising the protein and the crRNA, wherein the complex is directed to a target site comprised in the polynucleotide to modify the polynucleotide or its replication.
[0158] In the fifth aspect: -
[0159] A method of controlling the transcription or replication of a target DNA, which comprises contacting the target DNA with a complex, wherein the complex lacks a DNA nuclease and the complex binds to the target DNA, thereby controlling the transcription or replication of the target DNA.
[0160] In the sixth aspect: -
[0161] A method of controlling the replication of a target RNA, which comprises contacting the target RNA or DNA encoding the RNA with a complex, wherein the complex binds to the target RNA or DNA, thereby controlling the transcription of the target RNA.
[0162] In the seventh aspect: -
[0163] A method of editing a target DNA, comprising contacting the target DNA with a complex, wherein the complex binds to the target DNA so as to edit the target DNA.
[0164] In the eighth aspect: -
[0165] A method of cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with a complex, wherein the complex comprises a nuclease, wherein the dsDNA comprises a protospacer, with a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change at its 5'-flank, whereby the nuclease cleaves the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
[0166] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0167] In the ninth aspect: -
[0168] A method of cleaving single-stranded DNA (ssDNA), comprising contacting the DNA with a complex, wherein the complex comprises a nickase, wherein the DNA comprises a protospacer, with a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change at its 5'-flank, whereby the nuclease cleaves the single strand of the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
[0169] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0170] In the tenth aspect: -
[0171] A method of labeling or identifying a DNA region, comprising contacting the DNA with a complex, wherein the DNA comprises a protospacer, with a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change at its 5'-flank, whereby the complex binds to the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, and optionally wherein the complex comprises a detectable label.
[0172] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0173] In the eleventh aspect: -
[0174] A method of modifying the transcription of a DNA region, which comprises contacting the DNA with a complex, wherein the DNA contains a protospacer, and at its 5'-flank is a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the complex upregulates or downregulates the transcription of the DNA region or an adjacent gene.
[0175] The PAM is not CGG. Optionally,
[0176] - The PAM can be AAN, ANG, NAG; or
[0177] - The PAM can be AAG with no nucleotide change or with one nucleotide change.
[0178] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0179] In the twelfth aspect: -
[0180] A method of modifying a target dsDNA of a cell without introducing a dsDNA break, the method comprising generating a complex in the cell, wherein the complex targets the target DNA contained in the cell, the target DNA contains a protospacer, and at its 5'-flank is a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the DNA is modified without introducing a break in the DNA.
[0181] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0182] In the thirteenth aspect: -
[0183] A method of inhibiting the growth or proliferation of a cell without introducing a lethal dsDNA break, the method comprising generating a complex in the cell, wherein the complex targets the target DNA contained in the cell, the target DNA contains a protospacer, and at its 5'-flank is a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the replication of the DNA is inhibited without introducing a break in the DNA, thereby inhibiting the growth or proliferation of the cell.
[0184] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0185] In the fourteenth aspect: -
[0186] A method of treating or preventing a disease or condition in a human, animal, plant or fungal subject, the method comprising performing the method of any one of the fifth to thirteenth aspects, wherein the cells of the subject comprise the DNA or RNA and the cells mediate the disease or condition.
[0187] In the seventh configuration
[0188] A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence having at least (about) 70% or 80% identity to SEQ ID NO:12.
[0189] In the eighth configuration
[0190] In the first aspect: -
[0191] One or more nucleic acids encoding a plurality of Cas proteins and comprising at least one nucleotide sequence for generating crRNA, wherein the Cas proteins and the RNA are capable of forming a ribonucleoprotein CRISPR / Cas complex, wherein the RNA is capable of guiding the complex to a protospacer contained in a target DNA, wherein the 5' end of the protospacer is flanked by a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3', wherein
[0192] a) The complex lacks a DNA nuclease and is capable of modifying DNA without introducing a double-stranded DNA break;
[0193] b) The Cas protein does not comprise DinG, Cas3 and Cas10; and
[0194] c) The complex does not comprise all of the Cas proteins of a type I, II, III, V or VI CRISPR / Cas complex.
[0195] The PAM can alternatively be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0196] In the second aspect: -
[0197] A ribonucleoprotein CRISPR / Cas complex comprising a plurality of Cas proteins and crRNAs, wherein the RNA is capable of guiding the complex to a protospacer contained in a target DNA, wherein the 5'-end of the protospacer is flanked by a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3', wherein
[0198] a) The complex lacks a DNA nuclease and is capable of modifying DNA without introducing a double-stranded DNA break;
[0199] b) The complex lacks DinG, Cas3, and Cas10; and
[0200] c) The complex does not contain all of the Cas proteins of a type I, II, III, V, or VI CRISPR / Cas complex.
[0201] The PAM can alternatively be any of the PAMs described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0202] In the third aspect: -
[0203] A cell comprising a nucleic acid or complex described herein, wherein the DNA is contained in the chromosome of the cell.
[0204] In the fourth aspect: -
[0205] A cell comprising a nucleic acid or complex described herein, wherein the DNA is contained in a plasmid in the cell.
[0206] In the fifth aspect: -
[0207] A method of modifying DNA, the method comprising
[0208] a) contacting the DNA with a nucleic acid or vector described herein;
[0209] b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is guided to a target site contained in the DNA to modify the DNA.
[0210] The crRNA and Cas protein can be as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein.
[0211] In the sixth aspect: -
[0212] A method of inhibiting the replication of a plasmid containing DNA, the method comprising
[0213] a) contacting the DNA with a nucleic acid or vector described herein;
[0214] b) Permitting the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site comprised in DNA to inhibit plasmid replication.
[0215] The crRNA and Cas protein can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0216] In the seventh aspect: -
[0217] A method of inhibiting transcription of a nucleotide sequence comprised in DNA, the method comprising
[0218] a) contacting the DNA with a nucleic acid or vector as described herein;
[0219] b) Permitting the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site comprised in the nucleotide sequence to inhibit its transcription.
[0220] The crRNA and Cas protein can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0221] In the eighth aspect: -
[0222] A method of inhibiting the growth or proliferation of a cell (optionally a prokaryotic cell, such as a bacterial cell) comprising DNA, the method comprising
[0223] a) contacting the DNA with a nucleic acid or vector as described herein;
[0224] b) Permitting the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site comprised in the DNA to inhibit the growth or proliferation of the cell.
[0225] The crRNA and Cas protein can be as described in any configuration, concept, aspect, example, embodiment, option or other feature herein.
[0226] In the ninth aspect: -
[0227] A method of treating or preventing a disease or condition mediated by target cells in a human or animal subject, the method comprising performing any method described herein to modify the target cells, wherein the contacting comprises administering the nucleic acid to the subject, and wherein the modification treats or prevents the disease or condition.
[0228] In the tenth aspect: -
[0229] A method of using one or more nucleic acids of a first aspect for treating or preventing a target cell-mediated disease or condition in a human or animal subject, the method comprising performing any of the methods described herein to modify a target cell, wherein the contacting comprises administering the nucleic acid to the subject, and wherein the modification treats or prevents the disease or condition. BRIEF DESCRIPTION OF THE DRAWINGS
[0230] Figure 1 : Cloning components of the type S CRISPR system. A plasmid (p1624) was constructed that carried the pSC101 origin of replication and tetracycline resistance marker and a DNA sequence encoding an RNA-guided endonuclease (rge), two ORFs downstream of the rge gene, a homologous CRISPR array, and four ORFs immediately upstream of the CRISPR array. ORF 140 and ORF 423 were predicted to encode error-prone DNA polymerases, while ORF 624 and ORF 188 were predicted to encode a helicase and an RNA endonuclease, respectively. The target plasmid p1631 was constructed by first creating a synthetic DNA fragment that contained spacer-like sequences from the type S CRISPR / Cas system separated by a predicted PAM sequence (AAG), and then inserting this sequence into a plasmid that contained the p15A origin of replication, a chloramphenicol resistance marker (CmR), and the amilCP gene that produces purple.
[0231] Figure 2 : The CRISPR-Cas system inhibits the growth of colonies transformed with the p1631 target plasmid. The strain bSNP3127 was transformed with p1760 (right, for content see Figure 3 ) or p1624 (middle, for content see Figure 1 ) or a plasmid-free control (left). The resulting strains were each transformed with a 1:1 mixture of target (purple) and non-target (white) plasmids and streaked on LB plates supplemented with the appropriate antibiotics for selection. No purple colonies were observed in the presence of p1760.
[0232] Figure 3 : Schematic representation of a plasmid library constructed after a series of gene deletions in p1760 (for content see Figure 3 ). Boxes represent gene deletions. Constructs showing plasmid inhibition are indicated with an asterisk.
[0233] Figure 4 : Five genes and the CRISPR array constitute the type S CRISPR-Cas system. The bSNP3127 strain carrying some of the plasmids shown in Figure 3 was transformed with a 1:1 mixture of target p1631 (purple) and non-target (colorless) plasmids. Purple colonies with the test plasmids were observed only in the absence of the CRISPR array.
[0234] Figure 5 : The CRISPR-Cas system uses all five genes (S1-S5) and the CRISPR array for plasmid inhibition. Strains b5700 (ΔCas-S1), b5702 (ΔCas-S3), b5703 (ΔCas-S4), b5704 (ΔCas-S5) carry plasmids with single-gene deletions of S-type CRISPR-Cas genes. These strains, together with strain b4816 (negative control, without S-type CRISPR-Cas system) and b5408 (positive control, with intact S-type CRISPR-Cas system), were transformed with p1631 target plasmid or p144 non-target plasmid. The fold reduction in transformation efficiency of various strains is presented in this figure.
[0235] Figure 6 : The controlled induction of the S-type CRISPR-Cas system impedes the growth of colonies with p1631 target plasmid.
[0236] Figure 7 : Plasmid clearance assays were used to evaluate the targeting activity of each spacer in the S-type CRISPR array. All four tested spacers showed similar targeting efficiencies.
[0237] Figure 8 : The S-type CRISPR-Cas system does not induce lethal dsDNA breaks in the chromosome of Escherichia coli.
[0238] Figure 9 : Schematic representation of the position of the protospacer for S-type CRISPR-Cas binding assays.
[0239] Figure 10 : S-type CRISPR-Cas binds to chromosomal DNA. Growth curves of bSNP5810-derived strains in the presence of increasing chloramphenicol concentrations.
[0240] Figure 11 : Schematic representation of the S-type CRISPR-Cas PAM preference in the form of a 'PAM wheel' (see Leenay et al., Technology, 62(1), 137-147, 2016, doi:https: / / doi.org / 10.1016 / j.molcel.2016.02.031). The inner ring corresponds to the 3rd nucleotide position upstream (5') of the protospacer, the middle ring corresponds to the 2nd nucleotide position upstream (5') of the protospacer, and the outer ring corresponds to the 1st nucleotide position upstream (5') of the protospacer. In the wheel, each sequence occupies a sector, where the area is proportional to the relative enrichment of the sequence.
[0241] Figure 12 : Graphical representation of the targeting efficiency of the S-type CRISPR / Cas pair for the most preferred PAM. Based on the generated NGS data, the x-axis labels the PAM of the most targeted p2259 library members. The y-axis depicts the fold reduction of the sequencing reads for various PAMs when comparing the NGS data generated by strain b5408 (Cas-S targeted) with the NGS data generated by strain b6259 (control Cas-S).
[0242] Figure 13: Graphical representation of the transformation efficiency of each member of the p1935 plasmid library compared to the transformation efficiency of b5408 (wild-type S-type) with the negative control p2361 plasmid (no PAM). The wild-type protospacer is depicted at the top of each figure. The complementary spacer-protospacer portion of each tested protospacer is represented by dots. The mutated protospacer portion is represented by the corresponding single-letter nucleotide code. The tested protospacers have unique codes shown on the left side of the figure.
[0243] (A) PAM-proximal mismatch (segment); (B) PAM-distal mismatch (segment); (C) PAM-proximal single mismatch; (D) PAM-proximal double mismatch; (E) PAM-proximal triple mismatch; (F) PAM-proximal quadruple mismatch.
[0244] Note: The 'cross' symbol indicates the presence of a low (single cross) or high (double cross) number of uncountable "microcolonies" on the corresponding transformation plates. Microcolonies may indicate ineffective S-type CRISPR / Cas targeting that does not completely stop plasmid replication. As a result, colonies grow on the selection medium at a significantly slower rate.
[0245] Figure 14 : Graphical representation of the plasmids used in Examples 4 and 5. The S-type variants expressed by various plasmids are labeled on the left side of the figure. The names of various plasmids are labeled at the lower left of each plasmid. The S-type genes are presented as boxes with solid fills, and the corresponding gene names are labeled above each box. The dashed lines correspond to the missing S-type genes. The expression of the S-type genes is driven by the pBolA promoter, which is presented as an arrow pointing in the transcription direction. The expression of the crRNA module is driven by the 'leader' sequence, which is presented as a box filled with diagonal stripes and located immediately upstream (5') of the crRNA expression module. The CRISPR repeats (SEQ ID No: 6) of the crRNA expression module are presented as triangles, and the spacers are presented as circles. The localization of the BsmBI recognition site is indicated by an asterisk (*) symbol. The tetracycline resistance marker gene, the SC101 origin of replication, and the repA101 replication protein gene are presented as boxes filled with checkerboard, brick wall, and grid patterns, respectively.
[0246] Figure 15:Schematic representation of the positions of protospacers for S-type CRISPRi assays for downregulating GFP expression from strain b5815 containing chromosomally integrated gfp gene. Each protospacer position is represented by a triangle, and the ID of each protospacer is labeled above the triangle.
[0247] Figure 16: Graphical representation of chromosomal CRISPRi activity of the wild-type S-type system on GFP expression of strain b5815 over a 24-hour period. The targeted protospacer positions are A) at the start of the gfp orf and in the p70a promoter region (protospacers indicated in parentheses: S1F, S1R, S2F, S2R, S3F, S3R), B) at the end of the gfp orf and in the 200 bp regions upstream and downstream of gfp (protospacers indicated in parentheses: S4F, S4R, S5F, S6F, S6R), C) 1 kb upstream of the gfp orf (protospacers indicated in parentheses: U1kF, U1kR) and 2 kb upstream (protospacers indicated in parentheses: U2kF, U2kR).
[0248] Figure 17: Graphical representation of chromosomal CRISPRi activity of the ΔcasS3 S-type system on GFP expression of strain b5815 over a 24-hour period. The targeted protospacer positions are A) at the start of the gfp orf and in the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R), B) at the end of the gfp orf and in the 200 bp regions upstream and downstream of gfp (protospacers S4F, S4R, S5F, S6F, S6R).
[0249] Figure 18: Graphical representation of chromosomal CRISPRi activity of A) the ΔcasS3 S-type system, B) the ΔcasS1ΔcasS3 S-type system, and C) the ΔcasS1 S-type system on GFP expression of strain b5815 over a 24-hour period. The targeted protospacers are located at the start of the gfp orf and in the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R). The plasmids used for transforming strain b5815 are indicated by (:).
[0250] Figure 19: Graphical representation of plasmid-based CRISPRi activity of A) the ΔcasS3 S-type system, B) the ΔcasS4 S-type system, C) the Δcas S4 S-type system, and D) the ΔcasS1ΔcasS3 S-type system on GFP expression of strain b6386 over a 24-hour period. The targeted protospacers are located at the start of the gfp orf and in the p70a promoter region (protospacers S1F, S1R, S2F, S2R, S3F, S3R).
[0251] Figure 20 Schematic representation of plasmid combinations for A) CasS1, B) CasS1(ΔCasS3), C) CasS3, and D) CasS4 base editing assays. Asterisks (*) indicate that the spacer is against one of the following: non-targeting control, S1F, S1R, S2F, S2R, S3F, or S3R.
[0252] Figure 21: Heatmap-based visualization of representative sequencing results from S-type CRISPR / Cas base editing assays. The PmCDA1 cytidine deaminase is fused to the C-terminus of A) the Cas-S1 subunit of the wild-type S-type CRISPR / Cas system, B) the Cas-S1 subunit of the ΔCas-S3 S-type CRISPR / Cas system, C) the Cas-S3 subunit of the wild-type S-type CRISPR / Cas system, D) the Cas-S2 subunit of the wild-type S-type CRISPR / Cas system. i) The code of the tested protospacer and ii) the nucleotide sequence of the protospacer (emphasized with a square arrow additionally showing the direction of the protospacer) together with the sequence of the immediate genomic region where base editing modifications were detected are labeled on the top side of each box. i) The code of the unique Sanger sequencing results (each from a different colony) and ii) all C-to-T modifications detected in the sequencing results or G-to-A modifications when targeting and modifying the minus strand are labeled on the bottom side of each box. The color intensity of each heatmap cell represents the relative C-to-T (or G-to-A) abundance at each position according to the height ratio of the corresponding C and T peaks from the Sanger sequencing results. For example, black boxes indicate colonies with a 100% C-to-T mutant genotype at the corresponding position, while off-white boxes correspond to a 95% to 5% (the reliable detection limit of Sanger sequencing) C-to-T mixed mutant genotype at the corresponding position.
[0253] Figure 22 Schematic representation of plasmid combinations for TevCasS plasmid targeting assays.
[0254] Figure 23 Graphical representation of the results from the TevCasS plasmid targeting assay. The x-axis corresponds to the transformation efficiency of the MG1655 strain relative to that of the MG1655 strain expressing TevCasS1 but not targeting the target sequence of the TevCasS target plasmid, with a decrease in the transformation efficiency of the MG1655 strain expressing TevCasS1 and targeting the target sequence of the TevCasS target plasmid. The y-axis corresponds to the size of the I-TevI spacer sequence in the target sequence of the TevCasS target plasmid.
[0255] Figure 24: Schematic representation of results from the TevCasS chromosomal targeting assay. The x-axis corresponds to the decrease in transformation efficiency of the MG1655 strain expressing TevCasS1 and targeting the chromosomal TevCasS target sequence relative to the transformation efficiency of the MG1655 strain expressing TevCasS1 but not targeting the chromosomal TevCasS target sequence. The y-axis corresponds to the size of the I-TevI spacer sequence in the target sequence of the chromosomal TevCasS target sequence.
[0256] Figure 25 : Schematic representation of the TevCasS target cleavage site preference in the form of a 'Krona plot'. The sequence of the TevCasS target cleavage site is 5’CN 1 N 2 N 3 G-3’. The inner ring corresponds to N 1 and the middle ring corresponds to N 2 nucleotides, while the outer ring corresponds to N 3 . In the wheel, each sequence occupies a sector, the area of which is proportional to the relative abundance of the sequence. SUMMARY OF THE INVENTION
[0257] The techniques described herein relate to new and synthetic CRISPR / Cas systems and their components, vectors, proteins, ribonucleoprotein complexes, and methods of using the new systems and components.
[0258] We named the new system "CRISPR-S" and we refer to it as the "Type S CRISPR / Cas" system.
[0259] The basic components of this system do not rely on the characteristic genes of any of the already elucidated type I-III or V or VI systems. The S type does not use Cas3 (the characteristic protein of the type I system), Cas9 (the characteristic protein of the type II system), Cas10 (the characteristic protein of the type III system), Cas12 (the characteristic protein of the type V system), or Cas13 (the characteristic protein of the type VI system). The basic components do not include the IscB, IsrB, or IshB proteins (IscB and IsrB share a common evolutionary history with Cas9). We refer to the new system as the S type. Thus, in one embodiment, the vectors herein do not encode Cas3, 9, 10, 12, and / or 13. The nucleic acids herein may not encode Cas3, 9, 10, 12, and / or 13. The complexes herein may not contain Cas3, 9, 10, 12, and / or 13. The methods herein may not use Cas3, 9, 10, 12, and / or 13. In one embodiment, the vectors or nucleic acids herein do not encode TnpB. In one embodiment, the complexes described herein do not contain TnpB. In one embodiment, the vectors or nucleic acids herein do not encode the HD-nuclease domain. In one embodiment, the complexes herein do not contain the HD-nuclease domain. In one example, the vectors or nucleic acids herein encode Cas8. In one example, the complexes herein contain Cas8.
[0260] The complex may contain a protein for PAM identification, a protein for interacting with Cas8, a protein for processing precursor crRNA, and a scaffold protein, wherein the complex contains Cas-S1 to Cas-S5 proteins. The complex may contain a protein for PAM identification, a protein for interacting with Cas8, and a scaffold protein, wherein the complex contains Cas-S1, Cas-S2, Cas-S4, and Cas-S5 proteins.
[0261] In one example, the vector or nucleic acid contains a CRISPR array for generating crRNA. The array may contain one or more repeat sequences and one or more spacer sequences. The spacer sequence may be complementary to a first protospacer (as defined elsewhere herein). In some embodiments, the spacer sequences of the array may be complementary to second, third, fourth,... etc. protospacer regions in the target sequence, for example if multiple editing or modification is desired. For example, the repeat sequence is SEQ ID NO:6 or SEQ ID No:70 or a sequence identical except for 1, 2, 3, 4, or 5 variations, especially 1 or 2 variations, such as 1 variation.
[0262] The vector or nucleic acid of the present disclosure may comprise an S-type array, i.e., an array encoding crRNA, which can operate together with at least Cas-S1, 2, 3, 4, and 5 to target a first protospacer (or e.g., first and second protospacers) of a nucleic acid or target sequence. The vector or nucleic acid of the present disclosure may comprise an S-type array, i.e., an array encoding crRNA, which can operate together with at least Cas-S1, 2, 4, and 5 to target a first protospacer (or e.g., first and second protospacers) of a nucleic acid or target sequence. The vector or nucleic acid of the present disclosure may comprise an S-type array, i.e., an array encoding crRNA, which can operate together with at least Cas-S2, 3, 4, and 5 to target a first protospacer (or e.g., first and second protospacers) of a nucleic acid or target sequence. The vector or nucleic acid of the present disclosure may comprise an S-type array, i.e., an array encoding crRNA, which can operate together with at least Cas-S2, 4, and 5 to target a first protospacer (or e.g., first and second protospacers) of a nucleic acid or target sequence.
[0263] The protospacer is downstream of a 5'-AAG-3' PAM in the nucleic acid or target sequence. Alternatively,
[0264] -PAM may be AAN, ANG, NAG; or
[0265] -PAM may be AAG with no nucleotide changes or with one nucleotide change.
[0266] PAM is not CGG.
[0267] Alternatively, PAM may be AHN, KAG, AGG, GAC, or GTG.
[0268] Alternatively, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3', such as 5'-AAC-3' and 5'-ATG-3'. In particular, the PAM is 5'-AAC-3'.
[0269] Notably, we have identified that the components of the type S system (Cas-S1, Cas-S2, Cas-S3, Cas-S4 and Cas-S5) do not contain RNA-guided DNA nucleases. Nevertheless, we were surprised to find that the ribonucleoprotein complex of these components with crRNA can effectively and specifically target a predetermined site on DNA. Using this system, it is surprisingly also possible to modify cells without inducing dsDNA breaks, see Example 1.3.4 below. Armed with this knowledge, we can envision many applications of the new system. For example, it is possible to provide one or more proteins or domains with effector activity (such as nuclease (see Example 5 below), base editing (see Example 4 below) or prime editing activity) together with one or more components of the system to generate a synthetic system that can modify cells and nucleic acids in an RNA-guided or DNA-guided manner.
[0270] Thus, the type S system and its components provide new means for nucleic acid editing and cell targeting in environmental, medical and other contexts.
[0271] In the first aspect, the following aspects of a novel S-type system are provided: -
[0272] The Cas-S1 protein or a nucleic acid encoding Cas-S1.
[0273] The Cas-S2 protein or a nucleic acid encoding Cas-S2.
[0274] The Cas-S3 protein or nucleic acid encoding Cas-S3.
[0275] The Cas-S4 protein or nucleic acid encoding Cas-S4.
[0276] The Cas-S5 protein or nucleic acid encoding Cas-S5.
[0277] One or more nucleic acids encoding the Cas-S1, S2, S3, S4, and S5 proteins.
[0278] One or more nucleic acids encoding the Cas-S4 and S5 proteins.
[0279] One or more nucleic acids encoding the Cas-S3, S4, and S5 proteins.
[0280] One or more nucleic acids encoding the Cas-S2, S3, S4, and S5 proteins.
[0281] One or more nucleic acids encoding the Cas-S1, S2, S4, and S5 proteins.
[0282] One or more nucleic acids encoding the Cas-S2, S3, S4, and S5 proteins.
[0283] One or more nucleic acids encoding the Cas-S2, S4, and S5 proteins.
[0284] A vector or nucleic acid encoding the Cas-S1, S2, S3, S4, and S5 proteins, wherein the vector or nucleic acid comprises SEQ ID NO:12 or a DNA sequence having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:12.
[0285] A nucleic acid vector or nucleic acid comprising SEQ ID NO:12 or a DNA sequence having at least (about) 80% (or at least about 80%, 85%, 90%, 91, 92, 93, 94, 95, 96, 97, 98, or at least about 99%) identity to SEQ ID NO:12.
[0286] In one embodiment, each of the sequences is operably linked to a heterologous promoter (i.e., a promoter that is not naturally operably linked to the sequence). In one embodiment, each of the sequences is operably linked to a promoter that is a eukaryotic promoter or a viral (e.g., phage) promoter. The promoter can be a constitutive promoter. The promoter can be an inducible promoter. The promoter can be as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described herein.
[0287] Optionally, none of the nucleic acids or vectors encode a Cas nuclease. The nucleic acid can encode a non-Cas effector protein or domain, such as a nuclease that is not a Cas nuclease.
[0288] In the second aspect: -
[0289] A protein or ribonucleoprotein complex comprising 1, 2, 3, 4, or all of the Cas-S1, S2, S3, S4, and S5 proteins.
[0290] A protein or ribonucleoprotein complex comprising the Cas-S1, S2, S3, S4, and S5 proteins.
[0291] A protein or ribonucleoprotein complex comprising the Cas-S4 and S5 proteins.
[0292] A protein or ribonucleoprotein complex comprising the Cas-S3, S4, and S5 proteins.
[0293] A protein or ribonucleoprotein complex comprising the Cas-S2, S3, S4, and S5 proteins.
[0294] A protein or ribonucleoprotein complex comprising the Cas-S1, S2, S4, and S5 proteins.
[0295] A protein or ribonucleoprotein complex comprising the Cas-S2, S3, S4, and S5 proteins.
[0296] A protein or ribonucleoprotein complex comprising the Cas-S2, S4, and S5 proteins.
[0297] Optionally, the complex does not contain a Cas nuclease. The complex can contain a non-Cas effector protein or domain, such as a nuclease that is not a Cas nuclease.
[0298] The S-type protein is capable of performing RNA-guided DNA sequence modification, optionally without creating a break in the DNA. As used herein:
[0299] Cas-S1 protein: Cas-S1 is a Cas protein comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:1. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S1 is Cas-S1.1. Cas-S1.1 is a Cas protein comprising the amino acids of SEQ ID NO:1.
[0300] Cas-S2 protein: Cas-S2 is a Cas protein comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:2. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S2 is Cas-S2.1. Cas-S2.1 is a Cas protein comprising the amino acids of SEQ ID NO:2.
[0301] Cas-S3 protein: Cas-S3 is a Cas protein comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:3. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S3 is Cas-S3.1. Cas-S3.1 is a Cas protein comprising the amino acids of SEQ ID NO:3.
[0302] Cas-S4 protein: Cas-S4 is a Cas protein comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:4. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S4 is Cas-S4.1. Cas-S4.1 is a Cas protein comprising the amino acids of SEQ ID NO:4.
[0303] Cas-S5 protein: Cas-S5 is a Cas protein comprising amino acids having at least (about) 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% identity to SEQ ID NO:5. Preferably, the identity is at least (about) 80%. Even more preferably, the identity is at least (about) 90 or 95%. In another embodiment, the identity is at least about 96%. In another embodiment, the identity is at least about 97%. In another embodiment, the identity is at least about 98%. In another embodiment, the identity is at least about 99%. Preferably, Cas-S5 is Cas-S5.1. Cas-S5.1 is a Cas protein comprising the amino acids of SEQ ID NO:5.
[0304] For example, any percent identity herein is at least (about) 70%. For example, any percent identity herein is at least (about) 80%. For example, any percent identity herein is at least (about) 90%. For example, any percent identity herein is at least 95%. For example, any percent identity herein is at least about 96%. For example, any percent identity herein is at least about 97%. For example, any percent identity is at least about 98%. For example, any percent identity herein is at least about 99%.
[0305] The percent identity of amino acid sequences is determined using the blastP algorithm with the following parameters: -
[0306] The default parameters for short input sequences are adjusted, setting the expectation threshold to 0.05 and the seed sequence length for initiating alignment to 6. Regions of low compositional complexity are masked. The scoring matrix employed is 'BLOSUM62', where the scoring costs for creating and extending gaps are 11 and 1, respectively. Conditional compositional score matrix adjustment is used to compensate for the amino acid composition of the comparison sequences.
[0307] The percent identity of nucleotide sequences is determined using the blastn algorithm with the following parameters: -
[0308] The default parameters are automatically adjusted for short input sequences, setting the expectation threshold to 0.05 and the seed sequence length for initiating alignment to 28. Regions of low compositional complexity are masked. The query sequence is masked during the generation of the seed sequences for database scanning but not for extension. A match is scored as +1, while a mismatch is scored as -2.
[0309] For example, the amino acid sequences described herein are identical to the reference SEQ ID NO, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change. For example, the nucleic acid sequences described herein are identical to the reference SEQ ID NO, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotide changes. For example, the amino acid sequence of a polypeptide is identical to SEQ ID NO:1, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change. For example, the amino acid sequence of a polypeptide is identical to SEQ ID NO:2, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change. For example, the amino acid sequence of a polypeptide is identical to SEQ ID NO:3, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change. For example, the amino acid sequence of a polypeptide is identical to SEQ ID NO:4, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change. For example, the amino acid sequence of a polypeptide is identical to SEQ ID NO:5, except for 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid changes, particularly 1 to 5, such as 1 to 3, such as 1 or 2, such as 1 amino acid change.
[0310] For example, the amino acid sequences described herein are identical to a reference SEQ ID NO except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in the reference sequence. For example, the nucleotide sequences described herein are identical to a reference SEQ ID NO except for the total number of nucleotide changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of nucleotides in the reference sequence. Thus, in one instance, the amino acid sequence of a polypeptide is SEQ ID NO:1 or an amino acid sequence that is identical to SEQ ID NO:1 except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in SEQ ID NO:1, particularly no more than (about) 20%, such as no more than (about) 10%, of the number of amino acids in SEQ ID NO:1. Thus, in one instance, the amino acid sequence of a polypeptide is SEQ ID NO:2 or an amino acid sequence that is identical to SEQ ID NO:2 except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in SEQ ID NO:2, particularly no more than (about) 20%, such as no more than (about) 10%, of the number of amino acids in SEQ ID NO:2. Thus, in one instance, the amino acid sequence of a polypeptide is SEQ ID NO:3 or an amino acid sequence that is identical to SEQ ID NO:3 except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in SEQ ID NO:3, particularly no more than (about) 20%, such as no more than (about) 10%, of the number of amino acids in SEQ ID NO:3. Thus, in one instance, the amino acid sequence of a polypeptide is SEQ ID NO:4 or an amino acid sequence that is identical to SEQ ID NO:4 except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in SEQ ID NO:4, particularly no more than (about) 20%, such as no more than (about) 10%, of the number of amino acids in SEQ ID NO:4. Thus, in one instance, the amino acid sequence of a polypeptide is SEQ ID NO:5 or an amino acid sequence that is identical to SEQ ID NO:5 except for the total number of amino acid changes, where the total number is no more than (about) 5, 10, 15, 20, 25, or 30% of the number of amino acids in SEQ ID NO:5, particularly no more than (about) 20%, such as no more than (about) 10%, of the number of amino acids in SEQ ID NO:5.
[0311] In one embodiment, the vector herein encodes Cas-S1, 2, 3, 4, and 5.
[0312] In one embodiment, the vector herein encodes Cas-S4 and 5.
[0313] In one embodiment, the vector herein encodes Cas-S3, 4, and 5.
[0314] In one embodiment, the vector herein encodes Cas-S2, 3, 4, and 5.
[0315] In one embodiment, the vector herein encodes Cas-S1 and 2.
[0316] In one embodiment, the vector herein encodes Cas-S1, 2, and 3.
[0317] In one embodiment, the vector herein encodes Cas-S1, 2, 3, and 4.
[0318] In one embodiment, the vector herein encodes Cas-S1, 2, 4, and 5.
[0319] In one embodiment, the vector herein encodes Cas-S2, 4, and 5.
[0320] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, 3.1, 4.1, and 5.1.
[0321] In one embodiment, the vector herein encodes Cas-S4.1 and 5.1.
[0322] In one embodiment, the vector herein encodes Cas-S3.1, 4.1, and 5.1.
[0323] In one embodiment, the vector herein encodes Cas-S2.1, 3.1, 4.1, and 5.1.
[0324] In one embodiment, the vector herein encodes Cas-S1.1 and 2.1.
[0325] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, and 3.1.
[0326] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, 3.2, and 4.1.
[0327] In one embodiment, the vector herein encodes Cas-S1.1, 2.1, 4.1, and 5.1.
[0328] In one embodiment, the vector herein encodes Cas-S2.1, 4.1, and 5.1.
[0329] In one embodiment, the complexes herein comprise Cas-S1, 2, 3, 4, and 5.
[0330] In one embodiment, the complexes herein comprise Cas-S4 and 5.
[0331] In one embodiment, the complexes herein comprise Cas-S3, 4, and 5.
[0332] In one embodiment, the complexes herein comprise Cas-S2, 3, 4, and 5.
[0333] In one embodiment, the complexes herein comprise Cas-S1 and 2.
[0334] In one embodiment, the complexes herein comprise Cas-S1, 2, and 3.
[0335] In one embodiment, the complexes herein comprise Cas-S1, 2, 3, and 4.
[0336] In one embodiment, the complexes herein comprise Cas-S1, 2, 4, and 5.
[0337] In one embodiment, the complexes herein comprise Cas-S2, 4, and 5.
[0338] In one embodiment, the complexes herein comprise Cas-S1.1, 2.1, 3.1, 4.1, and 5.1.
[0339] In one embodiment, the complexes herein comprise Cas-S4.1 and 5.1.
[0340] In one embodiment, the complexes herein comprise Cas-S3.1, 4.1, and 5.1.
[0341] In one embodiment, the complexes herein comprise Cas-S2.1, 3.1, 4.1, and 5.1.
[0342] In one embodiment, the complexes herein comprise Cas-S1.1 and 2.1.
[0343] In one embodiment, the complexes herein comprise Cas-S1.1, 2.1, and 3.1.
[0344] In one embodiment, the complexes herein comprise Cas-S1.1, 2.1, 3.1, and 4.1.
[0345] In one embodiment, the complexes herein comprise Cas-S1.1, 2.1, 4.1, and 5.1.
[0346] In one embodiment, the complexes herein comprise Cas-S2.1, 4.1, and 5.1.
[0347] In one embodiment, the fusion proteins herein comprise Cas-S1.1, 2.1, 3.1, 4.1, or 5.1.
[0348] In one embodiment, the fusion proteins herein comprise Cas-S1.1, 2.1, 3.1, 4.1, and 5.1.
[0349] In one embodiment, the fusion proteins herein comprise Cas-S4.1 and 5.1.
[0350] In one embodiment, the fusion proteins herein comprise Cas-S3.1, 4.1, and 5.1.
[0351] In one embodiment, the fusion proteins herein comprise Cas-S2.1, 3.1, 4.1, and 5.1.
[0352] In one embodiment, the fusion proteins herein comprise Cas-S1.1 and 2.1.
[0353] In one embodiment, the fusion proteins herein comprise Cas-S1.1, 2.1, and 3.1.
[0354] In one embodiment, the fusion proteins herein comprise Cas-S1.1, 2.1, 3.1, and 4.1.
[0355] In one embodiment, the fusion proteins herein comprise Cas-S1.1, 2.1, 4.1, and 5.1.
[0356] In one embodiment, the fusion proteins herein comprise Cas-S2.1, 4.1, and 5.1.
[0357] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1, 4.1, or 5.1.
[0358] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1, 4.1, and 5.1.
[0359] In one embodiment, the methods herein use Cas-S4.1 and 5.1.
[0360] In one embodiment, the methods herein use Cas-S3.1, 4.1, and 5.1.
[0361] In one embodiment, the methods herein use Cas-S2.1, 3.1, 4.1, and 5.1.
[0362] In one embodiment, the methods herein use Cas-S1.1 and 2.1.
[0363] In one embodiment, the methods herein use Cas-S1.1, 2.1, and 3.1.
[0364] In one embodiment, the methods herein use Cas-S1.1, 2.1, 3.1, and 4.1.
[0365] In one embodiment, the methods herein use Cas-S1.1, 2.1, 4.1, and 5.1.
[0366] In one embodiment, the methods herein use Cas-S2.1, 4.1, and 5.1.
[0367] In one aspect, there is provided one or more vectors and nucleic acids comprising one or more nucleotide sequences selected from SEQ ID NO:7-11.
[0368] There is provided:-
[0369] One or more nucleic acids comprising SEQ ID NO:7-11.
[0370] One or more nucleic acids comprising SEQ ID NO:8-11.
[0371] One or more nucleic acids comprising SEQ ID NO:9-11.
[0372] One or more nucleic acids comprising SEQ ID NO:10 and 11.
[0373] One or more nucleic acids comprising SEQ ID No:7, 8, 10, and 11.
[0374] One or more nucleic acids comprising SEQ ID No:8, 9, 10, and 11.
[0375] One or more nucleic acids comprising SEQ ID No:8, 10, and 11.
[0376] A nucleic acid comprising SEQ ID NO:7.
[0377] A nucleic acid comprising SEQ ID NO:8.
[0378] A nucleic acid comprising SEQ ID NO:9.
[0379] A nucleic acid comprising SEQ ID NO:10.
[0380] A nucleic acid comprising SEQ ID NO:11.
[0381] Nucleic acids comprising SEQ ID No: 7 - 11.
[0382] Nucleic acids comprising SEQ ID No: 7, 8, 10 and 11.
[0383] Nucleic acids comprising SEQ ID No: 8, 9, 10 and 11.
[0384] Nucleic acids comprising SEQ ID No: 8, 10 and 11.
[0385] One or more nucleic acid carriers comprising SEQ ID NO: 7 - 11.
[0386] One or more nucleic acid carriers comprising SEQ ID NO: 8 - 11.
[0387] One or more nucleic acid carriers comprising SEQ ID NO: 9 - 11.
[0388] One or more nucleic acid carriers comprising SEQ ID NO: 10 and 11.
[0389] One or more nucleic acid carriers comprising SEQ ID No: 7, 8, 10 and 11.
[0390] One or more nucleic acid carriers comprising SEQ ID No: 8, 9, 10 and 11.
[0391] One or more nucleic acid carriers comprising SEQ ID No: 8, 10 and 11.
[0392] A nucleic acid carrier comprising SEQ ID NO: 7.
[0393] A nucleic acid carrier comprising SEQ ID NO: 8.
[0394] A nucleic acid carrier comprising SEQ ID NO: 9.
[0395] A nucleic acid carrier comprising SEQ ID NO: 10.
[0396] A nucleic acid carrier comprising SEQ ID NO: 11.
[0397] A nucleic acid carrier comprising SEQ ID NO: 7 - 11.
[0398] A nucleic acid carrier comprising SEQ ID No: 7, 8, 10 and 11.
[0399] A nucleic acid carrier comprising SEQ ID No: 8, 9, 10 and 11.
[0400] A nucleic acid carrier comprising SEQ ID No: 8, 10 and 11.
[0401] A vector comprising SEQ ID NO:12.
[0402] A nucleic acid comprising SEQ ID NO:12.
[0403] Optionally, the indicated nucleotide sequence or at least one of the indicated nucleotide sequences is operably linked to a promoter that
[0404] a) is heterologous to the nucleotide sequence;
[0405] b) is not an Escherichia coli or Klebsiella promoter;
[0406] c) is a eukaryotic promoter;
[0407] d) is an animal promoter (optionally a mammalian or human promoter);
[0408] e) is a plant promoter;
[0409] f) is a fungal promoter (optionally a yeast promoter);
[0410] g) is an insect promoter;
[0411] h) is a viral promoter (optionally a viral, AAV or lentiviral promoter); or
[0412] i) is a synthetic promoter.
[0413] Any promoter described herein can be used with the vectors and nucleic acids described herein.
[0414] In one embodiment, the vector or nucleic acid is contained in a cell that is not an Escherichia coli or Klebsiella cell (e.g., not in Klebsiella pneumoniae). In one embodiment, the vector or nucleic acid is contained in a cell that is not a Pseudomonas cell. In one embodiment, the vector or nucleic acid is contained in a cell that is not an Escherichia coli or Klebsiella cell containing an endogenous nucleotide sequence comprising SEQ ID NOs:7 - 11. In one embodiment, the vector or nucleic acid is contained in a cell (e.g., a bacterial cell) that does not contain an endogenous nucleotide sequence comprising SEQ ID NOs:7 - 11. In one embodiment, the vector or nucleic acid is contained in a eukaryotic cell. In one embodiment, the vector or nucleic acid is contained in a human cell. In one embodiment, the vector or nucleic acid is contained in an animal cell. In one embodiment, the vector or nucleic acid is contained in a plant cell. In one embodiment, the vector or nucleic acid is contained in a fungal cell.
[0415] Also provided:
[0416] One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NO: 7-11, wherein the nucleotide sequence
[0417] a) is operably linked to a heterologous promoter, synthetic promoter, eukaryotic promoter or non-bacterial promoter; and / or
[0418] b) is comprised within a cell that is a eukaryotic cell, non-bacterial cell or a bacterial cell (such as an Escherichia coli, Pseudomonas or Klebsiella cell) that does not contain an endogenous nucleotide sequence comprising SEQ ID NO: 7-11.
[0419] One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NO: 7-11, wherein the nucleotide sequence is comprised within a cell (such as a bacterial cell) that does not contain an endogenous nucleotide sequence comprising SEQ ID NO: 7-11.
[0420] Optionally, any of the vectors or nucleic acids or nucleotide sequences herein are codon-optimized for use in human, animal (such as mammalian, rodent, mouse or rat), plant or fungal cells. Optionally, any of the vectors or nucleic acids or nucleotide sequences herein are codon-optimized for use in eukaryotic cells. Vectors and nucleic acid sequences for use in eukaryotic cells may comprise a nuclear localization sequence (NLS). The NLS facilitates entry of the proteins and polypeptides described herein into eukaryotic cells.
[0421] Optionally, any of the vectors or nucleic acids or nucleotide sequences herein are codon-optimized for use in prokaryotic cells, such as for bacterial or archaeal cells. For example, the bacterial cell is not an Escherichia coli cell. For example, the bacterial cell is not a Klebsiella cell (such as Klebsiella pneumoniae). For example, the bacterial cell is not a Pseudomonas cell. In one instance, the cell comprises DNA or polynucleotide modified or targeted by the methods described herein.
[0422] In one instance, the cells herein are human, animal (e.g., mammalian, rodent, mouse or rat), plant or fungal cells. In one instance, the cells herein are Homo sapiens, Drosophila melanogaster, Mus musculus, Rattus norvegicus, Caenorhabditis elegans or Arabidopsis thaliana cells. In one instance, the cells herein are prokaryotic cells, such as bacterial or archaeal cells. For example, the bacterial cells are not Escherichia coli cells. For example, the bacterial cells are not Klebsiella cells. For example, the bacterial cells are not Pseudomonas cells.
[0423] In one instance, the animal or mammal herein is a vertebrate.
[0424] Optionally, the crRNA disclosed herein (see, e.g., Concept 10 below) is 15 - 100 nucleotides in length and comprises a sequence of at least 10 consecutive nucleotides (e.g., the spacer sequence) complementary to the target sequence. For example, the crRNA is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50 nucleotides in length. For example, the crRNA comprises a sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39 or 40 consecutive nucleotides (e.g., the spacer sequence) complementary to the target sequence.
[0425] In one embodiment, the crRNA is as described elsewhere herein. The crRNA can comprise two repeat sequences. The crRNA can comprise at least one repeat sequence having the nucleotide sequence of SEQ ID No:6. The crRNA can comprise two repeat sequences (e.g., having the nucleotide sequence of SEQ ID No:6) and one spacer sequence that hybridizes to the first protospacer sequence in the target sequence. The crRNA can comprise at least one repeat sequence having the nucleotide sequence of SEQ ID No:70. The crRNA can comprise two repeat sequences (e.g., having the nucleotide sequence of SEQ ID No:70) and one spacer sequence that hybridizes to the first protospacer sequence in the target sequence.
[0426] The length of the spacer sequence can be about 25 to 39 nucleotides. The length of the spacer sequence can be about 29 to 35 (such as about 30 to 34 or about 31 to 33) nucleotides. The length of the spacer sequence can be about 32 nucleotides. The length of the spacer sequence can be 32 nucleotides. The spacer region can be (about) 70% (such as (about) 80% or (about) 90% or (about) 95%) complementary to the protospacer in the target sequence. The spacer region can be 100% complementary to the protospacer in the target sequence. The spacer region and / or the protospacer region can be as described in Concept 10 herein.
[0427] The target sequence can be RNA. The target sequence can be single-stranded DNA.
[0428] In a particular embodiment, the target sequence can be the sequence of dsDNA. The target sequence can be a sequence in the genome of a mammal such as a human. Optionally, the target sequence contains a sequence associated with a disease or disorder in a human or animal subject.
[0429] Optionally, the target sequence contains a point mutation associated with a disease or disorder, and the complex edits the point mutation in the target sequence. Alternatively, the target sequence is contained in a gene whose expression is incorrect (such as present or overexpressed) in a disease or disorder, and the complex introduces a point mutation (or introduces one or more mutations in a region of the target sequence) that disrupts or prevents the expression of the gene. The (one or more) point mutations can be located between about 10 and about 20 nucleotides upstream (5') of the 5'-AAG-3' PAM in the target sequence. The (one or more) point mutations can be located upstream (5') or downstream (3') (such as downstream) of the PAM in the target sequence between about 10 and about 400 nucleotides or between about 10 and 350 nucleotides or between about 10 and 300 nucleotides or between about 10 and 250 nucleotides or between about 10 and 200 nucleotides or between about 10 and 150 nucleotides. The (one or more) point mutations can be located upstream (5') or downstream (3') (such as downstream) of the PAM in the target sequence between about 30 and 120 nucleotides (such as between about 50 and 100 nucleotides). Alternatively,
[0430] - The PAM can be AAN, ANG, NAG; or
[0431] - The PAM can be AAG with no nucleotide change or with one nucleotide change.
[0432] Alternatively, the PAM can be AHN, KAG, AGG, GAC or GTG.
[0433] Alternatively, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. For example, the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3', such as 5'-AAC-3' and 5'-ATG-3'. In particular, the PAM is 5'-AAC-3'.
[0434] For example, the mutation is a C to T point mutation (or one or more mutations are C to T point mutations), wherein the complex deaminates the target C point mutation, and wherein the deamination results in a sequence not associated with a disease or disorder. For example, the target C point mutation (or one or more mutations are target C point mutations) is present in a DNA strand that is not complementary to the crRNA. For example, the (one or more) C to T point mutations are introduced by a cytidine deaminase (such as the PmCDA1 cytidine deaminase described elsewhere herein), which is fused to a Cas-S1, Cas-S3 or Cas-S4 protein, particularly Cas-S1 or Cas-S3 (such as its C-terminus). The fusion protein may comprise all of Cas-S1, Cas-S2, Cas-S3, Cas-S4 and Cas-S5. Alternatively, the fusion protein comprising a cytidine deaminase fused to a Cas-S1 protein may lack the Cas-S3 protein. The fusion protein may further comprise a UGI protein.
[0435] Accordingly, there is provided a fusion protein complex, which comprises a fusion protein that forms a complex with another protein, and wherein the fusion protein comprises a protein of formula A-B-C-D-E (in the 5' to 3' orientation), wherein:
[0436] A is a CasS protein selected from Cas-S1, Cas-S3 or Cas-S4 (particularly Cas-S1);
[0437] B is optionally present and, when present, comprises a linker (such as any linker described herein, particularly an XTEN linker, such as having the amino acid sequence of SEQ ID No:17);
[0438] C is a base editor, such as a cytidine deaminase (such as the PmCDA1 cytidine deaminase described elsewhere herein, such as having the amino acid sequence of SEQ ID No:19);
[0439] D is optionally present and, when present, comprises a linker (such as any linker described herein, particularly the 10 - amino acid linker of SEQ ID No:23);
[0440] E is a uracil DNA glycosylase inhibitor (UGI) protein (more particularly described elsewhere herein, such as having the amino acid sequence of SEQ ID No:21); and
[0441] Additional proteins complexed with the fusion protein optionally include Cas - S2 and Cas - S5 proteins and Cas - S1 and / or Cas - S4 proteins, such that the fusion protein complex comprises at least Cas - S1, Cas - S2, Cas - S4, and Cas - S5 proteins.
[0442] In one embodiment, the fusion protein complex can further comprise a Cas - S3 protein. A first vector expressing the fusion protein of the expression A - B - C - D - E and a second vector expressing the additional proteins of the fusion protein complex are also provided.
[0443] In one embodiment, the fusion protein complex can further comprise a crRNA as described elsewhere herein.
[0444] In another example, the mutation is an A - to - G point mutation, where the complex deaminates the target A point mutation, and where the deamination results in a sequence not associated with a disease or disorder. For example, the target A point mutation is present in a DNA strand that is not complementary to the crRNA.
[0445] In one example, an isolated composition, vector, nucleic acid, or protein as described herein is provided. For example, isolated can mean providing a composition, vector, nucleic acid, or protein excluding Escherichia coli cells.
[0446] Optionally, the target sequence is contained in a gene whose expression is incorrect (e.g., present or overexpressed) in a disease or disorder, and the complex introduces a double-stranded (or single-stranded) break in the target sequence, thereby disrupting or preventing the expression of the gene. Optionally, the target sequence is contained in a pathogenic bacterium that causes a disease or disorder, and the complex introduces a double-stranded break (or single-stranded break) in the chromosome of the bacterium, thereby killing (e.g., selectively killing) the bacterium. Optionally, the target sequence is contained in an antibiotic resistance gene, which is contained in a pathogenic bacterium that causes a disease or disorder, and the complex introduces a double-stranded break (or single-stranded break) in the target sequence, thereby disrupting or preventing the expression of the antibiotic gene and rendering the bacterium sensitive to the antibiotic again. The double-stranded break (or single-stranded break) can be located upstream (5') of the PAM sequence in the target sequence (e.g., between about 29 and about 40 nucleotides upstream (5') of the PAM sequence (e.g., between 30 and 32 nucleotides). The PAM sequence can be any of those described herein, e.g., selected from AHN, KAG, AGG, GAC, and GTG (see also Concept 39 herein). The double-stranded break can be introduced by a fusion protein as described elsewhere herein, where Py is a nuclease (e.g., any nuclease described herein). The single-stranded break in the chromosome (or double-stranded DNA) can be introduced by a fusion protein as described elsewhere herein, where Py is a nickase (e.g., any nickase described herein).
[0447] For example, the double-stranded break is introduced by a nuclease. For example, the single-stranded break is introduced by a nickase. For example, the double-stranded break is introduced by an I-Tev nuclease (e.g., the I-TevI nuclease as described elsewhere herein), which is fused to a Cas-S1 or Cas-S4 protein, particularly Cas-S1 (e.g., its N-terminus). The fusion protein complex can comprise all of Cas-S1, Cas-S2, Cas-S4, and Cas-S5, and optionally can further comprise Cas-S3.
[0448] Accordingly, there is provided a fusion protein complex, which comprises a fusion protein that forms a complex with another protein, and wherein the fusion protein comprises a protein of the formula A-B-C (in the 5' to 3' orientation), wherein:
[0449] A is an I-Tev nuclease (e.g., any I-TevI protein described herein, e.g., having the amino acid sequence of SEQ ID No: 45);
[0450] B is optionally present, and when present comprises a linker (e.g., any linker described herein, particularly an XTEN linker, e.g., having the amino acid sequence of SEQ ID No: 17);
[0451] C is a CasS protein selected from Cas-S1 and Cas-S4; and
[0452] Additional proteins complexed with the fusion protein optionally include Cas-S2 and Cas-S5 proteins and Cas-S1 or Cas-S4 proteins, such that the fusion protein complex includes at least Cas-S1, Cas-S2, Cas-S4, and Cas-S5 proteins.
[0453] In one embodiment, the fusion protein complex can further include a Cas-S3 protein. Also provided are a first vector expressing a fusion protein of the formula A-B-C and a second vector expressing the additional proteins of the fusion protein complex.
[0454] In one embodiment, the fusion protein complex can further include a crRNA as described elsewhere herein.
[0455] In one embodiment, a fusion protein complex comprising an I-Tev nuclease recognizes and cleaves a target sequence that, in a 5' to 3' orientation, comprises:
[0456] a) an I-TevI cleavage site nucleotide sequence (as described elsewhere herein, but particularly where the cleavage site nucleotide sequence is 5'-CNNNG-3', for example where the I-TevI cleavage site nucleotide sequence has the nucleotide sequence of SEQ ID No:42);
[0457] b) an I-TevI spacer nucleotide sequence (as described elsewhere herein, but particularly where the I-TevI spacer nucleotide sequence is about 31 nucleotides in length, for example where the I-TevI spacer nucleotide sequence has at least (about) 80% (e.g., (about) 90%) identity (e.g., 100% identity) with the nucleotide sequence of SEQ ID No:43);
[0458] c) a PAM sequence (as described elsewhere herein, but particularly where the PAM is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', and 5'-ACA-3', for example 5'-AAC-3' and 5'-ATG-3', for example where the PAM is 5'-AAC-3'); and
[0459] d) a protospacer sequence (as described elsewhere herein, but particularly where the protospacer is a spacer sequence about 32 nucleotides in length,
[0460] and optionally where sequences a) to d) are each adjacent to one another.
[0461] In one embodiment, a fusion protein complex comprising an I-Tev nuclease introduces a double-strand break within a cleavage site nucleotide sequence. For example, when the cleavage site nucleotide sequence comprises the sequence of SEQ ID No: 42, the fusion protein introduces breaks after the second C on the sense strand of the target sequence and after the first T on the antisense strand of the target sequence (since the cleavage site nucleotide sequence on the antisense strand comprises the sequence 5'-CGTTG-3'). The I-TevI nuclease nicks the sense strand immediately after the second C and immediately after the first T on the antisense strand. The double nicking of the cleavage site results in a staggered double-stranded DNA break, and each site of the break carries a single-stranded dinucleotide overhang; 5'-AC-3' for the sense strand and 5'-GT-3' for the antisense strand.
[0462] In one example, the vector, complex, nucleic acid, or protein is not comprised within a cell that comprises an endogenous nucleotide sequence encoding Cas-S1, S2, S3, S4, and S5 proteins. In one example, the vector, complex, nucleic acid, or protein is not comprised within a cell of a species selected from the species in Table 3. For example, the vector, complex, nucleic acid, or protein is not comprised within an Escherichia coli cell. For example, the vector, complex, nucleic acid, or protein is not comprised within a Ktedonobacter, such as Ktedonobacter racemifer cell. For example, the vector, complex, nucleic acid, or protein is not comprised within an Allochromatium, such as Allochromatium warmingii cell. For example, the vector, complex, nucleic acid, or protein is not comprised within an Ignatius, such as Ignatius tetrasporus cell.
[0463] In one example, the vector, complex, nucleic acid, or protein is not comprised within a cell that comprises an endogenous nucleotide sequence encoding a polypeptide comprising SEQ ID NO:1, a polypeptide comprising SEQ ID NO:2, a polypeptide comprising SEQ IDNO:3, a polypeptide comprising SEQ ID NO:4, and a polypeptide comprising SE Q ID NO:5.
[0464] Concept:
[0465] Also provided are the following numbered concepts related to new systems, their components, and uses. Any concept may be combined with any configuration, aspect, example, embodiment, option, or other feature disclosed herein.
[0466] Concept 1A. A nucleic acid vector or nucleic acid vectors comprising an expressible nucleotide sequence, wherein the sequence comprises
[0467] a) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least (about) 80% (e.g., about 90%) identity to SEQ ID NO:1;
[0468] b) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least (about) 80% (e.g., about 90%) identity to SEQ ID NO:2;
[0469] c) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least (about) 80% (e.g., about 90%) identity to SEQ ID NO:3;
[0470] d) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least (about) 80% (e.g., about 90%) identity to SEQ ID NO:4; and
[0471] e) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least (about) 80% (e.g., about 90%) identity to SEQ ID NO:5.
[0472] Concept 1B. A nucleic acid vector or nucleic acid vectors comprising an expressible nucleotide sequence, wherein the sequence comprises
[0473] a) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least 94% (e.g., 95%) identity to SEQ ID NO:1;
[0474] b) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least 98% (e.g., 99%) identity to SEQ ID NO:2;
[0475] c) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least 99% (e.g., 100%) identity to SEQ ID NO:3;
[0476] d) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least 98% (e.g., 99%) identity to SEQ ID NO:4; and
[0477] e) A nucleotide sequence encoding a polypeptide that comprises an amino acid sequence having at least 88% (e.g., 89%) identity to SEQ ID NO:5.
[0478] Concept 1C. A nucleic acid vector or nucleic acid vectors comprising an expressible nucleotide sequence, wherein the sequence comprises
[0479] a) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least 98% (e.g., 99%) identity to SEQ ID NO:2;
[0480] b) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least 98% (e.g., 99%) identity to SEQ ID NO:4; and
[0481] c) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least 88% (e.g., 89%) identity to SEQ ID NO:5.
[0482] Concept 1C-1. A nucleic acid vector or nucleic acid vectors according to Concept 1C, wherein said vector further comprises a nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least 94% (e.g., 95%) identity to SEQ ID NO:1.
[0483] Concept 1C-2. A nucleic acid vector or nucleic acid vectors according to Concept 1C or 1C-1, wherein said vector further comprises a nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least 99% (e.g., 100%) identity to SEQ ID NO:3.
[0484] Concept 1D. A nucleic acid vector or nucleic acid vectors comprising an expressible nucleotide sequence, wherein said sequence comprises
[0485] a) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least about 80% (e.g., about 90%) identity to SEQ ID NO:2;
[0486] b) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least about 80% (e.g., about 90%) identity to SEQ ID NO:4; and
[0487] c) A nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least about 80% (e.g., about 90%) identity to SEQ ID NO:5.
[0488] Concept 1D-1. A nucleic acid vector or nucleic acid vectors according to Concept 1D, wherein said vector further comprises a nucleotide sequence encoding a polypeptide, said polypeptide comprising an amino acid sequence having at least about 80% (e.g., about 90%) identity to SEQ ID NO:1.
[0489] Concept 1D-2. A nucleic acid vector or multiple nucleic acid vectors according to Concept 1D or 1D-1, wherein the vector further comprises a nucleotide sequence encoding a polypeptide, and the polypeptide comprises an amino acid sequence having at least about 80% (e.g., about 90%) identity with SEQ ID NO:3.
[0490] Instead of the vectors in these concepts, one or more nucleic acids comprising the indicated components are alternatively provided. The present disclosure regarding the vectors herein should also be construed as being equally applicable to nucleic acids mutatis mutandis. For example, the nucleic acid can be incorporated into the chromosome of a cell, such as a prokaryotic cell. For example, the nucleic acid can be incorporated into a plasmid. For example, the modification can be a modification of the chromosome or plasmid, particularly a modification of the chromosome. Thus, the crRNA can comprise a spacer capable of hybridizing with a protospacer contained in the chromosome or plasmid, particularly a protospacer contained in the chromosome. Thus, the crRNA can comprise a spacer homologous to a protospacer contained in the chromosome or plasmid, particularly a protospacer contained in the chromosome.
[0491] Concept 2. The vector of Concept 1, wherein at least one nucleotide sequence is operably linked to a promoter, and the promoter
[0492] a) is heterologous to at least one nucleotide sequence;
[0493] b) is not an Escherichia coli or Klebsiella promoter;
[0494] c) is a eukaryotic cell promoter;
[0495] d) is an animal promoter (optionally a mammalian or human promoter);
[0496] e) is a plant promoter;
[0497] f) is a fungal promoter (optionally a yeast promoter);
[0498] g) is an insect promoter;
[0499] h) is a viral promoter (optionally a viral, AAV or lentiviral promoter); or
[0500] i) is a synthetic promoter.
[0501] In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure that involves a promoter, the promoter can be a constitutive promoter. In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure that involves a promoter, the promoter can be an inducible promoter. Suitable promoters are well known to those skilled in the art. Inducible promoters may be particularly useful when expression of the S-type protein and complex is desired only under certain defined parameters. In other cases where continuous production of the S-type protein and complex described herein is required, constitutive promoters can be employed.
[0502] In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure that involves a promoter, the promoter can be the BolA promoter, such as the pBolA promoter comprising the nucleotide sequence of SEQ ID No:14. In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure that involves a promoter, the promoter can be the p70a promoter, such as the p70a promoter comprising the nucleotide sequence of SEQ ID No:15. In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure that involves a promoter, the promoter can be an arabinose-inducible promoter, such as the pBAD promoter comprising the nucleotide sequence of SEQ ID No:24.
[0503] Concept 3. A vector of Concept 1 or Concept 2, wherein at least one nucleotide sequence is operably linked to a constitutive promoter.
[0504] In one embodiment, all nucleotide sequences are operably linked to a constitutive promoter. In one embodiment, all nucleotides are contained within a single operon under the control of a single constitutive promoter.
[0505] Concept 4. A vector of Concept 1 or Concept 2, wherein at least one nucleotide sequence is operably linked to an inducible promoter.
[0506] In one embodiment, all nucleotide sequences are operably linked to an inducible promoter. In one embodiment, all nucleotides are contained within a single operon under the control of a single inducible promoter.
[0507] Concept 5. A vector of any one of Concepts 1 to 4, wherein the vector lacks the nucleotide sequence encoding the polypeptide of part c) of Concept 1A or 1B.
[0508] Concept 6. A vector of any one of Concepts 1 to 5, wherein the vector lacks the nucleotide sequence encoding the polypeptide of part a) of Concept 1A or 1B.
[0509] Concept 7A: A nucleic acid vector comprising an expressible nucleotide sequence, said nucleotide sequence comprising the following
[0510] a) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 90% identity with SEQ ID NO: 2;
[0511] b) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 90% identity with SEQ ID NO: 4; and
[0512] c) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 90% identity with SEQ ID NO: 5.
[0513] Concept 7B. A nucleic acid vector comprising an expressible nucleotide sequence selected from the following
[0514] a) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., about 90%) identity with SEQ ID NO: 1;
[0515] b) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., about 90%) identity with SEQ ID NO: 2;
[0516] c) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., about 90%) identity with SEQ ID NO: 3;
[0517] d) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., about 90%) identity with SEQ ID NO: 4; and
[0518] e) A nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., about 90%) identity with SEQ ID NO: 5.
[0519] Concept 8A. The vector of Concept 7A, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 90% identity with SEQ ID NO: 1.
[0520] Concept 8B. The vector of Concept 7A or 7B, wherein the vector further comprises a nucleotide sequence encoding a Cas protein domain that can operate together with the polypeptides of parts a), b), and c) to bind to a nucleotide sequence selected from double-stranded DNA, single-stranded DNA, and RNA.
[0521] Such Cas proteins can be selected from Cas proteins of naturally occurring type I, II, III, IV, V or VI CRISPR / Cas complexes well-known to those skilled in the art. The methods provided herein can be used to identify proteins that can operate with Cas-S2, Cas-S4 and Cas-S5 proteins to provide additional activities.
[0522] Concept 9. A vector of Concept 7A or Concept 8A or 8B, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least (about) 90% identity to SEQ ID NO: 3.
[0523] Concept 10. A vector of any of the foregoing concepts, wherein the vector comprises one or more nucleotide sequences for generating a crRNA, wherein the crRNA comprises a spacer homologous to a first protospacer in a target sequence, optionally wherein the protospacer
[0524] a) is not found in Escherichia coli;
[0525] b) is a eukaryotic protospacer; or
[0526] c) is an animal (optionally human), plant, insect or fungal cell protospacer.
[0527] The protospacer may not be found in bacteria comprising an endogenous nucleotide sequence encoding a polypeptide as in a) to e) of Concept 1.
[0528] The crRNA optionally comprises a spacer having at least 10 consecutive nucleotide sequences complementary to the protospacer. The crRNA optionally comprises a spacer having at least 15 consecutive nucleotide sequences complementary to the protospacer. The crRNA optionally comprises a spacer having at least 20 consecutive nucleotide sequences complementary to the protospacer. The crRNA optionally comprises a spacer having at least 25 consecutive nucleotide sequences complementary to the protospacer. The crRNA optionally comprises a spacer having at least 28 consecutive nucleotide sequences complementary to the protospacer.
[0529] In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure involving a protospacer and a spacer, the length of the spacer and the protospacer can be from about 10 to 40 nucleotides. For example, the length of the spacer and the protospacer can be from about 25 to 40 nucleotides, such as from about 25 to 38 nucleotides, such as from about 28 to 35 nucleotides. The length of the spacer and the protospacer can be from about 25 to 39 nucleotides. The length of the spacer and the protospacer can be from about 29 to 35 (such as from about 30 to 34 or from about 31 to 33) nucleotides. For example, the length of the spacer and the protospacer can be about 32 nucleotides. For example, the length of the spacer and the protospacer can be about 28 nucleotides. For example, the length of the spacer and the protospacer can be 32 nucleotides. For example, the length of the spacer and the protospacer can be 28 nucleotides. The spacer and / or the protospacer can have any feature described in any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure.
[0530] In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure involving a protospacer and a spacer, the spacer can be (about) 90% complementary to the protospacer. The spacer can be (about) 70% complementary to the protospacer. The spacer can be (about) 80% complementary to the protospacer. The spacer can be (about) 95% complementary to the protospacer. The spacer can be 100% complementary to the protospacer across its entire length.
[0531] In one embodiment, when the length of the protospacer is about 32 (such as 32) nucleotides (or longer), nucleotides 1 to 28 of the spacer 5' of the PAM sequence are complementary to the protospacer sequence.
[0532] In any configuration, concept, aspect, instance, embodiment, option, or other feature of the present disclosure involving a crRNA, the crRNA comprises at least one repeat sequence. For example, the crRNA comprises two repeat sequences. The repeat sequence can be any sequence capable of forming a hairpin loop and being recognized by the type S system. The length of the repeat sequence can be from about 20 to 25 nucleotides (such as 22 or 23 nucleotides). The repeat sequence can be a nucleotide sequence having at least about 80%, such as about 90% (such as at least about 95, 96, 97, or 98%) identity to the nucleotide sequence of SEQ ID No: 6. The repeat sequence can be the nucleotide sequence of SEQ ID No: 6. The repeat sequence can be a nucleotide sequence having at least about 80%, such as about 90% (such as at least about 95, 96, 97, or 98%) identity to the nucleotide sequence of SEQ ID No: 70. The repeat sequence can be the nucleotide sequence of SEQ ID No: 70.
[0533] In any nucleic acid or vector described herein, the crRNA can be encoded by a CRISPR array (such as, for example, a CRISPR array described elsewhere herein).
[0534] According to any configuration, concept, aspect, example, embodiment, option, or other feature herein, the plant can be a monocotyledon or a dicotyledon. Optionally, the plant is selected from maize, soybean, cotton, wheat, canola, rapeseed, sorghum, rice, rye, barley, millet, oats, sugarcane, turfgrass, switchgrass, alfalfa, sunflower, tobacco, peanut, potato, Arabidopsis, safflower, and tomato.
[0535] As will be known to those skilled in the art, "cognate to" or "cognate with" or "complementary to" refers to components that can operate together. For example, it is known that the crRNA operates with the PAM to direct Cas to the protospacer in the target nucleic acid. In this sense, the crRNA is cognate with the PAM and cognate with the protospacer.
[0536] Concept 11. A vector of any of the foregoing concepts, wherein the vector comprises
[0537] a) a nuclear localization sequence (NLS);
[0538] b) a phage packaging sequence (optionally a pac or cos site);
[0539] c) an origin of plasmid replication;
[0540] d) an origin of plasmid transfer;
[0541] e) a bacterial plasmid backbone;
[0542] f) a eukaryotic plasmid backbone;
[0543] g) a human viral (optionally AAV or lentivirus) structural protein gene;
[0544] h) a viral (optionally AAV or lentivirus) rep and / or cap sequence;
[0545] i) a selectable marker or a sequence encoding a selectable marker;
[0546] j) a eukaryotic promoter; or
[0547] k) a nucleotide sequence of a human, animal, plant, or fungal gene.
[0548] When the target sequence is in a eukaryotic cell, in any configuration, concept, aspect, instance, embodiment, option, or other feature, the vectors, proteins, or complexes herein further comprise an NLS. NLSs are known to those of skill in the art and include, but are not limited to, monopartite or bipartite NLSs. The NLS sequence can be from the SV40 large T antigen (see Kalderon et al., Cell, 39(3), 499-509, 1984, doi: https: / / doi.org / 10.1016 / 0092-8674(84)90457-4, which is incorporated herein by reference in its entirety), from nucleoplamin, importin α, c-myc, EGL-13, the TUS protein, hnRNP A1, or the yeast transcriptional repressor Matα2.
[0549] Any vector as described herein can be delivered to a bacterial cell via a bacteriophage particle. Any vector as described herein can be delivered to a bacterial cell via a phagemid particle packaged within a bacteriophage particle. Any vector as described herein can be delivered to a bacterial cell via a conjugative plasmid.
[0550] Concept 12. A vector of any of the foregoing concepts, wherein each vector is
[0551] a) a plasmid vector (optionally a conjugative plasmid);
[0552] b) a transposon vector (optionally a conjugative transposon);
[0553] c) a viral vector (optionally a bacteriophage, AAV, or lentiviral vector);
[0554] d) a phagemid (optionally a packaged phagemid); or
[0555] e) a nanoparticle (optionally a lipid nanoparticle).
[0556] Concept 13A. One or more polypeptides as shown in any of the foregoing concepts.
[0557] Concept 13B. One or more polypeptides expressed from a vector as shown in any of the foregoing concepts.
[0558] Concept 13C. One or more polypeptides as shown in any of the foregoing claims expressed from a vector of any of the foregoing concepts.
[0559] Concept 14A. A fusion protein comprising a polypeptide (Px), wherein Px
[0560] a) comprises an amino acid sequence having at least (about) 80% identity (e.g., (about) 90% identity) to a sequence selected from SEQ ID No: 1-5; and
[0561] b) fused with a heterologous polypeptide (Py).
[0562] The 14th concept also provides: -
[0563] A fusion protein comprising Cas-S1 fused with a heterologous polypeptide (Py).
[0564] A fusion protein comprising Cas-S2 fused with a heterologous polypeptide (Py).
[0565] A fusion protein comprising Cas-S3 fused with a heterologous polypeptide (Py).
[0566] A fusion protein comprising Cas-S4 fused with a heterologous polypeptide (Py).
[0567] A fusion protein comprising Cas-S5 fused with a heterologous polypeptide (Py).
[0568] Concept 14B. A fusion protein complex, comprising:
[0569] (i) A fusion protein polypeptide (Px), where Px
[0570] a) comprises an amino acid sequence having at least (about) 80% identity (e.g., (about) 90% identity) with a sequence selected from SEQ ID No: 1, 3, and 4; and
[0571] b) is fused with a heterologous polypeptide (Py);
[0572] (ii) A polypeptide comprising an amino acid sequence having at least (about) 80% identity (e.g., (about) 90% identity) with SEQ ID No: 2;
[0573] (iii) A polypeptide comprising an amino acid sequence having at least (about) 80% identity (e.g., (about) 90% identity) with SEQ ID No: 5;
[0574] (iv) A crRNA comprising a spacer sequence homologous to the first protospacer in the target sequence; and
[0575] (v) Optionally, a polypeptide having at least (about) 80% identity (e.g., (about) 90% identity) with a sequence selected from SEQ ID No: 1, 3, and 4, and not based on the amino acid sequence of the polypeptide shown in part (i)a); and
[0576] (vi) An optional polypeptide that has at least (about) 80% identity (e.g., (about) 90% identity) with a sequence selected from SEQ ID No: 1, 3, and 4, and is not based on the amino acid sequence of the polypeptide shown in part (i) a), and if present, is not based on the amino acid sequence of the polypeptide shown in part (v).
[0577] Concept 14C. A fusion protein complex comprising:
[0578] (i) A fusion protein polypeptide (Px), where Px
[0579] a) comprises a Cas-S1, Cas-S3, or Cas-S4 protein; and
[0580] b) is fused to a heterologous polypeptide (Py);
[0581] (ii) A Cas-S2 protein;
[0582] (iii) A Cas-S5 protein;
[0583] (iv) A crRNA comprising a spacer sequence homologous to a first protospacer in a target sequence; and
[0584] (v) An optional protein selected from Cas-S1, Cas-S3, and Cas-S4 proteins and belonging to a different CasS protein class from the protein shown in part (i) a); and
[0585] (vi) An optional protein selected from Cas-S1, Cas-S3, and Cas-S4 proteins and belonging to a different CasS protein class from the protein shown in part (i) a), and if present, belonging to a different CasS protein class from the protein shown in part (v).
[0586] In one example, the fusion protein comprises two or more of Cas-S1 to S5, e.g., comprises Cas-S1 and S2 or S4 and S5.
[0587] In any configuration, concept, aspect, example, embodiment, option, or other feature involving the fusion protein, Py can be (directly or indirectly) fused to the N-terminus of the CasS protein. In any configuration, concept, aspect, example, embodiment, option, or other feature involving the fusion protein, Py can be (directly or indirectly) fused to the C-terminus of the CasS protein.
[0588] The fusion can be direct or indirect (e.g., via a peptide linker, e.g., (G 4 S) nLinker, where n = 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10). The length of the peptide linker can be about 10 to 20 amino acids (e.g., about 16 amino acids). The length of the peptide linker can be about 8 to 12 amino acids (e.g., about 10 amino acids). The peptide linker can have the amino acid sequence of SEQ ID No: 17. The peptide linker can have the amino acid sequence of SEQ ID No: 23. The linker can be as described elsewhere herein (e.g., in Concept 27 or 29 herein).
[0589] Here, "heterologous" refers to a polypeptide that is not naturally found fused to Px. In one example, Py comprises an I-Tev nuclease or a mutH protein.
[0590] In any configuration, concept, aspect, example, embodiment, option or other feature, Py is an I-Tev nuclease, such as an I-TevI nuclease, such as as described elsewhere herein. In one example, the I-TevI nuclease is a protein comprising the amino acid sequence of SEQ ID No: 45. In one embodiment, the I-TevI nuclease is encoded by a nucleic acid sequence that encodes an amino acid sequence comprising the amino acids of SEQ ID: No: 45. In one embodiment, the I-TevI nuclease is encoded by the nucleic acid sequence of SEQ ID: No: 46. In one example, the I-TevI nuclease is a protein that has an amino acid sequence that is at least about 80% identical (e.g., about 85%, 90%, 95%, 96%, 97% or 98% identical) to the amino acids of SEQ ID No: 45, and is capable of cleaving an I-TevI cleavage site nucleotide sequence having the nucleotide sequence of SEQ ID No: 42, and is capable of recognizing an I-TevI spacer nucleotide sequence having the nucleotide sequence of SEQ ID No: 43.
[0591] Generally, when one component is heterologous to another component, the components are not naturally present in nature or in the same cell.
[0592] For example, the selected sequence is SEQ ID NO: 1. For example, the selected sequence is SEQ ID NO: 2. For example, the selected sequence is SEQ ID NO: 3. For example, the selected sequence is SEQ ID NO: 4. For example, the selected sequence is SEQ ID NO: 5.
[0593] Py can comprise an effector domain. For example, the effector domain is a domain that comprises nuclease activity, nickase activity, recombinase activity, reverse transcriptase, helicase, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity or transcriptional repression activity.
[0594] Py can be naturally occurring. Py can be a synthetic sequence. Py can be a synthetic mutant derivative of a naturally occurring sequence.
[0595] Optionally, the effector domain is a nucleic acid editing domain. For example, the nucleic acid editing domain comprises a deaminase domain. The deaminase domain can be a cytidine deaminase domain. The cytidine deaminase domain can be a member of the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases. The cytidine deaminase domain can have at least (about) 80%, at least (about) 85%, at least (about) 90%, at least (about) 92%, at least (about) 95%, at least (about) 96%, at least (about) 97%, at least (about) 98%, at least (about) 99% or at least (about) 99.5% identity to any one of the cytidine deaminase domains of SEQ ID NOs: 350 - 389 disclosed in U.S. Application No. 16 / 976,047 or WO2019 / 168953, which sequences are hereby expressly incorporated by reference.
[0596] Optionally, Py comprises a uracil glycosylase inhibitor (UGI) domain. The UGI domain can comprise the amino acid sequence of SEQ ID NO: 500 disclosed in U.S. Application No. 16 / 976,047 or WO2019 / 168953, which sequence is hereby expressly incorporated by reference. The UGI domain can have the amino acid sequence of SEQ ID No: 21. The UGI domain can have the nucleotide sequence of SEQ ID No: 20.
[0597] Optionally, Py comprises the amino acid sequence of SEQ ID NO: 534. Optionally, Py can comprise the amino acid sequence of SEQ ID NO: 538. These sequences are disclosed in U.S. Application No. 16 / 976,047 or WO2019 / 168953, which sequences are hereby expressly incorporated by reference.
[0598] Optionally, Py comprises a fusion protein that further comprises a second UGI domain. Optionally, Py comprises the amino acid sequence of SEQ ID NO: 540. Optionally, Py comprises the amino acid sequence of SEQ ID NO: 541. These sequences are disclosed in U.S. Application No. 16 / 976,047 or WO2019 / 168953, which sequences are hereby expressly incorporated by reference.
[0599] Optionally, the deaminase domain is an adenosine deaminase domain. The fusion protein may comprise a second adenosine deaminase domain. For example, the first adenosine deaminase domain and the second adenosine deaminase domain comprise an ecT adA domain or a variant thereof. The first adenosine deaminase domain and the second adenosine deaminase domain may comprise the amino acid sequence of any one of SEQ ID NOs: 400-458. The first adenosine deaminase may comprise the amino acid sequence of SEQ ID NO: 400. The second adenosine deaminase may comprise the amino acid sequence of SEQ ID NO: 458. Optionally, Py comprises the amino acid sequence of SEQ ID NO: 535 or 539. These sequences are disclosed in US Application No. 16 / 976,047 or WO2019 / 168953, which are hereby incorporated by reference in their entirety.
[0600] Base editors that can be used with the Cas-S system are well known in the art. Base editors include systems in which a CasS protein is fused to a cytosine or adenosine deaminase domain and directed to a target sequence for a desired modification.
[0601] In one embodiment, the base editor is selected from:
[0602] 1) A cytosine base editor (CBE) that converts C:G to T:A (Komar et al., Nature, 533:420-4, 2016, which is hereby incorporated by reference in its entirety)
[0603] 2) An adenine base editor (ABE) that converts A:T to G:C (Gaudelli et al., Nature, 551(7681), 464-471, 2017, which is hereby incorporated by reference in its entirety)
[0604] 3) A cytosine-guanine base editor (CGBE) that converts C:G to G:C (both Chen et al., Biorxiv, 2020; Kurt et al., Nature Biotechnology, 2020, which are hereby incorporated by reference in their entirety)
[0605] 4) A cytosine-adenine base editor (CABE) that converts C:G to A:T (Zhao et al., Nature Biotechnology, 2020, which is hereby incorporated by reference in its entirety)
[0606] 5) An adenine-cytosine base editor (ACBE) that converts A:T to C:G (WO2020 / 181180, which is hereby incorporated by reference in its entirety)
[0607] 6) Adenine-thymine base editor (ATBE) that converts A:T to T:A (WO2020 / 181202, which is incorporated herein by reference in its entirety).
[0608] 7) Thymine-adenine base editor (TABE) that converts T:A to A:T (WO2020 / 181193 (or US2022 / 0170013); WO2020 / 181178; WO2020 / 181195, each of which is incorporated herein by reference in its entirety).
[0609] Base editors differ in base-modifying enzymes. CBEs rely on ssDNA cytidine deaminases, including: APOBEC1, rAPOBEC1, APOBEC1 mutants or evolved versions (evoAPOBEC1), and APOBEC homologs (APOBEC3A (eA3A), Anc689), cytidine deaminase 1 (CDA 1), evoCDA1, FERNY, evoFERNY.
[0610] ABEs rely on the deoxyadenosine deaminase activity of the tandem fusion TadA-TadA*, where TadA* is an evolved version of the Escherichia coli tRNA adenosine deaminase TadA that can convert adenosine on ssDNA to inosine. TadA* includes TadA-8a-e and TadA-7.10.
[0611] Examples of DNA-based editing proteins include, but are not limited to, BE1, BE2, BE3, BE4, BE4-GAM, HF-BE3, Sniper-BE3, Target-AID, Target-AID-NG, ABE, EE-BE3, YE1-BE3, YE2-BE3, YEE-BE3, BE-PLUS, SaBE3, SaBE4, SaBE4-GAM, Sa(KKH)-BE3, VQR-BE3, VRER-BE3, EQR-BE3, xBE3, Cas12a-BE, Ea3A-BE3, A3A-BE3, TAM, CRISPR-X, ABE7.9, ABE7.10, ABE7.10*, xABE, ABESa, VQR-ABE, VRER-ABE, Sa(KKH)-ABE, ABE8e, SpRY-ABE, SpRYCBE, SpG-CBE4, SpG-ABE, SpRY-CBE4, SpCas9-NG-ABE, SpCas9-NG-CBE4, enAs BE1.1, enAsBE1.2, enAsBE1.3, enAsBE1.4, AsBE1.1, AsBE1.4, CRISPR-Ab est, CRISPR-Cbest, eA3A-BE3, and AncBE4.
[0612] In one embodiment, Py is a cytidine deaminase. The cytidine deaminase can be the PmCDA1 cytidine deaminase. The cytidine deaminase can be from the sea lamprey. In one embodiment, the PmCDA1 cytidine deaminase comprises the amino acid sequence of SEQ ID No: 19. In one embodiment, the PmCDA1 cytidine deaminase is encoded by a nucleic acid sequence that encodes an amino acid sequence comprising the amino acids of SEQ ID No: 19. In one embodiment, the PmCDA1 cytidine deaminase is encoded by the nucleic acid sequence of SEQ ID No: 18. In one example, the PmCDA1 cytidine deaminase is a protein having an amino acid sequence that has at least about 80% identity (e.g., about 85%, 90%, 95%, 96%, 97%, or 98% identity) to the amino acids of SEQ ID No: 19 and is capable of generating a C to T mutation in a target nucleic acid.
[0613] Concept 15. The fusion protein of Concept 14, wherein the fusion protein is in the form of a fusion protein complex that comprises at least two additional polypeptides,
[0614] wherein each of the at least two additional polypeptides has an amino acid sequence that has at least 90% identity to a sequence selected from the sequences of SEQ ID No: 2, SEQ ID No: 4, and SEQ ID No: 5.
[0615] And wherein said fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2, SEQ ID No: 4, and SEQ ID No: 5.
[0616] Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 2, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 4 and SEQ ID No: 5. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 4, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2 and SEQ ID No: 5. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 5, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2 and SEQ ID No: 4.
[0617] Concept 16. The fusion protein complex of Concept 15, wherein said protein complex comprises at least 3 additional polypeptides, and wherein said fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2, SEQ ID No: 3, SEQ ID No: 4, and SEQ ID No: 5.
[0618] Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 2, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 3, SEQ ID No: 4, and SEQ ID No: 5. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 3, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2, SEQ ID No: 4, and SEQ ID No: 5. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 4, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2, SEQ ID No: 3, and SEQ ID No: 5. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 5, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 2, SEQ ID No: 3, and SEQ ID No: 4.
[0619] Concept 17A. The fusion protein complex of Concept 15, wherein the protein complex comprises at least 3 additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity to each of the sequences of SEQ ID No:1, SEQ ID No:2, SEQ ID No:4, and SEQ ID No:5.
[0620] Px may comprise an amino acid sequence having at least 90% identity to the sequence of SEQ ID No:1, and the additional polypeptides comprise amino acid sequences having at least 90% identity to the sequences of SEQ ID No:2, SEQ ID No:4, and SEQ ID No:5, respectively. Px may comprise an amino acid sequence having at least 90% identity to the sequence of SEQ ID No:2, and the additional polypeptides comprise amino acid sequences having at least 90% identity to the sequences of SEQ ID No:1, SEQ ID No:4, and SEQ ID No:5, respectively. Px may comprise an amino acid sequence having at least 90% identity to the sequence of SEQ ID No:4, and the additional polypeptides comprise amino acid sequences having at least 90% identity to the sequences of SEQ ID No:1, SEQ ID No:2, and SEQ ID No:5, respectively. Px may comprise an amino acid sequence having at least 90% identity to the sequence of SEQ ID No:5, and the additional polypeptides comprise amino acid sequences having at least 90% identity to the sequences of SEQ ID No:1, SEQ ID No:2, and SEQ ID No:4, respectively.
[0621] Concept 17B. The fusion protein complex of Concept 15, wherein the protein complex comprises at least 4 additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity to each of the sequences of SEQ ID No:1, SEQ ID No:2, SEQ ID No:3, SEQ ID No:4, and SEQ ID No:5.
[0622] Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 1, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 2, SEQ ID No: 3, SEQ ID No: 4, and SEQ ID No: 5, respectively. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 2, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 1, SEQ ID No: 3, SEQ ID No: 4, and SEQ ID No: 5, respectively. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 3, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 1, SEQ ID No: 2, SEQ ID No: 4, and SEQ ID No: 5, respectively. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 4, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 1, SEQ ID No: 2, SEQ ID No: 3, and SEQ ID No: 5, respectively. Px may comprise an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 5, and the additional polypeptide comprises an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 1, SEQ ID No: 2, SEQ ID No: 3, and SEQ ID No: 4, respectively.
[0623] Concept 18. The fusion protein of Concept 14, wherein Px comprises an amino acid sequence selected from the amino acid sequences having at least 90% identity with the sequences selected from SEQ ID No: 1, 3, and 4;
[0624] wherein the fusion protein is in the form of a fusion protein complex comprising at least two additional polypeptides,
[0625] wherein each of the two additional polypeptides has an amino acid sequence having at least 90% identity with the sequences selected from SEQ ID No: 2 and Seq ID No: 5, and
[0626] (i) an optional polypeptide having at least 90% identity with the sequences selected from SEQ ID No: 1, 3, and 4 and not based on the amino acid sequence of the polypeptide shown in Part I; and
[0627] (ii) An optional polypeptide that has at least 90% identity with a sequence selected from SEQ ID No: 1, 3, and 4, and is not based on the amino acid sequence of the polypeptide shown in Part I, and if present, is not based on the amino acid sequence of the polypeptide shown in part (i).
[0628] Concept 19. A fusion protein or fusion protein complex of any one of Concepts 14 to 18, wherein Py is fused to the C-terminus of the amino acid sequence of Part I.
[0629] Alternatively, Py is fused to the N-terminus of the amino acid sequence of Part I.
[0630] Concept 20. A fusion protein or fusion protein complex of any one of Concepts 14 to 19, wherein Py is an enzyme, optionally selected from deaminases (such as cytosine or adenine deaminases), base editors, prime editors, helicases, reverse transcriptases, methyltransferases, methylases, acetylases, acetyltransferases, transcriptional activators or deactivators, translational activators or deactivators, or nucleases (such as DNA nucleases, RNA nucleases, nickases, or inactivated (dead) nucleases such as dCas).
[0631] In any configuration, concept, aspect, example, embodiment, option, or other feature, Py is a nuclease, such as any nuclease disclosed herein, particularly the I-TevI nuclease. Py can be a DNA nuclease. Py can be an RNA nuclease. Py can be a nickase or an inactivated nuclease. In one embodiment, Py is a base editor, such as any base editor disclosed herein, particularly the PmCDA1 cytidine deaminase. In one embodiment, Py is a prime editor.
[0632] Concept 21. The fusion protein or fusion protein complex of Concept 20, wherein Py is a nuclease that is the I-TevI nuclease.
[0633] In any configuration, concept, aspect, example, embodiment, option, or other feature, in Part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence (such as any CasS protein described herein). In one embodiment, in Part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence that has at least (about) 80% (such as (about) 90%) identity with the sequence of SEQ ID No: 1. The I-TevI nuclease can be described elsewhere herein (such as in Concept 14).
[0634] Concept 22. The fusion protein or fusion protein complex of Concept 21, wherein the I-TevI nuclease does not contain a complete DNA binding domain.
[0635] I-TevI contains 245 amino acids and consists of an N-terminal catalytic domain and a C-terminal DNA-binding domain, which are connected by a long flexible linker. The crystal structure of the DNA-binding domain of I-TevI (residues 130 to 245) complexed with the 20-bp major binding region of its DNA target reveals the presence of a zinc finger (residues 151 to 167) that makes backbone contacts with DNA from the minor groove, an elongated segment containing a minor groove-binding α-helix (residues 183 to 194) and a helix-turn-helix (residues 204 to 245). The N-terminal catalytic domain is shown to contain residues 1-92. Biochemical data have shown that the zinc finger does not contribute to DNA-binding affinity or enzyme specificity, but rather it has a novel function and acts as a distance determinant that controls the relative positions of the catalytic and DNA-binding domains, see, e.g., Roey et al., 2005, DOI:10.1007 / 0-387-27421-9_7, which is incorporated herein by reference in its entirety.
[0636] Kleinstiver et al., G3 Genes|Genomes|Genetics, 4(6), 1155–1165, 2014, doi:https: / / doi.org / 10.1534 / g3.114.011445 (which is incorporated herein by reference in its entirety) showed that various truncations of the N-terminal catalytic domain retained activity when fused to a TALEN construct. The largest construct was residues 1 to 206 of I-TevI. The smallest construct that functioned when fused to a TALEN was residues 1 to 162.
[0637] Thus, in one embodiment, I-TevI does not contain the complete helix-turn-helix domain. I-TevI can contain the amino acid sequence of SEQ ID No:45. I-TevI can lack the helix-turn-helix domain. I-TevI can contain amino acids 1 to 195-203 of SEQ ID No:45. I-TevI can lack the minor groove-binding α-helix. I-TevI can contain amino acids 1 to 167 of SEQ ID No:45. I-TevI can lack the complete zinc finger domain. I-TevI can contain amino acids 1 to 162 of SEQ ID No:45.
[0638] Concept 23. A fusion protein or fusion protein complex of Concept 21 or 22, wherein the I-TevI nuclease contains an N-terminal catalytic domain.
[0639] Thus, I-TevI can contain amino acids 1 to 92 of SEQ ID No:45. I-TevI can contain the amino acid sequence of SEQ ID No:45.
[0640] Concept 24. A fusion protein or fusion protein complex of any one of Concepts 21 to 23, wherein the I-TevI nuclease is a bacteriophage I-TevI nuclease.
[0641] In any configuration, concept, aspect, instance, embodiment, option or other feature, the I-TevI nuclease can be T4 bacteriophage I-TevI. The I-TevI nuclease can be from T4 bacteriophage. The I-TevI nuclease can be encoded by an intron of bacteriophage T4. The I-TevI nuclease can comprise amino acids 1 to 206 of a naturally occurring I-TevI nuclease. The I-TevI nuclease can comprise the amino acid sequence of SEQ ID No: 45. The I-TevI nuclease can be encoded by the nucleotide sequence of SEQ ID No: 44.
[0642] Concept 25. The fusion protein or fusion protein complex according to claim 20, wherein Py comprises a base editor selected from the group consisting of adenine base editors (ABE, such as adenosine deaminase), cytosine base editors (CBE, such as cytidine deaminase or apolipoprotein B mRNA editing complex (APOBEC) family deaminases), cytosine-guanine base editors (CGBE), cytosine-adenine base editors (CABE), adenine-cytosine base editors (ACBE), adenine-thymine base editors (ATBE), thymine-adenine base editors (TABE), and uracil DNA glycosylase inhibitor (UGI) proteins.
[0643] In any configuration, concept, aspect, instance, embodiment, option, or other feature involving a fusion protein or fusion protein complex, Py is a CBE. In one embodiment, Py is the PmCDA1 cytidine deaminase (as described elsewhere herein, e.g., in Concept 14). In one embodiment, Py is the PmCDA1 cytidine deaminase in combination with the UGI protein. In one embodiment, Py is a CBE linked to the UGI protein via a linker (as described elsewhere herein). In one embodiment, Py is the PmCDA1 cytidine deaminase linked to the UGI protein via a linker (as described elsewhere herein). In one embodiment, Py is a CBE linked to the UGI protein via a linker (as described elsewhere herein), and the UGI protein is linked to the CasS protein described herein via a linker (as described elsewhere herein). In one embodiment, Py is the PmCDA1 cytidine deaminase linked to the UGI protein via a linker (as described elsewhere herein), and the UGI protein is linked to the CasS protein described herein via a linker (as described elsewhere herein). In one embodiment, Py is a CBE linked to the UGI protein via a linker (as described elsewhere herein), and the UGI protein is linked to the CasS protein described herein via a linker (as described elsewhere herein). In one embodiment, Py is the PmCDA1 cytidine deaminase linked to the UGI protein via a linker (as described elsewhere herein), and the UGI protein is linked to the CasS protein described herein via a linker (as described elsewhere herein). In one embodiment, the CasS protein is selected from the group consisting of CasS1 protein, CasS3 protein, and CasS4 protein.
[0644] Concept 26. The fusion protein or fusion protein complex of Concept 25, wherein the base editor is a lamprey base editor.
[0645] In any configuration, concept, aspect, instance, embodiment, option, or other feature, the base editor is the PmCDA1 cytidine deaminase having the amino acid sequence encoded by SEQ ID No: 19. In one embodiment, the base editor is from lamprey.
[0646] Concept 27. The fusion protein or fusion protein complex of Concept 25 or Concept 26, wherein the base editor is fused to the C-terminus of the amino acid sequence of Part I.
[0647] In any configuration, concept, aspect, instance, embodiment, option, or other feature, a base editor (such as the PmCDA1 cytidine deaminase) is fused to the amino acid sequence of Part I via a peptide linker. The peptide linker can be as described in Concept 14 or Concept 29. In any configuration, concept, aspect, instance, embodiment, option, or other feature, a base editor (such as the PmCDA1 cytidine deaminase) is fused to the CasS protein via a peptide linker. The length of the linker can be (about) 10 to 20 amino acids. The length of the linker can be (about) 10 to 22, (about) 11 to 21, (about) 12 to 20, (about) 13 to 19, (about) 14 to 18, (about) 15 to 17 amino acids. The length of the linker can be (about) 16 amino acids. The linker can be a linker having the amino acid sequence of SEQ ID No: 17. The linker can be any linker described herein.
[0648] Concept 28. The fusion protein or fusion protein complex of Concept 27, wherein the UGI protein is fused to the base editor.
[0649] In any configuration, concept, aspect, instance, embodiment, option, or other feature, the UGI protein has the amino acid sequence of SEQ ID No: 21. The UGI protein can be as described elsewhere herein (such as in Concept 14). The UGI protein can be directly fused to the base editor. The UGI protein can be fused to the base editor via a linker such as a peptide linker as described elsewhere herein.
[0650] Concept 29. The fusion protein or fusion protein complex of Concept 28, wherein the UGI protein is fused to the base editor via a peptide linker.
[0651] In any configuration, concept, aspect, instance, embodiment, option, or other feature, the UGI protein is fused to the base editor via a peptide linker having a length of (about) 8 to 12 amino acids. The length of the peptide linker can be (about) 9 to 11 amino acids. The length of the linker can be (about) 10 amino acids. The linker can have the amino acid sequence of SEQ ID No: 23. The linker can be fused to the C-terminus of the base editor, and the UGI protein is fused to the C-terminus of the linker. The linker can be fused to the N-terminus of the base editor, and the UGI protein is fused to the N-terminus of the linker. The linker can be as described elsewhere herein (such as in Concept 14 or 27).
[0652] Concept 30. The fusion protein or fusion protein complex of any one of Concepts 14 to 29, wherein the fusion protein or fusion protein complex is combined with a crRNA that comprises a spacer homologous to a first protospacer in a target sequence.
[0653] The crRNA, spacer, and / or protospacer may have any of the features described elsewhere herein (e.g., in Concept 10 or in the first aspect of the Summary of the Invention). In any configuration, concept, aspect, instance, embodiment, option, or other feature, the protospacer is not found in bacteria containing an endogenous nucleotide sequence encoding a polypeptide of a) to e) as shown in Concept 1A or 1B. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the protospacer is a eukaryotic protospacer. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the protospacer is heterologous to the source of Py. For example, when Py is a naturally occurring polypeptide, the protospacer is from a different species (or genus) than Py. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the protospacer is an animal (optionally human), plant, insect, fungal cell protospacer.
[0654] In one embodiment, the fusion protein complex forms a ribonucleoprotein complex with a crRNA that contains a spacer homologous to a first protospacer in a target sequence. In one embodiment, the fusion protein forms a ribonucleoprotein complex with a crRNA that contains a spacer homologous to a first protospacer in a target sequence.
[0655] Concept 31. A nucleic acid vector comprising one or more nucleotide sequences encoding a fusion protein or fusion protein complex as shown in any one of Concepts 14 to 30.
[0656] The vector may further encode a crRNA as described elsewhere herein (e.g., in the first aspect of the Summary of the Invention, in Concept 10, or in Concept 30). In one embodiment, a first and a second vector are provided that encode a fusion protein complex as described elsewhere herein. In one embodiment, the first vector encodes a fusion protein Px, and the second vector encodes an additional protein of the fusion protein complex.
[0657] The vector may encode second, third, fourth,... etc. crRNAs that contain spacers homologous to second, third, fourth,... etc. protospacers in a target sequence. Expression of multiple crRNAs may be useful for multiplex editing, e.g., for targeting multiple genes in a cell or organism.
[0658] Concept 32. The vector of Concept 31, wherein the fusion protein is encoded by a single first nucleotide sequence and wherein at least two additional polypeptides are encoded by a second nucleotide sequence.
[0659] In one embodiment, the fusion protein described herein is encoded by a single first nucleotide sequence contained within a first vector, and at least two additional CasS polypeptides of the complex are encoded by a second nucleotide sequence contained within a second vector. In one embodiment, all additional CasS polypeptides of the fusion protein complex described herein are encoded by the second nucleotide sequence. In one embodiment, all additional CasS polypeptides of the complex described herein are encoded by a second nucleotide sequence contained within a second vector.
[0660] Concept 33. The vector of Concept 31, wherein the fusion protein complex is encoded by:
[0661] a) a first nucleotide sequence encoding Py, which fuses Py to an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to the sequence of SEQ ID No: 1, SEQ ID No: 3, or Seq ID No: 4; and
[0662] b) a second nucleotide sequence encoding at least two additional polypeptides, the additional polypeptides encoding an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to the sequences of SEQ ID No: 1, 2, 3, 4, and / or 5.
[0663] In any configuration, concept, aspect, instance, embodiment, option, or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to each of the sequences of SEQ ID No: 1, 2, 3, 4, and 5. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to each of the sequences of SEQ ID No: 2, 4, and 5. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to each of the sequences of SEQ ID No: 2, 3, 4, and 5. In any configuration, concept, aspect, instance, embodiment, option, or other feature, the fusion protein complex described herein comprises a polypeptide comprising an amino acid sequence having at least (about) 80% (e.g., (about) 90%) identity to each of the sequences of SEQ ID No: 1, 2, 4, and 5.
[0664] Concept 34A. The vector of Concept 32 or Concept 33, wherein the crRNA is encoded by a first nucleotide sequence.
[0665] Concept 34B. A vector of Concept 32 or Concept 33, wherein the crRNA is encoded by or from a second nucleotide sequence.
[0666] In the case of encoding multiple crRNAs to target multiple protospacers, they can all be encoded by a first nucleotide sequence. In the case of encoding multiple crRNAs to target multiple protospacers, they can all be encoded by a second nucleotide sequence.
[0667] Concept 35A. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is heterologous to the at least one nucleotide sequence.
[0668] Concept 35B. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is a eukaryotic cell promoter.
[0669] The eukaryotic cell promoter can be any of those described herein.
[0670] Concept 35C. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is an animal promoter.
[0671] In one embodiment, the promoter is a mammalian or human promoter. The animal, mammalian or human promoter can be any of those described herein.
[0672] Concept 35D. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is a plant promoter.
[0673] The plant promoter can be any of those described herein.
[0674] Concept 35E. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is a fungal promoter.
[0675] In one embodiment, the promoter is a yeast promoter. The fungal or yeast promoter can be any of those described herein.
[0676] Concept 35F. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is an insect promoter.
[0677] The insect promoter can be any of those described herein.
[0678] Concept 35G. A vector of any one of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is a viral promoter.
[0679] In one embodiment, the promoter is a viral, AAV or lentiviral promoter. The viral, AAV or lentiviral promoter can be any of those described herein.
[0680] Concept 35H. A vector of any of Concepts 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter that is a synthetic promoter.
[0681] The synthetic promoter can be any of those described herein.
[0682] Concept 36. A vector of any of Concepts 31 to 35, wherein at least one nucleotide sequence is operably linked to a constitutive promoter.
[0683] Concept 37. A vector of any of Concepts 31 to 35, wherein at least one nucleotide sequence is operably linked to an inducible promoter.
[0684] Concept 38A. A vector, protein, fusion protein or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option or other feature that includes a crRNA, wherein the crRNA includes two repeat sequences (e.g., two repeat sequences).
[0685] The crRNA and other features of the repeat sequences can have other features as described elsewhere herein (e.g., in the first aspect of the Summary of the Invention or in Concept 10).
[0686] Concept 38B. A vector, protein, fusion protein or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option or other feature that includes a crRNA, wherein the crRNA includes at least one repeat sequence that includes the nucleotide sequence of SEQ ID No:6.
[0687] Concept 38C. A vector, protein, fusion protein or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option or other feature that includes a crRNA, wherein the crRNA includes two repeat sequences that include the nucleotide sequence of SEQ ID No:6.
[0688] Concept 38D. A vector, protein, fusion protein or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option or other feature that includes a crRNA, wherein the crRNA includes at least one repeat sequence that includes the nucleotide sequence of SEQ ID No:70.
[0689] Concept 38E. A vector, protein, fusion protein, or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option, or other feature that includes a crRNA, wherein the crRNA includes two repeat sequences that include the nucleotide sequence of SEQ ID No: 70.
[0690] Concept 39. A vector, protein, fusion protein, or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option, or other feature that includes a crRNA that includes a spacer that is homologous to a first protospacer in a target sequence, wherein the protospacer is adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, and the PAM sequence is selected from the group consisting of: 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3', and 5'-GTG-3'.
[0691] In any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA comprising a spacer homologous to a first protospacer in a target sequence, the protospacer can be adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, which is selected from the following: 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3'. In any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA comprising a spacer homologous to a first protospacer in a target sequence, the protospacer can be adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, which is selected from the following: 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3'. In any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA comprising a spacer homologous to a first protospacer in a target sequence, the protospacer can be adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, which is selected from the following: 5'-AAC-3' and 5'-ATG-3'. In any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA comprising a spacer homologous to a first protospacer in a target sequence, the protospacer can be adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, which is 5'-AAC-3'. In any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA comprising a spacer homologous to a first protospacer in a target sequence, the protospacer can be adjacent to a protospacer adjacent motif (PAM) sequence having any feature described elsewhere herein.
[0692] Concept 40. A vector, protein, fusion protein or fusion protein complex described in any configuration, concept, aspect, example, embodiment, option or other feature of a crRNA, which comprises a spacer homologous to a first protospacer in a target sequence, wherein the length of the spacer sequence is (about) 25 to 39 nucleotides.
[0693] The length of the spacer sequence can be (about) 28 to 32 nucleotides. The length of the spacer sequence can be about 32 nucleotides. The length of the spacer sequence can be 32 nucleotides.
[0694] In any configuration, concept, aspect, example, embodiment, option, or other feature that includes a crRNA that includes a spacer homologous to a first protospacer in a target sequence, the length of the spacer sequence can be 32 nucleotides. In any configuration, concept, aspect, example, embodiment, option, or other feature that includes a crRNA that includes a spacer homologous to a first protospacer in a target sequence, the spacer sequence can have any feature described elsewhere herein (e.g., in the first aspect of the Summary of the Invention, in Concepts 10 and 30).
[0695] Concept 41. A vector, protein, fusion protein, or fusion protein complex described in any configuration, concept (e.g., as in Concept 40), aspect, example, embodiment, option, or other feature that includes a crRNA that includes a spacer homologous to a first protospacer in a target sequence, wherein nucleotides 1 to 28 of the spacer are identical to the complement of nucleotides 1 to 28 of the protospacer sequence immediately 5' of the PAM sequence in the target sequence.
[0696] The protospacer can have any feature described elsewhere herein (e.g., in the first aspect of the Summary of the Invention, in Concept 10, and in Concept 30).
[0697] Concept 42. A vector, protein, fusion protein, or fusion protein complex described in any configuration, concept, aspect, example, embodiment, option, or other feature that includes a crRNA that includes a spacer homologous to a first protospacer in a target sequence, wherein the spacer is identical to the complement of the protospacer in the target sequence.
[0698] The spacer can be identical to the complement of the protospacer in the target sequence over its entire length. The spacer can have (about) 80% (e.g., (about) 90%) identity to the complement of the protospacer in the target sequence over its entire length. The spacer can be identical to the complement of the protospacer in the target sequence for the first 1 to 28 nucleotides immediately 5' of the PAM sequence in the target sequence and have (about) 80% (e.g., (about) 90%) identity for the remainder of the nucleotides in the protospacer.
[0699] Concept 43. A vector, fusion protein, or fusion protein complex described in any configuration, concept, aspect, example, embodiment, option, or other feature that includes a fusion protein, wherein Py is a nuclease of the I-TevI nuclease, and wherein the I-TevI nuclease recognizes an I-TevI cleavage site nucleotide sequence that is 5' or 3' (e.g., 5') of the protospacer in the target sequence.
[0700] The nucleotide sequence of the I-TevI cleavage site can be as described elsewhere herein. In one embodiment, the nucleotide sequence of the I-TevI cleavage site is 5' of the protospacer in the target sequence. The length of the nucleotide sequence of the I-TevI cleavage site can be (about) 3 to 7 nucleotides, such as (about) 4 to 6 nucleotides in length. The length of the nucleotide sequence of the I-TevI cleavage site can be about 5 nucleotides. The nucleotide sequence of the I-TevI cleavage site can comprise the motif 5'-CNNNG-3'. The nucleotide sequence of the I-TevI cleavage site can be selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', 5'-CAGCG-3', 5'-CACTG-3', 5'-CGACG-3', 5'-CGATG-3', 5'-CTCCG-3', 5'-CTACG-3', 5'-CTGTG-3', 5'-CTGCG-3', 5'-CCGTG-3', 5'-CCTAG-3', 5'-CCATG-3', 5'-CAAAG-3', 5'-CAGAG-3', 5'-CATTG-3', and 5'-CATCG-3'. The nucleotide sequence of the I-TevI cleavage site can be selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', and 5'-CAGCG-3 or selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', and 5'-CTGAG-3'. The nucleotide sequence of the I-TevI cleavage site can be selected from 5'-CCACG-3', 5'-CACCG-3', and 5'-CATAG-3'. The nucleotide sequence of the I-TevI cleavage site can be selected from 5'-CCACG-3' and 5'-CACCG-3'. The nucleotide sequence of the I-TevI cleavage site can have the nucleotide sequence of SEQ ID No:42.
[0701] Concept 44. A vector, fusion protein, or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option, or other feature, comprising a fusion protein, wherein Py is a nuclease of the I-TevI nuclease, and wherein the I-TevI nuclease recognizes an I-TevI spacer nucleotide sequence that is 5' or 3' (e.g., 5') of the protospacer in the target sequence,
[0702] The I-TevI spacer nucleotide sequence can be as described elsewhere herein. In one embodiment, the I-TevI spacer nucleotide sequence is 5' of the protospacer in the target sequence. The length of the I-TevI spacer nucleotide sequence can be (about) 10 to 35 nucleotides, such as 15 to 35 or 20 to 35 nucleotides in length. The length of the I-TevI spacer nucleotide sequence can be (about) 29 to 33 nucleotides. The length of the I-TevI spacer nucleotide sequence can be (about) 30 to 32 nucleotides. The length of the I-TevI spacer nucleotide sequence can be (about) 31 nucleotides. The I-TevI spacer nucleotide sequence can have at least (about) 80% (such as (about) 90%) identity with the nucleotide sequence of SEQ ID No: 43. The I-TevI spacer nucleotide sequence can have 100% identity with the nucleotide sequence of SEQ ID No: 43.
[0703] Concept 45. A vector, fusion protein, or fusion protein complex described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein, comprising a fusion protein, wherein Py is a nuclease that is an I-TevI nuclease, and wherein the target sequence comprises, in 5' to 3' orientation:
[0704] a) an I-TevI cleavage site nucleotide sequence;
[0705] b) an I-TevI spacer nucleotide sequence;
[0706] c) a PAM sequence; and
[0707] d) a protospacer sequence.
[0708] In one embodiment, sequences a) to d) are each adjacent to one another. The I-TevI cleavage site nucleotide sequence is as described elsewhere herein (such as as described in Concept 43). The I-TevI spacer nucleotide sequence is as described elsewhere herein (such as as described in Concept 44). The PAM sequence is as described elsewhere herein (such as as described in Concept 39). The protospacer sequence is as described elsewhere herein (such as as described in the first aspect of the Summary of the Invention or in Concept 10 or Concept 30).
[0709] Concept 46A. A cell comprising one or more vectors, fusion proteins, or fusion protein complexes described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein.
[0710] Concept 46B. A eukaryotic cell comprising one or more vectors, fusion proteins, or fusion protein complexes described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein.
[0711] Concept 46C. An animal, plant, insect, or fungal cell that contains one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein.
[0712] Concept 46D. A prokaryotic cell that contains one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein.
[0713] Concept 46E. A prokaryotic cell that contains one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein, wherein the cell is not a bacterial cell (such as an Escherichia coli, Pseudomonas, or Klebsiella cell) that contains an endogenous nucleotide sequence encoding a polypeptide as shown in Concept 1A, a) to e).
[0714] In one embodiment, the cell is a vertebrate, mammalian, or human cell.
[0715] Concept 47A. A composition that contains one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein.
[0716] The composition can be as described elsewhere herein. In one embodiment, the composition described herein can be an in vitro composition. In one embodiment, the composition described herein can be contained in a medical container.
[0717] There is also provided a pharmaceutical composition that contains one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein, and further contains a diluent, excipient, or carrier. The pharmaceutical composition is for use in any of the medical treatment methods described herein.
[0718] There is also provided one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein for use in therapy.
[0719] There is also provided one or more vectors, fusion proteins, or fusion protein complexes as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features described elsewhere herein for treating any of the diseases described herein by practicing any of the methods described herein.
[0720] The pharmaceutical composition can be contained within a medical device (such as an ampoule, syringe or inhaler). The pharmaceutical composition can be formulated in a tincture, capsule or sustained release formulation. The pharmaceutical composition can be an oral tablet contained within a blister pack.
[0721] In one embodiment, a pharmaceutical composition comprising one or more of the carriers, fusion proteins or fusion protein complexes described herein is formulated for oral or rectal administration. In one embodiment, the pharmaceutical composition is formulated for oral administration. In one embodiment, the pharmaceutical composition is formulated as a capsule or coated tablet.
[0722] Preparations comprising one or more of the carriers, fusion proteins or fusion protein complexes described herein can be lyophilized before encapsulation. In one embodiment, the pharmaceutical composition comprising one or more of the carriers, fusion proteins or fusion protein complexes described herein is a lyophilized preparation. In one embodiment, the pharmaceutical composition comprising one or more of the carriers, fusion proteins or fusion protein complexes described herein is an encapsulated preparation for release in the lower intestine of a subject.
[0723] Concept 47B. A composition comprising a crRNA or a nucleic acid encoding a crRNA (optionally an in vitro composition or a composition contained within a medical container); and one, more or all of the following
[0724] a) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:1;
[0725] b) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:2;
[0726] c) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:3;
[0727] d) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:4; and
[0728] e) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:5.
[0729] Where
[0730] A: The crRNA comprises a spacer homologous to a first protospacer, where the protospacer
[0731] f) Is not found in Escherichia coli;
[0732] g) Is a eukaryotic protospacer; or
[0733] h) a pro - spacer region of an animal (optionally a mammal or human), plant or fungal cell;
[0734] or
[0735] B: a pro - spacer region of Escherichia coli that lacks an endogenous nucleotide sequence encoding the polypeptides shown in parts a) to e) above.
[0736] In one instance, the composition comprises (a) - (e). In one instance, the composition comprises (a) - (d). In one instance, the composition comprises (a) - (c). In one instance, the composition comprises (a) - (b). In one instance, the composition comprises (d) - (e). In one instance, the composition comprises (c) - (e). In one instance, the composition comprises (b) - (e).
[0737] In one instance, the composition comprises (a). In one instance, the composition comprises (b). In one instance, the composition comprises (c). In one instance, the composition comprises (d). In one instance, the composition comprises (e).
[0738] The crRNA can be operable in a cell to direct the polypeptides, proteins, fusion proteins, and complexes described herein to a target sequence of a nucleic acid contained in the cell. A complex is provided that comprises one, more, or all of the polypeptides (a) - (e) and the crRNA, wherein the crRNA is capable of directing the complex to a pro - spacer region contained in the target DNA. For example, the complex comprises (a) - (e).
[0739] Concept 47C. A composition comprising a crRNA or a nucleic acid encoding a crRNA (optionally an in vitro composition or a composition contained in a medical container); and a nucleic acid encoding one, more, or all of the following
[0740] a) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO:1;
[0741] b) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO:2;
[0742] c) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO:3;
[0743] d) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO:4; and
[0744] e) a polypeptide comprising an amino acid having at least (about) 80% identity to SEQ ID NO:5.
[0745] wherein
[0746] A: The crRNA comprises a spacer homologous to a first protospacer, wherein the protospacer
[0747] f) is not found in Escherichia coli;
[0748] g) is a eukaryotic protospacer; or
[0749] h) a protospacer of an animal (optionally mammalian or human), plant or fungal cell;
[0750] or
[0751] B: A protospacer of Escherichia coli that lacks an endogenous nucleotide sequence encoding the polypeptide shown in parts a) to e) above.
[0752] In one instance, the composition comprises a vector encoding at least 2, 3 or 4 of the polypeptides. In one instance, (a)-(e) are encoded by the same vector. In one instance, the composition comprises a vector encoding all of (a)-(e).
[0753] In one instance, the composition nucleic acid encodes (a)-(e). In one instance, the composition nucleic acid encodes (a)-(d). In one instance, the composition nucleic acid encodes (a)-(c). In one instance, the composition nucleic acid encodes (a)-(b). In one instance, the composition nucleic acid encodes (d)-(e). In one instance, the composition nucleic acid encodes (c)-(e). In one instance, the composition nucleic acid encodes (b)-(e).
[0754] The crRNA can be operable in a cell to direct a polypeptide to a target sequence of a nucleic acid contained in the cell. A complex is provided that comprises one, multiple or all of the polypeptides (a)-(e) and the crRNA, wherein the crRNA is capable of directing the complex to a protospacer contained in the target DNA. For example, the complex comprises (a)-(e).
[0755] Methods are also provided. Any method herein can be an in vitro method. Optionally, the method is performed on cells contained in a subject, such as a human, an animal (such as a mammal, such as a rodent, a mouse or a rat), a plant, an insect or a fungus (such as yeast).
[0756] As used herein, a modification (such as modifying a target sequence of DNA, polynucleotide or cell) can be one in which the modification cuts, edits, blocks, labels or tags the target sequence.
[0757] Concept 48A. A method of modifying a nucleic acid target sequence in a cell, the method comprising
[0758] (I) contacting a cell with one or more vectors, proteins, fusion proteins, fusion protein complexes or compositions as described in any configuration, concept, aspect, instance, embodiment, option or other feature, which comprise a crRNA comprising a spacer homologous to a first protospacer in a target sequence;
[0759] (II) allowing introduction into the cell
[0760] (i) a vector, protein, fusion protein or fusion protein complex, whereby a polypeptide and the crRNA are expressed in the cell, or
[0761] (ii) the cRNA of the composition (or a nucleic acid encoding the crRNA, whereby the cRNA is expressed in the cell) and a nucleic acid of the composition encoding a polypeptide or protein, whereby the polypeptide or protein is expressed in the cell;
[0762] (III) wherein the crRNA forms a complex with the polypeptide or protein to direct the complex to the target sequence.
[0763] When the method comprises introducing a vector expressing a CasS protein or introducing a CasS protein, the modification can be downregulation of a gene comprising the target sequence. When the method comprises introducing a vector expressing a CasS protein or introducing a CasS protein, the modification can be downregulation of a gene adjacent to the target sequence (e.g., within 2 kb of the target sequence).
[0764] When the method comprises introducing a vector expressing a fusion protein or fusion protein complex or introducing a fusion protein or fusion protein complex, the modification can be cleavage of the target sequence (e.g., when Py comprises a nuclease or nickase).
[0765] When the method comprises introducing a vector expressing a fusion protein or fusion protein complex or introducing a fusion protein or fusion protein complex, the modification can be introduction of one or more mutations in the target sequence (e.g., when Py comprises a nuclease, nickase, prime editor or base editor).
[0766] When the method comprises introducing a vector expressing a fusion protein or fusion protein complex or introducing a fusion protein or fusion protein complex, the modification can be performing base editing in the target sequence (e.g., when Py is a base editor).
[0767] When the method comprises introducing a vector expressing a fusion protein or fusion protein complex or introducing a fusion protein or fusion protein complex, the modification can be performing prime editing in the target sequence (e.g., when Py is a prime editor).
[0768] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be deamination of one or more nucleotides in the target sequence (e.g., when Py is a deaminase, such as a cytosine or adenine deaminase).
[0769] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be methylation of one or more nucleotides in the target sequence (e.g., when Py is a methyltransferase).
[0770] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be methylation of one or more nucleotides in the target sequence (e.g., when Py is a methylase).
[0771] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be acetylation of one or more nucleotides in the target sequence (e.g., when Py is an acetylase).
[0772] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be subjecting the DNA of the target sequence to acetyltransferase activity (e.g., when Py is an acetyltransferase).
[0773] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be reverse transcribing the target sequence or RNA in a cell to produce DNA inserted at the target sequence (e.g., when Py is a reverse transcriptase).
[0774] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be activating gene transcription in the cell, the gene such as containing the target sequence or a gene adjacent to the target sequence (e.g., when Py is a transcriptional activator).
[0775] When the method involves introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be inactivating gene transcription in the cell, the gene such as containing the target sequence or a gene adjacent to the target sequence (e.g., when Py is a transcriptional deactivator).
[0776] When the method comprises introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be activation of gene translation in a cell, such as a gene containing a target sequence or a gene adjacent to the target sequence (e.g., when Py is a translation activator).
[0777] When the method comprises introducing a vector expressing a fusion protein or a fusion protein complex or introducing a fusion protein or a fusion protein complex, the modification can be inactivation of gene translation in a cell, such as a gene containing a target sequence or a gene adjacent to the target sequence (e.g., when Py is a translation deactivator).
[0778] Concept 48B. A method of modifying a nucleic acid target sequence in a cell, the method comprising
[0779] a) contacting the cell with a vector of any one of Concept 10 or Concepts 11 to 12 (when subordinated to Concept 10) or a composition according to Concept 47;
[0780] b) allowing introduction into the cell of
[0781] (i) a vector by which a polypeptide and a crRNA are expressed in the cell, or
[0782] (ii) the cRNA of the composition (or a nucleic acid encoding the crRNA, by which the cRNA is expressed in the cell) and the nucleic acid of the composition encoding the polypeptide, by which the polypeptide is expressed in the cell;
[0783] c) wherein the crRNA forms a complex with the polypeptide to direct the complex to the target sequence.
[0784] It is noted that we found that the orientation of the spacer does not affect the transcriptional efficiency of downregulating the targeted gene by the type S CRISPR-Cas system, which is not the case for other CRISPR-Cas systems that must target the coding or non-coding region of a gene to effectively downregulate its transcription. Thus, the complex, for example, modifies (e.g., cuts or edits) the coding or non-coding target sequence encoding the target sequence, optionally wherein (i) the coding strand of the DNA is modified rather than the non-coding strand, or (ii) the non-coding strand of the DNA is modified rather than the coding strand. The modifications herein can usefully modify (i) the coding strand of the DNA rather than the non-coding strand, or (ii) the non-coding strand of the DNA rather than the coding strand.
[0785] Concept 48C. The method of Concept 48B, wherein the complex is directed to the target sequence and modifies the target sequence, optionally wherein the complex comprises a component selected from the following:
[0786] a) a nuclease that cuts the target sequence;
[0787] b) a deaminase (such as a cytosine or adenine deaminase) that deaminates the target sequence;
[0788] c) a base editor that performs base editing on the target sequence;
[0789] d) a prime editor that performs prime editing on the target sequence;
[0790] e) a reverse transcriptase that reverse transcribes the target sequence or RNA in a cell to produce DNA inserted at the target sequence;
[0791] f) a methyltransferase that methylates DNA in a cell;
[0792] g) a methylase that methylates DNA in a cell;
[0793] h) an acetylase that acetylates DNA in a cell;
[0794] i) an acetyltransferase that subjects DNA in a cell to acetyltransferase activity;
[0795] j) a transcriptional activator that activates gene transcription in a cell, such as a gene containing or adjacent to the target sequence; and
[0796] k) a transcriptional deactivator that inactivates gene transcription in a cell, such as a gene containing or adjacent to the target sequence;
[0797] Optionally, wherein the component is encoded by a vector or nucleic acid of the composition.
[0798] In any configuration, concept, aspect, example, embodiment, option or other feature involving modifying target DNA or RNA, the complex further comprises a base editor (such as as described elsewhere herein).
[0799] In any configuration, concept, aspect, example, embodiment, option or other feature involving modifying target DNA or RNA, the complex further comprises a prime editor (such as as described elsewhere herein).
[0800] In any configuration, concept, aspect, example, embodiment, option or other feature involving modifying target DNA or RNA, the complex further comprises a nuclease (such as as described elsewhere herein).
[0801] Concept 49A. The method of Concept 48C, wherein the cell is a eukaryotic, animal, plant, insect or fungal cell.
[0802] Concept 49B. The method of Concept 48, wherein the cell is a prokaryotic cell.
[0803] In one embodiment, the cell is a prokaryotic cell, where the cell is not a prokaryotic cell (such as an Escherichia coli, Pseudomonas or Klebsiella cell) containing an endogenous nucleotide sequence encoding polypeptides a) to e) as shown in Concept 1.
[0804] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the animal herein is a non-human mammal or vertebrate. Optionally, the animal cell herein is a non-human mammal or vertebrate cell. Optionally, the cell is a plant cell, bacterial cell, fungal cell, mammalian cell, insect cell or archaeal cell. The cell herein can be ex vivo or in vivo.
[0805] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the cell herein is a pluripotent cell, such as a pluripotent stem cell. Optionally, the cell herein is a stem cell, such as an embryonic stem cell. Optionally, the cell herein is a totipotent cell. In embodiments of these options, the cell is a human cell. Alternatively, the cell is not a human cell and is not comprised in a human or human embryo.
[0806] Optionally, any method herein is not any of the following:
[0807] (a) A method for cloning humans;
[0808] (b) A method for modifying the germline genetic identity of humans;
[0809] (c) The use of a human embryo;
[0810] (d) A method for modifying the genetic identity of an animal that may cause the animal to suffer without any substantial medical benefit to a human or animal.
[0811] Optionally, any method herein is not a method of treatment of the human or animal body. Optionally, any method herein is not a surgical method.
[0812] Concept 50. The method of Concept 48 or 49, where the target sequence is comprised in a chromosome or episome in the cell.
[0813] In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is comprised in a chromosome. In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is comprised in a plasmid. In any configuration, concept, aspect, example, embodiment, option or other feature herein, the target sequence is not comprised in a plasmid.
[0814] Concept 51. The method of Concept 50, where the method inhibits chromosomal or episomal replication in the cell.
[0815] In one instance, the inhibition is at least (about) 20, 30, 40, 50, 60, 70, 80, 90, or 95% compared to replication in the same cells not exposed to the vector or composition. In one instance, the inhibition is 100%.
[0816] Concept 52. The method of any one of Concepts 48 to 51, wherein the target sequence is within (about) 2 kb upstream or downstream of the gene of interest.
[0817] In any configuration, concept, aspect, instance, embodiment, option, or other feature involving a target sequence, the target sequence is within (about) 2 kb upstream (5') or downstream (3') of the gene of interest. The target sequence can be upstream (5') of the gene of interest. The target sequence can be downstream (3') of the gene of interest. The target sequence can be within (about) 1.75 kb, (about) 1.5 kb, (about) 1.25 kb, (about) 1 kb, (about) 0.75 kb, (about) 0.5 kb, or (about) 0.25 kb upstream (5') or downstream (3') of the gene of interest. The target sequence can be within the gene of interest, for example, within the promoter, exon, and / or intron of the gene of interest. In one embodiment, the target sequence is within the promoter of the gene of interest.
[0818] Concept 53. The method of Concept 52, wherein the target sequence of the gene of interest is at least 2 kb upstream and downstream of any promoter, exon, and intron of any other gene.
[0819] As shown in Example 4 herein, the CasS3 protein can have a wide range of effects, where the binding effect is shown within about 2 kb on either side (i.e., 5' and 3') of the protospacer in the target sequence. Thus, for applications where the CasS system is used for CRISPRi (i.e., in any method of upregulating or downregulating a gene of interest described herein), it may be beneficial to avoid off-target effects if the protospacer in the target sequence is at least 2 kb upstream (5') and downstream (3') of any other gene (i.e., promoter, exon, or intron).
[0820] Concept 54. The method of any one of Concepts 48 to 53, wherein the modification is a substitution, deletion, or insertion of one or more nucleotides.
[0821] In one embodiment, the modification is a substitution of one or more nucleotides in the target sequence. In one embodiment, when Py is a CBE, the modification is one or more C to T substitutions in the nucleotides of the target sequence. In one embodiment, the CBE is PmCDA1 cytidine deaminase, as described elsewhere herein, for example.
[0822] Concept 55. A method of treating or preventing a disease or condition mediated by target cells in a subject, the method comprising performing a method as described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., as described in any one of Concepts 48 to 54) to modify the target cells, wherein the contacting comprises administering a vector, a protein (including a fusion protein or a fusion protein complex) or a composition to the subject, and wherein the modification treats or prevents the disease or condition.
[0823] The target cells can be, for example, pathogenic bacterial cells. They can be antibiotic-resistant bacterial cells, and the modification renders the bacterial cells sensitive to antibiotics again.
[0824] Concept 56. The method of Concept 55, wherein the subject is a human, an animal or a plant.
[0825] Concept 57. The method of Concept 56, wherein the target cells are cells of a subject comprising a nucleic acid defect, and the modification corrects the defect.
[0826] For example, certain diseases are known to be caused by single nucleotide polymorphisms (SNPs) (e.g., Hutchinson-Gilford progeria syndrome and mandibuloacral dysplasia are caused by the missense SNP c.1580G>T SNP in the LMNA gene encoding lamin A / C protein, while SNPs in the F5 gene result in factor V Leiden thrombophilia). Correcting these SNPs by targeted replacement can be used to ameliorate the diseases. In 2012, there were 124 genetic diseases confirmed to be caused by retrotransposon insertions, including cystic fibrosis (Alu), hemophilia A (L1) and X-linked dystonia-parkinsonism (SVA), see Hancks et al., Curr. Opin. Genet. Dev., 22, 191–203, 2012, doi:10.1016 / j.gde.2012.02.006, which is incorporated herein by reference in its entirety. There are certain targeted edits that can be used to correct these types of insertions and improve the related diseases.
[0827] Concept 58A. The method of Concept 56 or 57, wherein the modification adds a new function to the cell.
[0828] In one embodiment, the modification upregulates or downregulates gene expression in the cell. In one embodiment, the modification downregulates gene expression in the cell. In one embodiment, the modification inhibits gene expression in the cell. In one embodiment, the modification blocks gene expression in the cell.
[0829] In another embodiment, the modification is the insertion of a new gene or gene pathway (e.g., an operon) that can be used to produce substances beneficial to the cell. For example, the new gene can produce beneficial metabolites or therapeutic agents. In another embodiment, the modification is the insertion of a new gene or gene pathway (e.g., an operon) that can be used to remove substances harmful to the cell. For example, the new gene can metabolize harmful toxins.
[0830] Concept 58B. The method of Concept 56 or 57, wherein the modification adds a new nucleotide sequence into the genome of the cell for expressing a protein encoded by the new sequence.
[0831] For example, the protein is a heterologous protein. For example, the protein is a therapeutic protein, such as an antibody, an antibody fragment (e.g., an antibody single variable domain), a TCR (T cell receptor) binding site, a TCR variable domain, a hormone, an incretin, a growth factor, an anti-cancer agent, a neurotransmitter, or an enzyme.
[0832] Concept 58C. The method of Concept 56 or 57, wherein the modification regulates the expression of a nucleotide sequence contained in the chromosome or episome of the cell.
[0833] In one embodiment, the modification regulates the expression of a nucleotide sequence contained in a plasmid. In one embodiment, the modification regulates the expression of a nucleotide sequence not contained in a plasmid. For example, the expression is upregulated. For example, the expression is downregulated.
[0834] Concept 59. A vector, protein (including fusion proteins and fusion protein complexes), cell, or composition as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 1-58) for use in a method of treating or preventing a disease or condition in a human or animal, wherein the method is as described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 55 to 58).
[0835] Any vector herein (in any configuration, concept, aspect, example, embodiment, option, or other feature herein) can be
[0836] a) a plasmid vector (optionally a conjugative plasmid);
[0837] b) a transposon vector (optionally a conjugative transposon);
[0838] c) a viral vector (optionally a phage, AAV, or lentiviral vector);
[0839] d) a phagemid (optionally a packaged phagemid); or
[0840] e) a nanoparticle (optionally a lipid nanoparticle).
[0841] Concept 60. A method for introducing targeted editing in a target polynucleotide, the method comprising contacting the target polynucleotide with a fusion protein or fusion protein complex as described in any configuration, concept, aspect, example, embodiment, option or other feature herein, wherein the protein or fusion protein complex is combined with a crRNA, the crRNA comprising a spacer homologous to a first protospacer in the target sequence (e.g., as described in the first aspect of the Summary of the Invention, Concept 10 or in Concept 30 or in any one of Concepts 38 to 45 subordinate to Concept 30).
[0842] Wherein the first protospacer in the target sequence is contained in a polynucleotide, and the crRNA hybridizes with the protospacer to direct a protein (e.g., a ribonucleoprotein complex), whereby the protein (e.g., a ribonucleoprotein complex) edits the polynucleotide.
[0843] The method can be implemented in a parental cell containing a polynucleotide. Further provided are progeny cells or organisms derived from the edited parental cell, wherein the progeny cells or organisms retain the edit in their genome. The method can include one or more steps of culturing the edited cells to produce the progeny cells or a plurality of such progeny cells.
[0844] Concept 61A. The method of Concept 60, wherein the method inserts, deletes or replaces bases or nucleic acid sequences in a polynucleotide.
[0845] Concept 61B. The method of a vector of Concept 10 or any one of Concepts 11 or 12 subordinate to Concept 10, the method of any one of Concepts 48 - 54 or the method of Concept 60 or 61, wherein the crRNA is homologous to a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0846] Alternatively, the PAM can be any one of the PAMs described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., having one of the sequences shown in Concept 39).
[0847] Concept 62A. A container containing a plurality of proteins that are operable to be used with a crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide, the proteins comprising any protein described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., as described in Concept 13).
[0848] Wherein the complex can operate with any protospacer adjacent motif (PAM) (e.g., having one of the sequences shown in Concept 39) described in any configuration, concept, aspect, example, embodiment, option, or other feature herein, and the container is not a cell, and the protein is mixed with an in vitro buffer.
[0849] Concept 62B. A container comprising a plurality of proteins that are operable to be used with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the proteins comprising any fusion protein or fusion protein complex described in any configuration, concept, aspect, example, embodiment, option, or other feature herein (e.g., as described in any one of Concepts 14 to 29).
[0850] Wherein the complex can operate with any protospacer adjacent motif (PAM) (e.g., having one of the sequences shown in Concept 39) described in any configuration, concept, aspect, example, embodiment, option, or other feature herein, and the container is not a cell, and the protein is mixed with an in vitro buffer.
[0851] Concept 62C. A container comprising one or more proteins that are operable to be used with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the proteins comprising one, more, or all of the polypeptides selected from the following
[0852] a) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:1;
[0853] b) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:2;
[0854] c) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:3;
[0855] d) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:4; and
[0856] e) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:5.
[0857] Wherein the complex can operate with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3', the container is not a cell, and the protein is mixed with an in vitro buffer.
[0858] Alternatively, the PAM can be any of the PAMs described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein (e.g., having one of the sequences shown in Concept 39).
[0859] Also provided are: -
[0860] A protein or proteins that are operable with a crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide, the protein comprising one, more than one, or all of the polypeptides selected from the following
[0861] a) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:1;
[0862] b) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:2;
[0863] c) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:3;
[0864] d) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:4; and
[0865] e) A polypeptide comprising amino acids having at least (about) 80% identity to SEQ ID NO:5.
[0866] Wherein the complex is operable with a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3'.
[0867] Alternatively, the PAM can be any of the PAMs described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein (e.g., having one of the sequences shown in Concept 39).
[0868] Concept 63. A nucleic acid vector or nucleic acid vectors as described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein (e.g., according to any one of Concepts 1 to 12 or any one of Concepts 31 to 45), wherein the polypeptide or protein (including fusion proteins and fusion protein complexes) encoded by the vector is operable with a crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide, wherein the complex is operable with any protospacer adjacent motif (PAM) described in any of the configurations, concepts, aspects, examples, embodiments, options, or other features herein (e.g., having one of the sequences shown in Concept 39).
[0869] In one embodiment, the protospacer adjacent motif (PAM) has the sequence 5'-AAG-3'.
[0870] Concept 64. A method of targeting a polynucleotide, the method comprising
[0871] a) contacting the polynucleotide with a protein (including a fusion protein and a fusion protein complex) as described in any configuration, concept, aspect, example, embodiment, option or other feature herein, wherein the protein is combined with a crRNA comprising a spacer homologous to a first protospacer in a target sequence;
[0872] b) allowing the formation of a ribonucleoprotein complex comprising the protein and the crRNA, wherein the complex is directed to a target sequence comprised in the polynucleotide to modify the polynucleotide or its replication.
[0873] The targeting can be in a cell. The targeting can also be in vitro.
[0874] Concept 65. The method of Concept 64, wherein the polynucleotide is comprised in a chromosome or an episome.
[0875] In one embodiment, the polynucleotide is comprised in a chromosome. In one embodiment, the polynucleotide is comprised in a plasmid. In one embodiment, the polynucleotide is not comprised in a plasmid.
[0876] Concept 66A. The method of Concept 65, wherein the method inhibits the replication of the polynucleotide.
[0877] Concept 66B. The method of Concept 65, wherein the method edits the polynucleotide.
[0878] Editing can insert a base or a nucleic acid sequence into the polynucleotide. Editing can delete a base or a nucleic acid sequence in the polynucleotide. Editing can replace a base or a nucleic acid sequence in the polynucleotide.
[0879] Concept 67A. A kit comprising
[0880] a) one or more proteins as described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., as described in Concept 13); and
[0881] b) a crRNA, (or one or more nucleic acids encoding the crRNA), wherein the crRNA is homologous to any (PAM) as described in any configuration, concept, aspect, example, embodiment, option or other feature herein (e.g., a PAM having one of the sequences shown in Concept 39);
[0882] wherein the polypeptide is operable to be used with the crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide.
[0883] Concept 67B. A kit comprising
[0884] a) one or more fusion proteins or fusion protein complexes or one or more nucleic acids encoding such polypeptides described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., as described in any one of Concepts 14 to 29); and
[0885] b) a crRNA or one or more nucleic acids encoding a crRNA), wherein the crRNA is homologous to any (PAM) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., a PAM having one of the sequences shown in Concept 39);
[0886] wherein the polypeptide is operable to be used with the crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide.
[0887] Concept 67C. A kit comprising
[0888] a) one or more vectors or one or more nucleic acids encoding such polypeptides described as in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., as described in any one of Concepts 1 to 12 or Concepts 31 to 45); and
[0889] b) a crRNA or one or more nucleic acids encoding a crRNA), wherein the crRNA is homologous to any (PAM) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., a PAM having one of the sequences shown in Concept 39);
[0890] wherein the polypeptide is operable to be used with the crRNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide.
[0891] Concept 67D. A kit comprising
[0892] a) one or more polypeptides as shown in Concept 7 or one or more nucleic acids encoding such polypeptides; and
[0893] b) a crRNA, (or one or more nucleic acids encoding a crRNA, wherein the crRNA is homologous to a PAM having the sequence 5'-AAG-3';
[0894] wherein the polypeptide is operable to be used with the crRNA or a guide RNA to form a ribonucleoprotein complex for targeting a protospacer in a polynucleotide.
[0895] Provide a complex as described herein. In one instance, the complex further comprises one or more cascading Cas proteins, such as one, more, or all of the type I CasA-E proteins. The complex can further comprise Cas3 and optionally one or more homologous type I cascading proteins. The complex can further comprise Cas3, 9, 10, 12, or 13. Thus, the function of the S-type system or component can be provided together with the function of one or more proteins regarding different types of CRISPR / Cas systems. The complex can comprise a helicase fused to a nuclease. The nuclease can be, for example, Cas3, 9, 10, 12, or 13. For example, the nuclease is Cas3. For example, the nuclease is Cas9.
[0896] In one instance, the complex further comprises a TevI nuclease, such as an I-TevI nuclease domain. I-TevI can be as described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein (such as in concept 14 or concepts 21 to 24 or concepts 43 to 45).
[0897] In one instance, the complex further comprises a MutH protein.
[0898] In one instance, the complex is an isolated complex. The complex can be in vitro. The complex can be contained in a medical container, such as an intravenous bag, a medical vial, or a medical injection device.
[0899] Concept 68A. A ribonucleoprotein complex comprising
[0900] a) one or more proteins as described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein (such as as described in concept 13); and
[0901] b) a crRNA homologous to any (PAM) described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein (such as having one of the sequences shown in concept 39).
[0902] Concept 68B. A ribonucleoprotein complex comprising
[0903] a) one or more fusion proteins or fusion protein complexes as described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein (such as as described in any one of concepts 14 to 29); and
[0904] b) a crRNA homologous to any (PAM) described in any configuration, concept, aspect, instance, embodiment, option, or other feature herein (such as having one of the sequences shown in concept 39).
[0905] Concept 68C. Ribonucleoprotein complex, comprising
[0906] a) one or more polypeptides as shown in Concept 5 or one or more nucleic acids encoding such polypeptides; and
[0907] b) crRNA or one or more nucleic acids encoding crRNA), wherein the crRNA is homologous to a PAM having the sequence 5'-AAG-3'.
[0908] Concept 68D. The complex of Concept 68C, wherein the complex lacks
[0909] a) Cas nuclease;
[0910] b) DNA nuclease or RNA nuclease; or
[0911] c) nuclease.
[0912] Concept 69. Method for transcriptional control or replication of target DNA, comprising contacting the target DNA with any complex described herein (e.g., the complex of Concept 68), wherein the complex lacks a DNA nuclease and the complex binds to the target DNA, thereby controlling the transcription or replication of the target DNA.
[0913] Concept 70A. Method for controlling the replication of target RNA, comprising contacting the target RNA or DNA encoding RNA with any complex described herein (e.g., the complex of Concept 68), wherein the complex binds to the target RNA or DNA, thereby controlling the transcription of the target RNA.
[0914] Concept 70B. Method for controlling the transcription of target RNA, comprising contacting the target RNA or DNA encoding RNA with any complex described herein (e.g., the complex of Concept 68), wherein the complex binds to the target RNA or DNA, thereby controlling the transcription of the target RNA.
[0915] Concept 71. The method of Concept 69 or 70, wherein the DNA is contained in a chromosome or episome (optionally a plasmid).
[0916] In one embodiment, the DNA is contained in a chromosome. In one embodiment, the DNA is contained in a plasmid. In one embodiment, the DNA is not contained in a plasmid.
[0917] Concept 72. Method for editing target DNA, comprising contacting the target DNA with any complex described herein (e.g., the complex of Concept 68), wherein the complex binds to the target DNA, thereby editing the target DNA.
[0918] Concept 73A. A method for cleaving double-stranded DNA (dsDNA), which comprises contacting the dsDNA with any of the complexes described herein (e.g., the complex of Concept 68), wherein the complex comprises a fusion protein, wherein Py comprises a nuclease, wherein the dsDNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by any one of the PAMs (e.g., it has one of the sequences shown in Concept 39) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein, whereby the nuclease cleaves the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
[0919] Concept 73B. A method for cleaving RNA, which comprises contacting the RNA with any of the complexes described herein (e.g., the complex of Concept 68), wherein the complex comprises a fusion protein, wherein Py comprises an RNA nuclease, wherein the RNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by any one of the PAMs (e.g., it has one of the sequences shown in Concept 39) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein, whereby the nuclease cleaves the RNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
[0920] Concept 73C. A method for cleaving double-stranded DNA (dsDNA), which comprises contacting the dsDNA with the complex of Concept 68C or 68D, wherein the complex comprises a nuclease, wherein the dsDNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the nuclease cleaves the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
[0921] Alternatively,
[0922] - the PAM can be AAN, ANG, NAG; or
[0923] - the PAM can be AAG with no nucleotide change or with one nucleotide change.
[0924] The nuclease can be any nuclease described herein (e.g., the I-TevI nuclease described elsewhere herein). The nuclease can be a DNA nuclease. The nuclease can be an RNA nuclease. The nuclease can be a TALEN. The nuclease can be a zinc finger protein. The nuclease can be a Cas nuclease from any naturally occurring system. The nuclease can be a synthetic Cas nuclease.
[0925] Concept 74A. A method for cleaving single-stranded DNA (ssDNA), which comprises contacting the ssDNA with any of the complexes described herein (e.g., the complex of Concept 68), wherein the complex comprises a fusion protein, wherein Py comprises a nickase, wherein the ssDNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by any one of the configurations, concepts, aspects, examples, embodiments, options, or other features described herein (PAM) (e.g., it has one of the sequences shown in Concept 39), whereby the nickase cleaves the single strand of ssDNA in the region defined by the complementary binding of the spacer sequence of crRNA to the protospacer region.
[0926] Concept 74B. A method for cleaving a single strand in double-stranded DNA (dsDNA), which comprises contacting the dsDNA with any of the complexes described herein (e.g., the complex of Concept 68), wherein the complex comprises a fusion protein, wherein Py comprises a nickase, wherein the dsDNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by any one of the configurations, concepts, aspects, examples, embodiments, options, or other features described herein (PAM) (e.g., it has one of the sequences shown in Concept 39), whereby the nickase cleaves the single strand of dsDNA in the region defined by the complementary binding of the spacer sequence of crRNA to the protospacer region.
[0927] Concept 74C. A method for cleaving single-stranded DNA (ssDNA), which comprises contacting the DNA with the complex of Concept 68C or 68D, wherein the complex comprises a nickase, wherein the DNA comprises a protospacer sequence, and the protospacer sequence is flanked at its 5' end by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the nickase cleaves the single strand of DNA in the region defined by the complementary binding of the spacer sequence of crRNA to the protospacer region.
[0928] The nickase can be any nuclease described herein. The nickase can be a DNA nickase. The nickase can be selected from Nt.BstNBI, Nb.BsrDI, Nb.BtsI, Nt.AlwI, Nb.BbvCI, Nt.BbvCI, and Nb.BsmI (available from New England BioLabs). The nickase is a synthetic Cas nickase (e.g., nCas9 or Cas9D10A).
[0929] Concept 75A. A method of marking or identifying a DNA region, which includes contacting the DNA with any complex described herein (such as the complex of Concept 68), wherein the DNA contains a protospacer sequence, and the protospacer sequence is flanked by a protospacer adjacent motif (PAM) described in any one of any configuration, concept, aspect, example, embodiment, option or other feature herein at its 5' flank (for example, it has one of the sequences shown in Concept 39), whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, and optionally wherein the complex contains a fusion protein, and Py contains a detectable tag.
[0930] Concept 75B. A method of marking or identifying a DNA region, which includes contacting the DNA with the complex of Concept 68C or 68D, wherein the DNA contains a protospacer sequence, and the protospacer sequence is flanked by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change at its 5' flank, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, and optionally wherein the complex contains a detectable tag.
[0931] Concept 76A. A method of modifying the transcription of a DNA region, which includes contacting the DNA with any complex described herein (such as the complex of Concept 68), wherein the DNA contains a protospacer sequence, and the protospacer sequence is flanked by a protospacer adjacent motif (PAM) described in any one of any configuration, concept, aspect, example, embodiment, option or other feature herein at its 5' flank (for example, it has one of the sequences shown in Concept 39), whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the complex upregulates or downregulates the transcription of the DNA region or an adjacent gene.
[0932] Concept 76B. A method of modifying the transcription of a DNA region, which includes contacting the DNA with the complex of Concept 68C or 68D, wherein the DNA contains a protospacer sequence, and the protospacer sequence is flanked by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change at its 5' flank, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the complex upregulates or downregulates the transcription of the DNA region or an adjacent gene.
[0933] In one embodiment, the complex downregulates transcription. The downregulation can be at least about 50% compared to transcription in the absence of the complex. In one embodiment, the downregulation is at least about 60%, 70%, 80%, or 90%. In another embodiment, the downregulation is at least about 95%, 96%, 97%, 98%, or 99%. In another embodiment, the downregulation is 100%, i.e., transcription of the DNA region is completely blocked.
[0934] In some embodiments, when a gene is overexpressed and is the cause of a negative phenotype, downregulation may be desirable, but complete removal of the protein would be harmful. Downregulation of the target gene can be used to generate and study knockout phenotypes where generation of a complete knockout is lethal to the cell or organism.
[0935] In some embodiments, blocking transcription of an essential gene in a pathogenic bacterium can be used to kill the pathogenic bacterium while leaving other beneficial bacteria unaffected. Similarly, a pathogenic bacterium can be killed by blocking transcription of the origin of replication. Blocking non-essential genes in a cell or organism can be used to study knockout and phenotypic behavior.
[0936] Concept 77A. A method of modifying a target dsDNA of a cell without introducing a dsDNA break, the method comprising generating in the cell any complex described herein (e.g., the complex of Concept 68), wherein the complex targets a target dsDNA contained in the cell, the target dsDNA comprising a protospacer sequence, the protospacer sequence being flanked at its 5' end by any one of the configurations, concepts, aspects, examples, embodiments, options, or other features described in any of the herein (PAM) (e.g., having one of the sequences shown in Concept 39), whereby the complex binds to the dsDNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the dsDNA is modified without introducing a dsDNA break.
[0937] Concept 77B. A method of modifying a target dsDNA of a cell without introducing a dsDNA break, the method comprising generating in the cell a complex of Concept 68C or 68D, wherein the complex targets a target DNA contained in the cell, the target DNA comprising a protospacer sequence, the protospacer sequence being flanked at its 5' end by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the DNA is modified without introducing a break in the DNA.
[0938] Concept 78. The method of Concept 77, wherein the method edits the DNA, e.g., wherein the editing inserts, deletes, or replaces bases or nucleic acid sequences in the DNA.
[0939] When the complex is a complex of Concept 68B and Py is a base editor or a prime editor (e.g., any base editor or prime editor described elsewhere herein), the modification can be the introduction of one or more substitutions. When the complex is a complex of Concept 68B and Py is a nuclease or a nickase, and the cell is a cell capable of repairing dsDNA breaks or ssDNA breaks via the NHEJ mechanism, the modification can be the introduction of one or more insertions and / or deletions (INDELS).
[0940] Concept 79A. A method of inhibiting the growth or proliferation of a cell without introducing a lethal dsDNA break, the method comprising generating in the cell any complex described herein (e.g., a complex of Concept 68), wherein the complex targets a target DNA contained in the cell, the target DNA comprising a protospacer sequence, the protospacer sequence being flanked at its 5' end by any one of the configurations, concepts, aspects, examples, embodiments described herein (PAM), whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby DNA replication is inhibited without introducing a lethal dsDNA break in the DNA, thereby inhibiting the growth or proliferation of the cell.
[0941] Concept 79B. A method of inhibiting the growth or proliferation of a cell without introducing a lethal dsDNA break, the method comprising generating in the cell a complex of Concept 68C or 68D, wherein the complex targets a target DNA contained in the cell, the target DNA comprising a protospacer sequence, the protospacer sequence being flanked at its 5' end by a PAM having the sequence 5'-AAG-3' or a PAM identical except for one base change, whereby the complex binds to the DNA in the region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby DNA replication is inhibited without introducing a break in the DNA, thereby inhibiting the growth or proliferation of the cell.
[0942] When the complex is a complex of Concept 68B and Py is a nickase that introduces a single-strand break in the DNA, and the cell is a cell capable of repairing dsDNA breaks or ssDNA breaks via the NHEJ mechanism, the growth or proliferation of the cell can be inhibited by the introduction of one or more insertions and / or deletions (INDELS). When the complex is a complex of Concept 68A, the growth or proliferation of the cell can be inhibited by downregulating or blocking the transcription of an essential gene in the cell. Essential genes are well known to those of skill in the art. When the complex is a complex of Concept 68A, the growth or proliferation of the cell can be inhibited by downregulating or blocking the transcription of the chromosomal replication origin in the cell.
[0943] Concept 80. A method of treating or preventing a disease or condition in a human, animal, plant, or fungal subject, the method comprising practicing the method of any one of Concepts 69 to 79, wherein the cells of the subject contain the DNA or RNA and the cells mediate the disease or condition.
[0944] Concept 81A. The method of Concept 81, wherein the complex modifies a coding target sequence or a non-coding target sequence.
[0945] In one embodiment, the coding strand of the DNA is modified rather than the non-coding strand. In another embodiment, the non-coding strand of the DNA is modified rather than the coding strand.
[0946] Concept 81B. The method of any one of Concepts 69 to 80, wherein the complex modifies (e.g., cuts, edits, blocks, labels, or tags) a coding target sequence or a non-coding target sequence, optionally wherein (i) the coding strand of the DNA is modified rather than the non-coding strand or (ii) the non-coding strand of the DNA is modified rather than the coding strand.
[0947] Concept 82. The method of Concept 81, wherein the modification effects a cut, edit, block, label, or tag on the target sequence.
[0948] When the complex is the complex of Concept 68A, the modification can be a block. When the complex is the complex of Concept 68A, the modification can be a downregulation of the transcription of the coding target sequence. When the complex is the complex of Concept 68B, wherein Py is a nuclease or a nickase, the modification can be a cut of the coding target sequence. When the complex is the complex of Concept 68B, wherein Py is a base editor or a prime editor, the modification can be an edit of the coding target sequence. The base editor can introduce one or more substitutions in the coding target sequence.
[0949] Concept 83A. One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NO: 7-11, wherein the nucleotide sequence is operably linked to a heterologous promoter, a synthetic promoter, a eukaryotic promoter, or a non-bacterial promoter.
[0950] Concept 83B. One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NO: 7-11, wherein the nucleotide sequence is contained in a cell that is a eukaryotic cell, a non-bacterial cell, or a cell that is not a bacterial cell (e.g., an Escherichia coli, Pseudomonas, or Klebsiella cell) containing an endogenous nucleotide sequence comprising SEQ ID NO: 7-11.
[0951] Concept 83C. One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NOs: 7 - 11, wherein said nucleotide sequence is operably linked to a heterologous promoter, synthetic promoter, eukaryotic promoter or non - bacterial promoter, and wherein said nucleotide sequence is comprised within a cell, which is a eukaryotic cell, non - bacterial cell or a bacterial cell (such as Escherichia coli, Pseudomonas or Klebsiella cells) that does not contain an endogenous nucleotide sequence that comprises SEQ ID NOs: 7 - 11.
[0952] Concept 84. A nucleic acid vector or nucleic acid, comprising SEQ ID NO: 12 or a DNA sequence having at least (about) 80% identity with SEQ ID NO: 12.
[0953] Concept 85. One or more nucleic acids, encoding multiple Cas proteins and comprising at least one nucleotide sequence for generating crRNA, wherein said Cas proteins and RNA are capable of forming a ribonucleoprotein CRISPR / Cas complex, wherein said RNA is capable of guiding the complex to a protospacer comprised within a target DNA, wherein the 5' end of said protospacer is flanked by any one of the protospacer adjacent motifs (PAMs) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (such as having one of the sequences shown in Concept 39), wherein
[0954] a) the complex lacks a DNA nuclease and is capable of modifying DNA without introducing a double - strand DNA break;
[0955] b) the Cas proteins do not include DinG, Cas3 and Cas10; and
[0956] c) the complex does not include all of the Cas proteins of a type I, II, III, IV, V or VI CRISPR / Cas complex.
[0957] Concept 86. The nucleic acid of Concept 85, wherein said nucleic acid is comprised within a vector or said nucleic acid is comprised within one or more vectors.
[0958] Concept 87. The nucleic acid of Concept 85 or 86, wherein said DNA is comprised within a chromosome.
[0959] Concept 88. A ribonucleoprotein CRISPR / Cas complex comprising multiple Cas proteins and crRNAs, wherein the RNA is capable of guiding the complex to a protospacer contained in a target DNA, wherein the 5'-end of the protospacer is flanked by any one of the protospacer adjacent motifs (PAMs) described in any of the configurations, concepts, aspects, examples, embodiments, options or other features herein (e.g., having one of the sequences shown in Concept 39), wherein
[0960] a) the complex lacks a DNA nuclease and is capable of modifying DNA without introducing a double-stranded DNA break;
[0961] b) the complex lacks Cas3 and Cas10; and
[0962] c) the complex does not contain all the Cas proteins of a type I, II, III, IV, V or VI CRISPR / Cas complex.
[0963] Concept 89. A nucleic acid of any one of Concepts 85-87 or the complex of Concept 88, wherein the crRNA is not a type I, II, III, V or VI crRNA, respectively.
[0964] Concept 90. A nucleic acid or complex of any one of Concepts 85 to 89, wherein the complex is capable of modifying (i) the coding strand of DNA rather than the non-coding strand or (ii) the non-coding strand of DNA rather than the coding strand.
[0965] Concept 91. A nucleic acid or complex of any one of Concepts 85 to 90, wherein the complex modifies (e.g., cuts, edits, blocks, labels or tags) the coding target sequence or non-coding target sequence of DNA.
[0966] Concept 92. A nucleic acid or complex of any one of Concepts 85 to 91, wherein the multiple Cas proteins comprise Cas-S1, S2, S3, S4 and S5.
[0967] For example, the multiple Cas proteins comprise Cas-S1.1, S2.1, S3.1, S4.1 and S5.1.
[0968] Concept 93. A nucleic acid or complex of any one of Concepts 85 to 92, wherein the complex further comprises an effector protein domain for modifying DNA, optionally wherein the effector domain comprises nuclease activity, nickase activity, recombinase activity, reverse transcriptase, helicase, deaminase activity, methyltransferase activity, methylase activity, acetylase activity, acetyltransferase activity, transcriptional activation activity or transcriptional inhibition activity.
[0969] In one example, the complex comprises components selected from the following
[0970] a) A nuclease that cleaves a target site;
[0971] b) A deaminase (e.g., a cytosine or adenine deaminase) that deaminates the target site;
[0972] c) A base editor that performs base editing on the target site;
[0973] d) A prime editor that performs prime editing on the target site;
[0974] e) A reverse transcriptase that reverse transcribes the target site or an RNA in a cell to generate a DNA inserted at the target site;
[0975] f) A methyltransferase that methylates DNA in a cell;
[0976] g) A methylase that methylates DNA in a cell;
[0977] h) An acetylase that acetylates DNA in a cell;
[0978] i) An acetyltransferase that subjects DNA in a cell to acetyltransferase activity;
[0979] j) A transcriptional activator that activates gene transcription in a cell, such as a gene containing the target site or a gene adjacent to the target site; and
[0980] k) A transcriptional deactivator that inactivates gene transcription in a cell, such as a gene containing the target site or a gene adjacent to the target site.
[0981] Concept 94. A cell comprising a nucleic acid or complex of any one of Concepts 85 to 93, wherein the DNA is contained in the chromosome of the cell.
[0982] Concept 95. A cell comprising a nucleic acid or complex of any one of Concepts 85 to 93, wherein the DNA is contained in a plasmid in the cell.
[0983] Concept 96. The cell of Concept 95, wherein the complex is capable of inhibiting the growth or proliferation of the cell.
[0984] Concept 97. A method of modifying DNA, the method comprising
[0985] a) contacting the DNA with a nucleic acid defined in any one of Concepts 85 to 87 and 89 to 93;
[0986] b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site contained in the DNA to modify the DNA.
[0987] Concept 98. A method for inhibiting the replication of a plasmid containing DNA, the method comprising
[0988] a) contacting the DNA with a nucleic acid defined in any one of Concepts 85 to 87 and 89 to 93;
[0989] b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site contained in the DNA to inhibit plasmid replication.
[0990] Concept 99. A method for inhibiting the transcription of a nucleotide sequence contained in DNA, the method comprising
[0991] a) contacting the DNA with a nucleic acid defined in any one of Concepts 85 to 87 and 89 to 93;
[0992] b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site contained in the nucleotide sequence to inhibit its transcription.
[0993] Concept 100. A method for inhibiting the growth or proliferation of a cell (optionally a prokaryotic cell, such as a bacterial cell) containing DNA, the method comprising
[0994] a) contacting the DNA with a nucleic acid defined in any one of Concepts 85 to 87 and 89 to 93;
[0995] b) allowing the formation of a ribonucleoprotein complex comprising a Cas protein and a crRNA, wherein the complex is directed to a target site contained in the DNA to inhibit the growth or proliferation of the cell.
[0996] Concept 101. A method for treating or preventing a disease or condition mediated by target cells in a human or animal subject, the method comprising performing the method of any one of Concepts 97 to 100 to modify the target cells, wherein the contacting comprises administering the nucleic acid to the subject and wherein the modification treats or prevents the disease or condition.
[0997] Concept 102. The method of Concept 101, wherein the target cells are cells of a subject containing a nucleic acid defect and the modification corrects the defect.
[0998] Concept 103. The method of Concept 101 or 102, wherein the modification
[0999] a) adds a new function to the cell, optionally the modification upregulates or downregulates gene expression in the cell, or adds a new nucleotide sequence to the genome of the cell for expressing a protein encoded by the sequence; or
[1000] b) regulates the expression of a nucleotide sequence contained in a chromosome or episome (optionally a plasmid) of the cell; or
[1001] c) inhibiting the growth or proliferation of cells.
[1002] The cell can be a bacterial cell of any genus or species in Table 3. The method can modify cells of any genus or species in Table 3. The DNA or polynucleotide can be contained in cells of any genus or species in Table 3. The method can kill or inhibit the growth or proliferation of cells of any genus or species in Table 3. The cell can be a cell of a genus or species different from the genus or species in Table 3 respectively. For example, the cell is not an Escherichia coli cell or the cell is not a Klebsiella cell. The cell can be an archaeon, such as a methanogen. The method can kill or inhibit the growth or proliferation of archaea. The method can edit the genome of bacterial or archaeal cells. Any prokaryotic cell (such as a bacterial or archaeal cell) or fungal cell herein can be contained in a microbiome, such as the microbiome of a human, animal, plant or environment (such as soil or waterway). For example, the microbiome is contained in the gastrointestinal (GI) tract of a human or animal. For example, the microbiome is contained in the urinary system of a human or animal, such as contained in the bladder, kidney or urethra. For example, the microbiome is contained in the blood of a human or animal. For example, the microbiome is contained in the eyes, nose or ears of a human or animal. For example, the microbiome is contained in the hair of a human or animal. For example, the microbiome is contained in the skin of a human or animal. For example, the microbiome is contained in the leaves, roots, stems or seeds of a plant.
[1003] The disease or condition herein can be selected from any of the following:
[1004]
[1005]
[1006] Neurodegenerative or central nervous system diseases or conditions treated or prevented by the method
[1007] In one example, the neurodegenerative or central nervous system disease or condition is selected from Alzheimer's disease, senile psychosis, Down syndrome, Parkinson's disease, Creutzfeldt-Jakob disease, diabetic neuropathy, parkinsonism, Huntington's disease, Machado-Joseph disease, amyotrophic lateral sclerosis, diabetic neuropathy and Creutzfeldt-Jakob disease. For example, the disease is Alzheimer's disease. For example, the disease is parkinsonism.
[1008] In one instance, any method described herein is practiced on a human or animal subject for treating a central nervous system or neurodegenerative disease or condition, the method causing downregulation of Treg cells in the subject, thereby promoting systemic monocyte-derived macrophages and / or Treg cells to cross the choroid plexus into the brain of the subject, whereby the disease or condition (such as Alzheimer's disease) is treated, prevented or its progression is reduced. In one embodiment, the method causes an increase in IFN-γ in the central nervous system of the subject (such as in the brain and / or CSF). In one instance, the method restores nerve fibers and / or reduces the progression of nerve fiber damage. In one instance, the method restores myelin sheaths and / or reduces the progression of myelin sheath damage. In one instance, any method described herein treats or prevents the diseases or conditions disclosed in WO2015 / 136541, and / or the method can be used in combination with any method disclosed in WO2015 / 136541 (the disclosure of the document is incorporated herein by reference in its entirety, for example, to provide the disclosure of such methods, diseases, conditions and potential therapeutic agents, which potential therapeutic agents can be administered to a subject for achieving the treatment and / or prevention of central nervous system and neurodegenerative diseases and conditions, such as agents such as immune checkpoint inhibitors, such as anti-PD-1, anti-PD-L1, anti-TIM3 or other antibodies disclosed therein).
[1009] Cancer treated or prevented by the method
[1010] Cancers that can be treated include unvascularized or substantially unvascularized tumors, as well as vascularized tumors. The cancer can comprise non-solid tumors (such as blood cancers, such as leukemia and lymphoma) or can comprise solid tumors. Cancer types to be treated by the methods, proteins (including fusion proteins and fusion protein complexes), vectors, complexes and compositions described herein include, but are not limited to, carcinoma, blastoma and sarcoma, as well as certain leukemias or lymphoid malignancies, benign and malignant tumors and malignancies, such as sarcoma, carcinoma and melanoma. Also included are adult tumors / cancers and pediatric tumors / cancers.
[1011] Blood cancers are cancers of the blood or bone marrow. Examples of blood (or hematogenous) cancers include leukemia, including acute leukemia (such as acute lymphoblastic leukemia, acute myeloid leukemia, acute myelocytic leukemia and myeloblast, promyelocytic, myelomonocytic, monocytic and erythroleukemia), chronic leukemia (such as chronic myeloid (granulocytic) leukemia, chronic myelocytic leukemia and chronic lymphocytic leukemia), polycythemia vera, lymphoma, Hodgkin's disease, non-Hodgkin's lymphoma (indolent and high-grade forms), multiple myeloma, Waldenström's macroglobulinemia, heavy chain disease, myelodysplastic syndrome, hairy cell leukemia and myeloproliferation.
[1012] Solid tumors are abnormal masses of tissue that usually do not contain cysts or areas of fluid. Solid tumors can be benign or malignant. Different types of solid tumors are named for the cell types that form them (e.g., sarcomas, carcinomas, and lymphomas). Examples of solid tumors such as sarcomas and carcinomas include fibrosarcoma, myxosarcoma, liposarcoma, chondrosarcoma, osteosarcoma, and other sarcomas, synovioma, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, colon cancer, lymphoid malignancies, pancreatic cancer, breast cancer, lung cancer, ovarian cancer, prostate cancer, hepatocellular carcinoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, medullary thyroid carcinoma, papillary thyroid carcinoma, pheochromocytoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, liver cancer, cholangiocarcinoma, choriocarcinoma, Wilms' tumor, cervical cancer, testicular tumors, seminoma, bladder cancer, melanoma, and central nervous system tumors (e.g., gliomas (e.g., brainstem gliomas and mixed gliomas), glioblastoma (also known as glioblastoma multiforme), astrocytoma, central nervous system lymphoma, germ cell tumor, medulloblastoma, schwannoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, and brain metastases).
[1013] Autoimmune diseases treated or prevented by the method
[1014]
[1015]
[1016]
[1017] Inflammatory diseases treated or prevented by the method
[1018]
[1019] It should be understood that the specific embodiments described herein are shown by way of illustration and not as limitations of the invention. The main features of the invention can be used in various embodiments without departing from the scope of the invention. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific procedures described herein. Such equivalents are considered to be within the scope of the invention and are covered by the claims. All publications and patent applications mentioned in this specification represent the level of skill of those skilled in the art to which the invention pertains. All publications and patent applications, and all U.S. equivalent patent applications and patents, are hereby incorporated by reference into this application to the extent that each individual publication or patent application is specifically and individually indicated to be incorporated by reference. Reference is made to the publications mentioned herein and equivalent publications of the U.S. Patent and Trademark Office (USPTO) or WIPO, the disclosures of which are incorporated by reference herein to provide the disclosure that can be used in the invention and / or one or more features (such as carriers) that can be included in one or more of the claims herein.
[1020] When used in the claims and / or the specification in conjunction with the term “comprising,” the use of the word “a” or “an” can mean “one,” but it also is in accordance with the meaning of “one or more,” “at least one,” and “one or more than one.” The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer only to alternatives or the alternatives are mutually exclusive, although the present disclosure supports definitions that refer only to alternatives and “and / or.” Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the device, the method for determining the value, or the variation that exists among the subjects being studied.
[1021] As used in this specification and the claims, the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[1022] As used herein, the term "or combinations thereof" or similar refers to all permutations and combinations of the items listed before the term. For example, "A, B, C, or combinations thereof" is intended to include at least one of the following: A, B, C, AB, AC, BC, or ABC, and if order is important in a particular context, also BA, CA, CB, CBA, BCA, ACB, BAC, or CAB. Continuing with this example, combinations that contain repeats of one or more items or terms are explicitly included, such as BB, AAA, MB, BBC, AAABC CCC, CBBAAA, CABABB, and so forth. Those skilled in the art will understand that, unless otherwise apparent from the context, there is typically no limitation on the number of items or terms in any combination.
[1023] Unless otherwise apparent from the context, any part of the present disclosure may be read in combination with any other part of the present disclosure.
[1024] In view of the present disclosure, all of the compositions and / or methods disclosed and claimed herein can be made and implemented without undue experimentation. While the compositions and methods of the present invention have been described from the perspective of preferred embodiments, it will be apparent to those skilled in the art that changes can be applied to the compositions and / or methods described herein and to the steps or the order of steps of the methods without departing from the concept, spirit, and scope of the present invention. All such similar substitutions and modifications that are apparent to those skilled in the art are considered to be within the spirit, scope, and concept of the present invention as defined by the appended claims.
[1025] The present invention is described in more detail in the following non-limiting examples. Examples
[1026] Example 1
[1027] 1.1 Perform generalization
[1028] This example describes the in-silico identification and in vivo (in Escherichia coli) functional characterization of a previously uncharacterized CRISPR-Cas system whose targeting activity is independent of predicted RNA-guided Cas endonucleases. We confirmed that the system i) successfully targets plasmid DNA by an unknown mechanism, hindering plasmid-induced antibiotic resistance, and ii) effectively binds chromosomal DNA without inducing cell death. We named this system the "Type S CRISPR / Cas system".
[1029] 1.1.1 Objectives
[1030] Characterize CRISPR systems that are different from currently known functionally characterized systems.
[1031] 1.2 Materials and methods
[1032] 1.2.1 Plasmid and strain construction
[1033] The plasmids and strains used in this study are listed in the appendix (see Section 1.5 below). Plasmids were constructed by InFusion HD TM cloning (Takara) from DNA fragments generated by PCR. Using λ-Red-mediated recombineering, the target sequence was inserted into the chromosome of Escherichia coli strain C-1. The resulting strain was b5451 (ΔlacZ).
[1034] 1.2.2 Transformation assay for sequence-specific plasmid inhibition
[1035] First, cells were transformed with the p1624 plasmid shown in Figure 1 and the plasmid shown in Figure 3 that carry components of the novel CRISPR / Cas system (which we refer to herein as the type S CRISPR / Cas system). Next, to evaluate the function of the plasmids carrying components of the type S system, Escherichia coli cells were transformed with these plasmids individually and then with a 1:1 mixture of two plasmids, p1631 and p892. The target plasmid p1631 carries all the native target sequences (5 spacers) of the CRISPR array of the system and the amilCP gene that provides purple color to the colonies carrying the plasmid. Plasmid p892 serves as a non-target control, which has the same origin of replication (p15A) and antibiotic marker gene (CmR) as the target plasmid but does not carry the target sequence and the amil CP gene.
[1036] Cells not carrying any components of the type S system were also transformed with the p1631 / p892 plasmid mixture as a control for evaluating the ratio of the two plasmids in the mixture.
[1037] The transformed cells were plated on LB agar plates containing the appropriate antibiotics to select for the presence of the plasmids. The plates were incubated overnight at 37 °C, and the ratio of purple and white colonies carrying the target and non-target plasmids, respectively, was determined.
[1038] 1.2.3 Plasmid inhibition using controlled expression of S-type system components
[1039] Strain b5456 carrying plasmid p1826 was transformed by electroporation with a 1:1 mixture of two plasmids, p1631 and p892, and the plasmid p1826 is equivalent to p1793 ( Figure 3)However, it contains the S system under the control of an arabinose-inducible promoter (instead of pBolA). The cells were allowed to recover by standing for 1.5 hours and plated on LB agar containing the appropriate antibiotic with or without 0.2% arabinose.
[1040] 1.2.4 Growth inhibition assay
[1041] A 1:1 mixture of competent cells of Escherichia coli C-1 and b5451 was transformed with plasmid p1826 containing the required components of the S-type CRISPR-Cas system and allowed to recover by standing for 1.5 hours, then plated on LB agar containing 20 μg / mL tetracycline to select for plasmid-carrying transformants. The plates also contained X-gal (20 ng / mL) and IPTG (1 mM) to distinguish colonies of Escherichia coli C-1 (lacZ+, blue) and b5451 (lacZ-, white).
[1042] 1.2.5 Plasmid clearance assay
[1043] Strain b5408 was transformed separately with plasmids p1935, p1936, p1937, p1938, p1631, and p144 (see Table 2 and Figure 7 ). Serial dilutions of the recovered transformation solution were plated on LB agar plates supplemented with 10 μg / mL tetracycline and 20 μg / mL chloramphenicol. The plates were incubated overnight at 37 °C and the resulting colony-forming units were counted to calculate the effect of the S system on the transformation efficiency of strain b5408.
[1044] 1.2.6 Chromosome binding assay
[1045] Strain b5872 was transformed with plasmids p1985, p1993, p1994, p1995, and p1996 (see Table 2). The transformation mixture was inoculated and plated on LB agar plates supplemented with 10 ng / mL tetracycline and incubated overnight at 37 °C. Three randomly selected single colonies / transformations were used to inoculate LB cultures supplemented with 10 ng / mL tetracycline. The cultures were allowed to reach late stationary phase after incubation at 37 °C for 20 hours with vigorous shaking at 220 rpm. 1 μL from each of the various cultures was added to 99 μL of LB medium supplemented with 10 ng / mL tetracycline and 0 μg / mL, 20 μg / mL, 25 μg / mL, 30 μg / mL, or 35 μg / mL chloramphenicol. The cultures were transferred to the wells of a microtiter plate and incubated in a plate reader at 37 °C for 24 hours with vigorous shaking (900 rpm). The optical density of the cultures was monitored throughout the incubation period, with readings taken every 10 minutes at 600 nm.
[1046] 1.3 Results
[1047] 1.3.1 S-type CRISPR / Cas system
[1048] We cloned the DNA encoding the RNA-guided endonuclease (rge), two open reading frames (ORFs) downstream of the rge gene, the homologous CRISPR array, and four ORFs immediately upstream (5’) of the CRISPR array into a low-copy plasmid ( Figure 1 ). Separately, ORF 140 and 423 were predicted to encode error-prone DNA polymerases, while ORF 624 and 188 were predicted to encode a helicase and an RNA endonuclease. The resulting p1624 plasmid placed the inserted genes under the transcriptional control of a constitutive promoter. Meanwhile, the CRISPR array carried five spacers. The corresponding protospacers were preceded by the 5’-AAG-3’ protospacer adjacent motif (PAM) of the type S CRISPR-Cas system. We constructed a target plasmid (p1631) that carried four predicted protospacers (corresponding to spacer #1, #2, #4, and #5), flanked at their 5’ ends by the 5’-AAG-3’ PAM ( Figure 1 ).
[1049] We then tested the in vivo function of the system on p1624. The p1624 and p1631 plasmids were used for the sequence-specific plasmid inhibition assay described in Section 1.2.2 of ‘Materials and Methods’. Briefly, cells carrying plasmid p1624 were transformed with a 1:1 mixture of the p1631 target plasmid and a control (non-target) plasmid. p1631 expression promoted the screening of the purple protein. Both the target plasmid and the non-target plasmid were transformed with similar efficiencies, confirming that the current setup of the system did not induce plasmid targeting. Therefore, it was hypothesized that the system required additional subunits to be active. We extended our study by cloning six additional ORFs (ORF 83, 135, 136, 162, 145, and 286, see Figure 3 ) upstream (5’) of ORF 188 in p1624. All the additional ORFs were transcribed in the same orientation as the previously cloned ORFs and the CRISPR array. We repeated the plasmid inhibition assay using the resulting p1760 plasmid instead of p1624. p1760 inhibited the growth of colonies transformed with the p1631 target plasmid but allowed the growth of colonies transformed with the non-target plasmid, as all colonies appeared white ( Figure 2 ). The results indicated that p1760 carried an active CRISPR-Cas system.
[1050] 1.3.2 Deletion analysis for identifying the minimal functional system
[1051] Our goal was to identify the minimal number of elements required for sequence-specific plasmid inhibition / targeting by the CRISPR-Cas system. We created plasmid libraries, performed a series of gene deletions of p1760 ( Figure 3 ), and we repeated the plasmid inhibition assay (Materials and Methods, Section 1.2.3).
[1052] Eleven out of the 12 new constructs we tested showed target plasmid inhibition ( Figure 4 ). Plasmid inhibition was not observed in the absence of the CRISPR array, as confirmed by the assay using p1899 ( Figure 4 , bottom row, right panel). Plasmid inhibition was observed for the assay using p1793 ( Figure 4 , bottom row, third panel), which contains an operon of only five genes encoding Cas that we named Cas-S1, Cas-S2, Cas-S3, Cas-S4, and Cas-S5.
[1053] Each of these genes was individually deleted from p1793, generating plasmids p2009 (ΔCas-S1), p2011 (ΔCas-S3), p2012 (ΔCas-S4), and p2013 (ΔCas-S5) (see Table 2). The plasmids were transformed into Escherichia coli NEB10b cells, generating strains b5700 (ΔCas-S1), b5702 (ΔCas-S3), b5703 (ΔCas-S4), and b5704 (ΔCas-S5) (see Table 1). Each of these strains, together with the negative and positive targeting control strains b4816 and b5408, was transformed with the target plasmid p1631 and plated on agar plates with appropriate antibiotic selection. The transformation efficiency of b5408 decreased by approximately 3 orders of magnitude, while the transformation efficiency of all other strains did not decrease (see Figure 5 ). These results confirm that the S-type CRISPR-Cas system requires all 5 subunits for plasmid inhibition. Notably, there is no predicted RNA-guided DNA endonuclease gene among these five genes, suggesting that the mechanism of action of the CRISPR-Cas system probably does not rely on DNA cleavage.
[1054] To further confirm the function of the S-type CRISPR-Cas system, we cloned the DNA fragment containing the minimal system and the CRISPR array downstream of an arabinose-inducible promoter. In the absence of arabinose, the resulting plasmid (p1826) could be maintained in the same cells together with the p1631 target plasmid ( Figure 6, left small panel). In the presence of arabinose, expression of the CRISPR-Cas system causes repression of p1631 and, surprisingly, repression of bacterial growth ( Figure 6 , right small panel).
[1055] 1.3.3 Spacers in the S-type CRISPR array exhibit similar targeting efficiencies
[1056] We evaluated the targeting efficiency of each spacer in the CRISPR array used in the type S CRISPR-Cas system in these experiments. The array contains five spacers, and the p1631 target plasmid contains protospacers for 4 out of these 5 spacers. Using p1631 as our base, we constructed four additional target plasmids, namely p1935, p1936, p1937, and p1938, each of which contains only one protospacer (see Table 2 and Figure 7 , left small panel). We transformed all 5 target plasmids and the non-target p144 plasmid (see Table 2) individually in strain b5408, and plated serial dilutions of the transformation mixtures on appropriate selection plates ( Figure 7 , right small panel). The next day, colony-forming units were counted for analysis; all 4 tested spacers showed similar targeting efficiency, reducing the transformation efficiency of the corresponding target plasmids by 3 orders of magnitude relative to the non-target plasmid. In addition, the combined targeting efficiency of all 4 spacers was not significantly higher than the efficiency of each individual spacer.
[1057] 1.3.4 S-type CRISPR-Cas system binds to chromosomal DNA
[1058] We determined that the type S CRISPR-Cas system does not contain a predicted RNA-guided DNA endonuclease. Therefore, we decided to examine whether the subunits of the system contain a novel dsDNA cleavage domain that can efficiently introduce lethal dsDNA breaks into the Escherichia coli chromosome. By moving 4 previously tested protospacers from p1631 to the chromosome of Escherichia coli C-1 to inactivate the lacZ gene, we constructed strain b5451. The resulting strain can be distinguished from the original non-target Escherichia coli C-1 strain on X-gal plates because the latter strain metabolizes X-gal and converts it into a blue dye. Both the constitutive and arabinose-inducible type S CRISPR-Cas expression systems were tested as described in Section 1.2.6 of 'Materials and Methods'. Briefly, a 1:1 mixture of Escherichia coli C-1 and b5451 competent cells was transformed with the p1826 plasmid, which contains the type S CRISPR-Cas system under the control of an arabinose-inducible promoter. Regardless of the induction conditions, the ratio of white to blue colonies remained 1:1 under appropriate antibiotic selection conditions ( Figure 8) This result further supports the probable absence of a dsDNA endonuclease domain in the system.
[1059] Subsequently, we wanted to know whether the type S CRISPR-Cas system has DNA-binding function or its mode of action is at the RNA level. We constructed strain bSNP5810 by introducing the crm gene into the intergenic region of Escherichia coli C-1. We then constructed a negative control strain designated p1985 and four crm-targeting plasmids designated p1993, p1994, p1995, and p1996 (see Table 2). For this purpose, we replaced the spacer within the CRISPR array of p1793 with a non-targeting spacer or with a single spacer targeting four different positions in the promoter or coding region of the crm gene ( Figure 9 ). We transformed strain bSNP5810 with each of p1993, p1994, p1995, and p1996, and cultured the cells in media supplemented with different chloramphenicol concentrations. In the absence of chloramphenicol, the growth rates of all strains with crm-targeting spacers were similar to that of the control strain ( Figure 10 , the first figure). Nevertheless, increasing chloramphenicol concentrations decreased the growth rates of all targeting strains in a proportional manner ( Figure 10 , the second - fourth figures). The growth rate of the control strain remained unaffected by the presence of chloramphenicol in the medium. These results confirm that the type S CRISPR-Cas system binds stably to the chromosome and acts as a surprisingly efficient transcriptional inhibitor. Notably, the orientation of the crm-targeting spacer does not affect the binding efficiency of the type S CRISPR-Cas system and the subsequent downregulation of crm expression, which is not the case for other previously known CRISPR-Cas systems that must target the coding or non-coding regions of genes (system-dependent features) to efficiently downregulate their transcription (see Qi et al., Cell, 152(5), 1173 - 1183, 2013, doi:https: / / doi.org / 10.1016 / j.cell.2013.02.022, Rath et al., NAR, 43(1), 237 - 246, 2015, doi:https: / / doi.org / 10.1093 / nar / gku1257 and Luo et al., NAR, 43(1), 674 - 681, 2015, doi : https: / / doi.org / 10.1093 / nar / gku971).
[1060] 1.4 Conclusions
[1061] Characterizes a nuclease-deficient CRISPR-Cas system (which we refer to as the type S CRISPR / Cas) and determines its minimal functional elements. Our study reveals that the type S CRISPR / Cas contains five Cas proteins (which we call Cas-S1 to S5) and a homologous CRISPR array. None of the Cas proteins show DNA nuclease activity. The system surprisingly inhibits the growth of bacterial cells containing at least one target sequence adjacent to the AAG PAM. In addition, the system can be programmed to efficiently bind chromosomal targets and act as a transcriptional inhibitor without inducing lethal dsDNA breaks. With these findings, we have identified and characterized a functional, novel CRISPR / Cas system that can be used in its minimal form to inhibit plasmids, reduce bacterial cell growth, and inhibit transcription from the chromosome. We can see the application of individual Cas of the system, for example, by creating fusions of type S Cas with nuclease domains to generate targeted cleavage of DNA. Other fusion partners include base editors, prime editors, deaminases, and reverse transcriptases.
[1062] 1.5 Appendices
[1063] Table 1: Bacterial strains used in this study
[1064]
[1065] Table 2: Bacterial plasmids used in this study
[1066]
[1067]
[1068]
[1069] Example 2: PAM identification and characterization for the S-type CRISPR / Cas system
[1070] 2.1 Background
[1071] SNIPR Biome has identified and characterized the type S CRISPR / Cas system as described in Example 1.
[1072] An in silico analysis of public genomic databases was performed with the goal of identifying potential PAM sequences for the Type S CRISPR / Cas system of Example 1. This analysis was based on spacers from wild-type Type S CRISPR arrays. Several potential protospacers were found, all of which contained the 5'-AAG-3' sequence motif at their 5' ends but no common sequence motif at their 3' ends. The hypothesis was that this 5'-AAG-3' motif at the 5' end of the protospacer could act as an effective PAM for the Type S system. This hypothesis was confirmed using plasmid targeting / killing assays (as described in Examples 1.3.1 - 1.3.3). Nevertheless, the PAM sequences of most CRISPR-Cas systems are degenerate, and thus it was decided to further experimentally explore the diversity of sequences recognized as effective PAMs by the Type S CRISPR / Cas.
[1073] 2.2 Experimental design
[1074] The PAM identification process for the Type S CRISPR / Cas system was based on the following plasmids:
[1075] 1. PAM plasmid library (p2259); each member of this library encompasses the protospacer sequence present in the target plasmid p1935 (see Example 1.2.5) as well as a mixture of candidate 5-nucleotide-long PAMs at the 5' end of the protospacer.
[1076] 2. Type S targeting plasmid (p1793); this plasmid expresses the Type S CRISPR / Cas system (Cas-S1 to Cas-S5) as well as a spacer complementary to the protospacer of the PAM plasmid library.
[1077] 3. Non-targeting (control) Type S plasmid (p1985); this plasmid expresses the Type S CRISPR / Cas system (Cas-S1 to Cas-S5) as well as a spacer that is not complementary to the protospacer of the PAM plasmid library. Additionally, the spacer is not complementary to any part of the E. coli genome.
[1078] The steps of the PAM identification assay were as follows:
[1079] 1. Construct the PAM plasmid library designated p2259 (see the section below).
[1080] 2. Construct a Type S targeting E. coli NEBTurbo strain (b5408) by transforming the NEB Turbo b3127 strain (NEB) with the p1793 (Type S targeting) plasmid.
[1081] 3. The non-targeting E. coli NEB Turbo strain (b6259) was constructed by transforming the NEB Turbo b3127 strain (NEB) with the p1985 (control S-type) plasmid.
[1082] 4. The b5408 (S-type targeting) and b6259 (control S-type) were transformed with the p2259 plasmid library and cultured under appropriate antibiotic selection conditions.
[1083] 5. The p2259 plasmid library was isolated from both cultures and NGS fragments containing the PAM-protospacer were prepared.
[1084] 5. NGS and analysis of the sequencing results.
[1085] 2.2.1 Step 1: Construct a PAM library
[1086] The PAM library plasmid p2259 was designed to contain a 5-nucleotide-long variable PAM region at the 5' end of the protospacer. Thus, the library should contain 4 5 (1024) different PAMs.
[1087] The construction of the PAM library plasmid p2259 was based on the previously described target plasmid p1935 (see Example 1.2.5). The p1935 plasmid was used as the PCR amplification template, along with appropriately designed primers. CloneAMP PCR Premix (Takara) was used for the PCR reaction. The forward primer was designed to introduce a 5-nucleotide-long degenerate sequence at the 5’ end of the protospacer present on the p1935 plasmid. The gel-purified PCR product was used for In-Fusion cloning, using the 5× In-Fusion HD Enzyme Premix (Takara). 5 μL from the In-Fusion cloning reaction was transformed into 50 μL of NEB10b electrocompetent cells (NEB) by electroporation. The cells were then left in SOC medium (500 μL final volume) and recovered for 1 hour at 37 °C with shaking at 200 rpm. The recovered cells were transferred into 500 mL of LB liquid medium supplemented with 20 μg / mL chloramphenicol. 100 μL was taken from the 500 mL culture and plated onto an LB agar plate supplemented with 20 μg / mL chloramphenicol. The culture and the plate were incubated overnight at 37 °C with vigorous shaking (200 rpm). The next day, 536 colonies were counted on the plate, and it was then calculated that 2,680,000 successfully transformed cells were added to the 500 mL culture. For this calculation, Equation 1 was used.
[1088] Equation 1:
[1089] Number of transformed cells in the culture = culture volume / volume of the plated culture × number of colonies on the selection plate
[1090] According to Equation 2, the number of transformed cells in this culture corresponds to the coverage of 2,617 × PAM library members.
[1091] Equation 2:
[1092] Plasmid library coverage = Number of transformed cells in the culture / Number of different plasmid library members
[1093] After preparing glycerol stocks from 500 mL of the culture, the constructed p2259 plasmid library was isolated from the remaining culture using the Qiagen Plasmid Midi kit.
[1094] 2.2.2 Step 2: Transform the b5408 (CRIPSR-S targeted) and b6259 (control S-type) strains with the p2259 plasmid library
[1095] Some members of the p2259 plasmid library are expected to contain PAMs recognized by the type S CRISPR / Cas system. These PAM-positive plasmids will be successfully targeted, and their replication will be hindered by the type S system. Therefore, b5408 (type S-targeted) cells transformed with PAM-positive members of the p2259 plasmid library should not survive in medium supplemented with chloramphenicol because the p2259 plasmid library carries a chloramphenicol resistance marker gene. Meanwhile, b5408 (type S-targeted) cells transformed with PAM-negative members of the p2259 plasmid library should survive in medium supplemented with chloramphenicol. In contrast, none of the p2259 plasmids in the library will be targeted by the type S CRISPR / Cas system expressed in the b6259 (control type S) strain and will survive in medium supplemented with chloramphenicol.
[1096] The p2259 plasmid library was electroporated into b5408 (type S-targeted) and b6259 (control type S) electrocompetent cells, respectively. Electrocompetent cells from both strains were prepared using the protocol described at (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Five transformation reactions / strain were performed using 40 ng of the p2259 plasmid library / reaction. The cells were allowed to recover for 1 h at 37 °C and with shaking at 200 rpm in 500 μL of SOC medium. The recoveries from the various strains were pooled and the last two recoveries (one / strain) were each used to inoculate 500 mL of LB liquid medium cultures supplemented with 20 μg / mL chloramphenicol and 10 μg / mL tetracycline. The cultures were incubated overnight at 37 °C with vigorous shaking at 200 rpm. The next day, glycerol stocks were prepared from the various cultures and the p2259 plasmid library was isolated from the various strains using the Qiagen Plasmid Midi kit.
[1097] Plasmid isolation from cultures of b5408 (S-type targeted) transformed with p2259, followed by NGS, will show which PAM sequences are not recognized by the S-type CRISPR / Cas system and subsequently which are the PAM sequences required by the system. Plasmid isolation from cultures of b6259 (control S-type) transformed with p2259, followed by NGS, will indicate the abundances of the various plasmid members in the p2259 library. This information will be used for the normalization of the NGS results from strain b5408 (S-type targeted).
[1098] 2.2.3 Step 3: PAM identification via NGS
[1099] The p2259 isolated from b6259 (control S-type) should contain all members of the p2259 PAM library, providing quantitative information on the abundances of the various PAM plasmids in the library.
[1100] Subject the p2259 plasmids isolated from b5408 (S-type targeted) and b6259 (control S-type) to NGS. In a 50 μL KAPA HiFi HotStart ReadyMix reaction (Roche) (25 cycles, 30 s denaturation, 20 s annealing, 15 s extension), use appropriately designed primers to perform the first NGS PCR amplification step for both p2259 preparations. Purify the reaction products with AMPure XP reagent (Beckman Coulter). Subsequently, in a second NGS PCR amplification step, add the Illumina index and flow cell adapter sequences to the purified products. Use appropriately designed primers to subject the first NGS PCR product from strain b5408 (S-type targeted) to the second NGS PCR amplification step. Use appropriately designed primers to subject the first NGS PCR product from strain b6259 (control CRISPR-S) to the second NGS PCR amplification step. Both amplifications are 50 μL KAPA HiFi HotStart ReadyMix reactions (Roche) (10 cycles, 20 ng of template DNA, 30 s denaturation, 20 s annealing, 15 s extension).
[1101] Purify both reaction products with AMPure XP reagent (Beckman Coulter). Use Illumina with single-end 150 cycles Sequence the purified products.
[1102] 2.3 Results
[1103] NGS data were used to compare the relative abundances of the members of the p2259 library in the b5408 (S-type targeting) and b6259 (control) strains. Data analysis revealed a clear preference of the S-type CRISPR / Cas system for the following PAM sequences ( Figure 11 and 12 ):
[1104] 5’-AHN-3’
[1105] 5’-KAG-3’
[1106] 5’-AGG-3’
[1107] 5’-GAC-3’
[1108] 5’-GTG-3’
[1109] Example 3: Protospacer mismatch tolerance in the S-type CRISPR / Cas system
[1110] 3.1 Background
[1111] The CRISPR / Cas system poses various spacer-protospacer complementarity requirements for DNA / RNA binding and targeting. For example, the Escherichia coli type I-E CRISPR / Cas system tolerates spacer-protospacer mismatches outside the'seed' region without compromising the function of the type I-E CRISPR / Cas system as an immune system (see Semenova et al., PNAS, 108(25), 10098-10103, 2011, doi:https: / / doi.org / 10.1073 / pnas.110414410), while Streptococcus pyogenes Cas9 requires 15-nucleotide spacer-protospacer complementarity for in vitro DNA cleavage (see Jinek et al., Science, 337(6096), 816-821, 2012, doi:10.1126 / science.12258), but tolerates multiple (up to 15 nucleotides) spacer-protospacer mismatches for DNA binding in bacteria (see Ciu et al., Nature Comm., 9(1912), 2018, doi:https: / / doi.org / 10.1038 / s41467-018-04209-5).
[1112] The type S CRISPR / Cas system does not contain a predicted dsDNA nuclease domain. It was previously shown that the plasmid curing activity of the type S CRISPR / Cas system requires the presence of the Cas-S3 helicase subunit (see Example 1.2.5). It is expected that spacer-protospacer mismatches negatively affect the DNA target binding stability of the pre-helicase type S CRISPR / Cas system CRISPR-Cas complex (a type S CRISPR-Cas complex without a helicase subunit, which contains the following subunits: Cas-S1, -S2, -S4, and -S5). A decrease in the DNA binding efficiency of the type S CRISPR / Cas system is expected to result in a decrease in the plasmid curing rate. Therefore, the assay described in this example was designed to use plasmid curing / transformation efficiency as a measurable output of the effect of spacer-protospacer mismatches on the targeting efficiency of the type S CRISPR / Cas system CRISPR-Cas.
[1113] 3.2 Experimental design
[1114] Plasmid p1935 was previously designed (see Example 1.2.5) and constructed to contain a PAM sequence (5'-AAG-3') targeted by the type S CRISPR / Cas system and a 32 bp protospacer. In this example, different variants of the p1935 plasmid were designed and constructed, which contained various combinations of mutations in the protospacer. These mutations partially disrupted the complementarity between the protospacer and the targeting spacer. The attached transformation efficiency graph (Figure 13) depicts these mutations in detail.
[1115] Using appropriately designed primers, by site-directed mutagenesis (NEB), various protospacer mutations were introduced into the protospacer sequence of the p1935 plasmid. The resulting PCR products were circularized using a KLD enzyme mixture (NEB) reaction. The KLD reaction mixture was transformed into NEB-10β electrocompetent cells (NEB). Subsequently, the cells were allowed to recover for 1 hour in SOC medium (500 μL final volume) with shaking at 37 °C and 200 rpm. The recovered cells were plated on LB agar plates supplemented with 20 μg / mL chloramphenicol for plasmid selection. The plates were incubated overnight at 37 °C. Single colonies from each plate were inoculated into 10 mL LB liquid medium cultures supplemented with 20 μg / mL chloramphenicol for plasmid selection. The cultures were incubated overnight at 37 °C with vigorous shaking (200 rpm). Samples of the various cultures were taken for the preparation of glycerol stocks, which were stored at -80 °C. The remaining culture contents were used for plasmid isolation using a Qiagen plasmid miniprep kit.
[1116] b5408 is an NEB Turbo Escherichia coli strain that contains plasmid p1793 (which contains a wild-type type S CRISPR / Cas system, Cas-S1 to Cas-S5, including the type S CRISPR array). 50 ng of each member of the p1935 plasmid library was electroporated into electrocompetent b5408 (wild-type type S) cells. Electrocompetent cells from both strains were prepared using the protocol described in (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). The cells were allowed to recover for 1 hour in 500 μL of SOC medium with shaking at 37 °C and 200 rpm. The recoveries from the various strains were used to prepare serial dilutions, which were then spotted onto LB agar plates supplemented with the appropriate antibiotics for plasmid selection, the appropriate antibiotics being 20 μg / mL chloramphenicol and 10 μg / mL tetracycline. Transformations were performed in biological triplicate. Colonies formed from each transformation were counted, and a decrease in the transformation efficiency of the b5408 (wild-type type S) strain, when compared to the transformation efficiency when the b5408 (wild-type type S) strain was transformed with a negative (non-target) control plasmid, indicated how the corresponding spacer-protospacer mismatches affected the targeting efficiency of the type S CRISPR / Cas system (see Figure 13).
[1117] Control:
[1118] ● The original plasmid p1935 is a positive control plasmid that always generates very few transformants (only escape mutants).
[1119] ● Plasmid p2361 is a negative control plasmid. It is identical to p1935 except for a 5’-AAG-3’ PAM sequence upstream (5’) of the protospacer, which is exchanged with a 5’-CGG-3' sequence that is not predicted to be a type S CRISPR / Cas PAM (see Example 2). Thus, the transformation efficiency of p2361 (-ve control) should be the highest among all library members mentioned in this example.
[1120] 3.3 Results
[1121] Figure 13 shows the analysis of transformation efficiency regarding various mismatches.
[1122] Regarding how spacer-protospacer mismatches affect the plasmid clearance efficiency of the type S CRISPR / Cas system, careful examination of the figure leads to the following conclusions:
[1123] 1. Five consecutive PAM-distal mismatches are well tolerated.
[1124] 2. Eight consecutive PAM distal mismatches completely abolish targeting.
[1125] 3. A single mismatch at the 1st or 2nd PAM proximal position reduces targeting by three orders of magnitude.
[1126] 4. A single mismatch at the 3rd to 10th PAM proximal positions reduces targeting by 0 to 1 order of magnitude.
[1127] 5. Double mismatches at the 1st to 3rd PAM proximal positions reduce targeting by 3 to 4 orders of magnitude.
[1128] 6. Double mismatches at the 4th to 10th PAM proximal positions reduce targeting by 0 to 1 order of magnitude.
[1129] 7. Triple mismatches at the 1st to 5th PAM proximal positions reduce targeting by 3 to 4 orders of magnitude.
[1130] 8. Triple mismatches at the 4th to 9th PAM proximal positions reduce targeting by 0 to 1 order of magnitude.
[1131] 9. Triple mismatches at the 8th to 10th PAM proximal positions reduce targeting by 4 orders of magnitude.
[1132] 10. Quadruple mismatches at the 1st to 8th PAM proximal positions reduce targeting by 3 to 4 orders of magnitude.
[1133] 11. Quadruple mismatches at the 6th to 9th PAM proximal positions reduce targeting by 1 order of magnitude.
[1134] 12. Quadruple mismatches at the 7th to 10th PAM proximal positions reduce targeting by 5 orders of magnitude.
[1135] Example 4: Type S CRISPR / Cas System for CRISPRi Applications
[1136] 4.1 Purpose
[1137] This study evaluated i) the efficiency of the S-type CRISPR / Cas system acting as a CRISPRi tool for binding to Escherichia coli chromosomal and plasmid targets, and ii) which subunits are essential for DNA binding.
[1138] 4.2 Introduction
[1139] We previously demonstrated that the type S CRISPR / Cas system drives crRNA-dependent inhibition of plasmid replication (see Examples 1.3.2 and 1.3.3), while lacking a DNA nuclease domain. The crRNA-guided DNA-binding activity of the type S CRISPR / Cas system can be used to develop a CRISPR interference (CRISPRi) tool in Escherichia coli. We anticipate that it will also be able to be used in other cell types. In this example, the CRISPRi potential of the type S CRISPR / Cas system was evaluated via targeted downregulation of the transcription and subsequent translation of the reporter green fluorescent protein (gfp) gene. The targeted gfp gene was chromosomally integrated or plasmid-borne. Additionally, this example reveals which type S CRISPR / Cas system subunits are essential for the DNA-binding activity of the type S CRISPR / Cas system.
[1140] 4.3 Methods and Results
[1141] The CRISPRi potential of five variants of the type S CRISPR / Cas system was tested on them;
[1142] i) Wild-type (wt) type S CRISPR / Cas system
[1143] ii) Cas-S1 deletion mutant (ΔCas-S1) type S CRISPR / Cas system
[1144] iii) Cas-S3 deletion mutant (ΔCas-S3) type S CRISPR / Cas system
[1145] iv) Cas-S4 deletion mutant (ΔCas-S4) type S CRISPR / Cas system
[1146] v) Cas-S1 / Cas-S3 double deletion mutant (ΔCas-S1ΔCas-S3) type S CRISPR / Cas system.
[1147] All variants were placed under the transcriptional control of the pBolA promoter (SEQ ID No: 14) and cloned together with a non-targeting crRNA expression module to generate plasmids p1985 (wild-type), p2062 (ΔCas-S1), p2064 (ΔCas-S3), p2065 (ΔCas-S4), and p2160 (ΔCas-S1ΔCas-S3) ( Figure 14 ). The 5' and 3' ends of the spacer region of the crRNA module contain recognition sites for the BsmBI-v2 (NEB) restriction enzyme, which facilitates the cloning of the targeted spacer ( Figure 14)。All plasmids contain the repA101 replication protein gene, the SC101 low-copy replication origin, and the tetR tetracycline resistance marker gene( Figure 14 )。
[1148] 4.4 Chromosomal CRISPRi
[1149] The strains MG1655 and b5815 (gfp+ve strain) were used for chromosomal type I-F CRISPR / Cas system CRISPRi assays. b5815 is an Escherichia coli MG1655 derivative that was constructed to contain the gfp gene under the control of the p70a promoter (SEQ ID No:15) and is located between the ybcB and ybcC genes( Figure 15 )。Sixteen protospacers were selected and targeted within or near the gfp gene of the b5815 (gfp+ve) strain for chromosomal CRISPRi assays( Figure 15 , Table 4, Seq ID No:25 to 41). Eight protospacers were located on the gfp coding strand (i.e., protospacers U2kF, U1kF, S1F, S2F, S3F, S4F, S5F, S6F) and eight were located on the non-coding strand (i.e., protospacers U2kR, U1kR, S1R, S2R, S3R, S4R, S5R, S6R).
[1150] 4.4.1 Methods
[1151] Construction of plasmids expressing different type I-F CRISPR / Cas system variants and targeting the aforementioned protospacers was carried out by digesting the p1985 (wild type), p2062 (ΔCas-S1), p2064 (ΔCas-S3), p2065 (ΔCas-S4), and p2160 (ΔCas-S1ΔCas-S3) plasmids with the BsmBi-v2 (NEB) restriction enzyme, and then ligating the pre-annealed complementary ssDNA oligonucleotides encoding the respective spacers into the digested plasmids using T4 DNA ligase (NEB). 5 μL from each ligation reaction was transformed by electroporation into 50 μL of NEB 10β electrocompetent cells (NEB). The cells were then allowed to recover for 1 hour in SOC medium (500 μL final volume) at 37 °C with shaking at 200 rpm. 100 μL was taken from each recovery and plated onto LB agar plates supplemented with 20 μg / mL chloramphenicol. The plates were incubated overnight at 37 °C.
[1152] The next day, single colonies were randomly selected and used to inoculate 10 mL LB liquid culture supplemented with 10 μg / mL tetracycline. The cultures were incubated overnight at 37 °C with vigorous shaking (200 rpm).
[1153] The next day, glycerol stocks were prepared for each of the cultures, and the remaining cultures were used for plasmid isolation using the Qiagen plasmid miniprep kit. The various plasmids were sequence-verified by NGS. The various plasmids were transformed into electrocompetent b5815 (gfp+ve strain) cells. Electrocompetent cells were prepared using the protocol described in (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were carried out using 40 ng / reaction from the various plasmids. The cells were allowed to recover for 1 hour at 37 °C and 200 rpm shaking in 500 μL of SOC medium. 100 μL from each recovery was plated on LB agar plates supplemented with 10 μg / mL tetracycline. The plates were incubated overnight at 37 °C.
[1154] The next day, three individual colonies were randomly selected from each transformation and used to inoculate 5 mL of M9 minimal medium cultures supplemented with glycerol (0.4% w / v) and 10 μg / mL tetracycline. The cultures were incubated overnight at 37 °C with vigorous shaking (200 rpm).
[1155] The next day, 2 μL from each culture was transferred to 198 μL of M9 medium supplemented with glycerol (0.4% w / v) and 10 μg / mL tetracycline. The resulting 200 μL cultures were transferred to different wells in a Thermo Scientific TM Nunc microWell 96-well optical bottom plate. The plate was sealed with a Breathe- Seal membrane (Sigma-Aldrich) and loaded into a Synergy H1 microplate reader (Biotek) for incubation at 37 °C with continuous shaking. The optical density (absorbance at 600 nm) and green fluorescence emission (excitation at 485 nm, emission at 516 nm) of the cultures were measured at 10-minute intervals over a 24-hour period. The fluorescence emission data from each culture and each time point were normalized with the corresponding OD600 nm data. The normalized data from biological triplicate cultures were averaged, and the corresponding standard deviations were calculated. The final data were plotted in the graph presented in Figure 16.
[1156] 4.4.2 Results
[1157] When targeting the spacers (spacer S1F, S1R, S2F, S2R, S3F, S3R) at the start of the gfp orf and in the p70a promoter region ( Figure 16A) or at the end of the gfp orf and in the protospacers in the 200 bp regions upstream and downstream of gfp (protospacers S4F, S4R, S5F, S6F, S6R)( Figure 16B ) When this was the case, the S-type CRISPR / Cas system downregulated GFP expression from b5815 (gfp+ve strain) to the level of the non-fluorescent control strain (b230), confirming extremely efficient and robust CRISPRi performance. The only exception was the targeting of the protospacer S5R, which downregulated GFP expression by approximately 50%. When the S-type CRISPR / Cas system targeted protospacers located 1 kb upstream of the gfp orf (protospacers U1kF, U1kR), GFP expression was also downregulated, while for protospacers 2 kb upstream of the gfporf (U2kF, U2kR), GFP expression was completely unaffected( Figure 16C ).
[1158] When targeting the promoter p70a region or the sense strand in the gfp orf (protospacers S1F, S1R, S2F, S2R, S3F), the helicase-deficient ΔCas-S3 S-type CRISPR / Cas system widely or completely downregulated GFP expression from b5815 (gfp+ve strain), but when it targeted the antisense strand in the gfp orf (protospacer S3R), it did not affect GFP expression( Figure 17A ). It was confirmed here that Cas-S3 cannot unwind DNA and subsequently block the transcription of genes located 2 kb away from the targeted protospacer.
[1159] To confirm that the CRISPRi function of the S-type CRISPR / Cas system is Cas-S3-dependent when targeting protospacers at the very end or outside the orf of the targeted gene, the targeting assays for protospacers S4F, S4R, S5F, S5R, S6F and S6R were repeated using the ΔCas-S3 S-type CRISPR / Cas system( Figure 17B ) In fact, in the absence of Cas-S3, GFP expression from the selected target b5815 (gfp+ve strain) was not affected. These results indicate that Cas-S3 is not essential for the formation of the S-type CRISPR / Cas system RNP complex and its binding to the DNA target.
[1160] The ΔCas-S4 S-type CRISPR / Cas system did not affect GFP expression from b5815 (gfp+ve strain), indicating that the subunit is essential for the formation of the S-type CRISPR / Cas system RNP complex and / or its binding to the DNA target( Figure 18A ).
[1161] The ΔCas-S1 type S CRISPR / Cas system and the ΔCas-S1,ΔCas-S3 type S CRISPR / Cas system reduce GFP expression from b5815 (gfp+ve strain) in a similar manner and only when they (fully or partially) target the antisense strand of the p70a promoter( Figure 18B and 18C ). These results indicate that Cas-S1 is not essential for the formation of the type S CRISPR / Cas system RNP complex and its binding to the DNA target, although it is important for its stability and / or the efficiency of its DNA binding activity. In addition, these results indicate that Cas-S1 is crucial for loading Cas-S3 into the complex.
[1162] 4.5 Plasmid CRISPRi
[1163] Strain b6386 was used for plasmid type S CRISPR / Cas system CRISPRi assays. b6386 is an Escherichia coli MG1655 strain transformed with plasmid p2370, which contains the gfp gene (under the control of the p70a promoter (SEQ ID No:15)), the p15a origin of replication, and the cmR chloramphenicol resistance marker gene. Plasmid CRISPRi was also tested for only four of the five type S variants tested for chromosomal CRISPRi in Example 4.4. The wild-type type S CRISPR / Cas system containing Cas-S1-Cas-S5 was not tested due to its proven activity of inhibiting plasmid replication (see Examples 1.3.2 and 1.3.3). Six protospacers located within the p70a promoter region and the gfp gene in b6386 were selected and targeted for CRISPRi assays on the plasmid, the same as for the chromosomal CRISPRi assays; three protospacers were located on the gfp coding strand (i.e., protospacer S1F, S2F, S3F), and three were located on the non-coding strand (i.e., protospacer S1R, S2R, S3R) (see Figure 15 ).
[1164] 4.5.1 Methods
[1165] Construction of various type S expression plasmids targeting these protospacers was carried out as previously described in Example 4.4.1.
[1166] Transformation of the plasmids was carried out as described in Example 4.4.1, with the following two changes:
[1167] 1. 100 μL from each recovery was plated on LB agar plates supplemented with 20 μg / mL chloramphenicol instead of 10 μg / mL tetracycline
[1168] 2.2) The inoculum culture was supplemented with 20 μg / mL chloramphenicol instead of 10 μg / mL tetracycline.
[1169] 3. The final data were plotted in the graph presented in Figure 19.
[1170] 4.5.2 Results
[1171] ΔCas-S3 only fully or effectively (but not completely) downregulated GFP expression from b6386 (gfp +ve strain) when it targeted the protospacers (S1F and S1R) within the promoter p70a region, regardless of the orientation of the targeted strand ( Figure 19A ). None of the other tested variants downregulated GFP expression ( Figure 19B 、 19C 、19D). This was in complete contrast to the results obtained from chromosomal CRISPRi experiments.
[1172] Example 5: Type S CRISPR / Cas System in Base Editing Applications
[1173] 5.1 Introduction
[1174] Currently developed and reported CRISPR / Cas base editing tools rely on single-subunit effector modules of different type II CRISPR / Cas systems. Type I CRISPR / Cas systems have multi-subunit effector modules, which complicate the development of base editing tools. Therefore, the base editing capabilities of type I base editors, regarding, for example, editing specificity, efficiency, and window, remain unexplored.
[1175] In this example, the potential of the type S CRISPR / Cas system to be transformed into a flexible base editing platform was extensively studied. A collection of base editors was designed and constructed as follows: The PmCDA1 cytidine deaminase from sea lamprey (SEQ ID No: 18 and 19, see Nishida et al., Science, 353(6305), aaf8729, 2016, doi:10.1126 / science.aaf8729) was fused to the N-terminus or C-terminus of the Cas-S1, Cas-S3, or Cas-S4 subunit of the type S CRISPR / Cas system using a 16-amino acid long XTEN linker (SEQ ID No: 16 and 17, see Schellenberger et al., Nat. Biotechnol., 27(12), 1186-1190, 2009, doi:10.1038 / nbt.1588). The uracil DNA glycosylase inhibitor (UGI) protein (SEQ ID No: 20 and 21, see Mol et al., Cell, 82(5), 701-8, 1995, doi:10.1016 / 0092-8674(95)90467-0) was fused to the C-terminus of each chimera using a 10-amino acid long linker (SEQ ID No: 22 and 23). Subsequently, these base editors were tested for their ability to generate targeted chromosomal C to T modifications.
[1176] 5.2 Experimental Design
[1177] 5.2.1 Screening Strains and Plasmids Overview:
[1178] The gfp gene was inserted into the chromosome of Escherichia coli MG1655 under the transcriptional control of the p70a promoter (SEQ ID No: 15). The resulting b5815 (gfp+ve) strain was used as the test strain for all base editing assays. Six protospacer targets for the base editing assays were selected to have 5'-AAG-3' PAM at their 5' ends and were located within a 249-bp long region that contains the last 52 bp of the p70a promoter and the first 197 bp of the gfp sequence. Three protospacers were located on the template (with respect to transcription) strand and three on the non-template strand ( Figure 15 ).
[1179] Two classes of plasmids were designed and constructed for the base editing assays:
[1180] The first class contained 'base editor' plasmids responsible for expressing the pm_cda1 gene, which was fused to the N-terminus or C-terminus of the Cas-S1, Cas-S3, or Cas-S4 gene (see Figure 20, on the left - hand side). The expression of the chimeric body is placed under the transcriptional control of an arabinose - inducible promoter (p BAD , SEQ ID No:24, see Guzman et al., J. Bacteriol., 177(14), 4121–4130, 1995, doi:10.1128 / jb.177.14.4121 - 4130.1995). The backbone of the plasmid contains a chloramphenicol resistance marker gene (crm R ), for selection by means of the antibiotic chloramphenicol, and the p15a origin of replication.
[1181] All PCR reactions were carried out with Q5 High - Fidelity 2X Premix (Hi Fi 2X Master Mix) (NEB), with appropriately designed primers and PCR templates to construct the plasmids. The PCR fragments were used for High - Fidelity DNA Assembly Master Mix reactions. 5 μL from each assembly reaction was transformed by electroporation into 50 μL of NEB10β electrocompetent cells (NEB). The cells were then allowed to recover for 1 hour at 37 °C and 200 rpm shaking in SOC medium (500 μL final volume). 100 μL was taken from each recovery and plated on LB agar plates supplemented with 20 μg / mL chloramphenicol. The plates were incubated overnight at 37 °C.
[1182] The next day, single colonies were randomly selected and used to inoculate 10 mL of LB liquid medium cultures supplemented with 20 μg / mL chloramphenicol. The cultures were incubated overnight at 37 °C with vigorous shaking (200 rpm).
[1183] The next day, glycerol stock solutions were prepared for each culture, and the remaining cultures were used for plasmid isolation with the Qiagen plasmid miniprep kit. The various plasmids were sequence - verified by NGS.
[1184] The second category contains 'guide' plasmids, which are responsible for expressing i) the crRNA expression module (one for each of the non - targeting control, S1F, S1R, S2F, S2R, S3F, and S3R) and ii) different versions of the S - type Cas operon, in which the Cas - S1, Cas - S3, or Cas - S4 genes are deleted (see Figure 20 , the right - hand side plasmid). The expression of the crRNA is placed under the control of the native leader (promoter) sequence as it is found in the CRISPR locus of the S - type CRISPR / Cas system. The expression of the operon is placed under the constitutive promoter (p bolA) under the transcriptional control of. The backbone of the plasmid contains a tetracycline resistance marker gene (tet R ), for selection by means of the antibiotic tetracycline, and the pSC101 origin of replication.
[1185] The method for constructing the plasmid for 'guiding' is the same as that of the first category, except for the following: (i) 10 μg / mL tetracycline is used instead of 20 μg / mL chloramphenicol for plating on LB agar and (ii) 10 μg / mL tetracycline is used instead of 20 μg / mL chloramphenicol for inoculating LB liquid medium.
[1186] Subsequently, plasmids p2062 (non-targeting control, ΔCas-S1), p2064 (non-targeting control, ΔCas-S3), p2065 (non-targeting control, ΔCas-S4), and p2160 (non-targeting control, ΔCas-S1, ΔCas-S3) were subjected to restriction digestion with BsmBI-v2 (NEB), and the digestion products were gel-purified using the Zymoclean Gel DNA Recovery Kit (Zymoresearch). Appropriately designed oligonucleotides were annealed and ligated into the digested plasmids using T4 DNA ligase (NEB). 5 μL from each ligation reaction was transformed by electroporation into 50 μL of NEB10β electrocompetent cells (NEB). The cells were then allowed to recover for 1 hour in SOC medium (final volume 500 μL) with shaking at 37 °C and 200 rpm. 100 μL was taken from each recovery and plated on LB agar plates supplemented with 10 μg / mL tetracycline. The plates were incubated overnight at 37 °C.
[1187] The next day, a single colony was randomly selected and used to inoculate a 10 mL LB liquid medium culture supplemented with 10 μg / mL tetracycline. The culture was incubated overnight at 37 °C with vigorous shaking (200 rpm).
[1188] The next day, glycerol stocks were prepared for each culture, and the remaining cultures were used for plasmid isolation using the Qiagen plasmid miniprep kit. The various plasmids were sequence-verified by NGS.
[1189] 5.2.2 Base Editing Assay
[1190] Each of the 6 'base editor' plasmids was transformed into electrocompetent b5815 cells, and the plasmids have the PmCDA1 cytidine deaminase fused to the N-terminus or C-terminus of the wild-type S-type subunits Cas-S1, Cas-S3, and Cas-S4 respectively (see Figure 20, on the left - hand side). Electrocompetent cells were prepared using the protocol described in (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Transformation reactions were carried out using 40 ng of various plasmids per reaction. The cells were allowed to recover for 1 hour in 500 μL of SOC medium with shaking at 37 °C and 200 rpm. 100 μL from each recovery was plated on LB agar plates supplemented with 20 μg / mL chloramphenicol. The plates were incubated overnight at 37 °C.
[1191] The next day, single colonies were randomly selected and used to inoculate 10 mL of LB liquid medium cultures supplemented with 20 μg / mL chloramphenicol. The cultures were incubated overnight at 37 °C with vigorous shaking (200 rpm).
[1192] The next day, electrocompetent cells were prepared for each of the six resulting b5815 strains with base - editing plasmids according to the protocol described in (https: / / barricklab.org / twiki / bin / view / Lab / ProtocolsElectrocompetentCells). Subsequently, each of the six strains was electrotransformed with the corresponding set of guide plasmids, as presented on the Figure 20 right - hand side. These plasmids were responsible for expressing i) the S subunit not expressed by the paired base - editing plasmid, and ii) the targeting and non - targeting crRNAs. Transformation reactions were carried out using 40 ng per reaction from various plasmids. The cells were allowed to recover for 1 hour in 500 μL of SOC medium with shaking at 37 °C and 200 rpm. 100 μL from each recovery was plated on LB agar plates supplemented with 20 μg / mL chloramphenicol, 10 μg / mL tetracycline, and 0.2% v / w arabinose. The plates were incubated overnight at 37 °C.
[1193] The next day, single colonies were randomly selected and subjected to colony PCR using appropriately designed genomic - specific primers. The PCR products were sequenced by Sanger sequencing and the results were analyzed to detect C - to - T modifications or G - to - A modifications when targeting and modifying the antisense strand.
[1194] 5.3 Results
[1195] Representative results from the base - editing assays are presented in Figure 21.
[1196] A detailed analysis of the Sanger sequencing data from all base - editing assays is provided below:
[1197] 5.3.1 PmCDA1 Linked to the N-Terminus of the Following
[1198] ● Cas-S1 (wild-type S type): No base editing
[1199] ● Cas-S3 (wild-type S type): Single C to T mutation in only 1 target protospacer
[1200] ● Cas-S4 (wild-type S type): Single C to T mutation in only 1 target protospacer (same as for Cas-S3-N-terminal CDA)
[1201] 5.3.2 PmCDA1 Linked to the C-Terminus of the Following
[1202] ● Cas-S1:
[1203] ○ Protospacer S1F: No base editing
[1204] ○ Protospacer S1R: No base editing
[1205] ○ Protospacer S2F: Multiple positions base edited (105 nucleotide editing window), with different efficiency levels for each position. There are not only C to T conversions but also a few G to A conversions
[1206] ○ Protospacer S2R: 4 positions base edited (29 nucleotide editing window), with different efficiency levels for each position. There are not only G to A conversions but also 1 C to T conversion
[1207] ○ Protospacer S3F: 8 positions base edited (35 nucleotide editing window), with different efficiency levels for each position. There are not only C to T conversions but also 1 G to A conversion
[1208] ○ Protospacer S3R: A few positions base edited (outside the protospacer), with low efficiency. Streaking improves efficiency and expands the editing window (96 nucleotides). A mixture of C to T and G to A conversions was detected.
[1209] ● Cas-S1 (in the absence of Cas-S3: S type_ΔCas-S3):
[1210] ○ Protospacer S1F: No base editing
[1211] ○ Protospacer S1R: Single G to A modification within the protospacer
[1212] ○ Protospacer S2F: Multiple positions C to T modified within and 7 nucleotides downstream of the protospacer, with different efficiency levels. 2 additional C to T conversions were detected within a 102 nucleotide editing window upstream of the 5' end of the protospacer.
[1213] ○ Forward spacer S2R: 4 positions are base-edited (48-nucleotide editing window), with different efficiency levels for each position. Only 1 weak G-to-A modification at 3 PAM-proximal positions, and 3 effective C-to-T modifications downstream of the spacer
[1214] ○ Forward spacer S3F: Multiple very weak modifications
[1215] ○ Forward spacer S3R: A few very weak modifications are detected (a mixture of C-to-T and G-to-A conversions), and 1 very effective C-to-T modification at 40 nucleotides downstream of the 3'-end of the spacer.
[1216] ● Cas-S3:
[1217] ○ Forward spacer S1F: No base editing
[1218] ○ Forward spacer S1R: The dominant single position in the middle of the spacer is base-edited, and a few positions are ineffectively edited. Culturing with streaks expands the base-editing window (50 nucleotides), and while G-to-A conversions are detected, a few C-to-T conversions are also detected
[1219] ○ Forward spacer S2F: No base editing
[1220] ○ Forward spacer S2R: Multiple positions are ineffectively base-edited within a wide editing window (up to 208 nucleotides). Not only G-to-A conversions are detected, but mostly C-to-T conversions are also detected. Culturing with streaks improves the efficiency as it causes 3 clean conversions (both G-to-A and C-to-T) within a 178-nucleotide window.
[1221] ○ Forward spacer S3F: Multiple positions are base-edited (388-nucleotide editing window). The conversions within the spacer are the cleanest conversions. Only 2 positions with G-to-A conversions are detected.
[1222] ○ Forward spacer S3R: No base editing is initially detected, but culturing with streaks generates multiple G-to-A and C-to-T conversions within a 388-nucleotide editing window.
[1223] ● Cas-S4:
[1224] ○ All targeted forward spacers are edited at multiple positions within the spacer. Exceptions are S2R, which is also edited downstream of the spacer, and S3R, which has a single G-to-A mutation at the 2nd PAM-distal position of the spacer.
[1225] ○ Forward spacers S1F and S2R have a mixture of C-to-T and G-to-A modifications, while the remaining forward spacers have C-to-T or G-to-A modifications, depending on the direction of the spacer.
[1226] 5.4 Conclusions
[1227] This example generated the first reported class I CRISPR / Cas platform for efficient base editing. Seven developed base editors employed the PmCDA1 cytidine deaminase. Six base editors employed the wild-type type S CRISPR / Cas system, and one base editor employed the ΔCas-S3 type S CRISPR / Cas system.
[1228] The fusion of PmCDA1 to the N-terminus of Cas-S1 did not generate any detectable modification.
[1229] The fusion of PmCDA1 to the N-termini of Cas-S4 and Cas-S3 generated only a single detectable mixed (wild-type / mutant genotype) modification for one of the tested spacers. These systems can be used when very specific point mutations are needed. However, the low number of available spacers limits the applicability of these systems.
[1230] The fusion of PmCDA1 to the C-termini of Cas-S1, Cas-S4, and Cas-S3 generated three base editing systems with very different base editing outcomes (editing window, editing efficiency, and editing position) even for the same protospacer. In general, the Cas-S1- and Cas-S3-based base editors have very broad editing windows and may be very valuable in generating amino acid libraries of a single protein or introducing stop codons. In the absence of Cas-S3, the editing window of the Cas-S1-based base editor shrinks, and a shift in the position of the editing is noted. The Cas-S4-based base editor has a narrower base editing window and is the only base editor that introduces modifications throughout the targeted protospacer.
[1231] Example 6: Functionalizing the Type S CRISPR / Cas System for dsDNA Cleavage
[1232] 6.1 Objectives
[1233] Type S is a nuclease-deficient CRISPR-Cas system that does not carry a bioinformatics-predicted or experimentally determined (see Examples 1 and 4 herein) nuclease domain in any of its subunits. The goal of this study was to engineer a nuclease-deficient type S system (ΔCasS3, see Figure 22)'s N-terminal fusion of the CasS1 and CasS4 subunits to create synthetic crRNA-guided DNA nucleases to expand the repertoire of genome editing reagents. The developed chimeras were tested on chromosomal and plasmid-borne DNA targets, and a set of rules for efficient dsDNA cleavage was established.
[1234] 6.2 Introduction
[1235] I-TevI is a homing endonuclease containing the following: an N-terminal nuclease domain (1-92 amino acids (aa)), a zinc finger-containing spacer domain (92-169 aa), and a DNA binding domain (169-245 aa) (see Edgell et al., PNAS, 98(14), 7898-7903, 2001, doi:https: / / doi.org / 10.1073 / pnas.141222498 and Kleinstiver et al., PNAS, 109(21), 8061-8066, 2012, doi:https: / / doi.org / 10.1073 / pnas.111798410). I-TevI binds its DNA target as a monomer and introduces sequence-specific staggered dsDNA nicks in I-TevI cleavage via the GIY-YIG motif present in its N-terminal nuclease domain. The sequence of the naturally occurring I-TevI cleavage site is 5'-CAACG-3' (sense strand) / 5'-CGT TG-3' (antisense strand). Nevertheless, I-TevI has been reported to recognize a more flexible 5'-CNNNG-3' (sense strand) / 5'-CNNNG-3' (antisense strand) I-TevI cleavage site (see Edgell et al., 2001 and Kleinstiver et al., 2012, both supra). It is hypothesized that by fusing the N-terminal and zinc finger / linker domains of I-TevI to the N-terminal of CasS1 or CasS4, it is possible to create synthetic TevCasSx (x = 1 or 4) crRNA-guided DNA nucleases.
[1236] In Example 4, it was confirmed via CRISPRi experiments that the helicase-deficient ΔCasS3 S-type system retains strong DNA-binding activity. Additionally, in the same example, it was confirmed that the ΔCasS1 S-type system expressing helicase shows the same CRISPRi / DNA-binding efficiency as the helicase-deficient ΔCasS3 S-type system, while the ΔCasS1ΔCasS3 S-type system shows reduced CRISPRi / DNA-binding activity. These results support the model in which CasS1 is involved in recruiting CasS3 to the S-type complex. Therefore, it was decided to use the ΔCasS3 S-type helicase-deficient system for constructing the TevCasSx synthetic nuclease, with the goal of avoiding hindrance of the nuclease activity of the TevCasSx chimera by the recruitment or helicase activity of CasS3.
[1237] 6.3 Experimental Design
[1238] 6.3.1 TevCasS Plasmid-DNA Cleavage Assay
[1239] Three categories of plasmids were designed, constructed, and used for the TevCasS plasmid-DNA cleavage assay; i) the TevCasS target plasmid, ii) the TevCasSx expression plasmid, and iii) the S-type guide plasmid.
[1240] Twenty-three TevCasS target plasmids (named p2497–p2510, p2575–p2578, and p2655–p2659, respectively, see Figure 22, the right - hand plasmid) was designed to contain various TevCasS target sites (SEQ ID No: 46 to 68 respectively, see Table 4), which were individually cloned into a low - copy plasmid backbone containing a low - copy number pBBR1 replicon and a kanamycin resistance marker gene. The first TevCasS target plasmid, designated as p2497 (Table 4), was designed to cover a 71 - nucleotide - long target site with the following structure in the 5' to 3' direction: i) a 5 - nucleotide - long I - TevI 5′-CAACG-3’ (SEQ ID No: 42) cleavage site, followed by ii) a 31 - nucleotide - long native I - TevI DNA spacer sequence 5’-CTCAGTAGATGTTTTCTTGGGTCTACCGTTA-3’ (SEQ ID No: 46), which interacts with the I - TevI zinc - finger domain, followed by iii) a 3 - nucleotide - long S - type PAM sequence 5’-AAG-3’ (SEQ ID No: 13), followed by iv) a 32 - nucleotide - long protospacer sequence 5’-TCGTGTTGTCCACGGTTACCCGCTGGCTGGAA-3’ (SEQ ID No: 69), which is complementary to the sequence of the first spacer in the native CRISPR array of the S - type CRISPR - Cas system.The target sites of the additional TevCasS target plasmids (Table 4, SEQ ID No: 46 to 68) have the same I-TevI cleavage site, PAM, and protospacer portion as the targets on the p2497 plasmid, but their I-TevI DNA spacer portions are different; various numbers of nucleotides are removed from or added to the 3'-end of the I-TevI DNA spacer sequence from p2497 such that target plasmids are constructed with I-TevI DNA spacer sequences of the following lengths: 5 nucleotides (SEQ ID No: 59), 7 nucleotides (SEQ ID No: 58), 9 nucleotides (SEQ ID No: 57), 11 nucleotides (SEQ ID No: 56), 13 nucleotides (SEQ ID No...
Claims
1. A nucleic acid vector or nucleic acid vectors comprising an expressible nucleotide sequence, wherein said sequence comprises a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:1; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:2; c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:3; d) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:4; and e) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:
5.
2. The vector according to claim 1, wherein at least one nucleotide sequence is operably linked to a promoter which i. is heterologous to at least one nucleotide sequence; ii. is not an Escherichia coli, Pseudomonas or Klebsiella promoter; iii. is a eukaryotic cell promoter; iv. is an animal promoter (optionally a mammalian or human promoter); v. is a plant promoter; vi. is a fungal promoter (optionally a yeast promoter); vii. is an insect promoter; viii. is a viral promoter (optionally a viral, AAV or lentiviral promoter); or ix. is a synthetic promoter.
3. The vector according to claim 1 or 2, wherein at least one nucleotide sequence is operably linked to a constitutive promoter.
4. The vector according to claim 1 or 2, wherein at least one nucleotide sequence is operably linked to an inducible promoter.
5. The vector according to any one of claims 1 to 4, wherein the vector lacks the nucleotide sequence encoding the polypeptide as described in part c) of claim 1.
6. The vector according to any one of claims 1 to 5, wherein the vector lacks the nucleotide sequence encoding the polypeptide as described in part a) of claim 1.
7. A nucleic acid vector which comprises an expressible nucleotide sequence, said nucleotide sequence comprising a) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:2; b) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:4; and c) a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:
5.
8. The vector according to claim 7, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:
1.
9. The vector according to claim 7 or claim 8, wherein the vector further comprises a nucleotide sequence encoding a polypeptide comprising an amino acid sequence having at least 90% identity with SEQ ID NO:
3.
10. A vector as claimed in any one of the preceding claims, wherein the vector comprises one or more nucleotide sequences for generating a crRNA, wherein the crRNA comprises a spacer homologous to a first protospacer in a target sequence, optionally wherein the protospacer (a) is not found in Escherichia coli, Pseudomonas spp. and / or Klebsiella spp.; (b) is not found in bacteria containing an endogenous nucleotide sequence encoding a polypeptide of a) to e) as shown in claim 1; (c) is a eukaryotic protospacer; or (d) is a protospacer of an animal (optionally human), plant, insect, fungal cell.
11. A vector as claimed in any one of the preceding claims, wherein the vector comprises (i) a nuclear localization sequence (NLS); (ii) a phage packaging sequence (optionally a pac or cos site); (iii) an origin of plasmid replication; (iv) an origin of plasmid transfer; (v) a bacterial plasmid backbone; (vi) a eukaryotic plasmid backbone; (vii) a human virus (optionally an AAV or lentivirus) structural protein gene; (viii) a virus (optionally an AAV or lentivirus) rep and / or cap sequence; (ix) a selectable marker or a sequence encoding a selectable marker; (x) a eukaryotic promoter; or (xi) a nucleotide sequence of a human, animal, plant or fungal gene.
12. A vector as claimed in any one of the preceding claims, wherein the various vectors are i) a plasmid vector (optionally a conjugative plasmid); ii) a transposon vector (optionally a conjugative transposon); iii) a viral vector (optionally a phage, AAV or lentiviral vector); iv) a phagemid (optionally a packaged phagemid); or v) a nanoparticle (optionally a lipid nanoparticle).
13. One or more polypeptides as shown in any one of the preceding claims, which are expressed from the vector of any one of the preceding claims.
14. A fusion protein comprising a polypeptide (Px), wherein the Px I. comprises an amino acid sequence having at least 90% identity with a sequence selected from SEQ ID No: 1-5; and II. is fused to a heterologous polypeptide (Py).
15. The fusion protein as claimed in claim 14, wherein the fusion protein is in the form of a fusion protein complex comprising at least two additional polypeptides, wherein each of the two additional polypeptides has an amino acid sequence having at least 90% identity with a sequence selected from the sequences of SEQ ID No: 2, SEQ ID No: 4 and SEQ ID No: 5, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity with each sequence of SEQ ID No: 2, SEQ ID No: 4 and SEQ ID No:
5.
16. The fusion protein complex according to claim 15, wherein the protein complex comprises at least 3 additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No:2, SEQ ID No:3, SEQ ID No:4, and SEQ ID No:
5.
17. The fusion protein complex according to claim 15, wherein the protein complex comprises at least 3 additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No:1, SEQ ID No:2, SEQ ID No:4, and SEQ ID No:5; or wherein the protein complex comprises at least 4 additional polypeptides, and wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No:1, SEQ ID No:2, SEQ ID No:3, SEQ ID No:4, and SEQ ID No:
5.
18. The fusion protein according to claim 14, wherein Px comprises an amino acid sequence selected from amino acid sequences having at least 90% identity with the sequences selected from SEQ ID No:1, 3, and 4; wherein the fusion protein is in the form of a fusion protein complex comprising at least two additional polypeptides, wherein each of the two additional polypeptides has an amino acid sequence having at least 90% identity with the sequences selected from SEQ ID No:2 and Seq ID No:5, and (i) an optional polypeptide having at least 90% identity with the sequences selected from SEQ ID No:1, 3, and 4 and not based on the amino acid sequence of the polypeptide shown in Part I; and (ii) an optional polypeptide having at least 90% identity with the sequences selected from SEQ ID No:1, 3, and 4 and not based on the amino acid sequence of the polypeptide shown in Part I, and if present, not based on the amino acid sequence of the polypeptide shown in part (i).
19. The fusion protein or fusion protein complex according to any one of claims 14 to 18, wherein Py is fused to the C-terminus of the amino acid sequence of Part I.
20. The fusion protein or fusion protein complex according to any one of claims 14 to 19, wherein Py is an enzyme, optionally selected from deaminases (such as cytosine or adenine deaminase), base editors, prime editors, helicases, reverse transcriptases, methyltransferases, methylases, acetylases, acetyltransferases, transcriptional activators or deactivators, translational activators or deactivators, or nucleases (such as DNA nucleases, RNA nucleases, nickases, or inactivated nucleases, such as dCas9), optionally wherein Py is selected from base editors, prime editors, and nucleases.
21. The fusion protein or fusion protein complex according to claim 20, wherein Py is a nuclease of I-TevI nuclease, and optionally, in part I, the I-TevI nuclease is fused to the N-terminus of an amino acid sequence that has at least 90% identity with the sequence of SEQ ID No:
1.
22. The fusion protein or fusion protein complex according to claim 21, wherein the I-TevI nuclease lacks a complete DNA-binding domain.
23. The fusion protein or fusion protein complex according to claim 21 or claim 22, wherein the I-TevI nuclease comprises an N-terminal catalytic domain.
24. The fusion protein or fusion protein complex according to any one of claims 21 to 23, wherein the I-TevI nuclease is a bacteriophage I-TevI nuclease, particularly T4 bacteriophage I-TevI, optionally wherein the I-TevI nuclease comprises amino acids 1 to 206 of the naturally occurring I-TevI nuclease, for example wherein the I-TevI nuclease comprises the amino acid sequence of SEQ ID No:
45.
25. The fusion protein or fusion protein complex according to claim 20, wherein Py comprises a base editor selected from the group consisting of adenine base editors (ABE, such as adenosine deaminase), cytosine base editors (CBE, such as cytidine deaminase or apolipoprotein B mRNA editing complex (APOBEC) family deaminases), cytosine-guanine base editors (CGBE), cytosine-adenine base editors (CABE), adenine-cytosine base editors (ACBE), adenine-thymine base editors (ATBE), thymine-adenine base editors (TABE), and uracil DNA glycosylase inhibitor (UGI) proteins, particularly CBE, such as PmCDA1 cytidine deaminase (optionally in combination with a UGI protein).
26. The fusion protein or fusion protein complex according to claim 25, wherein the base editor is a lamprey base editor, for example wherein the base editor is PmCDA1 cytidine deaminase, which has the amino acid sequence encoded by SEQ ID No:
19.
27. The fusion protein or fusion protein complex according to claim 25 or claim 26, wherein the base editor is fused to the C-terminus of the amino acid sequence of part I, optionally wherein the base editor (such as PmCDA1 cytidine deaminase) is fused to the amino acid sequence via a peptide linker, such as a linker of 10 to 20 amino acids (such as about 16 amino acids), for example a linker having the amino acid sequence of SEQ ID No:
17.
28. The fusion protein or fusion protein complex according to claim 27, wherein the UGI protein is fused to the base editor, optionally wherein the UGI protein has the amino acid sequence of SEQ ID No:
21.
29. The fusion protein or fusion protein complex according to claim 28, wherein the UGI protein is fused to the base editor via a peptide linker, such as a linker of about 8 to 12 amino acids (such as about 10 amino acids), such as a linker having the amino acid sequence of SEQ ID No: 23, and optionally wherein the linker is fused to the C-terminus of the base editor, and the UGI protein is fused to the C-terminus of the linker.
30. The fusion protein or fusion protein complex according to any one of claims 14 to 29, wherein the fusion protein or fusion protein complex is combined with a crRNA (such as wherein the fusion protein or fusion protein complex forms a ribonucleoprotein complex with the crRNA), the crRNA comprising a spacer homologous to a first protospacer in a target sequence, optionally wherein the protospacer (a) is not found in a bacterium containing an endogenous nucleotide sequence encoding the polypeptide of a) to e) as shown in claim 1; (b) is a eukaryotic cell protospacer; (c) is heterologous to the source of Py (such as when Py is a naturally occurring polypeptide); or (d) is an animal (optionally human), plant, insect, fungal cell protospacer.
31. One or more nucleic acid vectors, which comprise one or more nucleotide sequences encoding the fusion protein or fusion protein complex as shown in any one of claims 14 to 30 and optionally the crRNA as shown in claim 30, particularly two vectors.
32. The vector according to claim 31, wherein the fusion protein is encoded by a single first nucleotide sequence (such as contained in a first vector), and wherein at least two additional polypeptides are encoded by a second nucleotide sequence (such as contained in a second vector), optionally wherein all additional polypeptides are encoded by the second nucleotide sequence (such as contained in a second vector).
33. The vector according to claim 31, wherein the fusion protein complex is encoded by: a) a first nucleotide sequence encoding Py, which is fused to an amino acid sequence having at least 90% identity with the sequence of SEQ ID No: 1, SEQ ID No: 3 or Seq ID No: 4; and b) a second nucleotide sequence encoding at least two additional polypeptides, the additional polypeptides encoding an amino acid sequence having at least 90% identity with the sequences of SEQ ID No: 1, 2, 3, 4 and / or 5, optionally wherein the fusion protein complex comprises an amino acid sequence having at least 90% identity with each of the sequences of SEQ ID No: 1, 2, 3, 4 and 5.
34. The vector according to claim 32 or 33, wherein the crRNA is encoded by the first nucleotide sequence or by the second nucleotide sequence.
35. The vector according to any one of claims 31 to 34, wherein at least one nucleotide sequence is operably linked to a promoter, the promoter i. is heterologous to at least one nucleotide sequence; ii. is a eukaryotic cell promoter; iii. is an animal promoter (optionally a mammalian or human promoter); iv. is a plant promoter; v. is a fungal promoter (optionally a yeast promoter); vi. is an insect promoter; vii. is a viral promoter (optionally a viral, AAV or lentiviral promoter); or viii. is a synthetic promoter.
36. The vector according to any one of claims 31 to 35, wherein at least one nucleotide sequence is operably linked to a constitutive promoter.
37. The vector according to any one of claims 31 to 35, wherein at least one nucleotide sequence is operably linked to an inducible promoter.
38. The vector according to claim 10 or the vector according to claim 11 or 12 when dependent on claim 10; or the vector according to claim 31 or any one of claims 32 to 37 when dependent on claim 31; or the fusion protein according to claim 30; or the protein according to claim 13 when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10, wherein the crRNA comprises two repeat sequences, and / or wherein the crRNA comprises a repeat sequence, and the repeat sequence comprises the nucleotide sequence of SEQ ID No: 6 or SEQ ID No:
70.
39. The vector as claimed in claim 10 or the vector as claimed in claim 11, 12 or 38 when dependent on claim 10; or the vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 38 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in claim 38 when dependent on claim 30; or the protein as claimed in claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10) or the protein as claimed in claim 38 when dependent on claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10), wherein the protospacer is adjacent to a protospacer adjacent motif (PAM) sequence in the target sequence, which is selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3', 5'-TAG-3', 5'-ACC-3', 5'-AGG-3', 5'-ATT-3', 5'-GAC-3' and 5'-GTG-3', for example selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3', 5'-ACA-3', 5'-ACT-3', 5'-ATC-3', 5'-ATA-3', 5'-GAG-3' and 5'-TAG-3', for example selected from 5'-AAC-3', 5'-ATG-3', 5'-AAA-3', 5'-AAG-3', 5'-ACG-3', 5'-AAT-3' and 5'-ACA-3', for example 5'-AAC-3' and 5'-ATG-3', and in particular the PAM is 5'-AAC-3'.
40. The vector as claimed in claim 10 or the vector as claimed in claim 11, 12, 38 or 39 when dependent on claim 10; or the vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 39 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in claim 38 or claim 39 when dependent on claim 30; or the protein as claimed in claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10) or the protein as claimed in claim 38 or claim 39 when dependent on claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10), wherein the length of the spacer sequence is 25 to 39 nucleotides (for example, the length is 28 to 32 nucleotides) or the length is about 32 nucleotides (for example, the length is 32 nucleotides).
41. The vector, fusion protein or protein as claimed in claim 40, wherein nucleotides 1 to 28 of the spacer region are identical to the complement of nucleotides 1 to 28 of the protospacer immediately 5' of the PAM sequence in the target sequence.
42. The vector as claimed in claim 10 or the vector as claimed in claim 11, 12 or 38 to 41 when dependent on claim 10; or the vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 41 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in any one of claims 38 to 41 when dependent on claim 30; or the protein as claimed in claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10) or the protein as claimed in any one of claims 38 to 41 when dependent on claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10), wherein the spacer region is identical to the complement of the protospacer in the target sequence.
43. The vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 42 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in any one of claims 38 to 42 when dependent on claim 30, wherein said Py is a nuclease for I-TevI nuclease, and the I-TevI nuclease recognizes an I-TevI cleavage site nucleotide sequence at the 5' or 3' (e.g., 5') of the protospacer in the target sequence, optionally wherein the length of the I-TevI cleavage site nucleotide sequence is about 3 to 7 nucleotides (e.g., about 5 nucleotides in length), for example wherein the I-TevI cleavage site nucleotide sequence is 5'-CNNNG-3', for example wherein the I-TevI cleavage site nucleotide sequence is selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3', 5'-CAGCG-3', 5'-CACTG-3', 5'-CGACG-3', 5'-CGATG-3', 5'-CTCCG-3', 5'-CTACG-3', 5'-CTGTG-3', 5'-CTGCG-3', 5'-CCGTG-3', 5'-CCTAG-3', 5'-CCATG-3', 5'-CAAAG-3', 5'-CAGAG-3', 5'-CATTG-3' and 5'-CATCG-3' (e.g., selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3', 5'-CTGAG-3', 5'-CTTAG-3', 5'-CCTCG-3' and 5'-CAGCG-3' or selected from 5'-CCACG-3', 5'-CACCG-3', 5'-CATAG-3', 5'-CTAAG-3' and 5'-CTGAG-3' or selected from 5'-CCACG-3', 5'-CACCG-3' and 5'-CATAG-3' or selected from 5'-CCACG-3' and 5'-CACCG-3'), particularly wherein the I-TevI cleavage site nucleotide sequence has the nucleotide sequence of SEQ ID No:
42.
44. The vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 43 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in any one of claims 38 to 43 when dependent on claim 30, wherein said Py is a nuclease of I-TevI nuclease, and the I-TevI nuclease recognizes an I-TevI spacer nucleotide sequence at the 5' or 3' (e.g., 5') of the protospacer in the target sequence, optionally wherein the length of the I-TevI spacer nucleotide sequence is about 10 to 35 nucleotides or e.g., about 29 to 33 nucleotides in length (e.g., about 31 nucleotides in length), e.g., wherein the I-TevI spacer nucleotide sequence has at least about 90% identity with the nucleotide sequence of SEQ ID No: 43 (e.g., 100% identity with the nucleotide sequence of SEQ ID No: 43).
45. The vector as claimed in claim 31 or the vector as claimed in any one of claims 32 to 44 when dependent on claim 31; or the fusion protein as claimed in claim 30 or the fusion protein as claimed in any one of claims 38 to 44 when dependent on claim 30, wherein said Py is a nuclease of I-TevI, and wherein the target sequence comprises, in a 5' to 3' orientation: a) an I-TevI cleavage site nucleotide sequence (optionally as shown in claim 43); b) an I-TevI spacer nucleotide sequence (optionally as shown in claim 44); c) a PAM sequence (optionally as shown in claim 39); and d) a protospacer sequence (optionally wherein the length of the protospacer sequence is about 32 base pairs (e.g., 32 base pairs in length) or as shown in claim 30, and optionally wherein sequences a) to d) are each adjacent to one another.
46. A cell, wherein the cell is A. a eukaryotic cell comprising the vector or protein as claimed in any one of the preceding claims; B. an animal, plant, insect or fungal cell comprising the vector or protein as claimed in any one of the preceding claims; or C. a prokaryotic cell comprising the vector or protein as claimed in any one of the preceding claims, optionally wherein the cell is not a bacterial cell (e.g., Escherichia coli, Pseudomonas or Klebsiella cells) comprising an endogenous nucleotide sequence encoding the polypeptide of a) to e) as claimed in claim 1.
47. A composition comprising the vector or protein as claimed in any one of claims 1 to 45 or the cell as claimed in claim 46 (optionally an in vitro composition or wherein the composition is contained in a medical container).
48. A method of modifying a nucleic acid target sequence in a cell, the method comprising I) contacting the cell with: a vector as claimed in claim 10 or a vector as claimed in any one of claims 11, 12 or 38 to 45 when dependent on claim 10; or a vector as claimed in claim 31 or a vector as claimed in any one of claims 32 to 45 when dependent on claim 31; or a fusion protein as claimed in claim 30 or a fusion protein as claimed in any one of claims 38 to 45 when dependent on claim 30; or a protein as claimed in claim 13 (when reciting claim 10 or when reciting claim 11 or 12 when dependent on claim 10) or a protein as claimed in any one of claims 38 to 45 when dependent on claim 13 (when reciting claim 10 or when reciting claim 11 or 12 dependent on claim 10); or a composition as claimed in claim 44; II) allowing introduction into the cell of (i) a vector or a protein, whereby the polypeptide or protein and the crRNA are expressed in the cell, or (ii) the cRNA of the composition (or a nucleic acid encoding the crRNA or the guide RNA, whereby the cRNA or the guide RNA is expressed in the cell) and the nucleic acid of the composition encoding the polypeptide or protein, whereby the polypeptide or protein is expressed in the cell; III) wherein the crRNA forms a complex with the polypeptide or protein to direct the complex to the target sequence.
49. The method according to claim 48, wherein the cell is a eukaryotic, animal, plant, insect or fungal cell; or a prokaryotic cell, optionally wherein the cell is not a prokaryotic cell containing an endogenous nucleotide sequence (such as an Escherichia coli, Pseudomonas or Klebsiella cell) encoding the polypeptide of a) to e) as claimed in claim 1.
50. The method according to claim 48 or 49, wherein the target sequence is contained in a chromosome or episome (optionally a plasmid) in the cell, particularly in the chromosome of the cell.
51. The method according to claim 50, wherein the method inhibits chromosomal or episomal replication in the cell.
52. The method according to any one of claims 48 to 51, wherein the target sequence is within about 2 kb upstream or downstream of the gene of interest, for example the target sequence is within about 1 kb upstream or downstream of the gene of interest, such as within the promoter, exon and / or intron of the gene of interest, such as within the promoter of the gene of interest.
53. The method according to claim 52, wherein the target sequence of the gene of interest is at least 2 kb upstream and downstream of any promoter, exon and intron of any other gene.
54. The method according to any one of claims 48 to 53, wherein the modification is a substitution, deletion or insertion of one or more nucleotides, for example a substitution (such as a C to T substitution when Py is the PmCDA1 cytidine deaminase).
55. A method of treating or preventing a target cell-mediated disease or condition in a subject, the method comprising performing the method according to any one of claims 48 to 54 to modify a target cell, wherein the contacting comprises administering a vector, protein or composition to the subject, and wherein the modification treats or prevents the disease or condition.
56. The method according to claim 55, wherein the subject is a human, animal or plant.
57. The method according to claim 56, wherein the target cell is a cell of a subject having a nucleic acid defect, and the modification corrects the defect.
58. The method according to claim 56 or 57, wherein the modification a) adds a new function to the cell, optionally the modification upregulates or downregulates (especially downregulates, e.g., blocks) gene expression in the cell, or adds a new nucleotide sequence into the genome of the cell for expression of a protein encoded by the new sequence; or b) regulates the expression of a nucleotide sequence contained in a chromosome or episome (optionally a plasmid) of the cell.
59. The vector, protein, cell or composition according to any one of claims 1 to 47, for use in a method of treating or preventing a disease or condition in a human or animal, wherein the method is according to any one of claims 55 to 58.
60. A method of introducing a targeted edit in a target polynucleotide, the method comprising contacting the target polynucleotide with a fusion protein according to claim 30 or a fusion protein according to any one of claims 38 to 45 when dependent on claim 30, wherein a first protospacer in the target sequence is contained in the polynucleotide, and wherein the crRNA hybridizes with the protospacer to direct a protein (e.g., a ribonucleoprotein complex), whereby the protein (e.g., a ribonucleoprotein complex) edits the polynucleotide.
61. The method according to claim 60, wherein the method inserts, deletes or replaces bases or nucleic acid sequences in the polynucleotide.
62. A container comprising a plurality of proteins that are operable with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, the proteins comprising a) any protein as shown in claim 13; or b) any fusion protein as shown in any one of claims 14 to 29; wherein the complex is operable with a protospacer adjacent motif (PAM) having one of the sequences as shown in claim 39, and the container is not a cell, and the proteins are mixed with an in vitro buffer.
63. A nucleic acid vector or a plurality of nucleic acid vectors according to any one of claims 1 to 12 or any one of claims 31 to 45, wherein the polypeptide or protein encoded by the vector is operable with a crRNA to form a ribonucleoprotein complex for protospacer targeting in a polynucleotide, wherein the complex is operable with a protospacer adjacent motif (PAM) having one of the sequences as shown in claim 39.
64. A method of targeting a polynucleotide (optionally in a cell or in vitro), the method comprising a) contacting the polynucleotide with: a protein as described in claim 13 (claim 11 or 12 when claim 10 is recited or when claims dependent on claim 10 are recited) and a crRNA, or a protein and a crRNA as described in any one of claims 38 to 45 when dependent on claim 13 (when claim 10 is recited or when claims dependent on claim 10 are recited); or a protein and a crRNA as described in claim 30 or a protein and a crRNA as described in any one of claims 38 to 45 when dependent on claim 30; b) allowing the formation of a ribonucleoprotein complex comprising the protein and the crRNA or guide RNA, wherein the complex is directed to a target sequence comprised in the polynucleotide to modify the polynucleotide or its replication.
65. The method according to claim 64, wherein the polynucleotide is comprised in a chromosome or episome (optionally a plasmid).
66. The method according to claim 65, wherein the method a) inhibits the replication of the polynucleotide; or b) edits the polynucleotide, for example wherein the editing inserts, deletes or substitutes bases or nucleic acid sequences in the polynucleotide.
67. A kit comprising a) one or more proteins as shown in claim 13; or a fusion protein as shown in any one of claims 14 to 29; or one or more vectors as described in any one of claims 1 to 12 or any one of claims 31 to 37 when dependent on any one of claims 1 to 12; and b) a crRNA or one or more nucleic acids encoding the crRNA, wherein the crRNA is homologous to a PAM having one of the sequences as shown in claim 39; wherein the polypeptide is operable to be used with the crRNA to form a ribonucleoprotein complex for targeting of a protospacer in a polynucleotide.
68. A ribonucleoprotein complex comprising a) one or more proteins as shown in claim 13; or a fusion protein as shown in any one of claims 14 to 29; and b) a crRNA, the crRNA being homologous to a PAM having one of the sequences as shown in claim 39.
69. A method of controlling the transcription or replication of a target DNA, the method comprising contacting the target DNA with the complex according to claim 68, wherein the complex lacks a DNA nuclease and the complex binds to the target DNA, thereby controlling the transcription or replication of the target DNA.
70. A method of controlling the transcription of a target RNA, the method comprising contacting the target RNA or the DNA encoding the RNA with the complex according to claim 69, wherein the complex binds to the target RNA or DNA, thereby controlling the transcription of the target RNA.
71. The method according to claim 69 or 70, wherein the DNA is comprised in a chromosome or episome (optionally a plasmid).
72. A method of editing a target DNA, comprising contacting the target DNA with a complex as claimed in claim 68, wherein the complex binds to the target DNA so as to edit the target DNA.
73. A method of cleaving double-stranded DNA (dsDNA), comprising contacting the dsDNA with a complex as claimed in claim 68, wherein the complex comprises a fusion protein, wherein Py comprises a nuclease, wherein the dsDNA comprises a protospacer sequence which is flanked at its 5'-end by a PAM having one of the sequences as claimed in claim 39, whereby the nuclease cleaves the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
74. A method of cleaving single-stranded DNA (ssDNA), comprising contacting the ssDNA with a complex as claimed in claim 68, wherein the complex comprises a fusion protein, wherein Py comprises a nickase, wherein the ssDNA comprises a protospacer sequence which is flanked at its 5'-end by a PAM having one of the sequences as claimed in claim 39, whereby the nickase cleaves the single strand of the ssDNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region.
75. A method of labeling or identifying a DNA region, comprising contacting the DNA with a complex as claimed in claim 68, wherein the DNA comprises a protospacer sequence which is flanked at its 5'-end by a PAM having one of the sequences as claimed in claim 39, whereby the complex binds to the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, and optionally wherein the complex comprises a fusion protein, wherein Py comprises a detectable tag.
76. A method of modifying the transcription of a DNA region, comprising contacting the DNA with a complex as claimed in claim 68, wherein the DNA comprises a protospacer sequence which is flanked at its 5'-end by a PAM having one of the sequences as claimed in claim 39, whereby the complex binds to the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the complex upregulates or downregulates the transcription of the DNA region or an adjacent gene, particularly downregulates transcription.
77. A method of modifying a target dsDNA of a cell without introducing a dsDNA break, the method comprising generating in the cell a complex as claimed in claim 68, wherein the complex targets a target dsDNA contained in the cell, the target dsDNA comprising a protospacer sequence which is flanked at its 5'-end by a PAM having one of the sequences as claimed in claim 39, whereby the complex binds to the dsDNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby the dsDNA is modified without introducing a dsDNA break.
78. The method as claimed in claim 77, wherein the method edits the DNA, for example wherein the editing inserts, deletes or substitutes bases or nucleic acid sequences in the DNA.
79. A method of inhibiting the growth or proliferation of cells without introducing lethal dsDNA breaks, the method comprising generating in a cell a complex as claimed in claim 68, wherein the complex targets a target DNA contained in the cell, the target DNA comprising a protospacer, the protospacer being flanked at its 5' end by a PAM having one of the sequences as shown in claim 39, whereby the complex binds to the DNA in a region defined by the complementary binding of the spacer sequence of the crRNA to the protospacer region, whereby DNA replication is inhibited without introducing lethal dsDNA breaks in the DNA, thereby inhibiting the growth or proliferation of the cells.
80. A method of treating or preventing a disease or condition in a human, animal, plant or fungal subject, the method comprising performing a method as claimed in any one of claims 69 to 79, wherein the cells of the subject contain the DNA or RNA and the cells mediate the disease or condition.
81. A method as claimed in any one of claims 69 to 80, wherein the complex modifies a coding or non-coding target sequence, optionally wherein (i) the coding strand of the DNA is modified rather than the non-coding strand or (ii) the non-coding strand of the DNA is modified rather than the coding strand.
82. The method as claimed in claim 81, wherein the modification effects cleavage, editing, blocking, labelling or tagging of the target sequence.
83. One or more nucleic acid vectors or one or more nucleic acids, comprising at least one nucleotide sequence selected from SEQ ID NO: 7 - 11, wherein the nucleotide sequence a) is operably linked to a heterologous promoter, synthetic promoter, eukaryotic promoter or non-bacterial promoter; and / or b) is contained in a cell which is a eukaryotic cell, non-bacterial cell or a bacterial cell (such as an Escherichia coli, Pseudomonas or Klebsiella cell) which does not contain an endogenous nucleotide sequence comprising SEQ ID NO: 7 - 11.
84. A nucleic acid vector or nucleic acid comprising SEQ ID NO: 12 or a DNA sequence having at least 80% identity to SEQ ID NO:
12.
85. A ribonucleoprotein CRISPR / Cas complex comprising a plurality of Cas proteins and a crRNA or guide RNA, wherein the RNA is capable of guiding the complex to a protospacer region contained in a target DNA, the protospacer region being flanked at its 5' end by a protospacer adjacent motif (PAM) having the sequence 5'-AAG-3' or being homologous to a PAM having one of the sequences as shown in claim 39, wherein a) the complex lacks a DNA nuclease and is capable of modifying DNA without introducing double-stranded DNA breaks; b) the complex lacks DinG, Cas3 and Cas10; and c) the complex does not contain all of the Cas proteins of a type I, II, III, IV, V or VI CRISPR / Cas complex.
Citation Information
Patent Citations
T:a to a:t base editing through adenosine methylation
US20220170013A1
Evolved CAS9 variants and uses thereof
US20220307001A1
Reducing systemic regulatory t cell levels or activity for treatment of disease and injury of the CNS
WO2015136541A2
Evolved CAS9 variants and uses thereof
WO2019168953A1
T:a to a:t base editing through thymine alkylation
WO2020181178A1
Cited By
Atherococcus turbinatus and application thereof
CN116622544A
Moraxella nasalis and uses thereof
CN116622544B
Enterobacter mori capable of resisting fish aeromonas and application of enterobacter mori
CN121518353A