Compositions and methods for microbial self-selection
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- RGT UNIV OF CALIFORNIA
- Filing Date
- 2025-10-23
- Publication Date
- 2026-07-23
AI Technical Summary
Current microbiome modulation approaches rely heavily on antibiotics, which indiscriminately affect microbial populations and are increasingly ineffective due to the rise of antibiotic-resistant bacteria, necessitating a need for compositions and methods that selectively enrich and maintain edited microbes without disrupting the broader microbial community.
A self-selecting system leveraging bacteriocins, naturally evolved antimicrobial peptides, is used to selectively enrich edited cells by providing a competitive advantage over unedited cells, engineered with a controllable external inducer for precise regulation, combined with genome editing tools like CRISPR Cas Transposase (CAST) to achieve enrichment across multiple orders of magnitude.
The system effectively increases the proportion of edited cells while sparing members of other microbial phyla, minimizing disruption of the indigenous microbiota, and is applicable for microbiome engineering in human health, agriculture, and environmental conservation.
Abstract
Description
COMPOSITIONS AND METHODS FOR ICROBIAL SELF-SELECTIONCross-Reference
[0001] This application claims the benefit of U.S. Provisional Patent Application No.63 / 711,895 filed October 25, 2024, which application is incorporated herein by reference in its entirety.Incorporation by reference of Sequence Listing provided as an XML File
[0002] A Sequence Listing is provided herewith as a Sequence Listing XML, “BERK- 545PRV_SeqList_v2.xml” created on October 24, 2024 and having a size of 114,684 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in their entirety.I. Introduction
[0003] Genetic manipulation of microbiomes is crucial for understanding microbial interactions within communities, with hosts, and with the environment, as well as for developing innovative solutions in human health, agriculture, and climate change mitigation. Current microbiome modulation approaches rely heavily on antibiotics, which indiscriminately affect microbial populations and are increasingly ineffective due to the rise of antibiotic-resistant bacteria.
[0004] There is a need for compositions and methods that selectively enrich (e.g., enrich and maintain) edited microbes such as bacteria, e.g., without disturbing the broader microbial community. Such is provided herein.II. SUMMARY
[0005] Precise microbiome engineering requires alternative selection systems to antibiotics that spare the microbiome. Bacteria have naturally evolved systems to outcompete competitors. Bacteriocins are an alternative selection strategy for microbiome engineering - they are: antimicrobial peptides, widespread across phyla, evolved to kill related strains, are encoded in small gene clusters (4-15kb), and have demonstrated safety and efficacy as food preservatives.
[0006] The inventors contemplated that Bacteriocins may be used to selectively enrich and increase penetrance of cells (e.g., edited cells) in microbial communities. Fig. 3 provides a schematic of how one can use the methods and compositions described herein to enrich for edited cells within a microbial community using bacteriocins for self-selection.
[0007] The inventors have developed a self-selecting system that leverages a natural bacterial antagonism mechanism to selectively enrich edited cells by providing them with a strong competitive advantage over unedited cells of the same or closely related species. This system can be engineered to be controllable by an external synthetic inducer, allowing for precise regulation. By using this self-selecting system in combination with genome editing tools such as RNA-guided CRISPR Cas Transposase (CAST), one can deliver a desired edit (e.g., gene edit) linked to the genes encoding for the selection system. The selection system significantly increases the proportion of edited cells over time (via selection after editing), achieving enrichment (self-selection) across multiple orders of magnitude. Provided systems are highly effective against strains of the target species, while sparing members of other major phyla in microbial communities. The findings disclosed herein demonstrate that the subject self-selection systems are a precise and effective approach for engineering microbial communities, minimizing disruption of the indigenous microbiota. These systems can revolutionize microbiome editing and engineering by enabling precise, efficient and sustainable interventions that can enhance human health, optimize agricultural practices, and contribute to environmental conservation efforts.
[0008] Provided are self-selection DNA editing systems that include (a) a gene editing tool;and (b) a donor DNA comprising a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein. In some cases, the bacteriocin unit further comprises factors involved in bacteriocin production and protection, modifying enzymes for modified peptides, export machinery, or any combination thereof. In some cases, the bacteriocin unit comprises a bacteriocin gene cluster - and in some such cases, all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon. In some cases, a bacteriocin unit comprises a nucleotide sequence encoding a bacteriocin immunity protein but does not include a nucleotide sequence encoding a bacteriocin protein.
[0009] The donor DNA (e.g., a transposon, an HDR template, and the like) is a DNA that is configured for integration of a nucleotide sequence of the DNA into a target DNA. The gene editing tool can be any tool that facilitates gene editing, e.g., a ZNF, a TALEN, a meganuclease, a CRISPR-Cas effector protein and guide RNA, a group II intron (e.g. ClosTron / TargeTron), a site-specific recombinase or integrase, etc. In some cases, the gene editing tool includes (i) nucleotide sequences encoding polypeptides that form a CRISPR-associated transposase (CAST) complex, and (ii) a nucleotide sequence encoding a guide RNA. In such cases, the donor DNA can be a transposon (e.g., flanked by recognition sites that are recognized by the CAST complex).
[0010] In some cases, the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter (e.g., an anhydrotetracycline (ATC) inducible promoter). In some cases, the components of a subject self-selection DNA editing system (gene editing tool and donor DNA) are present on the same nucleic acid (e.g., a phage DNA, a plasmid such as a conjugation plasmid, integrative conjugative element, and the like). In some cases, the bacteriocin is microcin V, B17, J25, N, or S. In some cases, the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA. In some cases, the donor DNA further comprises a therapeutic unit comprising: one or more nucleotide sequences encoding one or more expression products of interest, e.g., one or more proteins that confer antibiotic resistance on a bacterium, one or more enzymes of a biosynthetic pathway (e.g., a magnetosome, gas vesicle, B-vitamin, mevalonate, or polyketide biosynthetic pathway), one or more enzymes in a carbon utilization pathway (e.g., porphyran polysaccharide, glycosaminoglycan, non-caloric artificial sweetener, ethanolamine, or sucrose utilization pathway), one or more detectable markers (e.g., one or more fluorescent proteins, one or more fluorogenic aptamers, one or more proteins that provide for isolation of a target cell); one or more antibodies or nanobodies; one or more antiinflammatory factors; one or more immunomodulatory or antitumor factors (e.g. tumor- associated antigen, cytolysin A, synthetic CAR targets), or any combination thereof.
[0011] Also provided are prokaryotic cells that include a subject self-selection DNA editing system. Also provided are genetically modified prokaryotic cells that include a bacteriocin unit integrated into their genomic DNA or into a mobile DNA element (e.g., a conjugative transposon, a plasmid, a virus (an archaeal virus), a viral satellite (e.g., an archaeal viral satellite), a phage, a prophage, a phage satellite, a phage plasmid, and the like) inside of the prokaryotic cell, where the bacteriocin unit comprises anucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein (e.g., in some cases the bacteriocin unit comprises a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein).
[0012] Also provided are methods of modifying a microbial community, where the method includes introducing a genetically modified prokaryotic cell to the microbial community, or includes editing a target prokaryotic cell in situ. In some cases, the editing includes providing to the target prokaryotic cell, a DNA editing system and a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein, wherein the target prokaryotic cell is a member of the microbial community.
[0013] Also provided are methods of editing a prokaryotic cell to facilitate self-selection, where the methods include providing a bacteriocin unit that comprises a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein to a target prokaryotic cell, wherein the bacteriocin unit integrates into prokaryotic cell’s genomic DNA or into a mobile DNA element inside of the prokaryotic cell.
[0014] In some cases, any of the prokaryotic cells above (a prokaryotic cell, a genetically modified prokaryotic cell, a target prokaryotic cell) are a bacterial cell. In some cases, a prokaryotic cell is a proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longurri), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa or Enterobacter cell. In some cases, a prokaryotic cell is an E. coli cell.
[0015] Examples of use of the compositions and methods provided herein include, but are not limited to: enriching for edited cells (e.g., enriching and sustaining / maintaining edited cells) in microbial communities without causing major disruptions to the community (in situ niche replacement), editing bacterial strains that are resistant to the most commonly used antibiotics (which is valuable because the major approach to selecting for edits in isolates is based on antibiotic markers), controllable engraftment of edited probiotic strains in a natural community (e.g., a human microbiome such as the gut microbiome), in situ editing to facilitate colonization resistance, and uses involving editing with immunity genes (e.g., in some cases where the bacteriocin unitincludes a nucleotide sequence encoding one or more immunity genes but does include a nucleotide sequence encoding a bacteriocin) (see, e.g., FIG. 3 and FIG. 5).
[0016] Reagents, compositions, and kits / systems that find use in practicing the subject methods are provided.III. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The following detailed description of embodiments of the invention will be better understood when read in conjunction with the appended drawings. It should be understood that the invention is not limited to the precise arrangements and instrumentalities of the embodiments shown in the drawings.
[0018] FIG. 1A-1B. Refactored bacteriocin systems facilitate controllable antimicrobial activity. (FIG. 1A) Schematic of a gene cluster encoding at least a bacteriocin (toxin) and its corresponding resistance mechanism (self-immunity). The cluster may also include genes for secretion machinery and post-translational modification enzymes. The gene cluster can be refactored to remove internal regulatory elements, placing the toxin gene — and in some cases, the entire operon — under an inducible promoter for controlled expression. This modified gene cluster can be referred to as the "selfselecting unit" or "enrichment unit" or “bacteriocin unit”. The ‘bacteriocin unit’ can be encoded in the same DNA molecule as the ‘therapeutic unit’, which can encode various elements, such as: therapeutic genes (conferring therapeutic benefits), a DNA sequence necessary for disruption or deletion of target genes, markers such as antibiotic resistance genes, fluorescent markers, DNA barcodes, or therapeutically relevant genes. Together, the self-selecting unit and the therapeutic unit form the "genetic payload", intended for delivery into target bacterial cells. The genetic payload can be cloned into a plasmid, either suicide or replicative, which can also contain the necessary systems for delivery and genomic integration into the target bacteria. This delivery system can include, for example, a CRISPR-Associated Transposase (CAST) with a genome-targeting guide RNA. (FIG. 1B) In the absence of ATC, minimal or no antimicrobial activity was observed for all bacteriocin-encoding strains. However, as the concentration of ATC increased, clear zones of inhibition formed and expanded, matching or exceeding the diameter of those produced by natural bacteriocin systems.
[0019] FIG. 2A-2B. Bacteriocins facilitate self-selection of CAST-mediated genome edits in E. coli. (FIG. 2A) Schematics of delivery and self-selection of a geneticpayload using bacteriocins, including quantification of edited cells in the population. (FIG. 2B) After the first round of growth and bacteriocin induction with 200 nM ATC, the fraction of edited E. coii cells carrying both the therapeutic and self-selecting units increased 19-fold compared to the initial timepoint (t=Oh), and up to 97x after three passages. In contrast, the fraction of edited cells in the population that received only the therapeutic unit remained unchanged throughout the experiment.
[0020] FIG. 3. Schematic depicting a non-limiting example of enriching for edited cells within a microbial community using bacteriocins.
[0021] FIG. 4. Schematic of non-limiting examples of bacteriocin gene clusters (e.g.,bacteriocin V, S, N, J25, B17) that can be refactored to generate self-selecting unit of a genetic payload.
[0022] FIG. 5. Schematic of non-limiting examples of different uses for the subject methods and compositions, which include enrichment of genome-edited cells in culture (e.g., E. coli), niche replacement of cells (e.g., E. coli) with edited cells (e.g., E. coli) in a community, and microbiome editing (enrichment of genome-edited cells in a community).
[0023] FIG. 6A-6B. Bacteriocins collectively kill 85% of E. coli isolates while sparing other members of the microbiome. (Fig. 6A) Results demonstrating that three of the bacteriocins (V, B17 and S) killed at least 60% of the E. coli strains each. These data suggest that with just three bacteriocins, this selection strategy can be used in about 85% of E. coli strains. (FIG. 6B) Results demonstrating that all four bacteriocins selectively target E. coli strains, except for bacteriocin N that killed Salmonella enterica, but that is a very close relative of E. coli and this result is consistent with existing literature. Overall, these four bacteriocins exhibit a high degree of specificity compared to most antibiotics.
[0024] FIG. 7A-7C. Bacteriocins enable selection of CAST-mediated genome edits in E.coli across growth conditions in vitro. (Fig. 7A) Schematics of delivery and selfselection of a genetic payload using bacteriocins, including quantification of edited cells in the population. (FIG. 7B) Results demonstrating that after three rounds of growth and bacteriocin induction in LB broth, the fraction of edited E. coli cells carrying both the therapeutic and self-selecting units increased 77-fold compared to the control that only carries the therapeutic unit. This assay was then performed in different growth conditions to capture a natural environment for editing bacteria in-situ. In M9 broth, the fraction of edited cells that received the bacteriocin cluster increased by 5-orders of magnitude relative to edited cells that didn’t receive the bacteriocin,reaching 20% of total cells. In LB solid media, a similar trend was observed, with a 1000x increase in the fraction of cells edited that received the bacteriocin relative to the control edit. (FIG. 7C) Results demonstrating that after two rounds of growth and bacteriocin induction in M9 broth, the fraction of edited E. coli cells carrying both the therapeutic and self-selecting units increased 5 orders of magnitude compared to the control that only carries the therapeutic unit. At that timepoint, the fraction of edited cells is 24% of the total cells.
[0025] FIG. 8. Edited cells are maintained in the population after removing inducer.Results demonstrating that after self-selection in both LB and M9 broth, the fraction of edited cells didn’t decrease, hence edited cells persisted without bacteriocin inducer. This suggests that bacteriocin self-selection created a stable population of edited cells.
[0026] FIG. 9. Schematic one example embodiment in which bacteriocins are used for selfselection of edited cells in microbiomes.IV. DEFINITIONS
[0027] The terms “polynucleotide” and “nucleic acid,” used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0028] By "hybridizable" or “complementary” or “substantially complementary" it is meant that a nucleic acid (e.g. RNA, DNA) comprises a sequence of nucleotides that enables it to non-covalently bind, i.e. form Watson-Crick base pairs and / or G / U base pairs, “anneal”, or “hybridize,” to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. Standard Watson-Crick base-pairing includes: adenine (A) pairing with thymidine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C) [DNA, RNA], In addition, for hybridization between two RNA molecules (e.g., dsRNA), and for hybridization of a DNA molecule with an RNA molecule (e.g., when a DNA target nucleic acid base pairs with a guide RNA, etc.): guanine (G) can also base pair with uracil (U). For example, G / U base-pairing is at least partiallyresponsible for the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anti-codon base-pairing with codons in mRNA. Thus, in the context of this disclosure, a guanine (G) (e.g., of dsRNA duplex of a guide RNA molecule; of a guide RNA base pairing with a target nucleic acid, etc.) is considered complementary to both an uracil (U) and to an adenine (A). For example, when a G / U base-pair can be made at a given nucleotide position of a dsRNA duplex of a guide RNA molecule, the position is not considered to be non-complementary, but is instead considered to be complementary.
[0029] Hybridization and washing conditions are well known and exemplified in Sambrook, J., Fritsch, E. F. and Maniatis, T. Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (1989), particularly Chapter 11 and Table 11.1 therein; and Sambrook, J. and Russell, W., Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor (2001). The conditions of temperature and ionic strength determine the "stringency" of the hybridization.
[0030] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. The conditions appropriate for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementarity, variables well known in the art. The greater the degree of complementarity between two nucleotide sequences, the greater the value of the melting temperature (Tm) for hybrids of nucleic acids having those sequences. For hybridizations between nucleic acids with short stretches of complementarity (e.g. complementarity over 35 or less, 30 or less, 25 or less, 22 or less, 20 or less, or 18 or less nucleotides) the position of mismatches can become important (see Sambrook et al., supra, 11.7-11.8). Typically, the length for a hybridizable nucleic acid is 8 nucleotides or more (e.g., 10 nucleotides or more, 12 nucleotides or more, 15 nucleotides or more, 20 nucleotides or more, 22 nucleotides or more, 25 nucleotides or more, or 30 nucleotides or more). Temperature, wash solution salt concentration, and other conditions may be adjusted as necessary according to factors such as length of the region of complementation and the degree of complementation.
[0031] It is understood that the sequence of a polynucleotide need not be 100%complementary to that of its target nucleic acid to be specifically hybridizable or hybridizable. Moreover, a polynucleotide may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a bulge, a loop structure or hairpin structure, etc.). A polynucleotide cancomprise 60% or more, 65% or more, 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 98% or more, 99% or more, 99.5% or more, or 100% sequence complementarity to a target region within the target nucleic acid sequence to which it will hybridize. For example, an antisense nucleic acid in which 18 of 20 nucleotides of the antisense compound are complementary to a target region, and would therefore specifically hybridize, would represent 90 percent complementarity. In this example, the remaining noncomplementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous to each other or to complementary nucleotides. Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined using any convenient method. Example methods include BLAST programs (basic local alignment search tools) and PowerBLAST programs (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489), and the like.
[0032] The terms "peptide," "polypeptide," and "protein" are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.
[0033] "Binding" as used herein (e.g. with reference to binding of a guide RNA to a target nucleic acid, binding of an RNA-binding domain of a polypeptide to a guide RNA, and the like) refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid; between a CAST polypeptide / guide RNA complex and a target nucleic acid; and the like). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it is meant the molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction need be sequence-specific (e.g., contacts with phosphate residues in a DNA backbone), but some portions of a binding interaction may be sequence-specific. Binding interactions are generally characterized by a dissociation constant (KD) of less than 10'6M, less than 10-7M, less than 10-8M, less than 10'9M, less than 10-10M, less than 10-11M, less than 10-12M, less than 10-13M, less than 10-14M, or less than10-15M. "Affinity" refers to the strength of binding, increased binding affinity being correlated with a lower KD.
[0034] As used herein, a “promoter” or a "promoter sequence" is a DNA regulatory region capable of initiating transcription of a downstream (3' direction) coding or non-coding sequence. Various promoters, including inducible promoters, may be used to drive expression of the various components of the present disclosure (e.g., non-coding RNAs or coding RNAs, which in some cases lead to translation into protein).
[0035] "Operably linked" refers to a juxtaposition wherein the components so described are in a relationship permitting them to function in their intended manner. For instance, a promoter is operably linked to a coding sequence (or the coding sequence can also be said to be operably linked to the promoter) if the promoter affects its transcription / expression.
[0036] Several techniques exist to determine nucleic acid or polypeptide sequence homology and / or three-dimensional homology to polypeptides. These methods are routinely employed to discover the extent of identity that one sequence, domain, or model has to a target sequence, domain, or model. For example, percent identity may be determined using the BLAST software (Altschul, S. F., et al. (1990) “Basic local alignment search tool.” J. Mol. Biol. 215:403-410, accessible on the world wide web at blast.ncbi.nlm.nih.gov) with the default parameters.
[0037] Additional definitions are provided throughout this disclosure.
[0038] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.
[0039] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0040] Certain ranges are presented herein with numerical values being preceded by the term "about." The term "about" is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number.
[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative illustrative methods and materials are now described.
[0042] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates which may need to be independently confirmed.
[0043] It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. As such, the articles “a” and “an” are used herein to refer to one or to more than one (i.e. , to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element. Thus, for example, reference to “a cell” includes a plurality of such cells and reference to “the polypeptide” includes reference to one or more polypeptides and equivalents thereof known to those skilled in the art, and so forth. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.
[0044] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any ofthe other several embodiments without departing from the scope or spirit of the present invention. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. For example, it is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments pertaining to the invention are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0045] While the apparatus and method has or will be described for the sake of grammatical fluidity with functional explanations, it is to be expressly understood that the claims, unless expressly formulated under 35 U.S.C. §112, are not to be construed as necessarily limited in any way by the construction of "means" or "steps" limitations, but are to be accorded the full scope of the meaning and equivalents of the definition provided by the claims under the judicial doctrine of equivalents, and in the case where the claims are expressly formulated under 35 U.S.C. §112 are to be accorded full statutory equivalents under 35 U.S.C. §112.V. DETAILED DESCRIPTIONCompositions and Methods
[0046] As noted above, the present disclosure provides a self-selection DNA editing system.A subject self-selection DNA editing system includes: (a) a gene editing tool, and (b) a donor DNA comprising a bacteriocin unit, which includes a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein. In some cases, (a) and (b) are present on the same nucleic acid (e.g., plasmid DNA). The term “donor DNA” is used herein to describe a DNA that includes a bacteriocin unit, and is configured for integration of a nucleotide sequence of the DNA into a target DNA (e.g., a genomic DNA), e.g., inside of a prokaryotic cell. Thus, when a bacteriocin unit of the subject disclosure is present on (as part of) a donor DNA, the bacteriocin unit is configured for integration into a target DNA such asa genomic DNA. This integration provides for self-selection. For example, once the target DNA (e.g., inside of a targeted prokaryotic cell such as a bacterial cell) is modified (to include the integrated bacteriocin unit), organisms that include that modified target DNA will self-select / self-enrich because the bacteriocin can kill neighboring competition, in particular target bacterial cells that did not receive the donor DNA (i.e. , unedited cells), while at the same time the bacteriocin immunity protein will provide protection against the bacteriocin.
[0047] The present disclosure also provides methods of modifying a microbial community, where the method includes introducing a genetically modified prokaryotic cell (comprising a bacteriocin unit integrated into its genomic DNA or into a mobile DNA element inside of the prokaryotic cell) to the microbial community, or includes editing a target prokaryotic cell (where the target prokaryotic cell is a member of a microbial community) in situ. In some cases, the editing includes providing to the target prokaryotic cell, a DNA editing system and a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein.
[0048] The present disclosure also provides methods of editing a prokaryotic cell to facilitate self-selection, where the methods include providing a bacteriocin unit that comprises a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein to a target prokaryotic cell, where the bacteriocin unit integrates into a target DNA inside of the prokaryotic cell (e.g., the prokaryotic cell’s genomic DNA, e.g., at a locus for which disruption with have a beneficial effect, or into a mobile DNA element inside of the prokaryotic cell).Bacteriocins and immunity proteins
[0049] “Antibiotic,” and variations of this root term, refers to a metabolite, or an intermediate of a metabolic pathway which can kill or arrest the growth of at least one microbial cell. Some antibiotics can be produced by microbial cells, for example bacteria. Some antibiotics can be synthesized chemically. It is understood that “bacteriocins” can also kill or arrest the growth of at least one microbial cell, but are distinct from antibiotics, at least in that bacteriocins refer to gene products (e.g., proteins) (which, in some cases undergo additional post-translational processing) or synthetic analogs of the same, while antibiotics refer to intermediates or products of metabolic pathways or synthetic analogs of the same. Most bacteriocins have evolved to kill related strainsand therefore have a narrow activity spectrum, making them great candidates to selectively enrich for edited cells without impacting other species. To the contrary, most antibiotics have broad activity against bacteria.
[0050] A segment of DNA that encodes a bacteriocin-encoding gene (i.e. , a nucleotide sequence encoding a bacteriocin protein) and a bacteriocin immunity gene (i.e., a nucleotide sequence encoding an immunity protein) is referred to herein as a “bacteriocin unit” (also referred to as a “self-selecting bacteriocin unit” or a “selfselecting unit”).
[0051] In some embodiments, bacteriocins neutralize cells of a different species or strain from the host cell. In some embodiments, bacteriocins neutralize cells of the same species or strain as the host cell if these cells lack an appropriate immunity modulator (immunity gene / immunity protein). Neutralizing activity of bacteriocins can include arrest of microbial reproduction, or cytotoxicity. Some bacteriocins have cytotoxic activity (e.g. “bacteriocide” effects), and thus can kill microbial organisms, for example bacteria, yeast, algae, synthetic microorganisms, and the like. Without being limited by any particular theory, bacteriocins can affect neutralization of a target microbial cell in a variety of ways. For example, in some cases a bacteriocin can permeabilize a cell wall, thus depolarizing the cell wall and interfering with respiration. Some bacteriocins can inhibit the reproduction of microbial organisms (e.g. “bacteriostatic” effects), for example bacteria, yeast, algae, synthetic micoorganisms, and the like, for example by arresting the cell cycle.
[0052] Bacteriocins are diverse and widespread across the phyla that produce them, with heterogeneity in gene cluster organization, bio-synthetic machinery and peptide structures. Bacteriocins are usually classified into two major groups: those with post- translational modifications (class I) and unmodified peptides (class II). Bacteriocins are generally considered large molecule antimicrobials relative to small-molecule antibiotics but can range from bicyclic peptides, such as darobactin (965 Da) that is smaller than the clinically used colistin (~1,200 Da), to larger circular peptides such as pumilarin (7,083 Da).
[0053] Class I bacteriocins undergo extensive modification with enzymes that may dehydrate residues and form inter-amino acid bonds, leading to their intricate and stable structures. Such modifications can result in peptides that are extensively cyclized and dehydrated, such as polyheterocyclic klebsazolicin and corynaridin, the latter of which contains eight dehydrated residues. Lantibiotics (lanthionine-containing antibiotics) are well-established class I bacteriocins. Type V lanthipeptides (lanthidins), areglycosylated and N-terminally dimethylated, as evident in the founding member cacaoidin.
[0054] Class II bacteriocin structures include simple linear a-helical peptides such as plantaricin EF, peptides with conserved complex secondary structures such as the four-helix peptide BacSp222 or linear peptides with specific conserved disulfides such as pediocin-like maltaricin CPN and actifensin. Class II bacteriocins do not feature the substantial modifications, such as dehydration and formation of unusual amino acids or linkages of class I, and previously included circular bacteriocins that are subject to N-terminal to C-terminal covalent linkage resulting in their simple circular shape and the microcins that are post-translationally conjugated to the siderophore (iron binding) moiety. Circular and siderophore-modified peptides can be referred to as class I bacteriocins, as they undergo modification despite relatively simple mature structures.
[0055] As would be known to one of ordinary skill in the art, bacteriocin gene clusters encode factors involved in bacteriocin production and protection, typically incorporating genes encoding a core peptide, modifying enzymes for modified peptides, export machinery, self-immunity proteins (also referred to as ‘bacteriocin immunity proteins’ ‘immunity proteins’ or ‘immunity modulators’) that protect the producer, and many clusters also contain genes encoding regulatory proteins. For example, Colicin V (ColV; also referred to herein as bacteriocin V and Mircrosin V) is a peptide secreted by some members of the Enterobacteriaceae to kill closely related bacterial cells, thereby reducing competition for essential nutrients. Four plasmid-borne genes(cvaA, cvaB, cvaC, and cvi) are involved in ColV synthesis, export, and immunity (see, e.g., FIG. 4). ColV is encoded by the cvaC gene and synthesized as a 103- amino-acid primary translation product with a conserved double glycine leader peptide at its N terminus. It is secreted through a dedicated ABC exporter composed of three proteins — CvaA, CvaB, and TolC — and processed to an 88-amino-acid polypeptide. ColV kills sensitive cells by disrupting the membrane potential. The cvi gene (encoding an immunity protein) encodes the cognate immunity protein of 78 residues, with two transmembrane helices, which alone is sufficient to fully protect a cell from the bactericidal activity of ColV. For additional information related to bacteriocins, see, e.g., Cotter et al., Nat Rev Microbiol. 2005 Oct; 3(10): 777-88; Sugrue et al., Nat Rev Microbiol. 2024 Sep;22(9):556-571; and Gerard etal., J Bacteriol. 2005 Mar;187(6):1945-50.
[0056] Additional non-limiting examples of gene bacteriocin gene clusters include Microcin S (also referred to as bacteriocin S), Microcin N (also referred to as bacteriocin N),Microcin J25 (also referred to as bacteriocin J25), and Microcin B17 (also referred to as bacteriocin B17), and are depicted in FIG. 4. While these examples are from E. coli, bacteriocins come from diverse prokaryotes - any convenient bacteriocin, immunity protein, bacteriocin gene cluster, etc. can be used.
[0057] A large number of bacteriocins will be known to one of ordinary skill in the art, and any convenient bacteriocin can be used. For examples of bacteriocins that can be used with the methods and compositions disclosed herein, see, e.g., US patent application publications US20170266306, US20200263221, US20210070812, and US20220017573, as well as US Patent Number 9333227, all of which are incorporated herein by reference for their disclosures related to bacteriocins, including bacteriocin types and sequences.
[0058] In some embodiments, a subject bacteriocin is a Class I bacteriocin (e.g., Microcins J25 and B17 (also referred to herein as Bacteriocins J25 and B17). In some embodiments, a subject bacteriocin is a Class II bacteriocin (e.g., Microcins V, S, or N (also referred to herein as Bacteriocins V, S, or N).
[0059] Examples of bacteriocins that can be used with the methods and compositions disclosed herein include, but are not limited to: bacteriocin V, B17, J25, N, and S (See, e.g., FIG. 5):>Bacteriocin_V (also known as (colicin V and Microcin V) MRTLTLNELDSVSGGASGRDIAMAIGTLSGQFVAGGIGAAAGGVAGGAIYDYASTHK PNPAMSPSGLGGTIKQKPEGIPSEAWNYAAGRLCNWSPNNLSDVCL(SEQ ID NO: 51)>Bacteriocin_B17 MELKASEFGVVLSVDALKLSRQSPLGVGIGGGGGGGGGGSCGGQGGGCGGCSNG CSGGNGGSGGSGSHI(SEQ ID NO: 52)>Bacteriocin_J25 MIKHFHFNKLSSGKKNNVPSPAKGVIQIKKSASQLTKGGAGHVPEYFVGIGTPISFYG(SEQ ID NO: 53)>Bacteriocin_N MRELDREELNCVGGAGDPLADPNSQIVRQIMSNAAWGAAFGARGGLGGMAVGAA GGVTQTVLQGAAAHMPVNVPIPKVPMGPSWNGSKG(SEQ ID NO: 54)>Bacteriocin_S MSNIRELSFDEIALVSGGNANSNYEGGGSRSRNTGARNSLGRNAPTHIYSDPSTVKC ANAVFSGMVGGAIKGGPVGMTRGTIGGAVIGQCLSGGGNGNGGGNRAGSSNCSG SNVGGTCSR(SEQ ID NO: 55)
[0060] Thus, in some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, or S (i.e. , any one of SEQ ID NOs: 51-55). In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, or S (i.e., any one of SEQ ID NOs: 51-55). In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, or S (i.e., any one of SEQ ID NOs: 51-55). In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, or S (i.e., any one of SEQ ID NOs: 51-55). In some cases, a bacteriocin is bacteriocin V, B17, J25, N, or S (i.e., comprises the amino acid sequence of any one of SEQ ID NOs: 51-55).
[0061] In some embodiments, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 51. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 51. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 51. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 51). In some cases, a bacteriocin comprises the amino acid sequence of SEQ ID NO: 51.
[0062] In some embodiments, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more,98% or more, 99% or more, or 100%) identical to SEQ ID NO: 52. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 52. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 52. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 52). In some cases, a bacteriocin comprises the amino acid sequence of SEQ ID NO: 52.
[0063] In some embodiments, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 53. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 53. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 53. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 53). In some cases, a bacteriocin comprises the amino acid sequence of SEQ ID NO: 53.
[0064] In some embodiments, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 54. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 54. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 54. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 54). In some cases, a bacteriocin comprises the amino acid sequence of SEQ ID NO: 54.
[0065] In some embodiments, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 55. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% ormore, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 55. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 55. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 55). In some cases, a bacteriocin comprises the amino acid sequence of SEQ ID NO: 55.
[0066] Examples of bacteriocins that can be used with the methods and compositions disclosed herein also include, but are not limited to: bacteriocin 21 (from E. faecalis), bacteriocin 31 (from E. faecalis), bacteriocin 32 (from E. faecium), bacteriocin 43 (from E. faecium), bacteriocin 51 (from E. faecium), bacteriocin L50A / L50B (from E. faecium), bacteriocin EntQ (from E. faecium), bacteriocin EntA (from E. faecium), bacteriocin EntP (from E. faecium), bacteriocin EntB (from E. faecium), bacteriocin HirJM79 (from E. faecium), bacteriocin 35 (from Enterococcus mundtii), bacteriocin 41 (from Enterococcus faecalis), and bacteriocin RC714 (from E. faecium). Example bacteriocins also include any bacteriocin produced by Bifidobacterium. See, e.g., Yu et al., 2023, Appl Environ Microbiol 89:e00979-23. Thus, in some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin is bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced byBifidobacterium.
[0067] In some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin is bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714.
[0068] In some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin is bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium.
[0069] Additional examples of bacteriocins that can be used with the methods and compositions disclosed herein include, but are not limited to: Kawaguchipeptin B (Microcystis aeruginosa), Pallidocin (Aeribacillus pallidus), Kunkecin A (Apilactobacillus kunkeei), Roseocin (Streptomyces roseosporus), Cacaoidin (Streptomyces cacaoi), Ubonodin (Burkholderia ubonensis), Corynaridin (Corynebacterium lactis), Ruminococcin C (Ruminococcus gnavus), Pumilarin (Bacillus pumilus), MccH47 (E. coli), Maltaricin CPN (Carnobacterium maltaromaticum CRN), Plantaricin EF (Lactobacillus plantarum), Bactofencin (Lactobacillus salivarius), Actifensin (Actinomyces ruminicola), BacSp222 (Staphylococcus pseudintermedius), Epidermidin (Staphylococcus epidermidis), Circularin A (Clostridium beijerinckii ATCC 25752), Closticin 574 (Clostridium tyrobutyricum), Nisin (Lactococcus lactis), and Thuricin CD (Bacillus thuringiensis).
[0070] Thus, in some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, or Thuricin CD. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, or Thuricin CD. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, or Thuricin CD. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, or Thuricin CD. In some cases, a bacteriocin is Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, RuminococcinC, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, orThuricin CD.
[0071] In some cases, a bacteriocin comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, Thuricin CD, bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, Thuricin CD, bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, Thuricin CD, bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, Thuricin CD, bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin is Kawaguchipeptin B, Pallidocin, Kunkecin A, Roseocin, Cacaoidin, Ubonodin, Corynaridin, Ruminococcin C, Pumilarin, MccH47, Maltaricin CPN, Plantaricin EF, Bactofencin, Actifensin, BacSp222, Epidermidin, Circularin A, Closticin 574, Nisin, Thuricin CD, bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B,EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium.
[0072] Examples of DNA sequences encoding a bacteriocin unit include, but are not limited to those of SEQ ID NOs: 62-64. SEQ ID NO: 62 encodes a re-factored bacteriocin V operably linked to an ATC-inducible promoter; SEQ ID NO: 63 encodes a re-factored bacteriocin B17 operably linked to an ATC-inducible promoter; SEQ ID NO: 64 encodes a re-factored bacteriocin J25 operably linked to an ATC-inducible promoter; while SEQ ID NO: 61 encodes an example of a therapeutic unit, in this particular example a chloramphenicol resistance gene & GFP fluorescence marker.
[0073] Bacteriocin immunity proteins (also referred to herein simply as “immunity proteins”) are known in the art, and any convenient immunity protein can be used, e.g., which can depend on which bacteriocin protein will be used. Examples of immunity proteins that can be used with the methods and compositions disclosed herein include, but are not limited to those for bacteriocins V, B17, J25, N, and S:>Bacteriocin_V immunity protein (immunity gene eval):MDRKRTKLELLFAFIINATAIYIALAIYDCVFRGKDFLSMHTFCFSALMSAICYFVGDNY YSISDKIKRRSYENSDSK (SEQ ID NO: 71)>Bacteriocin_B17 immunity protein (immunity gene mcbE):MVTLKMAIIFLMEQIRTPFSLIWTIMSPTVLFFFLHFNEIELHYGDTAWLGKQISWFVG YISFSVVLFNYCLYLVGRRESGFIATFVHNMDGRLLFIRSQLIASLIMSILYVFFFILVVL TGFQASPDYQIVMIILKSIYINAFMMVSLTFMASFRVTFQTASTIYSVLITVCMVSGIVSL KYNEGIVYWINQVNPIAIYSTILQSDQELSLMTIFFYSIMLIISIISALTFKTEPVWSSQ(SEQ ID NO: 72)>Bacteriocin_B17 immunity protein (immunity gene mcbF):MTIPLLEINSLSFSYKVNLPPVFNNLSLKIEQGELIGLLGENGAGKTTLFNLIRGGVSNY EGTLKRNFSGGELVSLPQVINLSGTLRNEEVLDLICCFNKLTKKQAWTELNHKWNDN FFIRYDKIRRKRTYTVSYGEKRWLIISLMLTLCKNARLFLLDEPTVGIDIQYRMMLWELI NKITADGKTVFFSTHIFDELTRDKIPFYMLSKNSINRYSDMSDFIQSNNETTPEKAFIKE VMGTGD (SEQ ID NO: 73)>Bacteriocin_B17 immunity protein (immunity gene mcbG)MDIIEKRITKRHLSESELSGVNYYNCIFERIQLDNFNFRDCEFEKCRFVNCSIKNLKLN FFKLIDCEFKDCLLQGVNAADIMFPCTFSLVNCDLRFVDFISLRLQKSIFLSCRFRDCL FEETDLRKSDFTGSEFNNTEFRHSDLSHCDFSMTEGLDINPEINRILSIKIPQEAGLKIL KRMGVVVGG (SEQ ID NO: 74)>Bacteriocin_J25 immunity protein (immunity gene mcjD) MERKQKNSLFNYIYSLMDVRGKFLFFSMLFITSLSSIIISISPLILAKITDLLSGSLSNFSY EYLVLLACLYMFCVISNKASVFLFMILQSSLRINMQKKMSLKYLRELYNENITNLSKNN AGYTTQSLNQASNDIYILVRNVSQNILSPVIQLISTIVVVLSTKDWFSAGVFFLYILVFVI FNTRLTGSLASLRKHSMDITLNSYSLLSDTVDNMIAAKKNNALRLISERYEDALTQEN NAQKKYWLLSSKVLLLNSLLAVILFGSVFIYNILGVLNGVVSIGHFIMITSYIILLSTPVEN IGALLSEIRQSMSSLAGFIQRHAENKATSPSIPFLNMERKLNLSIRELSFSYSDDKKILN SVSLDLFTGKMYSLTGPSGSGKSTLVKIISGYYKNYFGDIYLNDISLRNISDEDLNDAIY YLTQDDYIFMDTLRFNLRLANYDASENEIFKVLKLANLSVVNNEPVSLDTHLINRGNN YSGGQKQRISLARLFLRKPAIIIIDEATSALDYINESEILSSIRTHFPDALIINISHRINLLEC SDCVYVLNEGNIVASGHFRDLMVSNEYISGLASVTE (SEQ ID NO: 75)>Bacteriocin_N immunity protein (immunity gene mcnl) MKRNKLTRMSFLNFAFSPVFFSIMACYFIVWRNKRNEFVCNRLLSIIIISFLICFIYPWL NYKIEVKYYIFEQFYLFCFLSSLVAVVINLIVYFILYRRCI (SEQ ID NO: 76)>Bacteriocin_S immunity protein (immunity gene mcsl):M DERSSQFRYSKYSAI I FLAVVI ISTI VTLSPTFTLRYVGLDIAFFI VFITEI LISTLVYLFYL KEFPECRI KI RTDSATVKFSALSFLI 11 LIQLAVYCYRDYLYHYEPSQI N WITVLVMTLVV PYYEEIVYRACAFGFLRSIFKENIIIPCVITSLFFSLMHFQYYNVLDQSVLFVVSMLLLG VRIKSRSLFYPMLIHSGMNTFVILLNIQNIL (SEQ ID NO: 77)
[0074] Thus, in some cases, a “bacteriocin immunity protein” (also referred to herein simply as an “immunity protein”) comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, orS immunity proteins (i.e., any one of SEQ ID NOs: 71-77). In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, or S immunity proteins (i.e., any one of SEQ ID NOs: 71-77). In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, orS immunity proteins (i.e. , any one of SEQ ID NOs: 71-77). In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to bacteriocin V, B17, J25, N, orS immunity proteins (i.e., any one of SEQ ID NOs: 71-77). In some cases, a bacteriocin immunity protein comprises the amino acid sequence of any one of SEQ ID NOs: 71-77.
[0075] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 71. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 71. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 71. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 71. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 71.
[0076] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 72. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 72. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 72. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 72. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 72.
[0077] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 73. Insome cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 73. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 73. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 73. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 73.
[0078] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 74. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 74. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 74. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 74. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 74.
[0079] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 75. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 75. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 75. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 75. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 75.
[0080] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more,97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 76. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 76. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 76. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 76. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 76.
[0081] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 77. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 77. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 77. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to SEQ ID NO: 77. In some cases, a bacteriocin immunity protein comprises the amino acid sequence of SEQ ID NO: 77.
[0082] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or anybacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein is an immunity protein for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium.
[0083] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714. In some cases, a bacteriocin immunity protein is an immunity protein for bacteriocin 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, or RC714.
[0084] In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 80% or more (e.g., 85% or more, 90% or more, 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 90% or more (e.g., 92% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB,HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein comprises an amino acid sequence that is 95% or more (e.g., 97% or more, 98% or more, 99% or more, or 100%) identical to the immunity protein(s) for bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium. In some cases, a bacteriocin immunity protein is an immunity protein for bacteriocin V, B17, J25, N, S, 21, 31, 32, 43, 51, L50A / L50B, EntQ, EntA, EntP, EntB, HirJM79, 35, 41, RC714, or any bacteriocin produced by Bifidobacterium.Nucleic Acids
[0085] In some embodiments, a nucleic acid encoding a bacteriocin and an immunity protein as described herein is provided. In some embodiments, the nucleic acid is an expression vector. In some embodiments, the nucleic acid or expression vector is in a microbial cell.
[0086] The skilled artisan will readily understand that more than one nucleotide sequence can encode a particular polypeptide. For example, the genetic code is degenerate, and moreover, codon usage can vary based on the particular organism in which the gene product is being expressed. In some embodiments, a nucleotide sequence encoding a bacteriocin and / or a nucleotide sequence encoding an immunity protein is selected based on the codon usage of the organism expressing the protein(s). In some embodiments, a nucleic acid encoding a bacteriocin and / or an immunity protein is codon optimized based on the particular organism expressing the bacteriocin and / or the immunity protein.
[0087] Example nucleic acids (e.g., those that include a bacteriocin unit, and therefore encode at least a bacteriocin and an immunity protein) include, but are not limited to, plasmids, viruses (including phages), transposable elements, conjugative transposons, phage satellites, and phage plasmids.Promoters
[0088] In some embodiments, the bacteriocin protein and the immunity protein are expressed from the same promoter (i.e. , the nucleotide sequence encoding the bacteriocin protein and the nucleotide sequence encoding the immunity protein are operably linked to the same promoter) (e.g., in some cases the same single promoter as part of an operon). In some embodiments, the bacteriocin protein and the immunity protein are expressed from different promoters (i.e., the nucleotide sequence encoding the bacteriocin protein and the nucleotide sequence encoding the immunity protein are operably linked to different promoters). In some embodiments, the bacteriocin protein and the immunity protein are expressed from an inducible promoter (e.g., an anhydrotetracycline (ATC) inducible promoter). As such, in some embodiments, the bacteriocin protein is expressed from an inducible promoter (e.g., an anhydrotetracycline (ATC) inducible promoter). In some embodiments, the immunity protein is expressed from an inducible promoter (e.g., an anhydrotetracycline (ATC) inducible promoter). In some embodiments, the bacteriocin protein and the immunity protein are expressed from an inducible promoter (e.g., an anhydrotetracycline (ATC) inducible promoter).
[0089] In some embodiments, a single “coding” polynucleotide is under the control of a single promoter. In some embodiments, two or more “coding” polynucleotides are under the control of a single promoter, for example two, three, four, five, six, seven, eight, nine, or ten polynucleotides. As such, in some embodiments, a “cocktail” of different bacteriocins can be produced by a single microbial organism. In some embodiments, a bacteriocin polynucleotide is under the control of a promoter. In some embodiments, an immunity modulator is under the control of a promoter. In some embodiments, a polynucleotide encoding a desired gene product is under the control of a promoter. In some embodiments, the bacteriocin polynucleotide and the polynucleotide encoding a desired gene product are under the control of the same promoter. In some embodiments, a bacteriocin polynucleotide and the polynucleotide encoding a desired gene product are under the control of different promoters. In some embodiments, the bacteriocin protein and the immunity protein are expressed from the same promoter. In some embodiments, the bacteriocin protein and the immunity protein are expressed from different promoters.
[0090] In some cases, a bacteriocin unit includes a bacteriocin gene cluster. For example,Colicin V (ColV; also referred to herein as bacteriocin V and Mircrosin V, is a peptidesecreted by some members of the Enterobacteriaceae to kill closely related bacterial cells, thereby reducing competition for essential nutrients. Four plasmid-borne genes cvaA, cvaB, cvaC, and cvi) are involved in ColV synthesis, export, and immunity (see, e.g., FIG. 4). In some cases, the genes that make up a bacteriocin gene cluster are expressed from different promoters, i.e., are not part of an operon. In some embodiments a nucleic acid is engineered so that the genes that make up a bacteriocin gene cluster are oriented in the same 5’ to 3’ direction and are operably linked to the same single promoter. Thus, such a gene cluster has been engineered so that all of the genes of the gene cluster are part of the same operon, i.e., a ribosome binding site (RBS) (such as a Shine-Dalgarno sequence) precedes each gene (each protein-coding sequence) of the cluster so that each encoded protein can be expressed from the same single transcript (see, e.g., FIG. 1A). An engineered gene cluster in which the genes have been engineered to be part of the same operon (all operably linked to the same single promoter, and including an RBS before each of the protein-coding sequences) are referred to herein as a “refactored bacteriocin cluster.” In some cases, the bacteriocin unit includes a bacteriocin gene cluster that is a refactored bacteriocin cluster (see, e.g., FIG. 4).
[0091] A variety of promoters are suitable for use (e.g., to drive expression of a bacteriocin protein, an immunity protein, any protein encoded by a bacteriocin gene cluster, or any expression product encoded as part of a therapeutic unit described herein). A promoter can be a constitutively active promoter (a constitutive promoter) (e.g., a promoter that is constitutively in an active / “ON” state), it may be an inducible promoter (e.g., a promoter whose state, active / “ON” or inactive / “OFF”), or is controlled by an external stimulus, e.g., the presence of a particular temperature, compound, or protein. Examples of inducible promoters include, but are not limited to promoters induced by anhydrotetracycline (ATC), isopropyl-p-d-thiogalactoside (IPTG), arabinose (Ara), rhamnose (Rha), cuminic acid (Cuma), 2,4-Diacetylphophloroglucinol (DAPG), vanillic acid (Van), arabinogalactan (AG), chondroitin sulfate (ChS), crystal Violet (CV), and other cationic dyes. Additional examples include those that respond to bacterial- or human-derived biomolecules such as gut inflammation biomarkers such as tetrathionate, thiosulfate, or extracellular adenosine triphosphate (eATP). See, e.g., Meyer et al., Nat Chem Biol 15, 196-204 (2019); Ruegg et al., Nat Commun 9, 3617 (2018); Mimee et al., Cell Syst. 2016 Mar 23;2(3):214; Zou et al., Cell Host Microbe. 2023 Feb 8;31 (2): 199-212. e5). Additional promoters can be found in, e.g., US Patent No. 9333227, which is incorporated herein by reference for itsteachings related to promoters. In some cases, a subject promoter is an ATC, IPTG, Ara, Rha, Cuma, DAPG, Van, AG, ChS, CV, tetrathionate, thiosulfate, or eATP inducible promoter. In some cases, a subject promoter is and ATC inducible promoter.
[0092] Examples of inducible promoters include, but are not limited to, heat shock promoter, tetracycline-regulated promoter, metal-regulated promoter, etc. Inducible promoters can therefore be regulated by molecules including, but not limited to, doxycycline; IPTG; etc. Inducible promoters suitable for use include any inducible promoter described herein or known to one of ordinary skill in the art. Examples of inducible promoters include, without limitation, chemically / biochemically-regulated and physically-regulated promoters such as alcohol-regulated promoters, tetracycline- regulated promoters (e.g., anhydrotetracycline (ATC)-responsive promoters and other tetracycline-responsive promoter systems, which include a tetracycline repressor protein (tetR), a tetracycline operator sequence (tetO) and a tetracycline transactivator fusion protein (tTA)), temperature / heat-inducible promoters (e.g., heat shock promoters). Inducible promoters include sugar-inducible promoters (e.g., lactose-inducible promoters, arabinose-inducible promoters); amino acid-inducible promoters; alcohol-inducible promoters; and the like. Suitable promoters include, e.g., lactose-regulated systems (e.g., lactose operon systems, sugar-regulated systems, isopropyl-beta-D-thiogalactopyranoside (IPTG) inducible systems, arabinose regulated systems (e.g., arabinose operon systems, e.g., an ARA operon promoter, pBAD, PARA, portions thereof, combinations thereof and the like), synthetic amino acid regulated systems, fructose repressors, a tac promoter / operator (pTac), tryptophan promoters, PhoA promoters, recA promoters, proll promoters, cst-1 promoters, tetA promoters, cadA promoters, nar promoters, PL promoters, espA promoters, and the like, or combinations thereof. In certain cases, a promoter comprises a Lac-Z, or portions thereof. In some cases, a promoter comprises a Lac operon, or portions thereof. In some cases, an inducible promoter comprises an ARA operon promoter, or portions thereof. In certain embodiments an inducible promoter comprises an arabinose promoter or portions thereof. An arabinose promoter can be obtained from any suitable bacteria. In some cases, an inducible promoter comprises an arabinose operon of E. coli or B. subtilis. In some cases, an inducible promoter is activated by the presence of a sugar or an analog thereof. Non-limiting examples of sugars and sugar analogs include lactose, arabinose (e.g., L-arabinose), glucose, sucrose, fructose, IPTG, and the like. Suitable promoters include a T7 promoter; a pBAD promoter; a laclQ promoter; and the like. In some cases, the promoter is aJ23119 promoter (e.g., this promoter can be used to drive expression of a guide RNA in prokaryotic cells). Many bacterial promoters are known in the art; bacterial promoters can be found on the internet at parts(dot)igem(dot)org / promoters (see, e.g., constitutive promoters, cell signaling promoters, metal sensitive promoters, phage promoters, I IT Madras Stresskit promoters, and USTC logic promoters).Gene Editing Tools
[0093] The term “gene editing tool” as used herein to describe a tool for editing target DNA, e.g., genomic DNA, plasmid DNA, phage DNA, etc. In some cases (e.g., CRISPR-Cas editing tools), the tool includes a protein component (e.g., a CRISPR-Cas effector protein) and a nucleic acid component (e.g., a CRISPR-Cas guide RNA). The term “gene editing tool” can be used to refer to the actual tool, e.g., a CRISPR-Cas protein and a guide RNA, or can be used to refer to one or more nucleic acids encoding the gene editing tool, e.g., a nucleic acid (e.g., DNA or RNA) encoding a CRISRPR-Cas protein and a nucleic acid (e.g., DNA) encoding a CRISPR-Cas guide RNA. For example, in some embodiments, a nucleic acid (e.g., a plasmid) includes a gene editing tool, in which case the nucleic acid can include a nucleotide sequence encoding a gene editing protein (e.g., CRISPR-Cas effector protein such as Cas9 or Cas12) - and in some cases (e.g., when the gene editing tool includes a CRISPR- Cas effector protein such as Cas9 or Cas12) can also include a nucleotide sequence encoding a guide RNA. As such, the term gene editing tool refers to both protein and nucleic acid versions of a gene editing system. In some embodiments, a CRISPR-Cas effector protein and a guide RNA are provided on different nucleic acids (e.g., different plasmids), in which case the gene editing tool includes more than one nucleic acid molecule. In some cases, a CRISPR-Cas effector protein and a guide RNA are provided on the same nucleic acid (e.g., plasmid) (i.e., the nucleic acid includes nucleotide sequence encoding a CRISPR-Cas protein and a nucleotide sequence encoding a guide RNA). In some cases, the gene editing tool includes nucleotide sequences encoding polypeptides that form a CRISPR-associated transposase (CAST) complex, and includes a nucleotide sequence encoding a guide RNA, and in some such cases the donor DNA includes a transposon, flanked by recognition sites that are recognized by the CAST complex. Additional information related to a Donor DNA and related to transposon systems can be found below.
[0094] In some embodiments, the gene editing tool is a CRISPR-Cas system, which includes a CRISPR-Cas effector protein and a CRISPR-Cas guide RNA.
[0095] Examples of gene editing proteins (which are considered gene editing tools) include, but are not necessarily limited to: a CRISPR-Cas effector protein (which functions in combination with a guide RNA) (e.g., a Class 2 CRISPR-Cas effector protein), a meganuclease, a site-specific recombinase, a resolvase I integrase, a transposase, a transposon, and the like. In some cases, the gene editing protein is a site-specific recombinase (e.g., Cre recombinase, Dre recombinase, Flp recombinase, KD recombinase, B2 recombinase, B3 recombinase, R recombinase, Hin recombinase, Tre recombinase, PhiC31 integrase, Bxb1 integrase, R4 integrase, lambda integrase, HK022 integrase, HP1 integrase, and the like). In some cases, the gene editing protein is a meganulease (homing endonuclease) (e.g., I-Scel, l-Scell, l-Scelll, I- ScelV, l-SceV, l-SceVI, l-SceVII, l-Ceul, l-CeuAIIP, l-Crel, l-CrepsblP, l-CrepsbllP, I- CrepsblllP, l-CrepsblVP, l-Tlil, l-Ppol, Pl-Pspl, F-Scel, F-Scell, F-Suvl, F-Tevl, F- Tevll, l-Amal, l-Anil, l-Chul, l-Cmoel, l-Cpal, l-Cpall, l-Csml, l-Cvul, l-CvuAIP, l-Ddil, l-Ddill, l-Dirl, l-Dmol, l-Hmul, l-Hmull, l-HsNIP, l-Llal, l-Msol, l-Naal, l-Nanl, l-NcllP, I- NgrlP, l-Nitl, l-Njal, l-Nsp236IP, l-Pakl, l-PbolP, l-PculP, l-PcuAI, l-PcuVI, l-PgrIP, I- PoblP, l-Porl, l-PorllP, l-PbpIP, 1-SpBetalP, l-Scal, l-SexlP, l-SnelP, l-Spoml, I- SpomCP, l-SpomlP, l-SpomllP, l-SqulP, l-Ssp6803l, 1-SthPhiJP, l-SthPhiST3P, I- SthPhiSTe3bP, l-TdelP, I-Tevl, l-Tevll, l-Tevlll, l-UarAP, l-UarHGPAIP, I- UarHGPA13P, l-VinIP, l-ZbilP, Pl-Mtul, PI-MtuHIP PI-MtuHIIP, Pl-Pful, Pl-Pfull, Pl- Pkol, Pl-Pkoll, PI-Rma43812IP, PI-SpBetalP, Pl-Scel, Pl-Tful, Pl-Tfull, Pl-Thyl, Pl- Tlil, Pl-Tlil I, and the like). In some cases, the gene editing protein is a resolvase and / or invertase (e.g., Gin, Hin, y53, Tn3, Sin, Beta, and the like). In some cases, the gene editing protein is a transposon (e.g., bacterial transposons such as Tn3, Tn5, Tn7, Tn9, Tn 10, Tn903, Tn 1681, and the like; eukaryotic transposons such as Tc1 / mariner super family transposons, PiggyBac superfamily transposons, hAT superfamily transposons, PiggyBac, Sleeping Beauty, Frog Prince, Minos, Himarl, and the like). In some cases, a gene editing protein is a zinc finger nuclease (ZFN) or a transcription activator like effector nuclease (TALEN).
[0096] With regard to CRISPR-Cas proteins, in class 2 CRISPR systems, the functions of the effector complex (e.g., the cleavage of target DNA) are carried out by a single protein (which can be referred to as a CRISPR-Cas effector protein) - where the natural protein is an endonuclease (e.g., see Zetsche etal, Cell. 2015 Oct 22;163(3):759-71; Makarova et al, Nat Rev Microbiol. 2015 Nov;13(11):722-36; Shmakov et al., Mol Cell.2015 Nov 5;60(3):385-97; Shmakov et al., Nat Rev Microbiol. 2017 Mar; 15(3): 169- 182: “Diversity and evolution of class 2 CRISPR-Cas systems”; Koonin et al., Curr Opin Microbiol. 2017 Jun:37:67-78; and Makarova et al., Nat Rev Microbiol. 2020 Feb;18(2):67-83;). As such, the term “class 2 CRISPR-Cas protein” or “CRISPR-Cas effector protein” is used herein to encompass the effector protein from class 2 CRISPR systems - for example, type II CRISPR-Cas proteins (e.g., Cas9) and type V CRISPR-Cas proteins (e g., Cpf1 / Cas12a, C2c1 / Cas12b, C2C3 / Cas12c, Cas12d / CasY, Cas12e / CasX). Class 2 CRISPR-Cas effector proteins include type II and type CRISPR-Cas proteins, but the term is also meant to encompass any class 2 CRISPR-Cas protein suitable for binding to a corresponding guide RNA and forming a ribonucleoprotein (RNP) complex.
[0097] In some cases, a CRISPR-Cas effector protein (e.g., a Cas9 protein) has reduced catalytic activity (e.g., a Cas9 protein with nickase activity, also referred to as nCas9)). For example, when a Cas9 protein has a mutation at one or more amino acid positions corresponding to D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or a A987 (e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), the variant Cas9 protein can still bind to target DNA in a site-specific manner (because it is still guided to a target DNA sequence by a guide RNA) as long as it retains the ability to interact with the guide RNA. In some cases, Cas9 protein of a subject Cas9 fusion polypeptide is a nickase (e.g., cleaves one strand of a double stranded target nucleic acid but not the other strand) (e.g., the Cas9 protein can be a nickase, e.g., can include one or more amino acid mutations that make it a nickase). For example, in some cases, Cas9 protein of a subject Cas9 fusion polypeptide has a mutation in a catalytic domain (e.g., a mutation in a RuvC or HNH domain).
[0098] For example, in some cases, a CRISPR-Cas effector protein (e.g., a Cas9 protein) can cleave the complementary strand of a target nucleic acid but has reduced ability to cleave the non-complementary strand of a target nucleic acid. For example, the Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some cases, a Cas9 protein has a mutation at residue D10 (e.g., D10A, aspartate to alanine) and can therefore cleave the complementary strand of a double stranded target nucleic acid but has reduced ability to cleave the non-complementary strand of a double stranded target nucleic acid (thus resulting in a single strand break (SSB) instead of a double strand break (DSB) when the variant Cas9 protein cleaves a double stranded target nucleic acid)(see, for example, Jinek et al., Science. 2012 Aug 17;337(6096):816-21). Examples of such amino acid positions in a RuvC domain can include: D10, G12, G17, E762, H982, H983, A984, D986, and / or A987 (e.g., D10A, G12A, G17A, E762A, H982A, H983A, A984A, and / or D986A).
[0099] In some cases, a CRISPR-Cas effector protein (e.g., a Cas9 protein) can cleave the non-complementary strand of a target nucleic acid but has reduced ability to cleave the complementary strand of the target nucleic acid. For example, the Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain. Thus, the Cas9 protein can be a nickase that cleaves the non- complementary strand, but does not cleave the complementary strand (e.g., does not cleave a single stranded target nucleic acid). As a non-limiting example, in some embodiments, the Cas9 protein has a mutation at position H840 (e.g., an H840A mutation, histidine to alanine) and can therefore cleave the non-complementary strand of the target nucleic acid but has reduced ability to cleave (e.g., does not cleave) the complementary strand of the target nucleic acid. Such a Cas9 protein has a reduced ability to cleave a target nucleic acid (e.g., a single stranded target nucleic acid). Examples of such amino acid positions in an HNH domain can include: H840, N854, and / or N863 (e.g., H840A, N854A, and / or N863A).
[0100] In some cases, a CRISPR-Cas effector protein (e.g., a Cas9 protein) is a variant. In some cases, such a variant is a high fidelity (HF) protein such as a HF Cas9 protein (also referred to as SpCas9-HF1 or HF1 - and also -HF2, -HF3, -HF4) (e.g., see Kleinstiver et al. (2016) Nature 529:490). For example, amino acids N497, R661, Q695, and Q926 can be substituted, e.g., with alanine. In some cases, a suitable parent Cas9 protein exhibits altered PAM specificity. See, e.g., Kleinstiver et al.(2015) Nature 523:481. Additional examples of Cas9 variants that can be used, include, but are not limited to: HiFiCas9 (e.g., R691A), eSpCas9 (e.g., K810A, K1003A, R1060A), eSpCas9 (e.g., D1135E), HypaCas9 (e.g., N692A, M694A, Q695A, H698A), xCas9 (e.g., E108G, S217A, A262T, S409I, E480K, E543D, M694I, E1219V), Sniper-Cas9 (e.g., F539S, M763I, K890N), evoCas9 (e.g., M495V, Y515N, K526E, R661Q), SpartaCas (e.g., D23A, T67L, Y128V, D1251G), LZ3Cas9 (e.g., N690C, T769I, G915M, N980K), miCas9 (e.g., SV40 NLS linker fused with brex27 motif), SuperFi-Cas9 (e.g., Y1010D, Y1013D, Y1016D, V1018D, R1019D, Q1027D, K1031D) (see, e.g., Allemailem et al, Int J Mol Sci. 2023 Apr 11;24(8):7052.
[0101] In some cases, the CRISPR-Cas effector protein is a Cas9. In some cases, the CRISPR-Cas effector protein is a Cas 12 (e.g., Cas12a).
[0102] In some cases, the Cas9 is an iGeoCas9 (see, e.g., international patent publication WO2024112479, which is incorporated herein by reference for such disclosure). In some cases, the CRISPR-Cas effector protein is an iGeoCas9, but with the following mutations: E149G, T182I, N206D, P466Q, Q817R, E843K, E884G, and K908R (see, e.g., Chen et al., bioRxiv. Preprint. 2023 Nov 15: doi: 10.1101 / 2023.11.15.566339).
[0103] In some cases, the CRISPR-Cas effector protein comprises an amino acid sequence that is iGeoCas9, but with the following mutations: E149G, T182I, N206D, P466Q, E843K, E884G, K908R, T1015A, and D1017N (see, e.g., Chen et al., bioRxiv.Preprint. 2023 Nov 15: doi: 10.1101 / 2023.11.15.566339).
[0104] In some cases, the CRISPR-Cas effector protein is a Type II CRISPR-Cas effector polypeptide. In some cases, the CRISPR-Cas effector protein is a Cas9 protein. In some cases, the Cas9 protein is a Streptococcus pyogenes Cas9 (spyCas9) protein. In some cases, the Cas9 protein is a Staphylococcus aureus Cas9 (saCas9) protein. In some cases, the Cas9 protein is a Geobacillus thermodenitrificans Cas9 polypeptide. In some cases, the Cas9 protein is a Neisseria meningitidis Cas9 (NmCas9) polypeptide.Guide RNA
[0105] A nucleic acid that binds to a class 2 CRISPR-Cas effector protein (e.g., a Cas9 protein; a Cas12 protein; etc.) (thereby forming a ribonucleoprotein complex (RNP)) and targets the complex to a specific location within a target nucleic acid is referred to herein as a “guide RNA” or “CRISPR-Cas guide nucleic acid” or “CRISPR-Cas guide RNA” or a “gRNA” or simply a “guide.” A guide RNA can be used to guide the protein to the target sequence. It is to be understood that in some cases, a hybrid DNA / RNA can be made such that a guide RNA includes DNA bases in addition to RNA bases - but the term “guide RNA” is still used herein to encompass such hybrid molecules.
[0106] A guide RNA provides target specificity to the complex (the RNP complex) by including a targeting segment, which includes a “guide sequence” (also referred to as a “targeting sequence” or a “spacer”), which is a nucleotide sequence that is complementary to (and hybridizes to) a sequence of a target nucleic acid, e.g., a target DNA (and thereby can be said to “target” a specific sequence or “target” a specific gene). The guide sequence can be changed each time a new target sequence is selected. As will be understood by one of ordinary skill in the art, a guide sequence of a guide RNA can be modified (e.g., by genetic engineering) / designed tohybridize to any desired target sequence (e.g., while taking the PAM into account, e.g., when targeting a dsDNA target) within a target nucleic acid.
[0107] A guide RNA also includes a portion that interacts with (binds to) the CRISPR-Cas effector protein. Because this region does not need to change each time a new target sequence is selected, this region is referred to as a “constant region” or “scaffold” (or “handle”). A guide RNA can be referred to by the protein to which it corresponds. For example, when a CRISPR-Cas effector protein is a Cas9 protein, the corresponding guide RNA can be referred to as a “Cas9 guide RNA.” Likewise, as another example, when a CRISPR-Cas effector protein is a Cas12a protein, the corresponding guide RNA can be referred to as a “Cas12a guide RNA.”
[0108] As will be known to one of ordinary skill in the art, in some embodiments, a guide RNA includes two separate nucleic acid molecules: an “activator” and a “targeter” (or tracrRNA and crRNA) and is referred to as a “dual guide RNA”, a “double-molecule guide RNA”, a “two-molecule guide RNA”, or a “dgRNA.” In some embodiments, the guide RNA is one molecule (e.g., for some class 2 CRISPR-Cas proteins, the corresponding natural guide RNA is a single molecule; and in some cases, an activator and targeter can be covalently linked to one another, e.g., via intervening nucleotides), and the guide RNA is referred to as a “single guide RNA”, a “singlemolecule guide RNA,” a “one-molecule guide RNA”, or simply “sgRNA.”
[0109] As would be understood to one of ordinary skill in the art, the guide RNA can be introduced into a cell as an RNA (or as a DNA / RNA hybrid) or can be introduced as a nucleic acid encoding the RNA (e.g., a DNA such as an expression vector such as a plasmid or phage), in which case the cell transcribes the RNA from the introduced DNA. In some cases, the nucleotide sequence encoding the guide RNA is operably linked to a promoter (e.g., a J23119 promoter for expression in prokaryotic cells - other examples of promoters include, but are not limited to: J23101, J23105, J23114, MAGIC-1, T7-lacO, Lac, and the like - see, e.g., US. Patent No. 10,947,534, which incorporated herein by reference). In some cases, one or more guide RNAs (e.g., 1, 2, 3, 4, 5, 6, 1-10, 1-8, 1-6, 1-5, 1-4, 1-3, 2-10, 2-8, 2-6, 2-5, 2-4, 3-10, 3-8, 3-6, 3-5, two or more, three or more, four or more, or five or more) (or nucleotide sequences that encode said guide RNAs) can be introduced into the same cell (e.g., to target different sequences of the same target nucleic, to target different target nucleic acids, etc.).
[0110] Scaffold sequences for various CRISPR-Cas guide RNAs are known in the art. For example, in some cases, the portion of the targeter-RNA (e.g., crRNA) thatcontributes to the scaffold (i.e., is 3’ of the guide sequence) (e.g., when using an S. pyogenes Cas9 protein) includes: 5’-GUUUUAGAGCUAUGCUGUUUUG-3' (SEQ ID NO: 81). In some cases, it includes: 5’-GUUUUAGAGCUA-3' (SEQ ID NO: 82). in some cases, the activator-RNA (e.g., tracrRNA) (e.g., when using an S. pyogenes Cas9 protein) includes: ’5- AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUG GCACCGAGUCGGUGCUU-3' (SEQ ID NO: 83). In some cases (e.g., when using an S. pyogenes Cas9 protein) a sgRNA includes 5’- GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCG-3' (SEQ ID NO: 84). In some cases (e.g., when using an S. pyogenes Cas9 protein) a sgRNA includes 5’- GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUG AAAAAGUGGCACCGAGUCGGUGCUU-3' (SEQ ID NO: 85). Mutations / variants of the above sequences can also be used and many suitable examples will be known to one of ordinary skill in the art.
[0111] Examples of crRNA repeat sequences (also known as the scaffold) for Cas12a proteins include:LbCas12a crRNA:5’ AAUUUCUACUAAGUGUAGAU 3’ (SEQ ID NO: 86) - [spacer] 3’ AsCas12a crRNA:5’ AAUUUCUACUCUUGUAGAU 3’ (SEQ ID NO: 87) - [spacer] 3’ FnCas12a crRNA:5’ AAUUUCUACUGUUGUAGAU 3’ (SEQ ID NO: 88) - [spacer] 3’ PmCas12a crRNA:5’ AAUUUCUACUAUUGUAGAU 3’ (SEQ ID NO: 89) - [spacer] 3’ M bCas 12a / M b2Cas 12a / M b3Cas 12a crR N A:5’ AAUUUCUACUGUUUGUAGAU 3’ (SEQ ID NO: 90) - [spacer] 3’TsCas12a crRNA5’ AAUUUCUACUGUUGUAGAU 3’ (SEQ ID NO: 91) - [spacer] 3’ BsCas12a crRNA5’ AAUUUCUACUAUUGUAGAU 3’ (SEQ ID NO: 92) - [spacer] 3’Protospacer adjacent motif (PAM)
[0112] A wild type CRISPR / Cas protein (e.g., Cas9, Cas12) normally has nuclease activity that cleaves a target nucleic acid (e.g., a double stranded DNA (dsDNA)) at a target site defined by the region of complementarity between the guide sequence of the guide RNA and the target nucleic acid. In some cases, site-specific targeting to the target nucleic acid occurs at locations determined by both (i) base-pairing complementarity between the guide nucleic acid and the target nucleic acid; and (ii) a short motif referred to as the “protospacer adjacent motif” (PAM) in the target nucleic acid. For example, when a Cas9 protein binds to (in some cases cleaves) a dsDNA target nucleic acid, the PAM sequence that is recognized (bound) by the Cas9 protein is present on the non-complementary strand (the strand that does not hybridize with the targeting segment of the guide nucleic acid) of the target DNA. In some cases, a PAM sequence has a length in a range of from 1 nt to 15 nt (e.g., 1 nt to 14 nt, 1 nt to 13 nt, 1 nt to 12 nt, 1 nt to 11 nt, 1 nt to 10 nt, 1 nt to 9 nt, 1 nt to 9 nt, 1 nt to 8 nt, 1 nt to 7 nt, 1 nt to 6 nt, 1 nt to 5 nt, 1 nt to 4 nt, 1 nt to 3 nt, 2 nt to 15 nt, 2 nt to 14 nt, 2 nt to 13 nt, 2 nt to 12 nt, 2 nt to 11 nt, 2 nt to 10 nt, 2 nt to 9 nt, 2 nt to 8 nt, 2 nt to 7 nt, 2 nt to 6 nt, 2 nt to 5 nt, 2 nt to 4 nt, 2 nt to 3 nt, 2 nt, or 3 nt).
[0113] As would be known by one of ordinary skill in the art, different CRISR-Cas effector proteins (e.g., Cas9, Cas12) from different species can have different PAM requirements (different PAM sequences, different location of the PAM relative to the target sequence, etc.), and PAM requirements for many different examples are known in the art. As an illustrative example, in some embodiments (e.g., when the Cas9 protein is derived from S. pyogenes or a closely related Cas9 is used; see for example, Chylinski et al., RNA Biol. 2013 May;10(5):726-37; and Jinek et al., Science.2012 Aug 17;337(6096):816-21; both of which are hereby incorporated by reference in their entirety), the PAM sequence can be NRG because the S. pyogenes Cas9 PAM (PAM sequence) is NAG or NGG (or NRG where “R” is A or G). For example, a Cas9 PAM sequence for S. pyogenes Cas9 can be: NGG, NAG, AGG, CGG, GGG, TGG, AAG, CAG, GAG, and TAG. In some cases, the PAM is NGG.
[0114] In some cases (e.g., when a Cas9 protein is derived from the Cas9 protein of Neisseria meningitidis or a closely related Cas9 is used), the PAM sequence (e.g., of a target nucleic acid) can be 5’-NNNNGANN-3’, 5’-NNNNGTTN-3’, 5’-NNNNGNNT-3’, 5’-NNNNGTNN-3’, 5’-NNNNGNTN-3’, or 5’-NNNNGATT-3’, where N is any nucleotide. In some embodiments (e.g., when a Cas9 protein is derived from Streptococcus thermophilus #1 ora closely related Cas9 is used), the PAM sequence(e.g., of a target nucleic acid) can be 5’-NNAGAA-3’, 5’-NNAGGA-3’, 5’-NNGGAA-3’, 5’-NNANAA-3’, or 5’-NNGGGA-3’ where N is any nucleotide. In some embodiments (e.g., when a Cas9 protein is derived from Treponema denticola (TD) or a closely related Cas9 is used), the PAM sequence (e.g., of a target nucleic acid) can be 5’- NAAAAN-3’, 5’-NAAAAC-3’, 5’-NAAANC-3’, 5’-NANAAC-3’, or 5’-NNAAAC-3’, where N is any nucleotide. As would be known by one of ordinary skill in the art, additional PAM sequences for other Cas9 proteins are known in the art and / or can readily be determined using bioinformatic analysis (e.g., analysis of genomic sequencing data), and / or routine experimentation. See Esvelt et al., Nat Methods. 2013 Nov;10(11):1116-21, for additional information. In some cases, the PAM sequence for S. aureus Cas9 is 5’-NNGRR(N)-3’ (in some cases 5’-NNGRRT-3’).
[0115] For additional information related to programmable gene editing tools (e.g., CRISPR- Cas RNA-guided proteins such as Cas9, CasX, CasY, Cas12, CRISPR-Cas guide RNAs, PAMs, and the like) refer to, for example, Dreier, et al., (2001) J Biol Chem 276:29466-78; Dreier, etal., (2000) J Mol Biol 303:489-502; Liu, et al., (2002) J Biol Chem 277:3850-6); Dreier, et al., (2005) J Biol Chem 280:35588-97; Jamieson, et al., (2003) Nature Rev Drug Discov 2:361-8; Durai, etal., (2005) Nucleic Acids Res 33:5978-90; Segal, (2002) Methods 26:76-83; Porteus and Carroll, (2005) Nat Biotechnol 23:967-73; Pabo, et al., (2001) Ann Rev Biochem 70:313-40; Wolfe, et al., (2000) Ann Rev Biophys Biomol Struct 29:183-212; Segal and Barbas, (2001) Curr Opin Biotechnol 12:632-7; Segal, et al., (2003) Biochemistry 42:2137-48; Beerli and Barbas, (2002) Nat Biotechnol 20:135-41; Carroll, et al., (2006) Nature Protocols 1:1329; Ordiz, et al., (2002) Proc Natl Acad Sci USA 99:13290-5; Guan, et al., (2002) Proc Natl Acad Sci USA 99:13296-301; Sanjana et al., Nature Protocols, 7:171-192 (2012); Zetsche et al, Cell. 2015 Oct 22;163(3):759-71 ; Makarova etal, Nat Rev Microbiol. 2015 Nov;13(11):722-36; Shmakov et al., Mol Cell. 2015 Nov 5;60(3):385- 97; Jinek et al., Science. 2012 Aug 17; 337(6096) :816-21 ; Chylinski et al., RNA Biol.2013 May;10(5):726-37; Ma et al., Biomed Res Int. 2013;2013:270805; Hou et al., Proc Natl Acad Sci U S A. 2013 Sep 24; 110(39): 15644-9; Jinek et al., Elife.2013;2:e00471; Pattanayak et al., Nat Biotechnol. 2013 Sep;31(9):839-43; Qi et al, Cell. 2013 Feb 28; 152(5): 1173-83; Wang et al., Cell. 2013 May 9;153(4):910-8; Auer et al., Genome Res. 2013 Oct 31; Chen et al., Nucleic Acids Res. 2013 Nov 1;41(20):e19; Cheng et al., Cell Res. 2013 Oct;23(10): 1163-71; Cho et al., Genetics.2013 Nov; 195(3): 1177-80; DiCarlo et al., Nucleic Acids Res. 2013 Apr;41(7):4336-43; Dickinson et al., Nat Methods. 2013 Oct; 10(10): 1028-34; Ebina et al., Sci Rep.2013;3:2510; Fujii et. al, Nucleic Acids Res. 2013 Nov 1;41(20):e187; Hu et al., Cell Res. 2013 Nov;23(11):1322-5; Jiang et al., Nucleic Acids Res. 2013 Nov 1;41(20):e188; Larson etal., Nat Protoc. 2013 Nov;8(11):2180-96; Mali et. at., Nat Methods. 2013 Oct;10(10):957-63; Nakayama et al., Genesis. 2013 Dec;51(12):835- 43; Ran etal., Nat Protoc. 2013 Nov;8(11):2281-308; Ran et al., Cell. 2013 Sep 12; 154(6): 1380-9; Upadhyay et al., G3 (Bethesda). 2013 Dec 9;3(12):2233-8; Walsh et al., Proc Natl Acad Sci U S A. 2013 Sep 24;110(39): 15514-5; Xie et al., Mol Plant.2013 Oct 9; Yang et al., Cell. 2013 Sep 12;154(6):1370-9; Briner et al., Mol Cell. 2014 Oct 23;56(2):333-9; Burstein et al., Nature. 2016 Dec 22 - Epub ahead of print; Gao et al., Nat Biotechnol. 2016 Jul 34(7):768-73; Shmakov etal., Nat Rev Microbiol. 2017 Mar;15(3):169-182; Makarova et al., Nat Rev Microbiol. 2020 Feb;18(2):67-83; as well as international patent application publication Nos. W02002099084; WOOO / 42219; WO02 / 42459; W02003062455; W003 / 080809; W005 / 014791; W005 / 084190;W008 / 021207; W009 / 042186; WO09 / 054985; and W010 / 065123; U.S. patent application publication Nos. 20030059767, 20030108880, 20140068797;20140170753; 20140179006; 20140179770; 20140186843; 20140186919;20140186958; 20140189896; 20140227787; 20140234972; 20140242664;20140242699; 20140242700; 20140242702; 20140248702; 20140256046;20140273037; 20140273226; 20140273230; 20140273231; 20140273232;20140273233; 20140273234; 20140273235; 20140287938; 20140295556;20140295557; 20140298547; 20140304853; 20140309487; 20140310828;20140310830; 20140315985; 20140335063; 20140335620; 20140342456;20140342457; 20140342458; 20140349400; 20140349405; 20140356867;20140356956; 20140356958; 20140356959; 20140357523; 20140357530;20140364333; 20140377868; 20150166983; and 20160208243; and U.S. Patent Nos.6,140,466; 6,511,808; 6,453,242 8,685,737; 8,906,616; 8,895,308; 8,889,418;8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; and 8,697,359; all of which are hereby incorporated by reference in their entirety.
[0116] In some cases, the CRISPR-Cas components (effector protein and guide RNA (gRNA) can be delivered as a nucleic acid encoding the components (e.g., the protein can be delivered as a protein or as a DNA or RNA encoding the protein, while a guide RNA can be delivered as the guide RNA or as a DNA encoding the guide RNA). Thus, an alternative mechanism to delivering the CRISPR-Cas components (effector protein and guide RNA (gRNA)) via nucleic acid (e.g., one or more plasmids), is the delivery via pre-made Cas9 / gRNA RNP complexes. For example, an active recombinantCRISPR-Cas effector protein (e.g., a Cas9 or a Cas12) can be purified from a heterologous system (e.g., E. coli) using known methods and guide RNA can be in vitro transcribed or chemically synthesized. These two elements can also be purchased from commercial vendors. The main advantage of this method is that it does not rely on the host transcription and translation machinery. The RNP complex is usually degraded shortly after transfection, avoiding the toxic effects of a continuous expression. An RNP can be delivered using any desired technique (e.g., electroporation), several of which will be known to one of ordinary skill in the art. In some cases, the CRISPR-Cas effector protein (e.g., a Cas9 or a Cas12) can be polymer-derivatize, e.g., via direct covalent modification of the protein with a cationic polymer (e.g., bPEI). The derivatized protein can then be complexed with guide RNA to generate nanosized complexes.
[0117] For additional information related to prokaryotic gene editing tools, see, e.g., Arroyo- Olarte et al., Microorganisms. 2021 Apr 15;9(4):844.
[0118] There are several methods other than CRISPR-based tools that can be used for DNA editing (e.g., genome editing) in prokaryotes, and some examples are listed above. Additional examples of non-CRISPR gene editing tools for use in prokaryotes include, but are not limited to suicide plasmids (for targeted gene disruption), a scarless system (e.g., using a homing endonuclease such as l-Scel), a recombineering system (e.g., a Lambda-Red system), and a Clostron system (AB resistance marker - disrupted) [see, e.g., Arroyo-Olarte etal., Microorganisms. 2021 Apr 15;9(4):844],
[0119] Suicide plasmids include a homologous sequence with the desired insertion, deletion or site-directed mutation coupled to a marker, usually an antibiotic resistance cassette -where the plasmid may harbor a transposon sequence that facilitates insertion into a target DNA in a recipient cell, e.g., after conjugation with a donor.
[0120] l-Scel is a homing endonuclease from Saccharomyces cerevisiae that targets an 18 bp asymmetric sequence (TAGGGATAACAGGGTAAT), cleaving both DNA strands to leave 3’-overhangs with a four base-pairs length, which can induce homologous recombination. Suicide plasmids can incorporate an l-Scel site between the mutant allele and the antibiotic resistance marker. The suicide plasmid is transformed into an E. coli strain already harboring an inducible plasmid for l-Scel expression. The suicide plasmid is then integrated into the genome and colonies are selected by their antibiotic resistance at the non-permissive temperature for plasmid replication.Induction of l-Scel cleaves the target gene locus, which is then repaired via native RecA-mediated homologous recombination, providing large enough homology arms(>500 bp). This can result either in a reversion to the wild-type chromosome or in a markerless allele replacement.
[0121] Lambda Red consists primarily of three proteins: a, [3, and y. a is an exonuclease (exo), which processively digests the 5’-ended strand of a dsDNA end. p (bet) binds to ssDNA and promotes strand annealing. Finally, y (gam) binds to the bacterial RecBCD enzyme (which degrades any linear DNA used as a template) and inhibits its activities. These proteins induce a “hyper-recombination” state in E. coli and other bacteria, in which recombination events between DNA species with as little as 35-50 bp of shared sequence occur at high frequency.
[0122] The ClosTron method utilizes an endogenous intron with transposon activity, a bacterial group II intron, as an insertional gene inactivation tool. These are broad host-range elements whose target specificity is determined largely by homology between intron RNA and target site DNA. Such introns can therefore be re-targeted by altering the sequence of an intron RNA-encoding plasmid. The ClosTron system uses an element derived from the broad host range LI.LtrB intron of Lactococcus lactis.Donor DNA
[0123] As noted above, the term “donor DNA” is used herein to describe a DNA that includes a bacteriocin unit, and is configured for integration of a nucleotide sequence of the DNA into a target DNA. Thus, when a bacteriocin unit of the subject disclosure is present on (as part of) a donor DNA, the bacteriocin unit is configured for integration into a target DNA. Examples of donor DNAs include, but are not limited to, transposons and HDR donor templates (also referred to herein as “donor templates”).Transposons and Transposon systems
[0124] One illustrative example of a donor DNA is one that includes a transposon, which integrates into a target DNA. For example, in some embodiments, a subject selfselection DNA editing system is a transposon system that includes: (i) a nucleotide sequence encoding polypeptides that form a CAST complex; (ii) a nucleotide sequence(s) encoding one or more guide RNAs; and (iii) a transposon (for insertion into a target DNA), which is flanked by CAST complex recognition sites. In such cases, parts (i) (CAST complex) and (ii) (guide RNA) together can be considered agene editing tool, while part (iii) (transposon) can be considered a donor DNA. As such, in some cases, a subject gene editing tool includes: (i) nucleotide sequences encoding polypeptides that form a CRISPR-associated transposase (CAST) complex, and (ii) a nucleotide sequence encoding a guide RNA (e.g., in some cases one that targets a target nucleotide sequence in a prokaryotic cell genome, i.e. , one that includes a nucleotide sequence that hybridizes to a target nucleotide sequence in a prokaryotic cell genome); while the donor DNA includes a transposon that is flanked by recognition sites recognized by the CAST complex.
[0125] In some cases, part (i) and part (ii) are present on different nucleic acids. In some cases, part (i) and part (ii) are present on the same nucleic acid. In some cases, part (i) and part (ii) are present on the same nucleic acid, and part (iii) is on a separate nucleic acid. In some cases, parts (i), (ii), and (iii) are all present on the same nucleic acid. In some embodiments, the nucleic acid is a plasmid, e.g., in some cases a conjugative plasmid.CAST
[0126] CRISPR-associated transposases (CASTs) include a CRISPR-associated polypeptide and one or more additional polypeptides that, in complex with one another, mediate transposition of a target transposon.ShCAST
[0127] In some cases, a CAST comprises: i) a Cas 12k polypeptide; ii) a TnsC polypeptide;iii) a TnsB polypeptide; and iv) a TniQ polypeptide. An example of such a CAST is a Scytonema hofmanni CAST (ShCAST).
[0128] A Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni Cas12k amino acid sequence of SEQ ID NO: 1. A Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from 500 amino acids to 639 amino acids (e.g., from 500 amino acids (aa) to 550 aa, from 550 aa to 575 aa, from 575 aa to 600 aa, from 600 aa to 625 aa, or from 625 aa to 639 aa) of the S. hofmanni Cas12k amino acid sequence of SEQ ID NO: 1. In some cases, the Cas12k polypeptide has a length of from about 600 amino acids to 650 amino acids (e.g., from 600 amino acids (aa) to 625 aa, orfrom 625 aa to 650 aa). In some cases, the Cas12k polypeptide has a length of 639 aa.
[0129] Non-limiting examples of other suitable Cas12k polypeptides are provided as SEQ ID NOs: 1 and 17-21. For example, a Cas12k polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas12k polypeptide amino acid sequences of SEQ ID NOs: 1 and 17-21.
[0130] A TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TnsB amino acid sequence of SEQ ID NO: 2. A TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 500 amino acids to 584 amino acids (e.g., from about 500 amino acids (aa) to 525 aa, from 525 aa to 550 aa, from 550 aa to 575 aa, or from 575 aa to 584 aa) of the S. hofmanni TnsB amino acid sequence of SEQ ID NO: 2. In some cases, the TnsB polypeptide has a length of from about 500 amino acids to about 600 amino acids (e.g., from about 500 amino acids (aa) to 525 aa, from 525 aa to 550 aa, from 550 aa to 575 aa, or from 575 aa to 600 aa). In some cases, the TnsB polypeptide has a length of 584 aa.
[0131] Non-limiting examples of other suitable TnsB polypeptides are provided as SEQ ID NOs: 2 and 12-18. For example, a TnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TnsB polypeptide amino acid sequences of SEQ ID NOs: 2 and 12-18.
[0132] A TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TnsC amino acid sequence of SEQ ID NO: 3. A TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 200 amino acids to 276 amino acids (e.g., from about 200 amino acids (aa) to 225 aa, from 225 aa to 250 aa or from 250 aa to 276 aa) of the S. hofmanni TnsC amino acid sequence of SEQ ID NO: 3. In some cases, the TnsC polypeptide has a length of from about 200 amino acids to 276 amino acids(e.g., from about 200 amino acids (aa) to 225 aa, from 225 aa to 250 aa or from 250 aa to 276 aa). In some cases, the TnsC polypeptide has a length of 276 aa.
[0133] Non-limiting examples of other suitable TnsC polypeptides are provided as SEQ ID NOs: 3 and 22-25. For example, a TnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TnsC polypeptide amino acid sequences of SEQ ID NOs: 3 and 22-25.
[0134] A TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the S. hofmanni TniQ amino acid sequence of SEQ ID NO: 4. A TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 100 amino acids to 167 amino acids (e.g., from 100 amino acids (aa) to 125 aa, from 125 aa to 150 aa, or from 150 aa to 167 aa) of the S. hofmanni TniQ amino acid sequence of SEQ ID NO: 4. In some cases, the TniQ polypeptide has a length of from about 100 amino acids to 167 amino acids (e.g., from 100 amino acids (aa) to 125 aa, from 125 aa to 150 aa, or from 150 aa to 167 aa). In some cases, the TniQ polypeptide has a length of 167 amino acids.
[0135] Non-limiting examples of other suitable TniQ polypeptides are provided as SEQ ID NOs: 4 and 26-29. For example, a TniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the TniQ polypeptide amino acid sequences of SEQ ID NOs: 4 and 26-29.VcCAST
[0136] In some cases, a CAST comprises: i) a Cas6 polypeptide; ii) a Cas7 polypeptide; iii) a Cas8 polypeptide; iv) a TnsA polypeptide; v) a TnsB polypeptide; vi) a TnsC polypeptide; and vii) a TniQ polypeptide. An example of such a CAST is a Vibrio cholerae CAST (VcCAST).
[0137] A Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas6 polypeptide amino acid sequence of SEQ ID NO: 11. A Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 125 amino acids to 199 amino acids (e.g., from about 125 amino acids (aa) to 150 aa, from 150 aa to 175 aa, or from 175 aa to 199 aa) of the V. cholera Cas6 polypeptide amino acid sequence of SEQ ID NO: 11. A Cas6 polypeptide can have a length of from about 125 amino acids to 199 amino acids (e.g., from about 125 amino acids (aa) to 150 aa, from 150 aa to 175 aa, or from 175 aa to 199 aa). A Cas6 polypeptide can have a length of 199 aa.
[0138] Non-limiting examples of other suitable Cas6 polypeptides are provided as SEQ ID NOs: 11 and 42-44. For example, a Cas6 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas6 polypeptide amino acid sequences of SEQ ID NOs: 11 and 42-44.
[0139] A Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas7 polypeptide amino acid sequence of SEQ ID NO: 10. A Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 275 amino acids to 352 amino acids (e.g., from about 275 amino acids (aa) to 300 aa, from 300 aa to 325 aa, or from 325 aa to 352 aa) of the V. cholerae Cas7 polypeptide amino acid sequence of SEQ ID NO: 10. A Cas7 polypeptide can have a length of from about 275 amino acids to 352 amino acids (e.g., from about 275 amino acids (aa) to 300 aa, from 300 aa to 325 aa, or from 325 aa to 352 aa). A Cas7 polypeptide can have a length of 352 aa.
[0140] Non-limiting examples of other suitable Cas7 polypeptides are provided as SEQ ID NOs: 10 and 45-47. For example, a Cas7 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas7 polypeptide amino acid sequences of SEQ ID NOs: 10 and 45-47.
[0141] A Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera Cas8 polypeptide amino acid sequence of SEQ ID NO: 9. A Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identityto a contiguous stretch of from about 575 amino acids to 640 amino acids (e.g., from about 575 amino acids (aa) to 600 aa, from 600 aa to 625 aa, or from 625 aa to 640 aa) of the V. cholerae Cas8 polypeptide amino acid sequence of SEQ ID NO: 9. A Cas8 polypeptide can have a length of from about 575 amino acids to 640 amino acids (e.g., from about 575 amino acids (aa) to 600 aa, from 600 aa to 625 aa, or from 625 aa to 640 aa). A Cas8 polypeptide can have a length of 640 aa.
[0142] Non-limiting examples of other suitable Cas8 polypeptides are provided as SEQ ID NOs: 9 and 48-50. For example, a Cas8 polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the Cas8 polypeptide amino acid sequences of SEQ ID NOs: 9 and 48-50.
[0143] A tnsA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsA polypeptide amino acid sequence of SEQ ID NO: 5. A tnsA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 150 amino acids to 222 amino acids (e.g., from about 150 amino acids (aa) to 175 aa, from 175 aa to 200 aa, or from 200 aa to 222 aa) of the tnsA amino acid sequence of SEQ ID NO: 5. A tnsA polypeptide can have a length of from about 150 amino acids to 222 amino acids (e.g., from about 150 amino acids (aa) to 175 aa, from 175 aa to 200 aa, or from 200 aa to 222 aa). A tnsA polypeptide can have a length of 222 amino acids.
[0144] Non-limiting examples of other suitable tnsA polypeptides are provided as SEQ ID NOs: 5 and 30-32. For example, a tnsA polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsA polypeptide amino acid sequences of SEQ ID NOs: 5 and 30-32.
[0145] A tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsB polypeptide amino acid sequence of SEQ ID NO: 6. A tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 525 amino acids to 603 amino acids (e.g., fromabout 525 amino acids (aa) to 550 aa, from 550 aa to 575 aa, or from 575 aa to 603 aa) of the tnsB amino acid sequence of SEQ ID NO: 6. A tnsB polypeptide can have a length of from about from about 525 amino acids to 603 amino acids (e.g., from about 525 amino acids (aa) to 550 aa, from 550 aa to 575 aa, or from 575 aa to 603 aa). A tnsB polypeptide can have a length of 603 amino acids.
[0146] Non-limiting examples of other suitable tnsB polypeptides are provided as SEQ ID NOs: 6 and 33-35. For example, a tnsB polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsB polypeptide amino acid sequences of SEQ ID NOs: 6 and 33-35.
[0147] A tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tnsC polypeptide amino acid sequence of SEQ ID NO: 7. A tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 225 amino acids to 330 amino acids (e.g., from about 225 amino acids (aa) to 250 aa, from 250 aa to 300 aa, or from 300 aa to 330 aa) of the tnsC amino acid sequence of SEQ ID NO: 7. A tnsC polypeptide can have a length of from about 225 amino acids to 330 amino acids (e.g., from about 225 amino acids (aa) to 250 aa, from 250 aa to 300 aa, or from 300 aa to 330 aa). A tnsC polypeptide can have a length of 330 amino acids.
[0148] Non-limiting examples of other suitable tnsC polypeptides are provided as SEQ ID NOs: 7 and 36-38. For example, a tnsC polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tnsC polypeptide amino acid sequences of SEQ ID NOs: 7 and 36-38.
[0149] A tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to the V. cholera tniQ polypeptide amino acid sequence of SEQ ID NO: 8. A tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to a contiguous stretch of from about 300 amino acids to 394 amino acids (e.g., from about 300 amino acids (aa) to 325 aa, from 325 aa to 350 aa, from 350 aa to 375 aa,or from 375 aa to 394 aa) of the tniQ amino acid sequence of SEQ ID NO: 8. A tniQ polypeptide can have a length of from about 300 amino acids to 394 amino acids (e.g., from about 300 amino acids (aa) to 325 aa, from 325 aa to 350 aa, from 350 aa to 375 aa, or from 375 aa to 394 aa). A tniQ polypeptide can have a length of 394 amino acids.
[0150] Non-limiting examples of other suitable tniQ polypeptides are provided as SEQ ID NOs: 8 and 3-41. For example, a tniQ polypeptide can comprise an amino acid sequence having at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% amino acid sequence identity to any one of the tniQ polypeptide amino acid sequences of SEQ ID NOs: 8 and 3-41.
[0151] The nucleotide sequence encoding the CAST complex polypeptides and / or the nucleotide sequence encoding the guide RNA can be operably linked to a promoter that is functional in a prokaryotic cell. In some cases, the nucleotide sequence encoding the CAST complex polypeptides is operably linked to a first promoter; and the nucleotide sequence encoding the guide RNA is operably linked to a second promoter. In some cases, the nucleotide sequence encoding the CAST complex polypeptides and the nucleotide sequence encoding the guide RNA are operably linked to the same promoter.Transposons
[0152] As noted elsewhere herein, in some cases a subject donor DNA includes a transposon. A transposon in a system of the present disclosure (e.g., in some cases as part of a CAST-complex transposon-based system) can have any convenient length, in some cases a length of up to about 100 kilobases (kb). For example, a transposon can have a length of from 0.1 kb to 0.5 kb, from 0.5 kb to 1 kb, from 1 kb to 5 kb, from 5 kb to 10 kb, from 10 kb to 15 kb, from 15 kb to 20 kb, from 20 kb to 25 kb, from 25 kb to 30 kb, from 30 kb to 35 kb, from 35 kb to 40 kb, from 40 kb to 45 kb, from 45 kb to 50 kb, from 50 kb to 55 kb, from 55 kb to 60 kb, from 60 kb to 65 kb, from 65 kb to 70 kb, from 70 kb to 75 kb, from 75 kb to 80 kb, from 80 kb to 85 kb, from 85 kb to 90 kb, from 90 kb to 95 kb, or from 95 kb to 100 kb (e.g., 0.1-100kb, 0.5- 100kb, 1-100kb, 5-100kb, 10-100kb, 50-100kb, 0.1-50kb, 0.5-50kb, 1-50kb, 5-50kb, 10-50kb, 0.1-25kb, 0.5-25kb, 1-25kb, 5-25kb, 10-25kb, 0.1-15kb, 0.5-15kb, 1-15kb, 5- 15kb, 10-15kb, , 0.1-10kb, 0.5-10kb, 1-10kb, 5-10kb, 0.1-5kb, 0.5-5kb, 1-5kb, 0.1-3kb, 0.5-3kb, or 1-3kb).
[0153] A transposon can function to knock out an endogenous nucleic acid in a target bacterium, e.g., to delete all or a portion of an endogenous nucleic acid in a target prokaryotic cell or to introduce a loss-of-function mutation in an endogenous nucleic acid in a target prokaryotic cell. A “knockout” includes deletion of all or a portion of a nucleic acid; and includes introduction of a loss-of-function mutation in a nucleic acid. For example, a transposon can function to delete all or a portion of an endogenous nucleic acid in a target prokaryotic cell (e g., target bacterium; target archaeon), or to introduce a loss-of-function mutation in an endogenous nucleic acid in a target prokaryotic cell, where the endogenous nucleic acid comprises one or more nucleotide sequences encoding one or more polypeptides that confer on a prokaryotic cell resistance to one or more antibiotics. Target loci are also discussed elsewhere herein.CAST-recognition sites
[0154] As noted above, a CAST transposon system of the present disclosure comprises a transposon flanked by recognition sites (nucleotide sequences) that are bound by and cleaved by a CAST complex. The recognition sites are referred to as “left end” and “right end.” Recognition sites bound by and cleaved by a CAST complex are known in the art.
[0155] For example, “left end” and “right end” recognition sites bound by and cleaved by a VcCAST are:TGTTGATGCAACCATAAAGTGATATTTAATAATTATTTATAATCAGCAACTTAACCA CAAAACAACCATATATTGATATCTCACAAAACAACCATAAGTTGATATTTT (left end; SEQ ID NO: 56); and GCAATATCAATTTATGGGTGTGATAATTATCAATTTATGGGTGTAATTATCATTTTA TGGTTGTATCAACA (right end; SEQ ID NO: 57).
[0156] As another example, “left end” and “right end” recognition sites bound by and cleaved by an ShCAST are:TGTACAGTGACAAATTATCTGTCGTCGGTGACAGATTAATGTCATTGTGACTATTT AATTGTCGTCGTGACCCATCAGCGTTGCTTAATTAATTGATGACAAATTAAATGTCA (left end; SEQ ID NO: 58); and CGACAGTCAATTTGTCATTATGAAAATACACAAAAGCTTTTTCCTATCTTGCAAAG CGACAGCTAATTTGTCACAATCACGGACAACGACATCTATTTTGTCACTGCAAAG AGGTTATGCTAAAACTGCCAAAGCGCTATAATCTATACTGTATAAGGATTTTACTGATGACAATAATTTGTCACAACGACATATAATTAGTCACTGTACA (right end; SEQ ID NO: 59).Transposon guide RNA
[0157] In some cases (e.g., for a CAST based transposon system), a transposon system of the present disclosure comprises a nucleotide sequence encoding one or more guide RNAs (also referred to herein as CAST guide RNAs). As with any CRISPR-Cas guide RNA, the CAST guide RNA comprises: i) a nucleotide sequence (referred to as a “guide sequence”, a “targeting sequence”, a “spacer”, and the like) that hybridizes to a target nucleotide sequence in a target DNA (e.g., in some cases a prokaryotic genome); and ii) a nucleotide sequence (referred to as a “constant region”, a “handle”, a “protein binding segment”, a “scaffold”, and the like) that binds to a polypeptide in the CAST complex. A CAST complex binds the guide RNA, thus forming a ribonucleoprotein complex (RNP). A CAST / guide RNA complex directs a transposon to a genomic site complementary to the guide RNA. See, e.g., Klompe et al. (2019) Nature 571:219; and Peters et al. (2019) Mol. Microbiol. 112:1635.
[0158] In some cases, a transposon system of the present disclosure comprises a nucleotide sequence encoding a single guide RNA. In some cases, a transposon system of the present disclosure comprises nucleotide sequences encoding two or more guide RNAs, each guide RNA comprising a nucleotide sequence that hybridizes to a target nucleotide sequence in a target DNA (e.g., in some cases a prokaryotic genome). For example, in some cases, a transposon system of the present disclosure comprises nucleotide sequences encoding 2, 3, 4, or 5 (or more than 5) different guide RNAs, each targeted to a different target sequence.
[0159] The guide sequence has complementarity with (hybridizes to) a target sequence of the target DNA. In some cases, the guide sequence is 16-35 nucleotides (nt) in length (e.g., 16-28, 16-26, 16-24, 16-22, 16-20, 16-18, 17-26, 17-24, 17-22, 17-20, 17-18, 18-26, 18-24,18-22, 18-20, 19-28, 19-26, 19-25, 29-22, 19-20, 20-25, 20-24, 20-23, 20-22, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nt in length). In some cases, the guide sequence is 18-24 nucleotides (nt) in length. In some cases, the guide sequence is at least 16 nt long (e.g., at least 17, 18, 20, or 22 nt long). In some cases, the guide sequence is at least 17 nt long. In some cases, the guide sequence is at least 18 nt long. In some cases, the guide sequence is at least 20 nt long. In some cases, the guide sequence is about 32 nt long. In some cases, VcCAST guides are included in aCRISPR array (repeat-spacer-repeat). In some cases, a guide RNA (e.g., a ShCAST guide RNA) includes a guide sequence that is about 23-nt long.
[0160] In some cases, the guide sequence has 80% or more (e.g., 85% or more, 90% or more, 95% or more, or 100% complementarity) with the target sequence of the target DNA. In some cases, the guide sequence is 100% complementary to the target sequence of the target DNA. In some cases, the target DNA includes at least 15 nucleotides (nt) of complementarity with the guide sequence of the guide RNA. In some cases, the target DNA includes at least 16 nucleotides (nt) of complementarity with the guide sequence of the guide RNA.
[0161] In some cases, the constant region of a guide RNA is 15 or more nucleotides (nt) in length (e.g., 16 or more, 17 or more, 18 or more, 20 or more, 21 or more, 22 or more, 23 or more, 24 or more, 25 or more, 26 or more, 27 or more, 28 or more, 29 or more, 30 or more, 31 or more nt, 32 or more, 33 or more, 34 or more, or 35 or more nt in length). In some cases, the constant region of a guide RNA is 18 or more nt in length.
[0162] In some cases, the guide RNA is a dual-molecule guide RNA. In some cases, the guide RNA is a single-molecule RNA (also referred to as a “single guide RNA” or “sgRNA”).
[0163] In some cases, VcCasTn gRNAs use a 5’-CC Type IF PAM. In some cases, ShCasTn gRNAs use a 5’-GTT Cas12k PAM.
[0164] One example of a crRNA for the VcCAST system is [It is to be understood that the Ts are Us when in RNA form]: gtgaactgccgagtaggtagctgataacGCGAGCAGAAGTGGCAGATGATGTTGTCAAAGgtga actgccgagtaggtagctgataac (SEQ ID NO:65).
[0165] One example of a sgRNA for the ShCas12k system is [It is to be understood that the Ts are Us when in RNA form]:ATATTAATAGCGCCGCAATTCATGCTGCTTGCAGCCTCTGAATTTTGTTAAATGAG GGTTAGTTTGACTGTATAAATACAGTCTTGCTTTCTGACCCTGGTAGCTGCTCAC CCTGATGCTGCTGTCAATAGACAGGATAGGTGCGCTCCCAGCAATAAGGGCGCG GATGTACTGCTGTAGTGGCTACTGAATCACCCCCGATCAAGGGGGAACCCTAAA TGGGTTGAAAGGTCAAAGAGTATGCGTCGTTAAT (SEQ ID NO:66).HDR donor template
[0166] Another non-limiting example of a donor DNA is a recombination template (also referred to as a “donor DNA template” or “donor template DNA” or “donor template”),which can be used to provide a template sequence for integrating a sequence from the donor template into a target DNA via homology directed repair (HDR), e.g., homologous recombination (HR), after cleavage of the target DNA (e.g., double strand breaks (DSBs) or single strand breaks (SSBs), e.g., when the protein is a nickase variant). Thus, in some cases, (e.g., when the gene editing tool is a CRISPR- Cas system), the donor DNA is a donor template.
[0167] A donor template includes flanking homology arms and the desired edit (insertion, deletion, or specific mutation). In some cases, the donor template also includes an internal sequence that disrupts the PAM at the target site (e.g., causes mutation of the PAM), thus preventing repeated targeting after successful recombination. Cleavage of unedited target genes by CRISPR nucleases is often lethal in bacteria because of the formation of a double-strand break (DSB), serving as a strong counterselection without the need of the insertion of a large resistance cassette marker into the genome. In this way the DSB drives editing through homologous recombination (HR) or, more rarely in prokaryotes, via non-homologous end joining (NHEJ). It is, therefore, the DNA repair systems of the host species / strain which actually perform the desired editing. In most bacterial organisms RecA-mediated HR is induced to repair DNA damage by DSB. This response however, can be error-prone and can insert undesired mutations, mainly through the recruitment of the mutagenic DNA polymerase IV (PollV) and inhibition of high-fidelity Poll II at the DSB site. In many cases where CRISPR nucleases have been used to achieve highly efficient genome editing, particularly in E. coli, they have been combined with an enhanced recombination system e.g., the Lambda Red phage to promote homology-directed repair (HDR).
[0168] A donor template will include a “donor sequence” (e.g., a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding an immunity protein), which is a nucleic acid sequence to be inserted at the cleavage site induced by a gene editing protein (e.g., a CRISPR-Cas protein). The donor polynucleotide can include sufficient homology to a genomic sequence at the cleavage site, e.g. 70%, 80%, 85%, 90%, 95%, or 100% homology with the nucleotide sequences flanking the cleavage site, e.g. within about 50 bases or less of the cleavage site, e.g. within about 30 bases, within about 15 bases, within about 10 bases, within about 5 bases, or immediately flanking the cleavage site, to support homology-directed repair between it and the genomic sequence to which it bears homology. Approximately 25, 50, 100, or 200 nucleotides, or more than 200 nucleotides, of sequence homology between adonor and a genomic sequence (or any integral value between 10 and 200 nucleotides, or more) will support homology-directed repair. Donor sequences can be of any length, e.g. 10 nucleotides or more, 50 nucleotides or more, 100 nucleotides or more, 250 nucleotides or more, 500 nucleotides or more, 1000 nucleotides or more, 5000 nucleotides or more, etc.
[0169] The donor sequence is typically not identical to the genomic sequence that it replaces. Rather, the donor sequence may contain at least one or more single base changes, insertions, deletions, inversions or rearrangements with respect to the genomic sequence, so long as sufficient homology is present to support homology- directed repair. In some embodiments, the donor sequence comprises a non- homologous sequence flanked by two regions of homology, such that homology- directed repair between the target DNA region and the two flanking sequences results in insertion of the non-homologous sequence at the target region. Donor sequences may also comprise a vector backbone containing sequences that are not homologous to the DNA region of interest and that are not intended for insertion into the DNA region of interest. Generally, the homologous region(s) of a donor sequence will have at least 50% sequence identity to a genomic sequence with which recombination is desired. In certain embodiments, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or 99.9% sequence identity is present.
[0170] In some cases, the donor template is single-stranded DNA (ssDNA). In some cases, the donor template is double-stranded DNA (dsDNA). It may be in linear or circular form. If introduced in linear form, the ends of the donor sequence may be protected (e.g., from exonucleolytic degradation) by methods known to those of skill in the art. As an alternative to protecting the termini of a linear donor sequence, additional lengths of sequence may be included outside of the regions of homology that can be degraded without impacting recombination. A donor sequence can be introduced into a cell as part of a vector molecule having additional sequences such as, for example, replication origins, promoters and genes encoding antibiotic resistance. Moreover, donor sequences can be introduced as naked nucleic acid, as nucleic acid complexed with an agent such as a liposome or poloxamer, or can be delivered by phage.
[0171] As alluded to above, because of the presence of a PAM sequence in the target DNA, it is contemplated herein that a CRISPR-Cas effector protein (e.g., Cas9, Cas12, and the like) could potentially cleave inserted sequence either prior to or following homology directed repair (e.g., homologous recombination), resulting in a possible repeated cleavage event. Therefore, to avoid cleavage of the donor sequence beforeand / or after CRISPR-Cas effector protein mediated homology directed repair, in some embodiments, alternate versions of the donor sequence may be used where mutations (e.g., silent mutations) are introduced. These mutations (e.g., silent mutations) may disrupt CRISPR-Cas effector protein binding and cleavage, but not disrupt the amino acid sequence of the repaired gene. For example, a donor sequence can in some cases include a mutated PAM (e.g., insertion of the donor sequence can result in mutation of the PAM sequence to reduce / eliminate additional cleavage events).Therapeutic Unit
[0172] In some embodiments, a donor DNA includes, in addition to the bacteriocin unit, a therapeutic unit that includes one or more nucleotide sequences encoding one or more expression products of interest. Examples of expression products of interest include, but are not limited to: one or more proteins that confer antibiotic resistance on a bacterium, one or more enzymes of a biosynthetic pathway (e.g., a magnetosome, gas vesicle, B-vitamin, mevalonate, or polyketide biosynthetic pathway), one or more proteins that confer resistance to a phage (e.g., bacteriophage), one or more enzymes of a biosynthetic pathway (e.g., a magnetosome, gas vesicle, B-vitamin, mevalonate, or polyketide biosynthetic pathway), one or more enzymes in a carbon utilization pathway (e.g., porphyran polysaccharide, glycosaminoglycan, non-caloric artificial sweetener, ethanolamine, or sucrose utilization pathway), one or more enzymes in a nitrogen utilization pathway, one or more enzymes in a sulfur utilization pathway, one or more detectable markers (e.g., a polypeptide that provides a detectable signal such as one or more fluorescent proteins, one or more fluorogenic aptamers, one or more proteins that provide for isolation of a target cell), one or more proteins that provide for detection of an analyte in a prokaryote, one or more antiinflammation factors (e.g., anti-inflammatory cytokines such as interleukin (IL)-1 receptor antagonist, IL-4, IL-6, IL-10, IL-11, and IL-13) (e.g., in some cases a subject self-selection DNA editing system is used with a human microbiome such as a gut microbiome, and the cell under self-selection, i.e., the cell that includes the therapeutic unit, can secrete anti-inflammation factors to the human); one or more antibodies or nanobodies; one or more immunomodulatory or antitumor factors (e.g., tumor-associated antigen, cytolysin A, synthetic CAR targets).
[0173] In some cases, a therapeutic unit comprises one or more nucleotide sequences encoding one or more polypeptides that provide for isolation of a target prokaryoticcell; e.g., a FLASH tag; FAST; iLOV; phiLOV; smURFP, IFP2.0; evoglow-Pp1; UnaG; a SNAP tag; a CLIP tag; a Halo tag; a spinach aptamer; mango aptamer; and the like. See, e.g., Thorn (2017) Mol. Biol. Cell 28:848; and Wang et al. (2017) Mol. Bhiochem. Parasitol. 216:1. A therapeutic unit can comprise one or more nucleotide sequences encoding one or more fluorescent proteins or tags that are detectable in anaerobic conditions, such as an anaerobic green fluorescent protein (GFP); see, e.g., Landete et al. ((2015) App. Microbiol. Biotechnol. 99:6865) and Streett et al. (2019) Appl. Environmental Microbiol. 85:e00622. Tagging surface exposed proteins with FLAG tag, His tag, Myc tag and the like, to be immunolabeled with fluorescence / magnetic- conjugated antibodies. Also suitable are tetracysteine tags to enable staining with biarsenical dyes (e.g., for staining with FIAsH and ReAsH dyes).
[0174] In some cases, a therapeutic unit comprises a nucleotide sequence encoding a fluorescent polypeptide. Examples of fluorescent proteins include, but are not limited to: green fluorescent protein (GFP) and variants thereof, blue fluorescent protein (BFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced GFP (EGFP), enhanced CFP (ECFP), enhanced YFP (EYFP), GFPS65T, Emerald, Topaz (TYFP), Venus, Citrine, mCitrine, GFPuv, destabilised EGFP (dEGFP), destabilized ECFP (dECFP), destabilized EYFP (dEYFP), mCFPm, Cerulean, T- Sapphire, CyPet, YPet, mKO, HcRed, t-HcRed, DsRed, DsRed2, DsRed-monomer, J- Red, dimer2, t-dimer2(12), mRFP1, pocilloporin, Renilla GFP, Monster GFP, paGFP, Kaede protein and kindling protein, Phycobiliproteins and Phycobiliprotein conjugates including B- Phycoerythrin, R-Phycoerythrin and Allophycocyanin. Other examples of fluorescent proteins include m Honeydew, m Banana, mOrange, dTomato, tdTomato, mTangerine, mStrawberry, mCherry, mGrapel, mRaspberry, mGrape2, mPlum (Shaner et al. (2005) Nat. Methods 2:905-909), and the like. See, e.g., Thorn (2017) Mol. Biol. Cell 28:848.Prokaryotic Cells, Target Prokaryotic Cells, and Genetically modified Prokaryotic cells
[0175] The present disclosure provides a prokaryotic cell comprising a subject self-selection DNA editing system (including a gene editing tool; and a donor DNA comprising a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding an immunity protein). In some cases, such a prokaryotic cell is a donor cell, e.g., to deliver a subject self-selection DNA editing system to a recipient cell. A donor cell is one that can deliver a subject system toanother cell, where the system is active. For example, a donor cell (a first cell) can transfer the self-selection DNA editing system (gene editing tool and the donor DNA) (e.g., in some cases where the editing system is a transposon system, e.g., a CAST system, in which the donor DNA includes a transposon) into a second cell (a “recipient cell”) (e.g., via conjugation), where the donor DNA then integrates into a target DNA in that second cell (recipient). The second cell will then self-select / enrich because it will express bacteriocin to kill neighboring cells, but will be immune to the bacteriocin. See, e.g., FIG. 2A. Any of the various embodiments described herein for self-selection DNA editing systems (e.g., various promoter arrangements, various types of gene editing tools, various donor DNA types, various bacteriocins, etc.) are equally applicable to cells that include such systems.
[0176] In some cases, the system is not active in the donor cell because the donor cell does not include the correct target sequence that the gene editing tool targets (e.g., does not include the target sequence targeted by a CRISPR-Cas guide RNA). Thus, no gene editing would occur and the bacteriocin unit would not integrate. However, if the self-selection DNA editing system is transferred (e.g., via conjugation) to a second cell (a recipient), and that second cell may have the targeted sequence, in which case the bacteriocin unit could integrate into the target DNA (e.g., genomic DNA). However, the second cell may not have the correct target sequence, and may itself become a donor cell, passing the system on to a third cell, etc.
[0177] The present disclosure also provides genetically modified prokaryotic cells that include a bacteriocin unit integrated into its genomic DNA or into a mobile DNA element (e.g., a conjugative transposon, a plasmid, a virus (an archaeal virus), a viral satellite (e.g., an archaeal viral satellite), a phage, a prophage, a phage satellite, a phage plasmid or a phage plasmid) inside of the cell. Such a cell can be a cell that was targeted by a subject self-selection DNA editing system (e.g., a cell that received the system from a donor cell). In other words, a genetically modified prokaryotic cell can be generated by targeting a prokaryotic cell with a subject self-selection DNA editing system (e.g., via direct delivery of the system to the target cell, or by contacting the target cell with a donor cell, which delivers the system) - resulting in integration of a bacteriocin unit into a target DNA inside of the targeted cell (genomic DNA or into a mobile DNA element such as a conjugative transposon, a plasmid, a virus (an archaeal virus), a viral satellite (e.g., an archaeal viral satellite), a phage, a prophage, a phage satellite, a phage plasmid or a phage plasmid). For example, in some cases, a subject genetically modified prokaryotic cell includes a bacteriocin unitintegrated into its genomic DNA. Any of the various embodiments described herein for bacteriocin units (e.g., various promoter arrangements, various donor DNA types, various bacteriocins, etc.) are equally applicable to genetically modified prokaryotic cells (e.g., cells that were generated by use of a subject self-selection DNA editing system).
[0178] For any of the cells discussed above in this section (e.g., Prokaryotic Cells, Target Prokaryotic Cells, and Genetically modified Prokaryotic cells), the cell can be any prokaryotic cell (e.g., any archaeal cell, any bacterial cell). In some cases, the cell is a bacterial cell. Bacteria of interest include, but are not limited to: proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, and ESKAPE pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa and Enterobacter spp.). For example, gram-negative bacteria of interest include, but are not limited to: proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bacteroides, and Prevotella. Gram-positive bacteria of interest include, but are not limited to: Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum), Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, and Streptomyces. Also of interest are ESKAPE pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella neumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa and Enterobacter spp.) (see, e.g., Miller et al., Nat Rev Microbiol 22, 598-616 (2024)).
[0179] In some cases, the cell is an E. coli cell. In some cases, the cell is a proteobacteria. In some cases, the cell is an Enterococcus (e.g., E. faecalis and E. faecium). In some cases, the cell is a Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum). In some cases, the cell is a Bacteroides. In some cases, the cell is a Prevotella. In some cases, the cell is a Clostridium (e.g., C. innocuum). In some cases, the cell is a Ruminococcus (e.g., R. gnavus). In some cases, the cell is a Lactococcus. In some cases, the cell is a Staphylococcus. In some cases, the cell is a Bacillus. In some cases, the cell is a Streptomyces. In some cases, the cell is an ESKAPE pathogen. In some cases, the cell is an Enterococcus faecium. In some cases, the cell is a Staphylococcus aureus. In some cases, the cell is a Klebsiella pneumoniae. In some cases, the cell is an Acinetobacter baumannii. In some cases, the cell is a Pseudomonas aeruginosa. In some cases, the cell is an Enterobacter spp.
[0180] See the below target cell section for additional example cells and a discussion of communities.Methods
[0181] The present disclosure provides methods of modifying a microbial community. The present disclosure also provides methods of modifying cells (e.g., editing a prokaryotic cell) to facilitate self-selection.
[0182] Techniques for introducing proteins and nucleic acids into cells will be known to one of ordinary skill in the art and any convenient technique can be used for introducing a subject self-selection DNA editing system into a prokaryotic cell (e.g., bacteria or archaea). In some cases, the self-selection DNA editing system is introduced via conditions that promote introduction of nucleic acid into prokaryotic cells including by electroporation, heat shock, use of chemically induced competence or other methods known in the art.
[0183] In some embodiments, plasmid conjugation can be used to introduce a desired plasmid from a “donor” prokaryotic cell (e.g., bacterium) to a recipient prokaryotic cell (e.g., bacterium). In some embodiments, plasmid conjugation can genetically modify a recipient microbial cell by introducing a conjugation plasmid from a donor microbial cell to a recipient microbial cell.
[0184] In some cases, a method of modifying a microbial community includes introducing a subject genetically modified prokaryotic cell (e.g., bacterial cell) to the microbial community. The genetically modified prokaryotic cell will then be self-selecting due to the presence of the bacteriocin unit. See elsewhere herein for details with respect to genetically modified prokaryotic cells - the various discussed embodiments are equally contemplated in the context of the methods described herein. Such a genetically modified prokaryotic cell can include a bacteriocin unit integrated into its genomic DNA (or in some embodiments integrated into a mobile DNA element inside of the cell, e.g., a conjugative transposon, a plasmid, a virus (an archaeal virus), a viral satellite (e.g., an archaeal viral satellite), a phage, a prophage, a phage satellite, a phage plasmid, and the like). The genetically modified prokaryotic cell will therefore be self-selective. If the bacteriocin protein and / or the immunity protein are under the control of an inducible promoter, then the self-selective advantage of the genetically modified prokaryotic cell will also be under the control of the inducible promoter in the sense that self-selective advantage will be present when the promoter is active. In some cases, the genetically modified prokaryotic cell will include an edit to itsgenomic DNA, e.g., at a locus that encodes a virulence factor, at a locus that encodes an antibiotic resistance gene, at a locus that affect metabolites that contribute to inflammation or risk of disease, and the like. See, e.g., the section herein discussing target loci.
[0185] In some cases, a method of modifying a microbial community includes editing a target prokaryotic cell in situ (within the community), which can lead to self-selection. Such editing includes providing to the target prokaryotic cell (e.g., bacterial cell), a DNA editing system and a subject bacteriocin unit (e.g., in some cases a subject selfselection DNA editing system is provided to the target cell). In some cases, a subject self-selection DNA editing system is provided to the target cell directly. In some cases, a subject self-selection DNA editing system is provided to the target cell using a donor cell (see discussion elsewhere herein). For example, a donor cell (that includes a subject self-selection DNA editing system) can be introduced to the microbial community. The donor cell will include the subject self-selection DNA editing system and will transfer the system to the target cell (directly or indirectly). In some cases, the donor cell will transfer directly to the target cell, and in some cases, the donor cell will transfer the system indirectly to the target cell, e.g., the donor cell might transfer the system to another cell that then becomes a donor cell - eventually, either from that donor cell or from yet another (or another etc.) donor cell, the system will be transferred to the target cell.
[0186] As noted above, the present disclosure also provides methods of modifying cells (e.g., editing a prokaryotic cell such as a bacterium) to facilitate self-selection. Such a method can occur in cells in culture (e.g., in a laboratory), or can occur in cells in their natural environment (e.g., cells that are part of a microbial community). In such a method, a bacteriocin unit is provided to a target prokaryotic cell, and the bacteriocin unit integrates into a target DNA inside of prokaryotic cell (e.g., the prokaryotic cell’s genomic DNA or into a mobile DNA element inside of the prokaryotic cell). In some cases, the bacteriocin unit integrates into the cell’s genome. The targeted prokaryotic cell will therefore become self-selective. If the bacteriocin protein and / or the immunity protein are under the control of an inducible promoter, then the self-selective advantage of the targeted cell will also be under the control of the inducible promoter in the sense that self-selective advantage will be present when the promoter is active. In some cases, the targeted prokaryotic cell will include an editto its genomic DNA (e.g., at a locus that encodes a virulence factor, at a locus that encodes an antibiotic resistance gene, at a locus that affect metabolites that contribute to inflammation orrisk of disease, and the like) - see, e.g., the section herein discussing target loci. In some cases, the cause of the edit to the target cell’s genomic DNA is the integration of the bacteriocin unit into a target locus of interest. As such, for example, the bacteriocin unit can be targeted for integration into a target locus of interest, simultaneously providing: (i) a self-selection advantage due to the bacteriocin unit and; (ii) an edit to the target DNA in the target cell, e.g., at a target locus of interest. See more details elsewhere herein with respect to target prokaryotic cells, bacteriocin units, and self-selection DNA editing systems - the various discussed embodiments are equally contemplated in the context of the methods described herein. For example, in some cases, that target cell is bacterial cell, and the self-selection DNA editing system is a CAST-based transposon system (e.g., where the gene editing tool of comprises: (1) nucleotide sequences encoding polypeptides that form a CRISPR- associated transposase (CAST) complex; and (2) a nucleotide sequence encoding a guide RNA; and where the donor DNA includes a transposon flanked by recognition sites that are recognized by the CAST complex.
[0187] With regard to any of the methods disclosed herein, see elsewhere herein for details with respect to target prokaryotic cells and self-selection DNA editing systems - the various discussed embodiments are equally contemplated in the context of the methods described herein. The act of providing a bacteriocin unit and an immunity protein to the cell will confer a self-selective advantage to the targeted cell (if an inducible promoter is used, the self-selective advantage will be realized in the presence of the inducer), in some cases, this will thus result in modifying a microbial community.
[0188] In some cases, a subject method comprises subjecting the heterogeneous population of prokaryotic cells to conditions for conjugation, transformation, or transduction, where such conditions permit conjugation, transformation, or transduction of a prokaryotic cell known to be susceptible to nucleic acid transfer via conjugation, transformation, or transduction. In some cases, the conditions comprise electroporation. For example, in some cases, a heterogeneous population of prokaryotic cells is electroporated in a liquid medium. In some cases, the conditions comprise chemically induced competence (e.g., calcium chloride; rubidium chloride; etc.).
[0189] A heterogeneous population of prokaryotic cells is referred to herein as a “community” or a “prokaryotic cell community” or a “microbial community.” Such a heterogeneous population can comprise from 5 to 109different prokaryotic cells; e.g., from 10 to 102,from 102to 103, from 103to 104, from 104to 105, from 105to 106, from 106to 107, from 107to 108, or from 108to 109different prokaryotic cells (e.g., from 5 to 108, 5 to 107, 5 to 106, 5 to 105, 5 to 104, 5 to 103, 5 to 102, 5 to 10, 102to 108, 102to 107, 102to 106, 102to 105, 102to 104, or 102to 103). In some cases, the population of prokaryotic cells are of the same genus. In some cases, the population of prokaryotic cells comprise prokaryotes (e.g., bacteria) of 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 (e.g., from 10 to 20, from 20 to 30, from 30 to 40, from 40 to 50, or more than 50), different genus and / or species. In some cases, a heterogeneous population of prokaryotic cells can include from 5 to 5000, or more than 5000, different species. For example, a heterogeneous population of prokaryotic cells can include from 5 to 25, from 25 to 50, from 50 to 100, from 100 to 250, from 250 to 500, from 500 to 1000, from 1000 to 2000, from 2000 to 3000, from 3000 to 4000, from 4000 to 5000, or more than 5000, different species.
[0190] Heterogeneous populations of prokaryotic cells can be found in a natural environment such as the gastrointestinal tract of a mammal (e.g., a human); the microbiome of a human; the microbiome of a non-human animal soil; hot springs; oceans; marshland; swamps; etc. Suitable heterogeneous populations of prokaryotic cells include prokaryotic cells found in wastewater, agricultural runoff, and the like. Suitable heterogeneous populations of prokaryotic cells include prokaryotic cells involved in food processing (e.g., fermentations to produce beverages or food that rely on a mixed community of cells such as with kimchi, soy sauce, or kombucha). Suitable heterogeneous populations of prokaryotic cells present in the rhizosphere. Suitable heterogeneous populations of prokaryotic cells present on the plant surface (the plant microbiome). Suitable heterogeneous populations of prokaryotic cells found in industrial processes relying on communities of microorganisms such as industrial wastewater treatment or bioreactors used for bioremediation of wastes (i.e. thiocyanate (SCN) degradation reactors used for gold mining runoff).
[0191] With regard to any of the methods disclosed herein, self-selection can result in an enriched population in which from 50% to more than 99% of the cells (e.g., from 50% to 60%, from 60% to 70%, from 70% to 80%, from 80% to 90%, from 90% to 95%, from 95% to 99%, e.g., 50-100%, 50-99%, 50-95%, 50-90%, 50-80%, 50-70%, 60- 100%, 60-99%, 60-95%, 60-90%, 60-80%, 60-70%, 70-100%, 70-99%, 70-95%, 70- 90%, 70-85%, 70-80%, 80-100%, 80-99%, 80-95%, 80-90%, 80-85%, or more than 99%) of the targeted cell type have a target DNA (e.g., genome) that has an integrated bacteriocin unit (and has been edited) as a result of providing a bacteriocinunit (and in some cases a therapeutic unit, and in some cases a self-selection DNA editing system) to a targeted cell.
[0192] One skilled in the art will appreciate that, for this and other functions, structures, and processes, disclosed herein, the functions, structures and steps may be implemented or performed in differing order or sequence. Furthermore, the outlined functions and structures are only provided as examples, and some of these functions and structures may be optional, combined into fewer functions and structures, or expanded into additional functions and structures without detracting from the essence of the disclosed embodiments.Targeted Cells (i.e., target cells)
[0193] For any of the methods discussed herein (e.g., methods of modifying a microbial community, methods of modifying cells to facility self-selection), the targeted cell (i.e., the target cell) can be any prokaryotic cell (e.g., any archaeal cell, any bacterial cell). In some cases, the cell is a bacterial cell. Bacteria of interest include, but are not limited to: proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifidum, B. infantis, and B. long urn), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, and ESKAPE pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa and Enterobacter spp.). For example, gramnegative bacteria of interest include, but are not limited to: proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bacteroides, and Prevotella. Grampositive bacteria of interest include, but are not limited to: Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum), Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, and Streptomyces. Aso of interest are ESKAPE pathogens (Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa and Enterobacter spp.) (see, e.g., Miller et al., Nat Rev Microbiol 22, 598- 616 (2024)). In some cases, the cell is an E. coli cell. In some cases, the cell is a proteobacteria.
[0194] In some cases, the cell is an Enterococcus (e.g., E. faecalis and E. faecium). In some cases, the cell is a Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum). In some cases, the cell is a Bacteroides. In some cases, the cell is a Prevotella. In somecases, the cell is a Clostridium (e.g., C. innocuurri). In some cases, the cell is a Ruminococcus (e.g., R. gnavus). In some cases, the cell is a Lactococcus. In some cases, the cell is a Staphylococcus. In some cases, the cell is a Bacillus. In some cases, the cell is a Streptomyces. In some cases, the cell is an ESKAPE pathogen. In some cases, the cell is an Enterococcus faecium. In some cases, the cell is a Staphylococcus aureus. In some cases, the cell is a Klebsiella pneumoniae. In some cases, the cell is an Acinetobacter baumannii. In some cases, the cell is a Pseudomonas aeruginosa. In some cases, the cell is an Enterobacter spp.
[0195] Target prokaryotic cells include bacteria and archaea. In some cases, the target prokaryotic cells are bacteria. In some cases, the target prokaryotic cells are archaea. In some cases, target prokaryotic cells include bacteria and / or archaea that have not yet been cultured or isolated in a laboratory in monoculture.
[0196] Target prokaryotic cells include prokaryotic cells found in a natural environment such as the gastrointestinal tract of a mammal (e.g., a human); the microbiome of a human; the microbiome of a non-human animal soil; hot springs; oceans; marshland; swamps; etc. Target prokaryotic cells include prokaryotic cells found in wastewater, agricultural runoff, and the like. Target prokaryotic cells include prokaryotic cells involved in food processing (e.g., fermentations to produce beverages or food that rely on a mixed community of cells such as with kimchi, soy sauce, or kombucha). Target prokaryotic cells include prokaryotic cells present in the rhizosphere. Target prokaryotic cells include prokaryotic cells present on the plant surface microbiome (the plant microbiome). Target prokaryotic cells include prokaryotic cells found in industrial processes relying on communities of microorganisms such as industrial wastewater treatment or bioreactors used for bioremediation of wastes (i.e. thiocyanate (SCN) degradation reactors used for gold mining runoff). Target prokaryotic cells include prokaryotic cells that find use in and / or are found in one or more of: the plant microbiome, food processing (e.g., wine, cheese, yogurt, etc.), bioremediation, and industrial processes.
[0197] Target bacteria include bacteria present in the human gastrointestinal tract. Target bacteria include bacteria of the phyla Firmicutes, Bacteroidetes, Actinobacteria, and Proteobacteria. Target bacteria include bacteria of the genera Lactobacillus, Bacteroides, Clostridium, Faecalibacterium, Eubacterium, Ruminococcus, Peptococcus, Roseburia, Peptostreptococcus, Bifidobacterium, Alistipes, Parabacteroides, Porphyromonas, Prevotella, Collinsalla, Escherichia, and Desulfovibrio. See, e.g., Rinninella et al. (2019) Microorganisms 7:14. Examples oftarget bacteria include, e g., Bacteroides fragilis ssp. vulgatus, Collinsella aerofaciens, Bacteroides fragilis ssp. thetaiotaomicron, Peptostreptococcus productus II, Parabacteroides distasonis, Faecalibacterium prausnitzii, Coprococcus eutactus, Peptostreptococcus productus I, Ruminococcus bromii, Bifidobacterium adolescentis, Gemmiger formicilis, Bifidobacterium longum, Eubacterium siraeum, Ruminococcus torques, Eubacterium rectale, Eubacterium eligens, Bacteroides eggerthii, Clostridium leptum, Bacteroides fragilis ssp. A, Eubacterium biforme, Bifidobacterium infantis, Eubacterium rectale, Coprococcus comes, Pseudoflavonifractor capillosus, Ruminococcus albus, Dorea formicigenerans, Eubacterium hallii, Eubacterium ventriosum, Fusobacterium russi, Ruminococcus obeum, Eubacterium rectale, Clostridium ramosum, Lactobacillus leichmannii, Ruminococcus callidus, Butyrivibrio crossotus, Acidaminococcus fermentans, Eubacterium ventriosum, Bacteroides fragilis ssp. fragilis, Coprococcus catus, Aerostipes hadrus, Eubacterium cylindroides, Eubacterium ruminantium. Staphylococcus epidermidis, Eubacterium limosum, Tissirella praeacuta, Fusobacterium mortiferum, Fusobacterium naviforme, Clostridium innocuum, Clostridium ramosum, Propionibacterium acnes, Ruminococcus flavefaciens, Bacteroides fragilis ssp. ovatus, Fusobacterium nucleatum, Fusobacterium mortiferum, Escherichia coli, Gemella morbillorum, Finegoldia magnus, Streptococcus intermedius, Ruminococcus lactaris, Eubacterium tenue, Eubacterium ramulus, Bacteroides clostridiiformis ssp. clostridliformis, Bacteroides coagulans, Prevotella oralis, Prevotella ruminicola, Odoribacter splanchnicus, and Desuifomonas pigra.
[0198] Target bacteria include bacteria present in the gastrointestinal tract of an ungulate (e.g., a bovine; an equine; an ovine; a caprine; etc.). Other target bacteria include, e.g., bacteria associated with nosocomial infections in humans. Other target bacteria include soil bacteria.
[0199] In some cases, a target prokaryotic cell is one that is refractory to genetic modification by electroporation. In some cases, a target prokaryotic cell is one that is refractory to genetic modification by chemically-induced competence (e.g., competence induced by calcium chloride, rubidium chloride, and the like). In some cases, a target prokaryotic cell is one that is refractory to genetic modification by heat shock. In some cases, a target prokaryotic cell is one that is refractory to natural transformation. In some cases, a target prokaryotic cell is one that is refractory to isolation. In some cases, a target prokaryotic cell is one that is refractory growth in monoculture (e.g., in an industrial setting, a research laboratory setting, or the like).
[0200] Archaea that are suitable target prokaryotic cells include, e.g., archaea any species in any of the phyla Aenigmarchaeota, Diapherotrites, Nanoarchaeota, Nanohaloarchaeota, Micrarchaeota, Pacearchaeota, Parvarchaeota, Woesearchaeota, Aigarchaeota, Bathyarchaeota, Crenarchaeota, Geoarchaeota, Korarchaeota, Thaumarchaeota, Lokiarchaeota, Thorarchaeota, Odinarchaeota, Heimdallarchaeota, and the like.
[0201] As noted above, in some cases, a target cell is a member of a microbial community.Target Loci and Target DNAs
[0202] Any DNA (e.g., inside of a prokaryotic cell) can serve as a target DNA (for integration of the donor DNA / donor DNA sequence). In some cases, the target DNA is genomic DNA (e.g., genomic DNA of a target prokaryotic cell). In some cases, the target DNA is a mobile DNA element inside of the prokaryotic cell. Examples of mobile DNA elements include, but are not limited to: integrative and conjugative elements (ICEs, also referred to as conjugative transposons), plasmids, viruses (e.g., archaeal viruses), viral satellites (e.g., archaeal viral satellites), phages, prophages, phage satellites, and phage plasmids.
[0203] Any locus of a target DNA (e.g., a genomic locus of a genomic DNA) can be targeted for integration of the donor DNA / donor DNA sequence (i.e. , integration of a bacteriocin unit). In some cases, the gene editing tool edits the target DNA (e.g., target prokaryotic cell’s genomic DNA) at a position other than the site of integration for the bacteriocin unit. In some cases, the gene editing tool edits the target DNA (e.g., target prokaryotic cell’s genomic DNA) at the site of integration for the bacteriocin unit (i.e., the edit to the target DNA is the integration of the bacteriocin unit).
[0204] In some cases, the target locus is chosen because disruption of the locus will have a beneficial effect. Examples include, but are not limited to loci that encode virulence factors, loci that encode microbial toxins, loci that encode antibiotic resistance genes, or loci that affect metabolites that contribute to inflammation or risk of disease.Specific examples of target loci include, but are not limited to: pduC, clbH, clbJ (colibactin genes), hlyD, iutA, cnf, ireA, fimK, and fimH (fimbria gene) (e.g., in E. coli). Additional examples of target loci include epoxide hydrolase (EH) genes in Enteroccous spp. and Bifidobacterium spp., loci that contribute to poor cardiac health (e.g., yBB utilization (gbu) gene cluster, see, e.g., Buffa et al., Nat Microbiol 7, 73-86(2022)), loci that contribute to degradation of medicines (e.g., cgr, see, e.g., Haiser et al., Science. 2013 Jul 19;341(6143):295-8), and loci that contribute to problematic antibiotic resistance (see, e.g., card.mcmaster.ca / prevalence).
[0205] As such, in some cases, a subject self-selection DNA editing system is configured to integrate the donor DNA into a target locus, where the target locus encodes a virulence factor, encodes an antibiotic resistance gene, or affects a metabolite that contributes to inflammation or risk of disease. For example, the gene editing tool can target such a desired locus, and sequence of the donor DNA can integrate into the targeted locus (e.g., an HDR donor template can include homology arms that target the locus, a transposon can integrate into the desired locus, e.g., in some cases due to the guide RNA used as part of a transposon system, and the like). For example, if a CRISPR-Cas gene editing tool is used (e.g., a CRISPR-Cas effector protein or a CAST complex), one or more guide RNAs can be used that target the locus of interest.Kits
[0206] Provided are kits / systems for carrying out a subject method. Such kits comprise various combinations of components useful in any of the methods described elsewhere herein. A kit can further include one or more additional reagents, where such additional reagents can be any convenient reagent. Components of a subject kit can be in separate containers; or can be combined in a single container. In some cases one or more of a kit’s components are pharmaceutically formulated for administration to a human.
[0207] In addition to above-mentioned components, a subject kit can further include instructions for using the components of the kit to practice the subject methods (e.g., dosing instructions, instructions to administer the component(s) to an individual. The instructions for practicing the subject methods are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e. , associated with the packaging or subpackaging) etc. In some embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g. CD-ROM, diskette, flash drive, etc. In some embodiments, the actual instructions are not present in the kit, but means forobtaining the instructions from a remote source, e.g. via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate.EXEMPLARY NON-LIMITING ASPECTS OF THE DISCLOSURE
[0208] Aspects, including embodiments, of the present subject matter described above may be beneficial alone or in combination, with one or more other aspects or embodiments. Without limiting the foregoing description, certain non-limiting aspects of the disclosure are provided below. As will be apparent to those of ordinary skill in the art upon reading this disclosure, each of the individually numbered aspects may be used or combined with any of the preceding or following individually numbered aspects. This is intended to provide support for all such combinations of aspects and is not limited to combinations of aspects explicitly provided below. It will be apparent to one of ordinary skills in the art that various changes and modifications can be made without departing from the spirit or scope of the invention.1. A self-selection DNA editing system, the system comprising:(a) a gene editing tool; and(b) a donor DNA comprising a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein.2. The system of 1, wherein the gene editing tool comprises a nucleotide sequence encoding a CRISPR-Cas effector protein and a nucleotide sequence encoding a guide RNA.3. The system of 1 or 2, wherein the gene editing tool and the donor DNA are present on the same nucleic acid.4. The system of 1, wherein:the gene editing tool of (a) comprises: (i) nucleotide sequences encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; and (ii) a nucleotide sequence encoding a guide RNA, andthe donor DNA of (b) comprises a transposon, flanked by recognition sites that are recognized by the CAST complex.The system of 4, wherein:(i) and (ii) are present on the same nucleic acid; or(i), (ii), and the donor DNA are all present on the same nucleic acid.The system of 3 or 5, wherein the nucleic acid is a conjugation plasmid. The system of any one of 1-3, wherein the donor DNA is a donor template for homology directed repair (HDR).The system of any one of 1-7, wherein the system is configured to integrate the donor DNA into a target locus that: encodes a virulence factor, encodes a microbial toxin, encodes an antibiotic resistance gene, contributes to poor cardiac health, contributes to degradation of medicines, that contributes to problematic antibiotic resistance, or affects a metabolite that contributes to inflammation or risk of disease.The system of any one of 1-8, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.The system of 9, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.The system of any one of 1-10, wherein the bacteriocin is microcin V, B17, J25, N, orS.The system of any one of 1-10, wherein the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA. The system of any one of 1-12, wherein the donor DNA further comprises a therapeutic unit comprising: one or more nucleotide sequences encoding one or more expression products of interest.The system of 13, wherein said one or more expression products of interest comprises: one or more proteins that confer antibiotic resistance on a bacterium, one or more enzymes of a biosynthetic pathway, one or more proteins that confer resistance to a phage, one or more enzymes of a biosynthetic pathway, one or more enzymes in a carbon utilization pathway, one or more enzymes in a nitrogen utilization pathway, one or more enzymes in a sulfur utilization pathway, one or more detectablemarkers, one or more proteins that provide for detection of an analyte in a prokaryote, one or more antibodies or nanobodies; one or more antiinflammatory factors; one or more immunomodulatory or antitumor factors, or any combination thereof.The system of any one of 4-14, wherein the CAST complex comprises: (i) a Cas6 polypeptide, a Cas7 polypeptide, a Cas8 polypeptide, a tnsA polypeptide, a tnsB polypeptide, a tnsC polypeptide, and a tniQ polypeptide; or (ii) a Cas12k polypeptide, a tnsC polypeptide, a tnsB polypeptide, and a tniQ polypeptide.The system of any one of 1-15, wherein the bacteriocin unit comprises a bacteriocin gene cluster.The system of any one of 1-16, wherein all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.A prokaryotic cell comprising the self-selection DNA editing system of any one of 1-17.The prokaryotic cell of 18, wherein the prokaryotic cell is a bacterial cell. The prokaryotic cell of 18, wherein the prokaryotic cell is a proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonasaeruginosa or Enterobacter cell.The prokaryotic cell of 18, wherein the prokaryotic cell is an E. coli cell. The prokaryotic cell of any one of 18-21, wherein the prokaryotic cell is a donor cell.A genetically modified prokaryotic cell, comprising a bacteriocin unit integrated into its genomic DNA or into a mobile DNA element inside of the prokaryotic cell, where the bacteriocin unit comprises a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein.The genetically modified prokaryotic cell of 23, wherein the mobile DNA element is a conjugative transposon, a plasmid, a phage, a prophage, a phage satellite, or a phage plasmid.The genetically modified prokaryotic cell of 23 or 24, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.The genetically modified prokaryotic cell of 25, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.The genetically modified prokaryotic cell of any one of 23-26, wherein the bacteriocin is microcin V, B17, J25, N, orS.The genetically modified prokaryotic cell of any one of 23-26, wherein the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA.The genetically modified prokaryotic cell of any one of 23-28, wherein the prokaryotic cell’s genomic DNA comprises a genomic DNA edit in addition to the integrated bacteriocin unit.The genetically modified prokaryotic cell of any one of 23-29, wherein the prokaryotic cell is a bacterial cell.The genetically modified prokaryotic cell of any one of 23-29, wherein the prokaryotic cell is a proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifidum, B. infantis, and B. longum), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonasaeruginosa or Enterobacter cell.The genetically modified prokaryotic cell of any one of 23-29, wherein the prokaryotic cell is an E. coli cell.The genetically modified prokaryotic cell of any one of 23-32, wherein the bacteriocin unit comprises a bacteriocin gene cluster.The genetically modified prokaryotic cell of any one of 23-33, wherein all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.A method of modifying a microbial community, the method comprising: (a) introducing the genetically modified prokaryotic cell of any one of 23-34 to the microbial community; or(b) editing a target prokaryotic cell in situ to facilitate selfselection, wherein said editing comprises providing to the target prokaryotic cell, a DNA editing system and a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein, wherein the target prokaryotic cell is a member of the microbial community. The method of 35, wherein the method comprises said editing, and wherein the bacteriocin unit integrates into the target prokaryotic cell’s genomic DNA.The method of 36, wherein said editing comprises introducing the selfselection DNA editing system of any one of 1-17 into the target prokaryotic cell.The method of 37, wherein introducing the self-selection DNA editing system into the target prokaryotic cell comprises introducing the donor cell of 22 to the microbial community, and the donor cell introduces the self-selection DNA editing system to the target prokaryotic cell.The method of any one of 35-38, wherein the DNA editing system is a CRISPR-Cas system.The method of any one of 35-39, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.The method of 40, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.The method of any one of 35-41 , wherein, in said editing of (b), the bacteriocin is microcin V, B17, J25, N, orS.The method of any one of 35-41 , wherein, in said editing of (b), the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA.The method of any one of 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is a bacterial cell.The method of any one of 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is a proteobacteria, Enterococcus (e.g., E. faecalis and E. faecium), Bifidobacterium (e.g., B. bifid urn, B. infantis, and B. longurri), Bacteroides, Prevotella, Clostridium (e.g., C. innocuum), Ruminococcus (e.g., R. gnavus), Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcusfaecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa or Enterobacter cell.The method of any one of 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is an E. coli cell.The method of any one of 35-46, wherein the bacteriocin unit comprises a bacteriocin gene cluster.The method of any one of 35-47, wherein all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.A method of editing a prokaryotic cell to facilitate self-selection, the method comprising:providing a bacteriocin unit that comprises a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein to a target prokaryotic cell, wherein the bacteriocin unit integrates into a target DNA inside of the prokaryotic cell. The method of 49, further comprising providing a DNA editing system to the target prokaryotic cell.The method of 50, wherein DNA editing system is a CRISPR-Cas system.The method of any one of 49-51 , wherein providing the bacteriocin unit comprises introducing the self-selection DNA editing system of any one of 1-17 into the target prokaryotic cell.The method of 52, wherein introducing the self-selection DNA editing system into the target prokaryotic cell comprises introducing the donor cell of 22 to the target prokaryotic cell, and the donor cell introduces the self-selection DNA editing system to the target prokaryotic cell.The method of any one of claims 49-53, wherein the target DNA is the prokaryotic cell’s genomic DNA or is a mobile DNA element inside of the prokaryotic cell.EXPERIMENTAL EXAMPLES
[0209] The following examples are provided for purposes of illustration only, and are not intended to be limiting unless otherwise specified. Thus, the invention should in no way be construed as being limited to the following examples, but rather should be construed to encompass any and all variations which become evident as a result of the teaching provided herein.
[0210] Without further description, it is believed that one of ordinary skill in the art can, using the preceding description and the following illustrative examples, make and utilize the present invention and practice the claimed methods. The following working examples therefore are not to be construed as limiting in any way the remainder of the disclosure.
[0211] General methods in molecular and cellular biochemistry can be found in such standard textbooks as Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., HaRBor Laboratory Press 2001); Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999); Protein Methods (Bollag et al., John Wiley & Sons 1996); Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999); Viral Vectors (Kaplift & Loewy eds., Academic Press 1995); Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997); and Cell and Tissue Culture: Laboratory Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.Reagents, cloning vectors, cells, and kits for methods referred to in, or related to, this disclosure are available from commercial vendors such as BioRad, Agilent Technologies, Thermo Fisher Scientific, Sigma-Aldrich, New England Biolabs (NEB), Takara Bio USA, Inc., and the like, as well as repositories such as e.g., Addgene, Inc., American Type Culture Collection (ATCC), and the like.Example 1: Refactored bacteriocin systems facilitate controllable antimicrobial activity
[0212] FIG. 1A provides schematics that include a gene cluster encoding at least a bacteriocin (toxin) and its corresponding resistance mechanism (self-immunity). The gene cluster can be refactored to remove internal regulatory elements, placing the toxin gene — and in some cases, the entire operon — under an inducible promoter for controlled expression. This modified gene cluster can be referred to as the "selfselecting unit" or "enrichment unit". The ‘self-selecting unit’ can be encoded in thesame DNA molecule as the ‘therapeutic unit’, which can encode various elements, such as: therapeutic genes (conferring therapeutic benefits), a DNA sequence necessary for disruption or deletion of target genes, markers such as antibiotic resistance genes, fluorescent markers, or DNA barcodes. Together, the self-selecting unit and the therapeutic unit form the "genetic payload", intended for delivery into target bacterial cells. The genetic payload can be cloned into a plasmid, either suicide or replicative, which can also contain the necessary systems for delivery and genomic integration into the target bacteria. This delivery system can include, for example, a CRISPR-Associated Transposase (CAST) with a genome-targeting guide RNA.
[0213] E. coli strains carrying either the natural or refactored bacteriocin gene cluster alongside a therapeutic unit were grown in 3 ml_ of LB media with appropriate antibiotics (as needed) at 37°C for 16 hours. Indicator strains (E coli MG 1655, Maehl , or BL21) were grown in 3 mL LB under the same conditions. For the antimicrobial spot assay, 200 pL of the indicator strain culture was mixed with 7 mL of 0.7% LB soft agar (pre-warmed to ~50°C) and poured onto a sterile LB agar plate. To test different induction conditions, the synthetic inducer Anhydrotetracycline (ATC) was added to both the soft agar and LB agar at final concentrations of 2 nM, 20 nM, or 200 nM. A 2 pL aliquot of the overnight culture from each bacteriocin-encoding strain was spotted onto the indicator strain lawn. Plates were incubated at 37°C for 16 hours, after which zones of growth inhibition were recorded around each spot. An E. coli strain carrying only the therapeutic unit was used as a negative control for bacteriocin production.
[0214] In the absence of ATC, minimal or no antimicrobial activity was observed for all bacteriocin-encoding strains. However, as the concentration of ATC increased, clear zones of inhibition formed and expanded, matching or exceeding the diameter of those produced by natural bacteriocin systems (FIG. 1B). These results demonstrated that the refactored gene cluster allowed for tunable expression of bacteriocins, resulting in controllable antimicrobial activity.Example 2: Bacteriocins enable self-selection of CAST-mediated genome edits in E coli
[0215] FIG. 2A provides schematics of delivery and self-selection of a genetic payload using bacteriocins, including quantification of edited cells in the population.
[0216] E. coli MFDpir strains carrying either the editing plasmid encoding both the therapeutic and self-selecting units, or only the therapeutic unit (negative control), were used as donor strains for conjugation into E. coli MG1655. Donor strains were grown in LB medium with diaminopimelic acid (DAP; 0.3 mM) and carbenicillin (100 pg / mL) at 37°C, 250 rpm, while the recipient strain was grown in LB under the same conditions. Donor and recipient cells were washed with LB and LB + DAP, mixed in equal numbers of cells, and spotted onto LB agar plates containing 0.3 mM DAP. The plates were incubated at 30°C for 20 hours, after which the cells were scraped off, resuspended in 1 mL LB, and serially diluted for spot-plating onto LB and LB with 34 pg / mL chloramphenicol to assess the initial fraction of edited E. coli MG1655 cells.
[0217] Each conjugation suspension was normalized to GD600 = 0.1 in LB or LB + 200 nM ATC and transferred (200 pL) into a sterile 96-well plate. The cultures were incubated at 37°C with constant shaking for 20 hours in a plate reader to induce bacteriocin expression and enable self-selection of edited cells. After incubation, an aliquot from each culture was serially diluted and spot-plated on LB and LB with chloramphenicol to quantify the fraction of edited cells. This process was repeated for three passages, with cultures diluted to GD600 = 0.1 and grown for another 16-20 hours after each passage. At each stage, the ratio of edited cells was determined via spot-plating.
[0218] After the first round of growth and bacteriocin induction with 200 nM ATC, the fraction of edited E. coli cells carrying both the therapeutic and self-selecting units increased 19-fold compared to the initial timepoint (t=Oh), and up to 97x after three passages (FIG. 2B). In contrast, the fraction of edited cells in the population that received only the therapeutic unit remained unchanged throughout the experiment (FIG. 2B).Example 3: Bacteriocins collectively kill 85% of E. coli isolates while sparing other members of the microbiome
[0219] FIG. 6A-6B. (Fig. 6A) E. coli strains carrying the refactored bacteriocin gene cluster alongside a therapeutic unit were grown in 3 mL of LB media with appropriate antibiotics (as needed) at 37°C for 16 hours. Indicator strains (a collection of 72 E. coli isolates that represent the genetic diversity in the species, known as the ECOR collection,) were grown in 1 mL LB under the same conditions. For the antimicrobial spot assay, 100 pL of the indicator strain culture was mixed with 4 mL of 0.7% LB soft agar (pre-warmed to ~50°C) and poured onto a sterile LB agar plate. The synthetic inducer Anhydrotetracycline (ATc) was added to both the soft agar and LB agar at a final concentration of 200 nM. A 2 pL aliquot of the overnight culture from eachbacteriocin-encoding strain was spotted onto the indicator strain lawn. Plates were incubated at 37°C for 16 hours, after which zones of growth inhibition were recorded around each spot. An E. coli strain carrying only the therapeutic unit was used as a negative control for bacteriocin production.
[0220] The results demonstrate that three of the bacteriocins (V, B17 and S) killed at least 60% of the E. coli strains each. These data suggest that with just three bacteriocins, this selection strategy can be used in about 85% of E. coli strains.
[0221] (Fig. 6B) E. coli strains carrying the natural bacteriocin gene cluster alongside a therapeutic unit were grown in 3 mL of LB media with appropriate antibiotics (as needed) at 37°C for 16 hours. Indicator strains (E. coli W, E. coli BL21, E. coli MG 1655, Salmonella enterica, Klebsiella michiganensis, Pseudomonas aeruginosa, Pseudomonas simiae, Pseudomonas putida, Vibrio natriegens, and Chryseobacterium joostei) were grown in 3 mL LB at 37°C or 30°C for 16 hours. For the antimicrobial spot assay, 200 pL of the indicator strain culture was mixed with 7 mL of 0.7% LB soft agar (pre-warmed to ~50°C) and poured onto a sterile LB agar plate. A 2 pL aliquot of the overnight culture from each bacteriocin-encoding strain was spotted onto the indicator strain lawn. Plates were incubated at 37°C or 30°C for 16 hours, after which zones of growth inhibition were recorded around each spot. An E. coli strain carrying only the therapeutic unit was used as a negative control for bacteriocin production.
[0222] The results demonstrate that all four bacteriocins selectively target E. coli strains, except for bacteriocin N that killed Salmonella enterica, but that is a very close relative of E. coli and this result is consistent with existing literature. Overall, these four bacteriocins exhibit a high degree of specificity compared to most antibiotics.Example 4: Bacteriocins enable selection of CAST-mediated genome edits in E. coli across growth conditions in vitro
[0223] FIG. 7A-7C. (Fig. 7A) Schematics of delivery and self-selection of a genetic payload using bacteriocins, including quantification of edited cells in the population.
[0224] (Fig. 7B) E. coli MFDpir strains carrying the editing plasmid encoding both the therapeutic and self-selecting units, or only the therapeutic unit (negative control), were used as donor strains for conjugation into E. coli MG 1655. Donor strains were grown in LB medium with diaminopimelic acid (DAP; 0.3 mM) and carbenicillin (100 pg / mL) at 37°C, 250 rpm, while the recipient strain was grown in LB under the same conditions. Donor and recipient cells were washed with LB and LB + DAP, mixed inequal numbers of cells, and spotted onto LB agar plates containing 0.3 mM DAP. The plates were incubated at 30°C for 20 hours, after which the cells were scraped off, resuspended in 1 mL PBS, and serially diluted for spot-plating onto LB and LB with 34 pg / mL chloramphenicol to assess the initial fraction of edited E. coli MG1655 cells.
[0225] For selection in liquid media, each conjugation suspension was normalized to OD600= 0.1 in a final volume of 1mL of LB, LB + 200 nM ATc, M9, or M9 + 2,000 nM ATc in a sterile 96-well deep plate. The cultures were incubated at 37°C with constant shaking for 20 hours to induce bacteriocin expression and enable self-selection of edited cells. For self-selection in solid media, 10M cells were harvest from each conjugation suspension and resuspended in 20uL of LB media with the 800 nM ATc. These were spotted onto an LB agar plate containing the same amount of ATc. Plates were incubated at 37°C for 20 hours to induce bacteriocin expression.
[0226] After incubation, plates were scraped and resuspended in 1mL of PBS. An aliquot from each resuspension, or an aliquot from culture (for self-selection in liquid media) was serially diluted and spot-plated on LB and LB with chloramphenicol to quantify the fraction of edited cells. This process was repeated for three passages, with cultures diluted to OD600 = 0.1 (for self-selection in liquid media), and spots harvested and diluted to 10M cells (for self-selection in solid media) and grown for another 20 hours after each passage. At each stage, the ratio of edited cells was determined via spotplating.
[0227] The results demonstrate that after three rounds of growth and bacteriocin induction in LB broth, the fraction of edited E. coli cells carrying both the therapeutic and selfselecting units increased 77-fold compared to the control that only carries the therapeutic unit. This assay was then performed in different growth conditions to capture a natural environment for editing bacteria in-situ. In M9 broth, the fraction of edited cells that received the bacteriocin cluster increased by 5-orders of magnitude relative to edited cells that didn’t receive the bacteriocin, reaching 20% of total cells. In LB solid media, a similar trend was observed, with a 1000x increase in the fraction of cells edited that received the bacteriocin relative to the control edit.
[0228] (Fig. 7C) E. coli FDpir strains carrying the editing plasmid encoding both the therapeutic and self-selecting units, or only the therapeutic unit (negative control), were used as donor strains for conjugation into E. coli ECOR-56. Donor strains were grown in LB medium with diaminopimelic acid (DAP; 0.3 mM) and carbenicillin (100 pg / mL) at 37°C, 250 rpm, while the recipient strain was grown in LB under the same conditions. Donor and recipient cells were washed with LB and LB + DAP, mixed inequal numbers of cells, and spotted onto LB agar plates containing 0.3 mM DAP. The plates were incubated at 30°C for 20 hours, after which the cells were scraped off, resuspended in 1 mL PBS, and serially diluted for spot-plating onto LB and LB with 34 pg / mL chloramphenicol to assess the initial fraction of edited E. coli ECOR-56 cells.
[0229] Each conjugation suspension was normalized to OD600 = 0.1 in a final volume of 1mL of M9 + 2,000 nM ATc in a sterile 96-well deep plate. The cultures were incubated at 37°C with constant shaking for 20 hours to induce bacteriocin expression and enable self-selection of edited cell.
[0230] After incubation, an aliquot from culture was serially diluted and spot-plated on LB and LB with chloramphenicol to quantify the fraction of edited cells. This process was repeated for two passages, with cultures diluted to OD600 = 0.1 and grown for another 20 hours after each passage. At each stage, the ratio of edited cells was determined via spot-plating.
[0231] The results demonstrate that after two rounds of growth and bacteriocin induction in M9 broth, the fraction of edited E. coli cells carrying both the therapeutic and selfselecting units increased 5 orders of magnitude compared to the control that only carries the therapeutic unit. At that timepoint, the fraction of edited cells is 24% of the total cells.Example 5: Edited cells are maintained in the population after removing inducer
[0232] FIG. 8. Editing plasmid encoding both the therapeutic and self-selecting units were delivered into E. coli MG1655 via conjugation as described in Figure 7B. After 3 rounds of growth and self-selection in LB + 200 nM ATc or M9 + 2,000 nM ATc, cultures were then passaged an additional two times in LB or M9, in absence of inducer (ATc). After each passage, the ratio of edited cells was determined via spotplating.
[0233] The results demonstrate that after self-selection in both LB and M9 broth, the fraction of edited cells didn’t decrease, hence edited cells persisted without bacteriocin inducer. This suggests that bacteriocin self-selection created a stable population of edited cells.
[0234] See FIG. 9 for a schematic of one example embodiment in which bacteriocins are used for self-selection of edited cells in microbiomes.
[0235] Although the foregoing invention has been described in some detail by way of illustration and example for purposes of clarity of understanding, it is readily apparentto those of ordinary skill in the art in light of the teachings of this invention that certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.
[0236] Accordingly, the preceding merely illustrates the principles of the invention. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e. , any elements developed that perform the same function, regardless of structure.Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.
[0237] The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of present invention is embodied by the appended claims. In the claims, 35 U.S.C. §112(f) or 35 U.S.C. §112(6) is expressly defined as being invoked for a limitation in the claim only when the exact phrase "means for" or the exact phrase "step for" is recited at the beginning of such limitation in the claim; if such exact phrase is not used in a limitation in the claim, then 35 U.S.C. § 112 (f) or 35 U.S.C. §112(6) is not invoked.
Claims
1. CLAIMS2.What is claimed is:
1. A self-selection DNA editing system, the system comprising:4.(a) a gene editing tool; and5.(b) a donor DNA comprising a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein.
2. The system of claim 1, wherein the gene editing tool comprises a nucleotide sequence encoding a CRISPR-Cas effector protein and a nucleotide sequence encoding a guide RNA.
3. The system of claim 1 or claim 2, wherein the gene editing tool and the donor DNA are present on the same nucleic acid.
4. The system of claim 1, wherein:9.the gene editing tool of (a) comprises: (i) nucleotide sequences encoding polypeptides that form a CRISPR-associated transposase (CAST) complex; and (ii) a nucleotide sequence encoding a guide RNA, and10.the donor DNA of (b) comprises a transposon, flanked by recognition sites that are recognized by the CAST complex.
5. The system of claim 4, wherein:12.(i) and (ii) are present on the same nucleic acid; or13.(i), (ii), and the donor DNA are all present on the same nucleic acid.
6. The system of claim 3 or claim 5, wherein the nucleic acid is a conjugation plasmid.
7. The system of any one of claims 1-3, wherein the donor DNA is a donor template for homology directed repair (HDR).
8. The system of any one of claims 1-7, wherein the system is configured to integrate the donor DNA into a target locus that: encodes a virulence factor, encodes a microbial toxin, encodes an antibiotic resistance gene, contributes to poor cardiac health, contributes to degradation of medicines, that contributes to problematic antibiotic resistance, or affects a metabolite that contributes to inflammation or risk of disease.
9. The system of any one of claims 1-8, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.
10. The system of claim 9, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.
11. The system of any one of claims 1-10, wherein the bacteriocin is microcin V, B17, J25, N, or S.
12. The system of any one of claims 1-10, wherein the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA.
13. The system of any one of claims 1-12, wherein the donor DNA further comprises a therapeutic unit comprising: one or more nucleotide sequences encoding one or more expression products of interest.
14. The system of claim 13, wherein said one or more expression products of interest comprises: one or more proteins that confer antibiotic resistance on a bacterium, one or more enzymes of a biosynthetic pathway, one or more proteins that confer resistance to a phage, one or more enzymes of a biosynthetic pathway, one or more enzymes in a carbon utilization pathway, one or more enzymes in a nitrogen utilization pathway, one or more enzymes in a sulfur utilization pathway, one or more detectable markers, one or more proteins that provide for detection of an analyte in a prokaryote, one or more antibodies or nanobodies; one or more anti-inflammatory factors; one or more immunomodulatory or antitumor factors, or any combination thereof.
15. The system of any one of claims 4-14, wherein the CAST complex comprises: (i) a Cas6 polypeptide, a Cas7 polypeptide, a Cas8 polypeptide, a tnsA polypeptide, a tnsBpolypeptide, a tnsC polypeptide, and a tniQ polypeptide; or (ii) a Cas12k polypeptide, a tnsC polypeptide, a tnsB polypeptide, and a tniQ polypeptide.
16. The system of any one of claims 1-15, wherein the bacteriocin unit comprises a bacteriocin gene cluster.
17. The system of any one of claims 1-16, wherein all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.
18. A prokaryotic cell comprising the self-selection DNA editing system of any one of claims 1-17.
19. The prokaryotic cell of claim 18, wherein the prokaryotic cell is a bacterial cell.
20. The prokaryotic cell of claim 18, wherein the prokaryotic cell is a proteobacteria, Enterococcus, Bifidobacterium, Bacteroides, Prevotella, Clostridium, Ruminococcus, Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus28.faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa or Enterobacter cell.
21. The prokaryotic cell of claim 18, wherein the prokaryotic cell is an E. coli cell.
22. The prokaryotic cell of any one of claims 18-21, wherein the prokaryotic cell is a donor cell.
23. A genetically modified prokaryotic cell, comprising a bacteriocin unit integrated into its genomic DNA or into a mobile DNA element inside of the prokaryotic cell, where the bacteriocin unit comprises a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein.
24. The genetically modified prokaryotic cell of claim 23, wherein the mobile DNA element is a conjugative transposon, a plasmid, an archaeal virus, an archaeal viral satellite, a phage, a prophage, a phage satellite, or a phage plasmid.
25. The genetically modified prokaryotic cell of claim 23 or claim 24, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.
26. The genetically modified prokaryotic cell of claim 25, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.
27. The genetically modified prokaryotic cell of any one of claims 23-26, wherein the bacteriocin is microcin V, B17, J25, N, orS.
28. The genetically modified prokaryotic cell of any one of claims 23-26, wherein the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA.
29. The genetically modified prokaryotic cell of any one of claims 23-28, wherein the prokaryotic cell’s genomic DNA comprises a genomic DNA edit in addition to the integrated bacteriocin unit.
30. The genetically modified prokaryotic cell of any one of claims 23-29, wherein the prokaryotic cell is a bacterial cell.
31. The genetically modified prokaryotic cell of any one of claims 23-29, wherein the prokaryotic cell is a proteobacteria, Enterococcus, Bifidobacterium, Bacteroides, Prevotella, Clostridium, Ruminococcus, Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa or Enterobacter cell.
32. The genetically modified prokaryotic cell of any one of claims 23-29, wherein the prokaryotic cell is an E. coli cell.
33. The genetically modified prokaryotic cell of any one of claims 23-32, wherein the bacteriocin unit comprises a bacteriocin gene cluster.
34. The genetically modified prokaryotic cell of any one of claims 23-33, wherein all proteincoding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.
35. A method of modifying a microbial community, the method comprising:42.(a) introducing the genetically modified prokaryotic cell of any one of claims 23- 34 to the microbial community; or43.(b) editing a target prokaryotic cell in situ to facilitate self-selection, wherein said editing comprises providing to the target prokaryotic cell, a DNA editing system and a bacteriocin unit comprising a nucleotide sequence encoding a bacteriocin protein and / or a nucleotide sequence encoding a bacteriocin immunity protein, wherein the target prokaryotic cell is a member of the microbial community.
36. The method of claim 35, wherein the method comprises said editing, and wherein the bacteriocin unit integrates into the target prokaryotic cell’s genomic DNA.
37. The method of claim 36, wherein said editing comprises introducing the self-selection DNA editing system of any one of claims 1-17 into the target prokaryotic cell.
38. The method of claim 37, wherein introducing the self-selection DNA editing system into the target prokaryotic cell comprises introducing the donor cell of claim 22 to the microbial community, and the donor cell introduces the self-selection DNA editing system to the target prokaryotic cell.
39. The method of any one of claims 35-38, wherein the DNA editing system is a CRISPR- Cas system.
40. The method of any one of claims 35-39, wherein the nucleotide sequence encoding a bacteriocin protein and / or the nucleotide sequence encoding a bacteriocin immunity protein are operably linked to an inducible promoter.
41. The method of claim 40, wherein the inducible promoter is Anhydrotetracycline (ATC) inducible.
42. The method of any one of claims 35-41, wherein, in said editing of (b), the bacteriocin is microcin V, B17, J25, N, or S.
43. The method of any one of claims 35-41, wherein, in said editing of (b), the bacteriocin is microcin V and the bacteriocin unit comprises the genes cvai, cvaC, cvaB and cvaA.
44. The method of any one of claims 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is a bacterial cell.
45. The method of any one of claims 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is a proteobacteria, Enterococcus, Bifidobacterium, Bacteroides, Prevotella, Clostridium, Ruminococcus, Lactococcus, Staphylococcus, Bacillus, Streptomyces, Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa or Enterobacter cell.
46. The method of any one of claims 35-43, wherein the genetically modified prokaryotic cell or the target prokaryotic cell is an E. coli cell.
47. The method of any one of claims 35-46, wherein the bacteriocin unit comprises a bacteriocin gene cluster.
48. The method of any one of claims 35-47, wherein all protein-coding sequences of the bacteriocin unit are operably linked to one single promoter, which is functional in a prokaryotic cell, as part of the same operon.
49. A method of editing a prokaryotic cell to facilitate self-selection, the method comprising:57.providing a bacteriocin unit that comprises a nucleotide sequence encoding a bacteriocin protein and a nucleotide sequence encoding a bacteriocin immunity protein to a target prokaryotic cell, wherein the bacteriocin unit integrates into a target DNA inside of the prokaryotic cell.
50. The method of claim 49, further comprising providing a DNA editing system to the target prokaryotic cell.
51. The method of claim 50, wherein DNA editing system is a CRISPR-Cas system.
52. The method of any one of claims 49-51, wherein providing the bacteriocin unit comprises introducing the self-selection DNA editing system of any one of claims 1-17 into the target prokaryotic cell.
53. The method of claim 52, wherein introducing the self-selection DNA editing system into the target prokaryotic cell comprises introducing the donor cell of claim 22 to the target prokaryotic cell, and the donor cell introduces the self-selection DNA editing system to the target prokaryotic cell.
54. The method of any one of claims 49-53, wherein the target DNA is the prokaryotic cell’s genomic DNA or is a mobile DNA element inside of the prokaryotic cell.