Apsab combines a nuclease / helicase protein and an argonaute-like protein to cleave DNA
The ApsAB system, comprising ApsA and ApsB proteins with helicase and nuclease domains, addresses antibiotic resistance by efficiently targeting and eliminating plasmids, offering a new tool for genome editing and plasmid curing.
Patent Information
- Application Number
- PCT/EP2025/059196
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-03
- Filing Date
- 2025-04-03
- Publication Date
- 2025-10-09
Smart Images

Figure IMGF000059_0001 
Figure IMGF000060_0001 
Figure IMGF000060_0002
Abstract
Description
1 APSAB COMBINES A NUCLEASE / HELICASE PROTEIN AND AN ARGONAUTE- LIKE PROTEIN TO CLEAVE DNA BACKGROUND OF THE INVENTION 5
[0001] Antibiotic resistance is of major public health concern, having been responsible for more than 1.2 million deaths in 20191. Mobile genetic elements, particularly plasmids carrying antibiotic resistance genes (ARG), are key contributors to antibiotic resistance spread within and between species2. While often carrying features beneficial under specific adverse environments, plasmids frequently confer a fitness cost to their bacterial host3,4. This is expected to lead to their 10 loss over time and to the integration of beneficial traits into the host chromosome5. However, some plasmids may persist over a long period even in the absence of selective pressure, a phenomenon known as the plasmid paradox6,7. Solutions to this paradox include compensatory mutations improving the fitness of plasmid-carrying bacteria8-10and high conjugation rates11. On the other hand, bacteria have developed defence mechanisms to counteract invasion by parasitic 15 DNAs, some of them able to limit plasmid diffusion12–14.
[0002] Carbapenemases are enzymes responsible for the hydrolysis of carbapenems, broad- spectrum antibiotics of major clinical importance. In western Europe, the class D enzyme OXA- 48 is the most frequent carbapenemase in Enterobacterales15–17. This is generally attributed to the high conjugation rate of the IncL group pOXA-48 plasmids18. These 56-76 kb-long plasmids 20 induce variable fitness costs to their hosts19. In addition, OXA-48 is also encoded on the chromosome15,20,21, mainly in E. coli isolates of sequence type ST38, a major cause of extra- intestinal infections22. These ST38 isolates belong to disseminated clones with the blaOXA-48 gene embedded in a complete or partial composite transposon, Tn6237, that likely derives from a pOXA-48 plasmid. 25
[0003] To counteract invasion by parasitic MGEs, bacteria have developed a plethora of defence mechanisms, a minority of them, such as prokaryotic Argonautes (pAgos), able to target plasmids at least in vitro (Koonin et al., Annu Rev Microbiol 71, 233-261, (2017), Makarova et al., BiolDirect 4, 29, (2009), Mayo-Munozet al., Cell Rep 42, 112672 (2023), Swarts et al. Nucleic Acids Res 43, 5120-5129 (2015)).
[0004] Eukaryotic Argonaute proteins (eAgos) bind small RNA guides to direct the RNA- induced silencing complex (RISC) towards complementary RNA, which either catalyzes endonucleolytic cleavage of the target RNA or indirectly silences the target RNA. Hegge et al., Nucleic Acids Research, 2019, Vol. 47, No. 11 5809-5821. Prokaryotes also encode Argonaute proteins (pAgos), which share a high degree of structural homology with eAgos in a two lobe and four domain (N-PAZ / MID-PIWI) architecture. Id. Several pAgos directly target DNA instead of RNA. Id. These pAgos use 5-end phosphorylated small interfering DNAs (siDNAs) for recognition and successive cleavage of complementary DNA targets. Id. These can be programmed with short synthetic siDNA which allows them to target and cleave dsDNA sequences of choice in vitro and to be used as a universal restriction endonuclease for in vitro molecular cloning. Furthermore, a diagnostic application termed NAVIGATER was developed which enables enhanced detection of rare nucleic acids with single nucleotide precision. Id. However, due to the thermophilic nature (optimum activity temperature > 65 °C) and low levels of endonuclease activity at the relevant temperatures (20-37°C), it is unlikely that the well-studied TtAgo, PfAgo and MjAgo are suitable for genome editing. Id.
[0005] There is a need in the art for unravelling the complexities of plasmid persistence, particularly in mesophilic bacteria, which can lead to the discovery of new DNA-modifying activities and development of molecular tools, particularly for genome editing. The current invention fulfils this need in the art.BRIEF SUMMARY OF THE INVENTION
[0006] The invention encompasses isolated ApsA or ApsB protein, compositions comprising an ApsA and / or ApsB protein.
[0007] The invention encompasses also an ApsAB system comprising the isolated ApsA and ApsB proteins.3
[0008] In some embodiments, the ApsA protein is a protein having the sequence of SEQ ID NO: 1, a homolog thereof, or a variant thereof; and / or the ApsB protein is a protein having the sequence of SEQ ID NO: 2, a homolog thereof, or a variant thereof.
[0009] In particular embodiments, the ApsA protein comprises a SF2 helicase domain and a 5 nuclease domain.
[0010] In more particular embodiments, the helicase domain contains six helicase motifs having the sequences: T(G / A)XGK; DE(E / L / V / Q)(H / E)XXY; (P / A)EXX(L / T / V)XXX(L / T / V); SAT; (S / T)X(Y / F)X(S / G)AXXG(L / I / V)(N / D); and Q(A / T)(I / V / L)GRXER, where X is any amino acid and the amino acids between parentheses indicate alternatives. In preferred embodiments, the first 10 helicase motif is TGFGK or TASGK; the second helicase motif is DELHEAY or DEEHEAY; the third helicase motif is PEVMLLRLL, PEAMLLRLL or PEIALLRVL; the fifth helicase motif is SSYKSAGTGLN or SHFQGAGTGLN; and / or the sixth helicase motif is QAVGRVER or QTIGRTER; preferably the six helicase motifs are chosen from: TGFGK, DELHEAY, PEVMLLRLL, SAT, SSYKSAGTGLN and QAVGRVER; TGFGK, DELHEAY, 15 PEAMLLRLL, SAT, SSYKSAGTGLN and QAVGRVER; TASGK, DEEHEAY, PEIALLRVL, SAT, SHFQGAGTGLN and QTIGRTER.
[0011] In more particular embodiments, the nuclease domain contains a nuclease motif having the sequence: (G / E)XXXE(X)20-50(Y / F)D(X)12-17(L / I / V / M)DXKX(W / Y), 20 where X is any amino acid; the amino acids between parentheses indicate alternatives; (X)12-17 and (X)20-50 represent a sequence of 12 to 17 and 20 to 50 amino acids, respectively, where each amino acid of the sequence is any amino acid. In preferred embodiments, the nuclease motif is chosen from: GNVGE(X)31FD(X)11IDVKRW and GNIGE(X)31FD(X)11IDVKNW.
[0012] In particular embodiments, the ApsB protein contains a nucleic acid guide 5’end binding 25 motif having the sequence: DXYXXXK(X)12-17Q(A / G) (X)35-90E(L / I)XXK,where X is any amino acid; the amino acids between parentheses indicate alternatives; (X)i2-i7 and (X)35-9O represent a sequence of 12 to 17 and 35 to 90 amino acids, respectively, where each amino acid of the sequence is any amino acid.
[0013] In particular embodiments, the ApsB protein has five B-sheets, ordered 32145, where B- strands 2 and 4 are anti-parallel to the other B-strands.
[0014] In particular embodiments, the ApsA and / or ApsB protein has at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the protein of SEQ ID NO: 1 or SEQ ID NO:2.
[0015] In particular embodiments, the ApsA and / or ApsB is a protein of any of SEQ ID NO: 3 to 6 or a variant thereof having at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 3 to 6.
[0016] In particular embodiments, the ApsA and / or ApsB protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 ammo acid changes from any of the proteins of SEQ ID NO: 1 to 6.
[0017] In particular embodiments, the ApsA or ApsB protein has the sequence of any of SEQ ID NO: 1 to 6.
[0018] In particular embodiments, the ApsA and / or ApsB is a protein of any of SEQ ID NO: 43 to 337.
[0019] In particular embodiments, the ApsA protein has endonuclease activity.
[0020] In particular embodiments, the ApsB protein binds a nucleic acid guide molecule; more particularly wherein the nucleic acid guide is single-stranded RNA, DNA or mixed RNA / DNA molecule; wherein the nucleic acid guide is small interfering DNA molecule, in particular 5-end phosphorylated small interfering DNA molecule; and / or wherein the guide molecule is complementary to the sequence of a DNA target.
[0021] In some embodiments, the composition or ApsAB system further comprises the nucleic acid guide molecule. In some embodiments, the composition or the ApsAB system comprises a complex of the ApsA and ApsB proteins. In particular embodiments, the nucleic acid guidemolecule is bound to the ApsB protein. In particular embodiments, the composition or ApsAB system has nucleic acid-guided endonuclease activity and / or cleaves the DNA target.
[0022] In some embodiments, the ApsA or ApsB protein is labelled.
[0023] The invention encompasses antibodies that bind to the ApsA or ApsB protein.
[0024] The invention encompasses an ApsAB complex.
[0025] The invention encompasses a recombinant nucleic acid encoding the ApsA or ApsB protein.
[0026] In some embodiments, the recombinant nucleic acid encodes the ApsA and ApsB protein.
[0027] In some embodiments, the recombinant nucleic acid has at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences SEQ ID NO: 14 to 19.
[0028] The invention encompasses a vector comprising a recombinant nucleic acid encoding the ApsA and / or ApsB protein.
[0029] In some embodiments, the vector is a mammalian expression vector.
[0030] In some embodiments, the vector is a bacterial expression vector.
[0031] In some embodiments, the vector has the sequence of any of SEQ ID NOs 7-13.
[0032] The invention encompasses a cell comprising a vector comprising a recombinant nucleic acid encoding the ApsA and / or ApsB protein.
[0033] The invention encompasses a method of producing an ApsA or ApsB protein comprising introducing an expression vector encoding an ApsA or ApsB protein into a host cell and expressing the ApsA or ApsB protein from the vector. In some embodiments, the method comprises isolating the expressed protein from the host cell.
[0034] The invention encompasses a method of cleaving a DNA target comprising providing an ApsB protein; binding the ApsB protein to a single stranded nucleic acid guide that is complementary to a dsDNA target to form an ApsB-guide complex; contacting the ApsB -guidecomplex with the dsDNA target; contacting the dsDNA target with an ApsA protein; and unwinding and cleaving the dsDNA target with the ApsA protein. In some embodiments, the single stranded nucleic acid guide that is complementary to a dsDNA target is 5 ’-phosphorylated. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in mammalian cells. In some embodiments, the method is performed in bacteria.
[0035] The invention encompasses a kit comprising a nucleic acid guide and, an ApsB protein, and optionally an ApsA protein. In some embodiments, the kit comprises a fluorescent reporter nucleic acid, which has a fluorescent group and a quenching group.DESCRIPTION OF THE DRAWINGS
[0036] Figure 1A-C: pOXA-48 plasmids induce variable fitness costs and are unstable in three ST38 strains Three different variants of pOXA-48 were introduced from three K. pneumoniae isolates into three ST38 E. coli strains. A. Alignment of the three pOXA-48 plasmids used in the study and displayed with Easyfig59. IS 7s are indicated in red and the bkuyx \. 48 gene is in pink. The two composite transposons are indicated by double headed arrows. Tn6237 carrying / > / <2OXA-48 is only present in the pOXA-48_l variant. Percentages of nucleotide identities are indicated by a grey gradient as indicated in the Figure key B. Comparison of the relative doubling time (DT) of plasmid- free strains and TCs calculated at exponential phase in the absence of meropenem in LB or M9 glucose minimum media. The relative DT is the ratio of means of technical replicates to one of the plasmid-free strain mean values. Three to five independent TCs for each condition were tested in three independent experiments with five technical replicates. C. Plasmid loss during serial passages in the absence of meropenem. The ratio of plasmidcarrying colonies over the whole population at days 0, 5 and 10 was quantified by using the number of colonies growing on meropenem containing plates as a proxy for the number of plasmid-carrying colonies. For boxplots in b and c, the median is indicated by the line, the box bounds the 25thand 75thquartiles and whiskers bound the minimum and maximum values excluding outliers. Outliers are values > 1.5 x interquartile range. For b, normal distribution of data was assessed with Shapiro-Wilk normality test. Statistical analysis was performed withpairwise two sample t-test and / ?- values (Source Data) were FDR-adjusted. Ns: no significant difference; *p<0.05, **p<0.01, ***p<0.001, ****£><0.0001.
[0037] Figure 2A-F: Experimental evolution of pOXA-48 transconjugants selects Tn6237 transposition and pOXA-48 loss. Experimental evolution was performed by 28-day passages of TCs from the three ST38 E. coli strains in LB or in M9 minimal medium. A. Evolution of the relative DT of ST38_l / pOXA-48_l TCs population and of the parental strain as control. Meropenem was added every day (MEM / day) or every three days (MEM / 3 days) in TC lineage cultures at subinhibitory concentration. B. Quantification of plasmid loss by WGS of whole populations. Plasmid loss was estimated by calculating the ratio of the number of reads on three pOXA-48 regions of same size, two outside Tn6237 and one inside. C. Comparison of the relative DT of three WaoxA-48 integrants and of the original transconjugant. The WaoxA-48 integrants, one with Tn6237 integrated in the chromosome at gadW and two with Tn6237 integrated in the IncFII endogenous plasmid. Of note, the integrant at gadW had a mutation in malT frequently present in TC and control evolved colonies (Chr: chromosome). D. Evolution of the relative DT of ST38 1 pOXA-48_l, pOXA-48_2 or pOXA-48_3 TC populations in the presence of meropenem at subinhibitory concentration and of the parental strain. E. Evolution of the relative DT from ST38 2 and ST38 3 pOXA-48_l TCs in the presence of meropenem at subinhibitory concentration. F. Evolution of the relative DT of ST38 2 pOXA-48_l TCs in M9 medium in the presence of meropenem at subinhibitory concentration. Boxplots show median, box bounds 25thand 75thquartiles, whisker bounds minimum and maximum excluding outliers and outliers are value > 1.5 x interquartile range. Each point represents an independent lineage. DTs were determined during exponential growth phase in the absence of meropenem. The relative DT is calculated as the ratio of means of technical replicates of plasmid-free strains and TCs compared to the mean value of one of the evolved plasmid-free lineages collected the same day for a, d, e and f. For c, the relative DT is calculated as the ratio of means of technical replicates of plasmid- free strains, TCs and WaoxA-48 integrants compared to the mean value of one of the plasmid-free strain. Normal distribution of data was assessed with Shapiro-Wilk normality test. Statistical analysis was performed with pairwise two sample t-test and / ?- values (Source Data) were FDR- adjusted. Ns: no significant difference; *j><0.05, **p<0.01, ***p<0.001, ****p<0.0001.
[0038] Figure 3A-F: Mutations in the F3141-F3140 operon led to pOXA-48 stabilization.A. Organization of the genomic island containing the F3141-F3140 operon. The organization of the genomic island and its insertion at tRNA-leu is identical for the three ST38 strains used in the study. Positions of the mutations of the analysed mutants named by letters are indicated by vertical dash. In black, IS 7 insertion, in green, a deletion in F3141 and in red, SNPs (strain J, L1444* and strain N, V1425W). B. DT comparison of ten F3141-F3140 mutant strains as indicated in (a) and of the original TC measured during exponential growth phase. Five replicates of three independent experiments were performed in the absence of meropenem in LB medium for ST38 1 strain and in M9 medium for ST38 2. The relative DT is calculated as the ratio of means of technical replicates of plasmid-free strains or TCs compared to the mean value of one of the plasmid-free strains. C, D and E. pOXA-48 stability in F3141 -F3140 mutants after 10-day passages in LB medium in the absence of meropenem. C. Quantification of pOXA-48_l carriage in three ST38 1 and three ST38 2 mutants. Letters define the mutant strain as in (A) and (B). D. Quantification of the carriage of the three pOXA-48 plasmids in ST38 1 and ST38 3 TCs deleted for the F3141-F3140 operon. E. Quantification of pOXA-48 plasmids carriage following complementation of F3141-F3140 operon deletion. A plasmid copy of the F3141-F3140 operon was expressed under the control of the PBAD inducible promotor in ST38 1AF3141-F3140 TC. TO refers to the value obtained following one 24-hour passage in LB glucose 0.4% and tl to the value determined after two additional passages in LB arabinose 0.02%. pHV7-empty is used as negative control. Each dot represents an independent experiment. Boxplots show median, box bounds 25thand 75thquartiles, whiskers bound minimum and maximum values excluding outliers and outliers are value > 1.5 * interquartile range. For B, normal distribution of data was assessed with Shapiro- Wilk normality test. Statistical analysis was performed with pairwise two sample t-test and p- values (Source Data) were FDR-adjusted. Ns no significant difference; *p<Q.Q5, **p<0.01, ***7?<0.001, ****7?<0.0001.
[0039] Figure 4A-D: F3141-F3140 (ApsAB) acts as an antiplasmid system active against low and high copy number plasmids. A. Activity of F3141-F3140 against non-IncL plasmids with different replicons and copy numbers: pACYt: pl5A ori, ca. 10 copies; pBbE8c: ColEl ori, ca. 20-30 copies; pUC19: pMBl ori, >100 copies; pBbS8c: pSClOl ori, ca. 5 copies. CNR146C9-9 pKCP: incFII / IncFIB, ca.1-2 copies; B. Comparison of pOXA-48 conjugation frequency between ST38_1 WT (wild type) and ST38_1∆F3141-F3140 performed in three independent experiments. C. Comparison of transformation efficiency to ST38_1 WT and ST38_1∆F3141-F3140 by using electroporation of pACYt and pBbE8c D. Quantification of the ColE1 derivative pBbE8c 5 following induction of the F3141-F3140 operon cloned under the pBAD promotor and inserted into MG1655 E. coli chromosome (TnF3141-F3140). TnzeoR used as control carries the zeocin resistance gene in place of the F3141-F3140 operon. The plasmid copy number was determined on DNA preparations by qPCR by using the ∆Ct method with primers in pBbE8c replication origin and rpsI (chromosomal gene) as reference gene (Table 5). Each dot represents an 10 independent experiment. Boxplots show median, box bounds 25thand 75thquartiles, whisker bounds minimum and maximum excluding outliers and outliers are value > 1.5 × interquartile range. For b, c and d, normal distribution of data was assessed with Shapiro-Wilk normality test. Statistical analysis was performed with Log10-transformed data by using a pairwise two sample t-test and p-values (Source Data) were FDR-adjusted. Ns no significant difference; *p<0.05, 15 **p<0.01, ***p<0.001, ****p<0.0001.
[0040] Figure 5: ApsAB does not target eight tested phages. Serial dilutions of high titre lysates of phages lambda, T4, P1, 186cIts, CLB_P2, LF82_P8, AL505_P2, T5 and T7 spotted on MG1655 strain and its isogenic derivatives chromosomally carrying the miniTnapsAB (apsAB under the control of the arabinose inducible promoter pBAD) or the miniTnzeoR as a negative 20 control. Experiments were performed in three independent biological replicates and showed reproducible results. Dilution factor is indicated in Log10.
[0041] Figure 6A-D: blaOXA-48 integration and pOXA-48_1 loss require the apsAB operon A. Comparison of the relative DT between the original TC and ST38_1Z103∆apsB TC B. Quantification of plasmid carrying strains after apsAB deletion in the original TC. C. Evolution 25 of the relative DT of ST38_1Z103∆apsAB whole population and of the parental strain. D. Quantification of plasmid loss by whole genome sequencing of whole populations. Plasmid loss was estimated by calculating the ratio of the number of reads on two pOXA-48 regions of same size, one outside Tn6237 and one inside. Each point represents an independent experiment. Boxplots show median, box bounds 25thand 75thquartiles, whisker bounds minimum andmaximum excluding outliers and outliers are value > 1.5 x interquartile range. Calculated at exponential phase growth in the absence of meropenem. DTs were determined during exponential growth phase in the absence of meropenem. The relative DT is calculated as the ratio of means of technical replicates of plasmid-free strains and TCs compared to the mean value of one of the evolved plasmid-free lineages from the same day for c. For a, the same method was used. Normal distribution of data was assessed with Shapiro-Wilk normality test. Statistical analysis was performed with LoglO-transformed data by using a pairwise two sample t-test and p- values (Source data) were FDR-adjusted. Ns no significant difference.
[0042] Figure 7A-D: ApsAB associates a protein with helicase and nuclease domains and a novel Argonaute-like protein. A. Predicted structure of ApsA as determined with Alphafold2. The conserved PD-(D / E)XK nuclease domain with active site residues (1420D...1433DV1435K) in red and the four motifs of the helicase domain in cyan218GFGKS222, blue532DELH535, grey823SAT825and purple1140QAVGRVER1147are indicated. B. Effect of mutations replacing key residues of the putative helicase (K221A, E533A) or nuclease (K1435A) domains of ApsA and of the putative 5’ end guide binding motif (K413A) of ApsB on pOXA-48_l maintenance. Effect on pOXA-48_2 and _3 are shown in Fig. 10. apsAB and variants were expressed under the control of the PBAD inducible promotor in ST38 1AF3141-F3140 TCs. TO refers to the value obtained following one 24-hour passage in LB glucose 0.4% and tl to the value determined after two additional passages in LB arabinose 0.02%. pHV7-empty is a negative control. C. Predicted structure of ApsB as determined with Alphafold2. The conserved motif acting as 5 ’-end guide binding motif in the MID domain of pAgos,409Y,413K,431Q,489K, is indicated as red balls; the RNAse H fold of the PIWI domain and the PAZ domain of pAgos are not conserved in ApsB. D. Sequence logo generated by WebLogo from the alignment of ApsB homologs in the region corresponding to the MID domain highlighting the conserved residues. A zoom on the ApsB MID domain structure at the409Y,413K,431Q,489K residues is provided in the lower panel. The C- terminal743N residue possibly interacting with the putative guide nucleic acid in the binding pocket is also indicated.
[0043] Figure 8: ApsAB belongs to a broad family of putative antiplasmid systems.Quantification of the ColEl derivative pBbE8c following induction of two homologs of theapsAB operon encoding proteins with 92 / 92% and 28 / 25% aa sequence identity with ApsA / B respectively. Operon were cloned in Tn7 following PCR amplification from the ST219 CNR36C9 and ST10 CNR81D10 isolates under the PBAD promotor and inserted in MG1655 E. coli chromosome. TnZeoR used as control carries the zeocin resistance gene in place of the apsAB operon. The plasmid copy number was determined by qPCR by using the ACt method with primers in pBbE8c replication origin and rpsl (chromosomal gene) as reference gene., Each dot represents an independent experiment. Boxplots show median, box bounds 25thand 75thquartiles, whisker bounds minimum and maximum excluding outliers and outliers are value > 1.5 x interquartile range. Statistical analysis was performed with LoglO-transformed data by using a pairwise two sample t-test and / ?- values (Source Data) were FDR-adjusted. ***j><0.001, ****p<Q 0001.
[0044] Figure 9: Tn6237 insertion is associated with pOXA-48 plasmid loss. A. B. and D. Quantification of plasmid loss by population WGS. Plasmid loss was estimated by calculating the ratio of the number of reads on three pOXA-48_l regions of same size, two outside and one inside Tn6237 respectively. A. pOXA-48_l in ST38 1 evolved in LB medium is gradually lost from day 7 (D7) to day 28 (D28). B. pOXA-48_3 in ST38 1 TCs evolved in LB medium is partially lost at D28. C. Comparison of the relative DT of WaoxA-48 integrants and the original transconjugant calculated at exponential phase growth in the absence of meropenem. The relative DT is calculated as the ratio of means of technical replicates of plasmid-free strains, TCs or / > / <2OXA-48 integrants compared to the mean value of one of the plasmid-free strains. Chr for chromosome. D. pOXA-48_l in ST38 2 evolved in M9 medium remains unchanged between D7 and D28. Each Point represents an independent experiment. Boxplots show median, box bounds 25thand 75thquartiles, whisker bounds minimum and maximum excluding outliers and outliers are value > 1.5 x interquartile range. For c, normal distribution of data was assessed with Shapiro- Wilk normality test. Statistical analysis was performed by using a pairwise two sample t-test and / ?- values (Source data) were FDR-adjusted. ****p<0.0001.
[0045] Figure 10: Effect of mutations replacing key residues of the putative helicase (K221A, E533A) or nuclease (K1435A) domains of ApsA and of the 5’ nucleic acid binding domains (K413A) of ApsB on pOXA-48_2 and pOXA-48_3 maintenance. apsAB and variantswere expressed under the control of the PBAD inducible promotor in ST38 1AF3141-F3140 TCs. tO refers to the value obtained following one 24h-passage in LB glucose 0.4% and tl to the value determined after two additional passages in LB arabinose 0.02%. pHV7-emty is a negative control.DETAILED DESCRIPTION OF THE INVENTION
[0046] To investigate the factors involved in blaOXA-48 gene integration into another replicon, the inventors first transferred three different pOXA-48 plasmids into three different ST38 E. coli strains by conjugation. The inventors observed that the three plasmids were unstable in ST38 strains and were rapidly lost when transconjugants (TC) were grown without carbapenem pressure (Fig. 1C). The inventors then performed experimental evolution of the pOXA-48 TC in the three ST38 strains. Subinhibitory concentration of carbapenem was added every passage or every three passages, leading to alternate selection for growth and for resistance. Rapidly, for one of the five strain / plasmid pairs tested (ST38_l / pOXA48_l), blaOXA-48 integrated lineages emerged and became dominant at day 28 (Fig. 2B).
[0047] In addition, in bacteria that kept the pOXA-48 plasmid during experimental evolution, the inventors observed convergent mutations (mainly inactivation) of the apsAB operon (also named operon F3140-3141} (Fig. 3A). The inventors showed that apsAB inactivation increased the stability of pOXA-48 plasmid and decreased the fitness cost associated with these plasmids in ST38. (Fig. 3B and 3C). Conjugation of the pOXA-48 plasmids to genetically modified derivatives of ST38 1 deleted for F3141-F3140 confirmed the stabilization of the plasmid in the absence of F3141-F3140 (Fig. 3D).
[0048] Furthermore, complementation in ST38 1AF3747-F3740 strain with a pHV7-derivative plasmid expressing F3141-F3140 under the control of the arabinose inducible promoter led to more than 80% loss of each pOXA-48 plasmid at 48 hours following induction (Fig. 3E). ApsAB was also found to decrease the pOXA-48 conjugation efficiency (Fig. 4B). Finally, the inventors showed that pOXA-48 destabilization, in the presence of an active copy of apsAB, was needed for the selection of blaOXA-48 integration (Fig. 6D). In addition, the inventors observed that apsAB inactivation in ST38 strains also increased the stability of diverse plasmids differing bytheir origins of replication and copy-numbers, such as pl5A, ColEl and pMBl-type plasmids (Fig. 4A and 4D). This confirmed that apsAB was a defense system able to target and eliminate various plasmids.
[0049] ApsAB location, embedded in a genomic island, and its distribution among ST38 isolates suggested that it has been horizontally acquired. Structural modelling of ApsB by using AlphaFold2 revealed a bi-lobe structure reminiscent of that of Argonautes (Fig. 7C) and the conservation of the amino-acids involved in the guide 5 ’end binding site in prokaryotic long-B Argonautes (Fig. 7D).
[0050] However, ApsB lacks the typical folding of argonaute nuclease PIWI domain and thus is predicted to have no nuclease activity per se (Fig. 7C). On the other side ApsA was predicted to bring nuclease and helicase activities to the ApsAB complex (Fig. 7A). Targeted mutations of amino-acids involved in the guide 5’ end binding site of ApsB as well as in the helicase and nuclease active sites of ApsA decreased the efficiency of pOXA-48 destabilization by the complex (Fig. 7B), strengthening inventors functional hypotheses.
[0051] In recent years, a plethora of bacterial defence systems have been discovered. But, only a few of them, such as CRISPR, restriction-modification, pAgos and Wadjet were also described to have antiplasmid activity36,37. Recently a novel system, DdmDE, was shown to have contributed to plasmid elimination from the seventh pandemic Vibrio cholerae13. By combining different in silico search strategies, the inventors uncovered a broad family of ApsA homologs that combine helicase and nuclease domains (Tables 4 and 6). Strikingly, this family was subdivided into two classes associated with two families of argonaute-like proteins, whose prototypes were ApsB and DdmE. However, ApsB had no sequence homology with DdmE and the conservation of the guide 5’ end binding site characteristic of the MID domain of long pAgos was unique to ApsB-like proteins. Unique to ApsB-like proteins is the conservation of a MID- like domain of long pAgos, which the inventors showed to be essential for its activity. At least three ApsAB related systems were identified in E. coir. ApsAB from ST38, and two ApsAB systems sharing 90% and 25 % sequence identity, respectively, with that of ST38 E. coli. The three systems showed an antiplasmid activity (Fig. 8).
[0052] Altogether, this suggests that ApsB is involved in a guide-dependent recognition of the target, ApsA bringing the nuclease activity, similarly to what has recently been described for a group of catalytically inactive long true pAgos38. While these Argonaute systems confer immunity also via abortive infection38, the inventors do not have evidence that ApsAB kills plasmid invaded cells. Despite the absence of similarity between DdmE and ApsB, both DdmDE and ApsAB share an antiplasmid activity, like a second E. coli system only 28 / 25% similar to ST38 ApsAB (Fig. 8). This suggests that the compendium of systems the inventors brought out represents a new diverse family of defence systems acting on plasmids. These systems can provide new tools for genomic engineering and plasmid curing, including those carrying ARGs.
[0053] The invention encompasses isolated ApsA or ApsB proteins, including ApsA-like and ApsB -like proteins, isolated ApsAB complexes, compositions comprising ApsA and ApsB proteins, antibodies against these proteins, nucleic acids encoding these proteins, compositions comprising these nucleic acids, proteins, complex and / or antibodies and cells and kits comprising these nucleic acids, proteins and / or antibodies. The invention further encompasses methods for making and using these compositions comprising nucleic acids, proteins and / or antibodies.Isolated or purified ApsA and ApsB proteins
[0054] The invention encompasses “isolated or purified” ApsA and ApsB proteins.
[0055] As used herein the term “ApsA protein” means a protein that has the sequence of any of SEQ ID NOs 1, 3, and 5, variants, mutants, homologs, and functional fragments thereof, and ApsA-like proteins.
[0056] As used herein the term “ApsB protein” means a protein that has the sequence of any of SEQ ID NOs 2, 4, and 6, variants, mutants, homologs, and functional fragments thereof, and ApsB-like proteins.
[0057] The terms “isolated or purified” mean modified “by the hand of humans” from the natural state; in other words if an object exists in nature, it is said to be isolated or purified if it is modified or extracted from its natural environment or both. For example, a polynucleotide or a protein / peptide naturally present in a living organism is neither isolated nor purified; on the other hand, the same polynucleotide or protein / peptide separated from coexisting molecules in itsnatural environment, obtained by cloning, amplification and / or chemical synthesis is isolated for the purposes of the present invention. Furthermore, a polynucleotide or a protein / peptide which is introduced into an organism by transformation, genetic manipulation or by any other method, is “isolated” even if it is present in said organism. The term “purified” as used in the present invention means that the proteins / peptides according to the invention are essentially free of association with the other proteins or polypeptides, as is for example the product purified from the culture of recombinant host cells or the product purified from a non-recombinant source. Various techniques can be used to obtain purified protein according to the invention, for example affinity chromatography or gel filtration.
[0058] When using recombinant techniques, the ApsA and ApsB proteins can be produced intracellularly, in the periplasmic space, or directly secreted into the medium. The ApsA and ApsB proteins may be produced simultaneously (co-expression) to form ApsAB complexes or separately. If the ApsA and / or ApsB protein is produced intracellularly, as a first step, the particulate debris, either host cells or lysed fragments, are removed, for example, by centrifugation or ultrafiltration. Carter et al., Bio / Technology 10: 163-167 (1992) describe a procedure for isolating proteins which are secreted to the periplasmic space of E. coli. Briefly, cell paste is thawed in the presence of sodium acetate (pH 3.5), EDTA, and phenylmethylsulfonylfluoride (PMSF) over about 30 min. Cell debris can be removed by centrifugation.
[0059] Where the ApsA and / or ApsB protein is secreted into the medium, supernatants from such expression systems are generally first concentrated using a commercially available protein concentration filter, for example, an Amicon or Millipore Pellicon ultrafiltration unit. A protease inhibitor such as PMSF may be included in any of the foregoing steps to inhibit proteolysis and antibiotics may be included to prevent the growth of adventitious contaminants. The protein composition prepared from the cells can be purified using, for example, hydroxylapatite chromatography, gel electrophoresis, dialysis, and affinity chromatography, with affinity chromatography being the preferred purification technique.
[0060] Affinity chromatography can be used to purify the ApsA and / or ApsB proteins using antibodies against ApsA and / or antibodies against ApsB proteins. In some embodiments the ApsA protein and / or the ApsB protein can be fused to an amino acid sequence encoding a protein or a fragment of a protein that bind DNA molecules or that bind to other cellular molecules, such as poly-histidine tag, glutathione S-transferase (GST), maltose binding protein (MBP), S-tag. In that case affinity chromotography targeting these amino-sequences can be used, such as immobilized metal affinity chromatography, glutathione affinity chromatography, chromatography on amylose resin, chromatography on S-agarose.
[0061] The matrix to which the affinity ligand is attached is most often agarose, but other matrices are available. Mechanically stable matrices such as controlled pore glass or poly(styrene- divinyl) benzene allow for faster flow rates and shorter processing times than can be achieved with agarose. Other techniques for protein purification such as fractionation on an ion-exchange column, ethanol precipitation, Reverse Phase HPLC, chromatography on silica, chromatography on heparin SEPHAROSE™ chromatography on an anion or cation exchange resin (such as a polyaspartic acid column), chromatofocusing, SDS-PAGE, hydrophobic interaction chromatography, and ammonium sulphate precipitation are also available.
[0062] ApsAB complexes may be isolated by co-expressing ApsA and ApsB proteins using recombinant techniques and isolating the complex. Alternatively, ApsA and ApsB proteins may be produced separately using recombinant techniques. The recombinant ApsA and ApsB proteins are then mixed together to form ApsAB complexes. ApsAB complexes may be purified using standard purification methods that are well-known in the art and disclosed herein.
[0063] As used herein, the terms ApsA or ApsB “variant” or “mutant” are used interchangeably herein to designate a polypeptide (protein or peptide) having an amino acid sequence that differs from the native sequence (AspsA sequence of SEQ ID NO: 1, 3 or 5 or ApsB sequence of SEQ ID No: 2, 4 or 6) by the mutation (substitution, insertion and / or deletion) of one or more amino acids. A homolog is a particular type of variant or mutant, that has structural similarity due to a common evolutionary ancestor.
[0064] ApsA and ApsB proteins (including variants, mutants, homologs, ApsA-like and ApsB- like proteins) and fragments thereof according to the invention are functional proteins or fragments having the activity of the native sequence (ApsA sequence of SEQ ID NO: 1 or ApsB sequence of SEQ ID NO: 2) as disclosed herein, which means that they have retained ApsA or ApsB activity. ApsA is an helicase and a nuclease having a SF2 helicase domain with conserved residues as disclosed herein and a nuclease domain containing a a PD-(D / E)XK- superfamily nuclease motif as disclosed herein. ApsB is a novel argonaute-like protein involved in a guidedependent recognition of a nucleic acid target. ApsB has a conserved motif characteristic of the MID domain of long prokaryotic Argonaute proteins (pAgos) that interacts with the 5 ’-end of guide nucleic acids as disclosed herein. ApsB has five B-sheets, ordered 32145, where B-strands 2 and 4 are anti-parallel to the other B-strands. ApsB does not have the nuclease catalytic tetrad DEDX. The nucleic acid guide molecule and ApsB protein form a complex. The guide molecule provides target specificity to the complex by comprising a nucleotide sequence that is complementary to a sequence of a target dsDNA. ApsA protein is guided to the target dsDNA molecule by virtue of the association of ApsB with the nucleic acid guide molecule. The ApsA protein and the ApsB protein form a complex and ApsA protein cleaves the target dsDNA. The nucleic acid-guided DNA cleavage activity of the ApsAB system may be measured using a method of cleaving a DNA target as disclosed herein, such as for example using the plasmid stability assay disclosed in the examples. For example, the activity of an ApsA or ApsB variant, mutant, homolog, ApsA-like or ApsB-like protein according to the invention may be determined by introducing a multicopy plasmid carrying a resistance gene for an antibiotic into bacteria expressing inducible-copies of ApsA and ApsB; and testing the decrease in the number of resistant colonies after ApsAB induction. Alternatively, the activity of ApsAB may be tested in an vitro plasmid stability test by following the plasmid DNA degradation by electrophoresis after plasmid DNA incubation with purified or semi-purified ApsA and ApsB and a guide sequence complementary to the plasmid DNA.
[0065] The invention encompasses proteins and peptides based on the ApsA and ApsB amino acid sequences. In one embodiment, the protein or peptide is a recombinant ApsA and / or ApsB protein or peptide. In one embodiment, the protein or peptide is an ApsA and / or ApsB protein orpeptide that comprises or consists of at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, or 1500 amino acids of any of the sequences of SEQ ID NO: 1 to 6. As used herein, the term “ApsAB system” encompasses ApsAB (E. coli ST38) composed of ApsA protein of SEQ ID NO: 1 and ApsB protein of SEQ ID NO: 2 and ApsAB related systems. ApsAB related systems in E. coli include: CNR36C9_ApsA-like protein of SEQ ID NO: 3 having 90% identity with SEQ ID NO: 1 and CNR36C9_ApsB-like protein of SEQ ID NO: 4 having 90% identity with SEQ ID NO: 2; CNR81D10_ApsA-like protein of SEQ ID NO: 5 having 25% identity with SEQ ID NO: 1 and CNR81D10_ApsB-like protein of SEQ ID NO: 6 having 25% identity with SEQ ID NO: 2.
[0066] The invention encompasses mutants and homologs of any of the proteins of SEQ ID NOs 1-6. The mutant or homolog can contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50 ,60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 amino acid changes from any of the proteins of SEQ ID NOs 1-6. Preferably, the changes are conservative amino acid changes.
[0067] Conservative substitutions are substitutions of one amino acid with another having similar chemical or physical properties (size, charge or polarity), which substitution generally does not adversely affect the biochemical, biophysical and / or biological properties of the protein. Examples of conservative substitutions may be within the following groups : Group 1 -small aliphatic, non-polar or slightly polar residues (A, S, T, P, G); Group 2-polar, negatively charged residues and their amides (D, N, E, Q); Group 3 -polar, positively charged residues (H, R, K); Group 4-large aliphatic, nonpolar residues (M, L, I, V, C); and Group 5-large, aromatic residues (F, Y, W); or within the groups of basic amino acids (R, K, H), acidic amino acids (D, E), polar amino acids (Q, N), hydrophobic amino acids (M, L, I, V), aromatic amino acids (F, W, Y), and small amino acids (G, A, S, T).
[0068] The invention encompasses a variant ApsA and / or ApsB protein having at least at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of SEQ ID NO: 1 or SEQ ID NO:2.
[0069] The percent amino acid sequence or nucleotide sequence identity is defined as the percent of amino acid residues or nucleotides in a Compared Sequence that are identical to theReference Sequence after aligning the sequences and introducing gaps if necessary, to achieve the maximum sequence identity and not considering any conservative substitutions for amino acid sequences as part of the sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways known to a person of skill in the art, for instance using publicly available computer software such as the GCG (Genetics Computer Group, Program Manual for the GCG Package, Version 7, Madison, Wisconsin) pileup program, or any of sequence comparison algorithms such as BLAST (Altschul etal., J. Mol. Biol., 1990, 215, 403, FASTA or CLUSTALW. When using such software, the default parameters, are preferably used. The BLASTP program uses as default a word length (W) of 3 and an expectation (E) of 10.
[0070] The invention encompasses a variant of the proteins of SEQ ID NOs 3-6 having at least at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NOs 3-6.
[0071] The invention encompasses the proteins in the Figures and Tables and variants having at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with these proteins. In various embodiments, the ApsA and ApsB proteins may be chosen from the proteins in Table 3 and Table 4. In some embodiments, the ApsA protein is chosen from a protein of any of SEQ ID NO: 43 to 147 or SEQ ID NO: 253 to 295, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 43 to 147 or SEQ ID NO: 253 to 295. In particular embodiments, the ApsA protein is chosen from a protein of any of SEQ ID NO: 43-46, 48-54, 56-79, 81-84, 86-87, 89 to 102, 104 to 113, 115, 117 to 124, 126, 128 to 132, 134 to 136, 138, 140 to 147, or SEQ ID NO: 253 to 295, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 43 to 102, 104 to 113, 115, 117 to 124, 126, 128 to 132, 134 to 136, 138, 140 to 147, or SEQ ID NO: 253 to 295. In more particular embodiments, the ApsA protein is chosen from a protein of any of SEQ ID NO: 253 to 261; preferably a protein of any of SEQ ID NO: 253 a 255, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 253 to 261; preferably a protein of any of SEQ ID NO: 253 to 255. In some embodiments, the ApsB protein is chosen from a protein of any of SEQ ID NO: 148 to 252 or SEQ ID NO: 296 to 337, or a variant thereof having at least80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 148 to 252 or SEQ ID NO: 296 to 337. In particular embodiments, the ApsB protein is chosen from a protein of any of SEQ ID NO: 148 to 252, SEQ ID NO: 296 to 334, 336 or 337, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 148 to 252, SEQ ID NO: 296 to 334, 336 or 337. In particular embodiments, the ApsB protein is chosen from a protein of any of SEQ ID NO: 148 to 171, 173, 175, 177 to 178, 190, 193 to 197, 199 to 201, 204, 205, 207, 213, 216, 218, 223, 225, 226, 232, 234, 236, 238, 296 to 336 or 337, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 148 to 171 , 173, 175, 177 to 178, 190, 193 to 197, 199 to 201, 204, 205, 207, 213, 216, 218, 223, 225, 226, 232, 234, 236, 238, 296 to 336, or 337. In more particular embodiments, the ApsB protein is chosen from a protein of any of SEQ ID NO: 148 to 152, 154 to 156, 158 or 160; preferably a protein of any of SEQ ID NO: 148, 149, 151, 152, 155, 158 or 160, or a variant thereof having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any SEQ ID NO: 148 to 152, 154 to 156, 158 or 160; preferably a protein of any of SEQ ID NO: 148, 149, 151, 152, 155, 158 or 160.
[0072] The invention encompasses systems related to the ApsAB system, and their ApsA and ApsB counterparts.
[0073] In various embodiments, these systems are constituted by two proteins phylogenetically related to the ApsAB proteins described in E. coli ST38. These relationships may be strong or loose (less than 25% amino-acid identity).
[0074] In various embodiments, the ApsA and ApsB proteins may be from Gamma- proteobacteria, Beta-proteobacteria or Cyanobacteria. The ApsA and ApsB proteins may be from any one of the bacteria species disclosed in Table 3 and Table 4. In particular embodiments, ApsA and ApsB proteins are from Enterobacterales (i.e., enterobacteria) such as E.coli, in particular E.coli ST38 isolates, Salmonella enterica, Escherichia albertii, Enterobacter sp. 120016, and others.
[0075] In various embodiments, the system has a protein related to ApsB containing a binding pocket for the guide 5’ end.
[0076] The invention encompasses “ApsB-like proteins.” As used herein the term “ApsB-like protein” means a protein having at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%,95%, 96%, 97%, 98%, or 99% identity with the protein of SEQ ID NO:2, containing the motif:DxYxxxK(x)12-17Q(A / G) (x)35-90E(L / I)xxK, having five B-sheets, ordered 32145, where 13- strands 2 and 4 are anti-parallel to the other 13-strands and not having the nuclease catalytic tetradDEDX.
[0077] In particular, ApsB-like proteins share with ApsB the DxYxxxK(x)12-17Q(A / G) (x)35-90E(L / I)xxK motif where the BOLD amino-acids are involved in the binding pocket for the guide5’ end and are identical to amino-acids found in Long-B pAgos:- ApsBGuide 5’end binding motif:407DTYTQL413K(X)i7431Q432G(X)52485ELWF489KMIPNLNELTDTPIARTNLIKLEEDQLTIIQRLLAPVSNIYTIDFMVQRFTKERKEKSADYY ARIHQEVI<SC VRQI<LGLEAGQEVI<YELHC LPNYHHVFFFLAPATNPNSQAHRTLAERI ETLCQQLTAENYDLSRLIQGLFSLHLKMVMLEKASERFSVSPTYFNSTLFLNARLTQPV TQKSGTGVMEAFELDIYASEYNELAFTLHKRKFLVEPEDEMLLSLDDTSVWFNINNRR LKARRKLDARDSKLDFFRERSGYGECQAYTYNWMNAACERLSELEIPHQPIPFQATH EVNQFATDLDQQLANTLLWNNGVKFSATQEAYFFDALAIQFPGYQLWPLASLKHSQQTGFSELPADTSILVLNAVEGERSNSISQQDNASVEYNDFYAAFADARKQPELNWDTY TQLKLDRLQGWLNQQPLPWLQGMNIDRKLLD AIDLINERS KRDPAQYEIDLTKPHSR LKSAVTLLNSKIRRTKTELWFKESLLNQHHILLPDLADGHYTAYAVRKTKNNPFLLGY VEMI<TEC GQLRVVDTGITEGEFEYLSVDHPALGRLI<I<LFDI<SFYLYDHTTDVLLTTYN SSRVPRLIGPAQFNIVDSYVYQEQEKALAERQGDKFSGYTITRSAKPEQNVLPYLISPAR SKYDSLAKTQKMKHfflHYLQPHENGVFVLVSDAQPTNPTIAHPNLVENLLIWDAQGKA VDVFNHPLTGAYLNSFTLDMLRSGESSKSSIFTKLARLMVEN (SEQ ID NO:2)- CNR36C9ApsB-like (90% identity)Guide 5’end binding motif:407DTYTQL413K(X)i7431Q432G(X)52485ELWF489KMIPNLNELTDTPIARTNLIKLEEDQLTTIQRLLAPVSNIYTIDFMVQHFTKERKEKSADY YARIHQEVKTCVRQKLGLEAGQEVKYELHCLPNYHHVFFFLAPAAAPNSLAHRTLAER lETLCQRLTAENYDLSRLIQGLFSLHLKMVMLEQASERFSVPPTYFNSTFYLNARLSQPV TQKSGTGVMEAFELDIYASEYNELAFTLHKRKFLVEPEDELHLSLDDTCVWFNIDNRRL KARRKLDARDSKLDFFRERSGYGECQAYTYNWMNAACERLSELEIPHQPIAFQATHE VNQFATDLDQQLTNTLLWNNGVEFSATQEAYFFDTLAIQFPGYQLWPLASLKHSQQT GFSELPASTSILVLNAVDEERSNSIRQQDNESVEYNDFYAAFADARKQPELNWDTYTQLKLDRLQGWLNQQPLPVVLQGMNIDHKLLDAIDFINEQLTSNPTQYEIDLTKPHSRLKS AVTLLNSKVRRTKTELWFKESLLNQHHIPLPDLADGHYTAYAVRKTKSYLPLLGYVEL KIEHGQLRVVDTGIAEGKLDYLSVDPPSLGRLKKLFDKSFYLYDHTADVLLTTYNSSRV1PRLIGPAQFNIVDSYAYQEQEKTLAERKGDKFNGYAITRSAKPDQNVLPYLISPGRSKYDSLTI<AQI<MI<HHHIYLQPHENGVFVLVSDAQPTNPTIARPNLVENLLIWDAQGI<AVDVFSHPLTGVYLNSFTLDMLRSGESSKCSIFAKLARLMVEN (SEQ ID NO:4)- CNR81D10 ApsB-like (25% identity)Guide 5’end binding motif:406DYYSRL412K(X)i5428Q429G(X)35465ELWL469KMTDVTESQQHNETDWMAAEWGLQPPDYAPKARTQLFDIHDTADVMLEKLLQLPSTN GWGIDAICITKVKEYSMKWDDFFIELHREIQNILSLKLGTKNNREFQYQTLRKGNSYLV MFLSPDARPYHQWVNELKIKGFDQYKIAQTVLGLRLKLHFLDNAGKYFEVSHGAYNS EIFLGAWQKEEPDKRGNYYVDALQSDLFYSPDYGDLTFTLKGICFRAKAEELGVLSD YGRILIKGKKSLQLLEKVDGRKYRKMPYMRFASKAYSGCKNHAENLVQNVLGDILSQ ADIAFSPRVFCATHSQLDFLTTNIKTLQRPLVIIDNLDNGEEPEAHQRFIHYLKKELDAV DIISPRELIDLDKDVNYLVLNQSKKINGSSISLGEDTFLNSTWQALNEAKAGKKGLLDY YSRLKVDFILDFEKNHQRVFQGCNIEAIEQNKKDPVTLWKKVLNPVSVAKLEMIRSE LWLKEAVLVRKKLTGFSINDGEYILVSIRCSSSGNRRGNDFYGVATPWFSHGVMNIA EPKMYTDENDFRNNVPGMNINCIDKLRDDNFYLVDTINQVVIKRYNTVRIPRLLGSAER CSLQVAEDNHDKINKFGAADKTCLPYYLTPKPKSHAVSQYQQLFIEKRGNEALLFCTK SNQPNQIMDRATLVYNLVATTFDGKEADVMAQPATELLLRTLTYDTTRYNEVSQTSLL EKISRLMLHN (SEQ ID NO:6).
[0078] However, ApsB-like proteins do not share with long-B pAgos the typical RNAseH fold of the PIWI domain characterized by five B-sheets, ordered 32145, where the B-strand 2 is antiparallel to other B-strands. Instead, ApsB-like proteins have five B-sheets, ordered 32145, where B-strands 2 and 4 are anti-parallel to the other B-strands. ApsB-like proteins do not have the nuclease catalytic tetrad DEDX.
[0079] In various embodiments, the ApsB homolog or ApsB-like protein has a sequence as shown in Table 3 or Table 4; in particular selected from the group consisting of SEQ ID NO: 148 to 252 and SEQ ID NO: 296 to 337; more particularly selected from the group consisting of SEQ ID NO: 148 to 252, SEQ ID NO: 296 to 334 and SEQ ID NO: 336 or 337. In some embodiments, the ApsB homolog or ApsB-like protein has a sequence selected from the group consisting of SEQ ID NO: 148 to 171, 173, 175, 177 to 178, 190, 193 to 197, 199 to 201, 204, 205, 207, 213, 216, 218, 223, 225, 226, 232, 234, 236 and 238; particularly selected from the group consisting of SEQ ID NO: 148 to 152, 154 to 156, 158 and 160; more particularly selected from the group consisting of SEQ ID NO: 148, 149, 151, 152, 155, 158 and 160.
[0080] The invention encompasses “ApsA-like proteins.” As used herein the term “ApsA-like protein” means a protein having at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%,95%, 96%, 97%, 98%, or 99% identity with the protein of SEQ ID NO: 1, containing a SF2 helicase domain and a nuclease domain. The helicase domain has conserved motifs : 1) aT(G / A)xGK motif corresponding to helicase motif I; 2) a DE(E / L / V / Q)(H / E)xxY motif corresponding to helicase motif II; 3) a (P / A)Exx(L / T / V)xxx(L / T / V) motif 4) a SAT motif corresponding to helicase motif III 5) a (S / T)x(Y / F)x(S / G)AxxG(L / I / V)(N / D) motif encompassing the helicase motif Va; and 6) a Q(A / T)(I / V / L)GRXER motif corresponding to helicase motif VI. The nuclease domain contains the conserved (G / E)xxxE(X)20- 50(Y / F)D(X)12-17(L / I / V / M)DxKx(W / Y) motif encompassing the catalytic motif of the PD-(D / E)XK nuclease superfamily.
[0081] In various embodiments, the system has a protein related to ApsA, which contains a SF2 helicase domain and a nuclease domain.-ApsAHelicase motifs1)217-221:TGFGK2)532-538: DELHEAY3)804-812:PEVMLLRLL4)823-825: SAT5)1034-1044: SSYKSAGTGLN6)1140-1147: QAVGRVER motif1383GNVG1387E (X)311419F142OD (X)II1432IDVKR1437WMRKSWTIEEDC KLLTLVRRYFSTLVSHNRLNATMPFSQQLHDAFDSPDRDAAALLYRL EQAKILGFASRPGGDPTKQLFRCLINNDLALYDYSLTFPTLRKSLHPDTIAAALNLFTISN PHEPLSNTINEIATALHLAPMQVEKILLESGQIIINSYRNYERVGEKIINNNLQDLISRQIP DITLVKDINACRAQVSQLYHVHERDGAEVIFSSDGTGFGKSYGVLQGYVEYLERFAKK TKSNDLFPEGGFTNLLFMSPQKSQIDLDNSQKDKILASGGEFICILSRQDIADLDFRDWA TGLKNRDRYIQWYKGAKDSKYIGSTMRSLYYHVSQIDRCEEQLKKLTTYGSQDTSYER DILEEQLKSCRHSIRSTIESACKLLFGPDSEKASIKEYIRRGLQARQKRMQNAETARKTV KPELSVYEVYFELIKQVLPFEVCQYRPSVLLMTTNKFDTSTYRLVPCQRGEGVRFQSISF DLLIGGKLLPKAPQISSVATAGHTVQVDYLRNEHFRRNPDCPFRQKNIRFTVIIDELHEA YTRLEEACHVKLITQENNLAHVISVVGRIHNAVLSLERRNKPKEAQTTFEQEMVKFITIL RDLLAKKCELSPGTTLGSLLEMFRDQLGAFEVNGDAAERIIAITRNVFSFNPKMYVNEE GLKRIRMRNSEGDITRTELYYEVENDTSDTNPTLHDLFQLISVILAACSEITNRHFKRWVKNGGQDNS S S QNTPLGQFVD AANNVAGWRHIFDRTTDKNLLIDHFYTYLQPKTVFTM TPIAELNYVNSGAERTIILAFEMDLVQELPEVMLLRLLTGTHNKVIGLSATSGFSHTKN GNFNRHFLARYSRDLGYRVVERQKTDIDTLRALRGLRASIRNVDFRVFDDNQLKLTDI YQNCEIYRRTYDNFFDALKKPLEYVLKNTYKQRQCMRELEALLLAAWEGKNSLILSLS GTFKRAFISAWRTHQSAWRQQYGMHSRCDEKTDINKKHDQILTFTPFKGRHTVHLVFF DSSLANVEDIRQETYLQNSNTVLVFMSSYKSAGTGLNYFVKYHDGDINDINAPHLDVD FERLVLINSPFYSEVKDNSGNLNTLPNYVTVLKHYADDDINVHRLADFSVNFAHGENY RLLMAEHDMSLFKVWQAVGRVERRDTLLKTEIFLPRDVFRNVAFQFAALSEDRGNE WSESMSLLNHRLMEECEKLSQSQSFSDAEQRHAFEQAIVENGRRIDAIHKRVLKKDWI NQVRAGNLEYLELCNLFRDSDSFTDPQRWLEKLQANSLYAANRQMQFIHHSLFIDRQQ GNQMILLCHKRGPDGLVHRDYSALSDFAGGAREYRPELTLFPQYRNDVDFTPGNLVGE LIRECNNIQETAFKKWVPNPRLVPLLKGNVGEYLFDKVLKHYRVTPLTDQQVFGRLAP LVYEFFDRFIEVGDDLLCIDVKRWSTQLDDLTRAENTLEKSKNKIRQIRNIASQKADTD GQKQILAALSGRYERIRFVYLNVAYSQNPNNLMWQDNVDHTIHYLNLLQTDYQYYQP I<NRESGRAQGISI<LRMALDINPMLLTLLGVEI<LLTI<GI<IS (SEQ ID NO:1)- CNR36C9 ApsA-like (90% identity)Helicase motifs1)217-221:TGFGK2)534-540: DELHEAY3)806-814:PEAMLLRLL4)825-828: SAT5)1036-1046: SSYKSAGTGLN6)1142-1149: QAVGRVER motif1385GNVG1389E (X)311421F1422D (X)II1434IDVKR1439WMRKSWnEEDCKLLTLVRQLFSALVSHNRLNATMPFSQQLHDAFDSPDRDAAALLYRL EQAKILGFASRPGGDPTKQLFRCLISNDLALYDYSLTFPTLRKALHPDTVAAALNHFTIS NPHEPLSNTTNEIATALHLAPMQVEKILIDSGQITINSYRKCERVGEKNINNNLQDLISRQI PDITLIKEINACRAQVSQLYHVHERDGAEVIFSSDGTGFGKSYGVIQGYVEYLERFAKT QKSDDLFPEGGFTNLLFMSPQKSQIDLDSSQKEKILAASGEFICVLSRKDVADLDFMDW ASGLKNRDRYIQWYEGAKGSKYIGGAMRSLNYHVLQIDRCEEQLKKLTTYGSQDTNY EREILEEQLKNCRHSIRNTIESACKLLFGPDSEKASIKEYIRRGLQARQERMQNAETARK PGKLEPKISVHEVYFELIKQVLPFEVCQYRPSVLLMTTNKFDTSTYRLAPRQRGEGVRF ESVGFDLLIGGKLTPKDPQISTVAAAGHTGQVTYLRDEHFRRNPDCPFRQKNIRFTVIID ELHEAYTRLEETCHVKLITQENNLAHVISVAGRIHNAVLSLERRNKPKEAQTTFEQEM VI<FITTLRNLLAEI<C ELSPGTRLGSILEMFRDQLGAFEVNGDAAERIISITRNVFSFNPI<M YVNEEGLKRIRMRNSEGDITRTELYYEVENDANDTNPTLHDLFQLVSVILAACSEITNR HFKRWVKNGGQDNSSSQNTPLGQFVDAANNVAGWRHIFDRTTDKNLLIDHFYTYLQ PKTVFTMTPIAELNYVNRGAERTIILAFEMDLVQELPEAMLLRLLTGTHNKVIGLSATS GFSHTI<NGNFNRHFLAHYSRDLGYRVVEREI<ADIDTLI<ALRGLRASIRNVDFRVFDDI< QLKLTDIYQNCEIYRRTYDNFFDALKKPLEYDLKNTYKRRQCQRELEALLLAAWEGKNSLILSLSGTFKRAFISAWRTHQTTWRQQYGMHSRCDEKTDNGKKHDQILTFTPFKGRH TVHLVFFDSPLANVEDIRQETYLQNSNTVLVFMSSYKSAGTGLNYFVKYHDGDINDIN APRLDVDFERLVLINSSFYSEVI<DNSGNLNTLPNYVTVLI<HYADDDITVHI<LADFNVNFAHGENYRLLMAEHDMSLFKVWQAVGRVERRDTLLKTEIFLPRDVFRNVAFQFAALS EDSGNEWSESMSLLNHRLMEECEKLSQGQSFNNAEQRLTFEQAIVANGRRIDEIHKRV LKTDWINKVRAGNLDYLEICNLFRDPDSFTDPQRWLAKLQANPLYTANRQMQSVHDA LFIDRQQGNQTILLCHKRGPDGLAHRDYSALSDFAGGAREYRPELTLFPQYRNDVDFTP GNLVGELIRECDNIQEKAFKKWVPNPRLVPLLKGNVGEYLFDKVLKSYGVTPLSDQQV FERLEPLVYEFFDRFIEVGDDLLCIDVKRWATQLDDLTRAEETLEKSDNKIRQIRNIASQ KADTEGQKQLQTALAGRYERIRFIYLNVAYSQNPNNLMWQDNVDHTIHYLNLLQTDY QYYQPKNRESGRAQENSKLRMTLDINPMLLTLLGVEKLPTKGKVS (SEQ ID NO:3)- CNR81D10 ApsA-like (25% identity)Helicase motifs1)50-54: TASGK2)339-345: DEEHEAY3)602-610: PEIALLRVL4)621-623: SAT5)830-840: SHFQGAGTGLN6)933-940: QTIGRTER motif1166GNIG1170E (X)3I1202F1203D (X)H1215IDVKN1220WMQSELNHNIRHLPLPQSVREAITAARKQAGTFYAFTEADTPLLIQSQDGTASGKSYNVF AQYIESVNPAINPGNSHRNLVFITPLKSQIDIPEQLIARARDKKIELLPYLSQGDIADLGFT SWVKSPEGKTPTNRERYRRWIKQANELMSGEIQLIFQRLKAEVDARNSLEVRLETEIKV GATFEQKRLEDELRLNSFHLAGRLQEATVAMLNALNQDLAVLLDGEATNPRYALATE MVKHLFPFQI ALYRRCILLATTKKFDGS VLLLKRDKEGD YHRS S QQLDHVLGQI<I<GLI< VNPLGEIANQTNPQRIESLKEQYWITDEDSPFAQRDIRFTLVLDEEHEAYRIFQQNVCQT LMNDDQQLPSLLAILHRVLEYVDECSSDERAPGYTPLKNYVDEINKALLQCDLRPDFD LKKLTRLFRNSLLDIWDSRNAEQVIALIGNIFSYQIQRFYRESDLRRIHLRSHGRHSYTE AYVSEDDKKVPDEFNLYELYQTLTALLLASARLETPKHVCDLLAQQEMANQNNPLSV FIRKAQSIKGDIVHLSSSEPQKEQLIDDFFTYFQPKTLFSLSPREQVRIEAEQLRGWVLLK FRMELIKELPEIALLRVLYNTRNTIILLSATAGFPTTYNGQYSRPFLDKYARELNYRIRQ RDVDSASPLSTIRDNRNLHRPVALEVFDDQVLEFLPQNEESEFRNAYRFWLKMLEPYSK AVQFNRYHKREFHRQLQCMLLAAYTGRHILCIGISSRFFAIIGNFLRANKLSANPYRGVT ILDNETRRGNDLPRVFEITPFAERHRLRVVMFDSKLNREDPVRDYLQIDHQNLAICMVS HFQGAGTGLNYYVTYPAPSEPVEADLEQIDFDELAMVCGSYWSQINATPSRNTLENYI TLLKHYAHGSVPRQVGDFDTDLVDSDAAQLLNTEHTVELHKIAMQTIGRTERRDAQ MNGIIRLPSGVHHNALCVFRDLDRHPNAQSLLASLSLHNHLWFKRSKEDLLKASFSSDS QRSEFENKVAKAIDMYSDFEAELKNKILPLARQGDRDAIELNEALRHRDSFTDPQSYIK RLKRNAIIKKNAYLSDCVSHFYLERTADWENVILAKTLDGYGLTDISAGANPYKPEYCL PQYHEAMAEESSGTEQRIFAKILGLDATPLQQYIPIPSLLPLLIGNIGEWQLHLILQEMNITLIPPQELSYYLDSSCYELFDVYCINKNRIVAIDVKNWRMQGNNRQLAKKMHNNSLGKVSELQKIITAKTQFDGVDWYLNTRYALNSLNIRAEHSNHKGVCYYNLFKNISSYEKD NGKNPYDAKIKSELRVNQYLLNLLGANYD (SEQ ID NO: 5).
[0082] In various embodiments, the ApsA homolog or ApsA-like protein has a sequence as shown in Table 3 or Table 4; in particular selected from the group consisting of SEQ ID NO: 43 to 147 and SEQ ID NO: 253 to 295; more particularly selected from the group consisting of SEQ ID NO: 43-46, 48-54, 56-79, 81-84, 86-87, 89 to 102, 104 to 113, 115, 117 to 124, 126, 128 to 132, 134 to 136, 138, 140 to 147 and SEQ ID NO: 253 to 295. In some embodiments, the ApsA homolog or ApsA-like protein has a sequence selected from the group consisting of SEQ ID NO: 253 to 261; particularly selected from the group consisting of SEQ ID NO: 253 a 255.
[0083] In some embodiments, the protein or peptide of the invention is labelled. In one embodiment, the protein or peptide is labelled with a visualizing molecule, such as a radioactive atom, a dye, a fluorescent molecule, a fluorophore, an enzyme, colloidal gold, a magnetic particle, or a latex bead.
[0084] Examples of tags or labels include histidine (His) tags, VS tags, FLAG tags, influenza hemagglutinin (HA) tags, Myc tags, VSV-G tags, and thioredoxin (Trx) tags. Examples of reporter genes include glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT) beta-galactosidase, beta-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP). The protein or peptide can be fused to an amino acid sequence encoding a protein or a fragment of a protein that bind DNA molecules or that bind to other cellular molecules, such as maltose binding protein (MBP), S-tag, Lex A DNA binding domain (DBD) fusions, GAL4 DNA binding domain fusions, and herpes simplex virus (HSV) BP 16 protein fusions.
[0085] The invention encompasses a composition comprising ApsA and / or ApsB proteins with a nucleic acid guide, which can be an RNA, DNA, or RNA / DNA guide. Preferably, the guide is a 5 -end phosphorylated small interfering DNA (siDNA) for recognition and successive cleavage of complementary DNA targets.Antibodies
[0086] In some embodiments, purified ApsA and / or ApsB proteins, including ApsA-like and ApsB-like proteins, are used to produce antibodies by conventional techniques. In some embodiments, recombinant or synthetic proteins or peptides of the invention are used to produce antibodies by conventional techniques.
[0087] Antibodies can be synthetic, semi-synthetic, monoclonal, or polyclonal and can be made by techniques well known in the art. Such antibodies specifically bind to proteins and polypeptides of the invention via the antigen-binding sites of the antibody (as opposed to nonspecific binding). Purified or synthetic proteins and peptides can be employed as immunogens in producing antibodies immunoreactive therewith. The proteins and peptides contain antigenic determinants or epitopes that elicit the formation of antibodies.
[0088] These antigenic determinants or epitopes can be either linear or conformational (discontinuous). Linear epitopes are composed of a single section of amino acids of the polypeptide, while conformational or discontinuous epitopes are composed of amino acids sections from different regions of the polypeptide chain that are brought into close proximity upon protein folding (C. A. Janeway, Jr. and P. Travers, Immuno Biology 3:9 (Garland Publishing Inc.,2nd ed. 1996)). Because folded proteins have complex surfaces, the number of epitopes available is quite numerous; however, due to the conformation of the protein and steric hinderances, the number of antibodies that actually bind to the epitopes is less than the number of available epitopes (C. A. Janeway, Jr. and P. Travers, Immuno Biology 2: 14 (Garland Publishing Inc.,2nd ed. 1996)). Epitopes can be identified by any of the methods known in the art. Such epitopes or variants thereof can be produced using techniques well known in the art such as solid-phase synthesis, chemical or enzymatic cleavage of a polypeptide, or using recombinant DNA technology.
[0089] Antibodies are defined to be specifically binding if they bind proteins or polypeptides of the invention with a Ka of greater than or equal to about 107M-l. Affinities of binding partners or antibodies can be readily determined using conventional techniques, for example those described by Scatchard et al., Ann. N.Y. Acad. Sci., 51:660 (1949).
[0090] Polyclonal antibodies can be readily generated from a variety of sources, for example, horses, cows, goats, sheep, dogs, chickens, alpaca, camels, rabbits, mice, or rats, using procedures that are well known in the art. In general, a purified protein or polypeptide of the invention that is appropriately conjugated is administered to the host animal typically through parenteral injection. The immunogenicity can be enhanced through the use of an adjuvant, for example, Freund's complete or incomplete adjuvant. Following booster immunizations, small samples of serum are collected and tested for reactivity to proteins or polypeptides. Examples of various assays useful for such determination include those described in Antibodies: A Laboratory Manual, Harlow and Lane (eds.), Cold Spring Harbor Laboratory Press, 1988; as well as procedures, such as countercurrent immuno-electrophoresis (CIEP), radioimmunoassay, radioimmunoprecipitation, enzyme-linked immunosorbent assays (ELISA), dot blot assays, and sandwich assays. See U.S. Pat. Nos. 4,376,110 and 4,486,530.
[0091] Monoclonal antibodies can be readily prepared using well known procedures. See, for example, the procedures described in U.S. Pat. Nos. RE 32,011, 4,902,614, 4,543,439, and 4,411,993; Monoclonal Antibodies, Hybridomas: A New Dimension in Biological Analyses, Plenum Press, Kennett, McKeam, and Bechtol (eds.), 1980.
[0092] For example, the host animals, such as mice, can be injected intraperitoneally at least once and preferably at least twice at about 3 week intervals with isolated and purified proteins or conjugated polypeptides of the invention, for example a peptide comprising or consisting of the specific amino acids set forth above. Mouse sera are then assayed by conventional dot blot technique or antibody capture (ABC) to determine which animal is best to fuse. Approximately two to three weeks later, the mice are given an intravenous boost of the protein or polypeptide. Mice are later sacrificed, and spleen cells fused with commercially available myeloma cells, such as Ag8.653 (ATCC), following established protocols. Briefly, the myeloma cells are washed several times in media and fused to mouse spleen cells at a ratio of about three spleen cells to one myeloma cell. The fusing agent can be any suitable agent used in the art, for example, polyethylene glycol (PEG). Fusion is plated out into plates containing media that allows for the selective growth of the fused cells. The fused cells can then be allowed to grow for approximately eight days. Supernatants from resultant hybridomas are collected and added to a plate that is firstcoated with goat anti-mouse Ig. Following washes, a label, such as a labeled protein or polypeptide, is added to each well followed by incubation. Positive wells can be subsequently detected. Positive clones can be grown in bulk culture and supernatants are subsequently purified over a Protein A column (Pharmacia).
[0093] The monoclonal antibodies of the invention can be produced using alternative techniques, such as those described by Alting-Mees et al., “Monoclonal Antibody Expression Libraries: A Rapid Alternative to Hybridomas”, Strategies in Molecular Biology 3:1-9 (1990), which is incorporated herein by reference. Similarly, binding partners can be constructed using recombinant DNA techniques to incorporate the variable regions of a gene that encodes a specific binding antibody. Such a technique is described in Larrick et al., Biotechnology, 7:394 (1989).
[0094] Antigen-binding fragments of such antibodies, which can be produced by conventional techniques, are also encompassed by the present invention. Examples of such fragments include, but are not limited to, Fab and F(ab’)2 fragments. Antibody fragments and derivatives produced by genetic engineering techniques are also provided.
[0095] The monoclonal antibodies of the present invention include chimeric antibodies, e.g., humanized versions of murine monoclonal antibodies. Such humanized antibodies can be prepared by known techniques, and offer the advantage of reduced immunogenicity when the antibodies are administered to humans. In one embodiment, a humanized monoclonal antibody comprises the variable region of a murine antibody (or just the antigen binding site thereof) and a constant region derived from a human antibody. Alternatively, a humanized antibody fragment can comprise the antigen binding site of a murine monoclonal antibody and a variable region fragment (lacking the antigen-binding site) derived from a human antibody. Procedures for the production of chimeric and further engineered monoclonal antibodies include those described in Riechmann et al. (Nature 332:323, 1988), Liu et al. (PNAS 84:3439, 1987), Larrick et al. (Bio / Technology 7:934, 1989), and Winter and Harris (TIPS 14: 139, May, 1993). Procedures to generate antibodies transgenically can be found in GB 2,272,440, U.S. Pat. Nos. 5,569,825 and 5,545,806.
[0096] Antibodies produced by genetic engineering methods, such as chimeric and humanized monoclonal antibodies, comprising both human and non-human portions, which can be made using standard recombinant DNA techniques, can be used. Such chimeric and humanized monoclonal antibodies can be produced by genetic engineering using standard DNA techniques known in the art, for example using methods described in Robinson et al. International Publication No. WO 87 / 02671; Akira, et al. European Patent Application 0184187; Taniguchi, M., European Patent Application 0171496; Morrison et al. European Patent Application 0173494; Neuberger et al. PCT International Publication No. WO 86 / 01533; Cabilly et al. U.S. Pat. No. 4,816,567; Cabilly et al. European Patent Application 0125023; Better et al., Science 240:1041 1043, 1988; Liu et al., PNAS 84:3439 3443, 1987; Liu et al., J. Immunol. 139:3521 3526, 1987; Sun et al. PNAS 84:214 218, 1987; Nishimura et al., Cane. Res. 47:999 1005, 1987; Wood et al., Nature 314:446 449, 1985; and Shaw et al., J. Natl. Cancer Inst. 80:1553 1559, 1988); Morrison, S. L., Science 229:1202 1207, 1985; Oi et al., BioTechniques 4:214, 1986; Winter U.S. Pat. No. 5,225,539; Jones et al., Nature 321 :552525, 1986; Verhoeyan et al., Science 239: 1534, 1988; and Beidler et al., J. Immunol. 141:4053 4060, 1988.
[0097] In connection with synthetic and semi-synthetic antibodies, such terms are intended to cover but are not limited to antibody fragments, isotype switched antibodies, humanized antibodies (e.g., mouse-human, human-mouse), hybrids, antibodies having plural specificities, and fully synthetic antibody-like molecules.
[0098] In one embodiment, the invention encompasses single-domain antibodies (sdAb), also known as nanobodies. A sdAb is a fragment consisting of a single monomeric variable antibody domain that can bind selectively to a specific antigen.
[0099] In one embodiment, the sdAbs are from heavy-chain antibodies found in camelids (VHH fragments), or cartilaginous fishes (VNAR fragments), or are obtained by splitting dimeric variable domains into monomers.
[0100] Preferably, the antibody is labelled. In one embodiment, the antibody is labelled with a visualizing molecule, such as a radioactive atom, a dye, a fluorescent molecule, a fluorophore, an enzyme, colloidal gold, a magnetic particle, or a latex bead.
[0101] The antibodies of the invention can be used for affinity purification of ApsA and ApsB proteins. They can also be used for immunoprecipitation, flow cytometry, western blot, ELISA, ELISPOT, antibodies microarrays, or tissue microarrays coupled to immunohistochemistry. Other suitable techniques include FRET or BRET, single cell microscopic or histochemistry methods.Nucleic acids encoding proteins (RNA & DNA)
[0102] The invention encompasses recombinant nucleic acids, RNA or DNA, encoding ApsA or ApsB proteins, including ApsA-like and ApsB-like proteins according to the invention. The recombinant nucleic acid comprises a nucleic acid encoding ApsA or ApsB proteins sequences linked to a heterologous nucleic acid sequence. In one embodiment, the recombinant nucleic acid comprises or consists of a DNA encoding any of the proteins of SEQ ID NO: 1-6.
[0103] In some embodiments, the recombinant nucleic acid comprises at least 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 100, or 150 sequential nucleotides of a DNA encoding any of the proteins of SEQ ID NO: 1-6.
[0104] The recombinant nucleic acid can comprise all or at least 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 100, or 150 sequential nucleotides identical to any of the nucleotide sequences encoding the proteins of any of the Figures or Tables herein.
[0105] In one embodiment, the recombinant nucleic acid comprises an origin of replication for replication in bacteria or yeast.
[0106] In one embodiment, the recombinant nucleic acid is contained in a plasmid, cosmid, or phage.
[0107] In one embodiment, the recombinant nucleic acid comprises heterologous sequences allowing expression, such as a heterologous promoter or enhancer.
[0108] In one embodiment, the recombinant nucleic acid encodes a protein with at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of SEQ ID NO: 1 to SEQ ID NO: 6.
[0109] In one embodiment, the recombinant nucleic acid encodes an ApsA or ApsB protein and has at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity to the sequences within the SEQ ID Nos 7-13 encoding these proteins. Preferably, the recombinant nucleic acid has at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity to the complement of nucleotides 296-5071 of SEQ ID NO:9 (ApsA) or the complement of joined nucleotides 10303-12235 / 1-299 of SEQ ID NO: 9 (ApsB).
[0110] In one embodiment, the recombinant nucleic acid encodes a protein with at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences in the Figures or Tables herein, in particular the sequences in Table 3 or Table 4. In particular embodiment, the recombinant nucleic acid encodes a protein with at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences SEQ ID NO: 43 to 337; particularly the sequences SEQ ID NO: 43 to 102, 104 to 113, 115, 117 to 124, 126, 128 to 132, 134 to 136, 138, 140 to 147, 148 to 171, 173, 175, 177 to 178, 190, 193 to 197, 199 to 201, 204, 205, 207, 213, 216, 218, 223, 225, 226, 232, 234, 236, 238, 253 to 334, 336, and 337; more particularly the sequences SEQ ID NO: 148 to 152, 154 to 156, 158, 160, and 253 to 261; even more particularly the sequences SEQ ID NO: 148, 149, 151, 152, 155, 158, 160 and 253 to 255. In one embodiment, the recombinant nucleic acid has at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences SEQ ID NO: 14 to 19. SEQ ID NO: 14 and 15 are the coding sequences for the ApsA and ApsB proteins of E. coli ST38 ; SEQ ID NO : 16 and 17 are the coding sequences for the ApsA-like and ApsB-like proteins of CNR36C9; SEQ ID NO: 18 and 19 are the coding sequences for the ApsA-like and ApsB-like proteins of CNR81D10.
[0111] Preferably, the recombinant nucleic acid has at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of SEQ ID NOs 14-19.
[0112] The invention encompasses a recombinant vector for expression of an ApsA or ApsB protein, including ApsA-like and ApsB-like proteins. The recombinant vector can be a vector for eukaryotic or prokaryotic expression, such as a plasmid, a phage for bacterium introduction, a YAC able to transform yeast, a viral vector and especially a retroviral vector, or any expressionvector. An expression vector as defined herein is chosen to enable the production of an ApsA and / or ApsB protein, either in vitro or in vivo.
[0113] In some embodiments, an expression vector as defined herein is chosen to enable the production of an ApsA protein and an ApsB protein, either in vitro or in vivo. In some embodiments, an expression vector as defined herein is chosen to enable the production of an ApsA protein, either in vitro or in vivo. In some embodiments, an expression vector as defined herein is chosen to enable the production of an ApsB protein, either in vitro or in vivo.
[0114] The expression vector can comprise an inducible or constitutive promoter operably linked to a sequence encoding an ApsA or ApsB protein. In various embodiments, the expression vector comprises the pBAD inducible promoter, for example in a pHV7 vector coding for apramycin resistance.
[0115] In one embodiment, the expression vector encodes a protein with at least 50%, 60%, 70%, 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of SEQ ID NO: 1 to SEQ ID NO:6.
[0116] In one embodiment, the expression vector encodes a protein with 50%, 60%, 70%, 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences in the Figures or Tables herein.
[0117] In one embodiment, the expression vector has the nucleotide sequence of any of SEQ ID Nos 7-13.
[0118] In one embodiment, the expression vector encodes a protein purification tag. In one embodiment the expression vector encodes ApsA and ApsB proteins. In one embodiment, the expression vector encodes a protease cleavage site, such as TEV cleavage site, inserted between the ApsA or ApsB protein coding sequence and a protein purification tag, such as polyHis tag. In a preferred embodiment, the expression vector encodes a His tag. In one embodiment, a protease cleavage site is positioned to remove the His tag, for example, after purification.
[0119] The expression vector can comprise transcription regulation regions (including promoter, enhancer, ribosome binding site (RBS), polyA signal), a termination signal, aprokaryotic or eukaryotic origin of replication and / or a selection gene. The features of the promoter can be easily determined by the person skilled in the art in view of the expression needed, i.e., constitutive, transitory or inducible (e.g. IPTG), strong or weak, tissue-specific and / or developmental stage-specific promoter. The vector can also comprise sequence enabling conditional expression, such as sequences of the Cre / Lox system or analogue systems.
[0120] In various embodiments, the expression vector is a plasmid, a phage for bacterium introduction, a YAC able to transform yeast, a viral vector, or any expression vector. An expression vector as defined herein is chosen to enable the production of a protein or polyepitope, either in vitro or in vivo.
[0121] The nucleic acid molecules according to the invention can be obtained by conventional methods, known per se, following standard protocols such as those described in Current Protocols in Molecular Biology (Frederick M. AUSUBEL, 2000, Wiley and son Inc., Library of Congress, USA). For example, they may be obtained by amplification of a nucleic sequence by PCR or RT- PCR or alternatively by total or partial chemical synthesis.
[0122] The vectors are constructed and introduced into host cells by conventional recombinant DNA and genetic engineering methods which are known. Numerous vectors into which a nucleic acid molecule of interest may be inserted in order to introduce it and to maintain it in a host cell are known; the choice of an appropriate vector depends on the use envisaged for this vector (for example replication of the sequence of interest, expression of this sequence, maintenance of the sequence in extrachromosomal form or alternatively integration into the chromosomal material of the host), and on the nature of the host cell.Cells comprising vectors
[0123] The invention further encompasses cells comprising the vectors of the invention. In some embodiments, the cells are recombinant cells.
[0124] Suitable host cells for cloning or expressing DNAs encoding ApsA protein and / or ApsB protein, including ApsA-like and ApsB-like proteins, in the vectors herein include prokaryotic cells, yeast cells, insect cells. These cells can be bacterial cells, such as E. coli cells.
[0125] Additional host cells for cloning or expressing the DNA in the vectors herein include higher eukaryote cells, including vertebrate host cells. Propagation of vertebrate cells in culture (tissue culture) has become a routine procedure. Examples of useful mammalian host cell lines are monkey kidney CV1 line transformed by SV40 (COS-7, ATCC CRL 1651); human embryonic kidney line (293 or 293 cells subcloned for growth in suspension culture, Graham et al., J. Gen Virol. 36:59 (1977)); baby hamster kidney cells (BHK, ATCC CCL 10); Chinese hamster ovary cells / iDHFR (CHO, Urlaub et al., Proc. Natl. Acad. Sci. USA 77:4216 (1980)); mouse sertoli cells (TM4, Mather, Biol. Reprod. 23:243-251 (1980)); monkey kidney cells (CV1 ATCC CCL 70); African green monkey kidney cells (VERO-76, ATCC CRL-1587); human cervical carcinoma cells (HELA, ATCC CCL 2); canine kidney cells (MDCK, ATCC CCL 34); buffalo rat liver cells (BRL 3A, ATCC CRL 1442); human lung cells (W138, ATCC CCL 75); human liver cells (Hep G2, HB 8065); mouse mammary tumor (MMT 060562, ATCC CCL51); TRI cells (Mather et al., Annals N.Y. Acad. Sci.383:44-68 (1982)); MRC 5 cells; FS4 cells; and a human hepatoma line (Hep G2). Host cells are transformed with the above-described expression or cloning vectors for ApsA or ApsB protein production and cultured in conventional nutrient media modified as appropriate for inducing promoters, selecting transformants, or amplifying the genes encoding the desired sequences.
[0126] The host cells used to produce the ApsA protein and / or ApsB protein including ApsA- like and ApsB-like proteins of the present application may be cultured in a variety of media. Commercially available media such as Ham's L10 (Sigma), Minimal Essential Medium ((MEM), (Sigma), RPMI-1640 (Sigma), and Dulbecco's Modified Eagle's Medium ((DMEM), Sigma) are suitable for culturing the host cells. In addition, any of the media described in Ham et al., Meth. Enz. 58:44 (1979), Barnes et al., Anal. Biochem. 102:255 (1980), U.S. Pat. No. 4,767,704; 4,657,866; 4,927,762; 4,560,655; or 5,122,469; WO 90 / 03430; WO 87 / 00195; or U.S. Pat. Re. 30,985 may be used as culture media for the host cells. Any of these media may be supplemented as necessary with hormones and / or other growth factors (such as insulin, transferrin, or epidermal growth factor), salts (such as sodium chloride, calcium, magnesium, and phosphate), buffers (such as HEPES), nucleotides (such as adenosine and thymidine), antibiotics (such as GENTAMYCIN™ drug), trace elements (defined as inorganic compounds usually present at finalconcentrations in the micromolar range), and glucose or an equivalent energy source. Any other necessary supplements may also be included at appropriate concentrations that would be known to those skilled in the art. The culture conditions, such as temperature, pH, and the like, are those previously used with the host cell selected for expression, and will be apparent to the ordinarily skilled artisan.
[0127] In some embodiments, the vectors encoding ApsA and / or ApsB proteins including ApsA-like and ApsB-like proteins, are integrated into a host cell chromosome.
[0128] In some embodiments, the vectors encoding ApsA and / or ApsB proteins including ApsA-like and ApsB-like proteins, are not integrated into a host cell chromosome, and exist, for example, as an extrachromosomal plasmid.
[0129] In some embodiments, the vectors encoding ApsA and / or ApsB proteins including ApsA-like and ApsB-like proteins, are present at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, or more copies per host cell.Methods for expression
[0130] The invention encompasses methods for expressing ApsA and / or ApsB proteins, including ApsA-like and ApsB-like proteins, in cells in vivo or in vitro.
[0131] In various embodiments, plasmids coding for expression of ApsA and / or ApsB proteins and RNAs coding for expression of ApsA and / or ApsB can be subjected to in vitro transcription and / or translation.
[0132] For example, a linearized SP6 ApsA or ApsB plasmid, which contains the SP6 promoter, followed by a 5' untranslated region, and the coding region for ApsA or ApsB can be in vitro transcribed with SP6 RNA polymerase. See e.g., Lingappa et al., Methods Mol Biol. 2009, 485: 185-195, which is hereby incorporated by reference.
[0133] A wheat germ extract can be used for in vitro translation to generate ApsA and ApsB proteins from the ApsA or ApsB RNAs. See e.g., Lingappa et al., Methods Mol Biol. 2009, 485: 185-195, which is hereby incorporated by reference.31
[0134] A rabbit reticulocyte lysate can also be used. See e.g., Spearman et al., J. Virol., 1996, p. 8187-8194, which is hereby incorporated by reference.
[0135] In various embodiments, vectors coding for expression of ApsA and / or ApsB proteins are transfected or transduced into host cells under conditions that allow expression of the proteins from the vectors. In some embodiments, the vector is a plasmid, a phage, a YAC able to transform yeast, a viral vector, especially a retroviral vector, or any expression vector.
[0136] Different types of vectors can be used for ApsAB expression, for example, vectors of the pET28 type (ex: pET28_b) for large-scale expression and purification of the system (expression of N-terminally 6xHis-tagged proteins facilitating the purification) or for expression and purification of the system (pET28_Aps_AB (SEQ ID NO: 9) and pET28_Aps_AB_CNR36C9 (SEQ ID NO: 10) or of one of the proteins (pET28_ApsA (SEQ ID NO: 7), pET28_ApsB (SEQ ID NO: 8)). The presence of the His-tag allows the purification of the proteins on affinity columns.
[0137] Purified proteins of the invention can be used for in vitro applications, such as the utilization as a programmable artificial restriction enzyme (cf US patent US 11,261,483 B2) or as a biosensor (Emerging Argonaute-based nucleic acid biosensors. Qin et al. Trends Biotechnol. 2022 Aug;40(8):910-914.; Highly specific enrichment of rare nucleic acid fractions using Thermus thermophilus argonaute with applications in cancer diagnostics. Song et al., Nucleic Acids Res. 2020 Feb 28;48(4):el9.), which are hereby incorporated by reference.
[0138] In addition, vectors of the miniTn7 type: (ex miniTn-Not-zeo-bad-RK6) can be used to generate strains able to defend against the invasion by some plasmids. This vector allows the insertion of the two genes at a precise position (downstream glmS) in the chromosome of a given strain (for instance MG1655). Examples of such plasmids include pminiTnApsAB (SEQ ID NO:11) comprising apsA and apsB genes of ST38 and pMiniTn_ApsAB_CNR36C9 (SEQ ID NO:12) comprising apsA and apsB genes of CNR36C9.
[0139] In addition, vectors of the phV7 type can be used for short-term induction of ApsAB expression under arabinose control in a bacterial cell and potential gene-editing in bacteria. An examples of such plasmid is TpHV7-apsAB (SEQ ID NO: 13).
[0140] The invention also encompasses a method of preparing a protein comprising culturing cells comprising an expression vector of the invention and recovering the expressed protein.
[0141] The invention further encompasses the proteins produced by these methods from the nucleic acids of the invention.Method for cleaving
[0142] ApsA and ApsB proteins, including ApsA-like and ApsB-like proteins, can be used to cleave dsDNA targets in vivo and also in vitro. The ApsB protein is bound to a single stranded nucleic acid guide, preferably DNA, that is complementary to a dsDNA target. The single stranded nucleic guide that is complementary to a dsDNA target, preferably DNA, may be 5’- phosphorylated or non-phosphorylated. In some embodiments, the single stranded nucleic guide that is complementary to a dsDNA target, preferably DNA, is 5 ’-phosphorylated The single stranded nucleic acid guide brings the ApsB protein to its corresponding target. The ApsA protein is then able to unwind and cleave each target at a specific position. In some embodiments, the method is performed at 37°C.
[0143] Thus, the invention encompasses a method of cleaving a DNA target comprising: providing an ApsB protein; binding the ApsB to single stranded nucleic acid guide, preferably DNA, that is complementary to a dsDNA target to form an ApsB-guide complex; contacting the ApsB -guide complex with the dsDNA target; contacting the dsDNA target with an ApsA protein to form an ApsAB complex; and unwinding and cleaving the dsDNA target with the ApsA protein. In some embodiments, the method is performed at 37°C. The single stranded nucleic acid guide that is complementary to a dsDNA target, preferably DNA, may be 5 '-phosphorylated or non-phosphorylated. In some embodiments, the single stranded nucleic acid guide that is complementary to a dsDNA target, preferably DNA, is 5 '-phosphorylated.
[0144] The ability of a guide molecule to direct sequence-specific binding of ApsB protein to a target DNA molecule can be tested using any suitable assay. For example, cleavage of a target DNA molecule can be evaluated in vitro by contacting a test target molecule, ApsB protein and ApsA protein, one or more test guide molecules having homology to the test target molecule, and a control guide sequence different from the test guide sequence(s), and comparing binding and / orrate of cleavage of the target DNA molecule between the test guide and control guide sequence reactions. A guide molecule can be selected or designed to target and hybridize to any target molecule at a specific target sequence. DNA guide target sequences can be unique in the target DNA molecule or can occur at two or more locations with the target DNA molecule.
[0145] In an embodiment, a protein according to the invention, preferably ApsB and more preferably ApsA and ApsB, can be used in combination with one or more single-stranded nucleic acid guide sequences, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or more guide sequences that target sense and / or antisense strands of target DNA molecules. This system allows both target DNA strands of the double-stranded DNA molecule to be nicked. Alternatively, only one target DNA strand can be nicked. In some embodiments, the method is performed at 37°C.
[0146] In some embodiments, a guide molecule is designed to reduce the degree of secondary structure within the guide molecule. Secondary structure can be determined by any suitable polynucleotide folding algorithm including, for example mFold (Zuker and Stiegler (Nucleic Acids Res 9 (1981), 133-148)). In an embodiment, the guide molecules are designed so that a first guide molecule hybridizes to a sense strand of a denatured or partially denatured DNA molecule and a second guide molecule hybridizes to an anti-sense strand of a denatured or partially denatured DNA molecule. The first and second guide molecules can have the same sequence or, more often, different sequences. The first and second guide molecules can be designed so that when the ApsA protein cleaves the sense and antisense strand of the target DNA the target molecule has overlapping or sticky ends (ends of DNA that can hybridize to one another). In another embodiment, the guide molecules are designed so that when the ApsA protein cleaves the sense and antisense strand of the target DNA, the target molecule has blunt ends. In some embodiments, the method is performed at 37°C.
[0147] DNA guide molecule and ApsB protein form a complex. The guide molecule provides target specificity to the complex by comprising a nucleotide sequence that is complementary to a sequence of a target DNA. The guide of the complex provides the site-specific targeting for the nuclease activity carried by ApsA. In other words, ApsA protein is guided to the target DNA molecule (e.g. a chromosomal sequence or an extrachromosomal sequence, e.g. an episomalsequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by virtue of the association of ApsB with the nucleic acid guide molecule. The ApsA protein and the ApsB protein form a complex and ApsA protein cleaves target DNA. In some cases, the ApsA protein and / or ApsB protein is a naturally-occurring polypeptide. In other cases, the ApsA protein and / or ApsB protein is not a naturally-occurring polypeptide (e.g., a chimeric polypeptide or a naturally- occurring polypeptide that is modified by, e.g., amino acid mutation, deletion, insertion, or a recombinant polypeptide).
[0148] A target deoxyribonucleotide molecule can be any prokaryotic, eukaryotic, or synthetic polynucleotide. The target molecule can be a genomic DNA, such as a chromosome, or can be an extrachromosomal molecule, such as a plasmid. The target molecule can comprise a sequence coding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory polynucleotide or a non-coding DNA). A target DNA is a polynucleotide that comprises a target site or target sequence. The terms target site or target sequence refer to a nucleic acid sequence present in a target DNA to which a nucleic acid guide can hybridize. For example, the target site can be a sequence of 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more consecutive nucleotides to which the nucleic acid guide can hybridize. Suitable DNA binding conditions include physiological conditions normally present in a cell. Other suitable DNA binding conditions are known in the art; see, e.g., Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0149] By performing the cleavage in cells, the ApsA protein and ApsB protein can be used for gene editing of cells. Since the proteins are active at 37°C, the proteins can be expressed from an expression vector in cells, and together with the nucleic acid guide, they can specifically cleave a DNA target.
[0150] The invention provides methods of genome editing or modifying sequences associated with or at a target locus of interest wherein the method comprises introducing ApsA and ApsB proteins or expression vector(s) expressing the same into any desired cell type, prokaryotic or eukaryotic cell, whereby the ApsA and ApsB proteins functions to create a break at a DNA target site in the genome of the eukaryotic or prokaryotic cell, which can then be repaired by the cell tocreate a change in the nucleotide and / or amino acid sequence at the cleavage site. In some embodiments, the method further comprises introducing at least one DNA guide molecule.
[0151] The invention also provides methods of genome editing or modifying sequences associated with or at a target locus of interest wherein the method comprises introducing ApsA and ApsB proteins or expression vector(s) expressing the same into any desired cell type, prokaryotic or eukaryotic cell, whereby the ApsA and ApsB proteins function to integrate a DNA insert into the genome or another nucleic acid sequence of the eukaryotic or prokaryotic cell. In preferred embodiments, the cell is a eukaryotic cell and the genome is a mammalian genome. In preferred embodiments the integration of the DNA insert is facilitated by non-homologous end joining (NHEJ)-based gene insertion mechanisms. In some embodiments, the method further comprises introducing at least one DNA guide molecule. In some embodiments, the method is performed at 37°C.
[0152] In some embodiment, the method further comprises introducing the DNA insert to be integrated. In preferred embodiments, the DNA insert is an exogenously introduced DNA template or repair template. In one preferred embodiment, the exogenously introduced DNA template or repair template is delivered with the ApsA and / or ApsB protein or a polynucleotide vector for expression of a ApsA and / or ApsB protein. In one embodiment, the eukaryotic cell is a non-dividing cell. In some embodiments, the method is performed at 37°C.Assay methods
[0153] The invention encompasses ApsA and / or ApsB proteins, including ApsA-like and ApsB -like proteins, for use in diagnostic assays and their use in these assays. In some embodiments, the method is performed at 37°C.
[0154] In an embodiment, a tagged or labeled ApsB can be used to identify the location of a target DNA sequence. For example, a single-stranded nucleic acid guide having complementarity to a target DNA molecule can be used with a labeled ApsB protein or labelled ApsA protein and labelled ApsB protein to identify a target DNA sequence. See US 2021 / 0164024, which is hereby incorporated by reference in its entirety.
[0155] The present invention provides a detection system for detecting target nucleic acid molecules, which comprises: (a) a nucleic acid guide; (b) an ApsB protein and optionally an ApsA protein; and (c) a fluorescent reporter nucleic acid. In some embodiments, the method is performed at 37°C.
[0156] In some embodiments, the present invention provides a detection system for detecting target nucleic acid molecules, which comprises: (a) a nucleic acid guide; (b) an ApsB protein and ApsA protein; and (c) a fluorescent reporter nucleic acid, which has a fluorescent group and a quenching group; wherein, the target nucleic acid molecule is target DNA. See US 2021 / 0164024. In some embodiments, the method is performed at 37°C.
[0157] The present invention also provides a nucleic acid detection method based on the gene editing enzyme ApsB protein and ApsA protein. In the method of the present invention, based on the cleavage activity of the complex formed by the ApsB protein and ApsA protein, a series of nucleic acid guides can be designed hybridizing to the different target nucleic acid sequences. These nucleic acid guides bind to ApsB to target the nucleic acid to be detected and mediate the ApsA protein enzyme to cleave the target fragment, to form a new secondary guide nucleic acid. In the presence of the ApsB enzyme, the secondary guide nucleic acid continues to guide the ApsA enzyme to cleave the fluorescent reporter nucleic acid complementary to the secondary guide ssDNA, so as to achieve the detection of the target nucleic acid. In some embodiments, the method is performed at 37°C.Kits
[0158] The invention encompasses kits comprising an ApsA and / or ApsB protein, including ApsA-like and ApsB-like proteins, or expression vector(s) expressing the same. In some embodiments, the kit further comprises at least one nucleic acid guide, preferable a DNA guide molecule.
[0159] In some embodiments, the kit comprises a tagged or labelled ApsA protein and / or ApsB protein or expression vector(s) expressing the same. The kit can be used to identify the location of a target DNA sequence. The kit can also be used to determine the presence of a target sequence. The kit can also be used to quantitate a target sequence.
[0160] In some embodiments, the kit comprises an ApsB protein and at least one nucleic acid guide. In some embodiments, the kit comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or more guide sequences. In some embodiments, the kit comprises an ApsA protein.
[0161] In some embodiments, the kit comprises at least one nucleic acid guide and, an ApsB protein, and ApsA protein. The kit can further comprise a fluorescent reporter nucleic acid, especially which has a fluorescent group and a quenching group.
[0162] In some embodiments, the kit comprises a nucleic acid encoding an ApsA and / or ApsB protein. In some embodiments, the kit comprises at least one nucleic acid guide. In some embodiments, the kit comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 25, 30, or more guide sequences.EXAMPLESMATERIALS AND METHODS1. Bacterial strains, growth conditions and plasmids
[0163] Characteristics of the strains used in this work are summarized in Table 1. All strains were submitted to Illumina sequencing and for some of them to long-read sequencing. Donor strains in conjugation experiments were three clinical K. pneumoniae isolates carrying either pOXA-48_l, pOXA-48_2 or pOXA-48_3. The organization of IS 7 sequences in the three plasmids was confirmed by PCR. Recipient strains were three A. coli ST38 environmental isolates characterized by different O-types / H-types. ST38 1 carried an endogenous IncFII, 70.8 kb-long plasmid deprived of any ARG while ST38 3 contained a 161 kb-long IncFIC(FII) plasmid. E. coli strains CNR36C9 (ST219) and CNR81D10 (ST10) encoding closely related (91% / 92% amino acid (a.a.) identity) or distantly related (28% / 25% a.a. identity) a / zs^ / l-like systems respectively were used for PCR-cloning of apsAB homologs. E. coli K-12 strain MG1655 was used for apsAB expression experiments from a mini-Tn7 inserted downstream glmS. E. coli strains DH5a and XL1 blue were used for cloning and MFD pit39for propagation of plasmids with RK6 origin of replication and for bacterial mating.
[0164] Bacterial growths were performed at 37°C. Liquid cultures were performed in LB Miller or in M9 medium supplemented with glucose 0.4%, MgSCh 1 mM and CaCh 0.1 mM with shaking. Where appropriate, antibiotics were added: meropenem (0.1 pg ml’1), apramycin (40 pg ml’1), zeocin (30 pg ml’1) and chloramphenicol (10 pg ml’1). Growth and mating experiments with MFDpir derivatives were performed in LB or LB agar complemented with 0.3 mM diaminopimelic acid (DAP, ThermoS cientific). Induction from PBAD promoter was performed in LB Miller supplemented with L-arabinose to a final concentration of 0.02 or 0.2% as indicated. E-test were performed on Mueller Hinton agar (MHA) medium. A list of plasmids used in the study and their main characteristics is provided in (Table 5). All inserted fragments were verified by Sanger sequencing (Eurofins Genomics) and expression plasmids carrying apsAB were WGS (PlasmidSaurus or Eurofins Genomics).2. Conjugation assay
[0165] Overnight precultures in LB of donors and recipient strains were diluted 1 : 100 into fresh LB and grown to an optical density at 600 nm (ODeoo) of 0.6. After mixing at 1: 1 ratio, 200 pl were spread on a filter (MILLIPORE type HAEP 0.45 pM) placed on LB agar and incubated at 37°C overnight or for one hour. Bacteria were harvested in physiological water followed by serial dilutions and plated on LB agar containing two antibiotics: meropenem (MEM) 0.1 pg ml’1and tetracycline (TET) 10 pg ml’1to select TCs. TCs were WGS and those devoid of mutations were selected for further analyses. Conjugation frequency was calculated as the ratio of transconjugants over donor after a Ih-mating followed by selection of TCs, donors and recipients (MEM 0.1 pg ml’1and TET 10 pg ml’1; MEM 0.1 pg ml’1, TET 10 pg ml’1respectively).3. Growth curves and relative fitness assessments
[0166] Overnight precultures in LB or M9 of the three ST38 E. coli plasmid-free isolates and their isogenic transconjugants were inoculated without and with antibiotic (MEM 0.1 pg ml’1) respectively. 96-well plates were inoculated with 100 pL of precultures diluted to 5*105colony forming unit (CFU) per millilitre in LB or M9 medium and incubated at 37° C with shaking for 6 (LB) or 16 hours (M9) in an automatic plate reader (Tecan infinite M Nano under i-control 2.0.10). Three independent experiments were carried out for each strain, with three to fivereplicates each. The doubling time was calculated using the formula: G = In (2) / max. max is the maximum growth rate, corresponding to the slope of the curve at exponential phase. The relative fitness was calculated using the formula: W Gtransconjugant / Gpiasmid-free4. Experimental evolution
[0167] Overnight precultures in LB or in M9 of TCs and plasmid-free strains were diluted 1 :200 into 10 ml fresh LB medium or into complemented M9 medium with and without MEM 0.1 pg ml’1respectively and incubated at 37°C with shaking (220 r.p.m, INFORS HT Minitron). For each experiment, five independent biological replicate cultures were evolved. Serial transfers were achieved for 28 days using 1 :200 dilutions every day into 10 ml of fresh medium (ca. 8 generations per day). For evolution of transconjugant lineages, MEM was added every day or every three days at 0.1 pg ml’1as indicated. Each seven days, whole populations were collected and frozen at -80°C and bacterial pellets for DNA sequencing were obtained by centrifuging 1 ml of culture. Relative fitness of whole population through experimental evolution was monitored by growth curves. PCR using primers targeting / > / <2OXA-48 and repA (Table 5) were performed on isolated colonies selected on MEM 0.1 pg ml’1to identify potential integration of WaoxA-48 (Tn6237) and pOXA-48 loss at day 7, day 14, day 21 and day 28.5. Plasmid stability assay
[0168] Overnight precultures of pOXA-48 TCs in LB with MEM 0.1 pg ml’1were set as day 0. Plasmid stability was assessed by serial passages using 1 :200 dilutions into 10 ml fresh LB without MEM for ten days. Bacteria were collected and frozen at -80°C at days 0, 5 and 10 and CFU were determined by plating dilutions on LB agar with and without MEM 0.05 pg ml’1. Colonies growing on MEM were used as a proxy for plasmid-carrying bacteria as Tn6237 integration was estimated as a rare event under these conditions. Automatic bacterial colony counting was done with scan4000 (INTERSCIENCE) and plasmid stability was calculated as the ratio of plasmid-carrying bacteria over total population. Stability of other plasmids was similarly tested after introduction by electroporation (pBbS8c, pBbE8c, pACYt, pUC19) or conjugation (pKPC, using CNR146C9 as donor) with adequate antibiotic selection (Table 5).6. Plasmid constructions for expression of wild-type and mutated copies of apsAB
[0169] F3141-F3140 operon, F3140 (apsB) and F3141 (apsA) were cloned under the control of a PBAD inducible promotor in a pHV7 vector coding for apramycin resistance. All growing steps of pHV7-F3141-F3140 transformants were performed in the presence of glucose 0.2 % to repress apsAB induction. Two mutated versions of this operon where a stop codon was introduced in F3141 or F3140 sequence were generated during this cloning and the corresponding plasmids were used for activity testing. For site-directed mutagenesis, selected codons were modified in the pHV7-F3141-F3140 plasmid by using the Q5 Site-Directed Mutagenesis kit (New England Biolabs) according to manufacturer’s recommendations. pHV7 plasmid mutants were WGS (Eurofins). For chromosomal expression of apsAB (F3141-F3140) in MG1655, apsAB or apsAB homologs were cloned under the control of a PBAD promoter in a miniTn7 transposon (miniTna / ?.s^ / l) using primers described in Table 5. miniTna / ?.s^ / l was integrated downstream glmS, a neutral chromosomal position of A’, coli K12 MG1655 strain following triparental mating as previously described40. Antiplasmid activity of these constructs was analysed by electroporating the ColEl plasmid pBbE8c. An MG1655 strain harbouring an empty miniTn7 (miniTnzeoA) integrated at glmS was used as control. When stated, chloramphenicol was added at 10 pg ml’1to arrest bacterial growth 2.5 hours after addition of arabinose 0.2%.7. F3140-3141 (apsAB) chromosomal deletion and complementation
[0170] Complete deletion of the F3140-3141 (apsAB) operon was obtained by > Red recombination41by using an in-lab pl5red vector carrying the lambda recombinase under the control of an inducible pBAD promotor (Table 5). Primers used for PCR amplification of the ZeoR cassette between F3141-F3140 homology sequences are shown in Table 5. zeoR marker was eliminated by using an in-lab pl 5Flip plasmid encoding the Flippase. SacB counterselection was used to eliminate recombineering plasmids. F3141-F3140 (apsAB) operon deletion were confirmed by PCR. The absence of mutation in the ST38_lAap&4B clone used for experimental evolution was determined by WGS. For complementation, recipient cells were electroporated with the plasmids pHV7-empty, pHV7-F3141, pHV7-F3140, pHV7-F3141-F3140 and pHV7- F3141-F3140-mutants (Table 5). pHV7-derivatives carrying strains were incubated 24 h in LB,apramycin 50 µg ml-1, 0.4% glucose to repress pBADactivity (t0); two serial passages using 1:200 dilutions were then performed in LB, apramycin 50 µg ml-1and 0.02% arabinose to induce F3141- F3140 transcription. Cultures were diluted and plated on LB agar with or without MEM 0.05 µg ml-1and pOXA-48 plasmids stability assessed as the ratio of resistant colonies over whole 5 population. 8. Carbapenem susceptibility testing
[0171] MEM Minimal Inhibitory concentration (MIC) was determined by Etest (Biomerieux). The plates were inoculated by flooding 2 ml of bacterial culture (106bacteria ml-1), spread by a gentle rocking motion, and excessive liquid was removed leaving 0.5 ml of culture (+ / - 10%). 10 The flooding method was selected over swab streaking as it provides a more accurate reading of the Etest result. 9. Bacteriophages plaque assays
[0172] Phage plaque assays have been performed as previously described and by using the same collection of bacteriophages (lambda, T4, P1, 186cIts, CLB_P2, LF82_P8, AL505_P2, and T5, 15 Table 1)28. Phages were obtained as active cultures. The preys, E. coli K12 MG1655 strain and its isogenic derivatives chromosomally encoding the miniTnapsAB or the miniTnzeoR as control were grown overnight. Overnight cultures were diluted to 1:20 in the presence of arabinose 0.2% in LB medium and incubated at 37°C for 3h to induce apsAB expression. Bacterial lawns were prepared by mixing 200 μL of the induced culture with 100 μL of CaCl21 M and 20 ml of LB + 20 0.5% agar and poured onto 12 x 12 cm square plates of LB containing 0.2% arabinose. High-titer (>108pfu ml-1) stocks of phages lambda, T4, P1, 186cIts, CLB_P2, LF82_P8, AL505_P2, and T5 serially diluted were spotted on each plate and incubated at 37°C overnight except for phage T7 incubated overnight at room temperature. 10. Whole genome sequencing and mutation identification 25
[0173] ST38_1, ST38_2 and ST38_3 strains were fully sequenced and used as reference sequences by combining Illumina sequencing and PacBio sequencing (ST38_1, ST38_2) or Oxford Nanopore technology (ST38_3). DNA was extracted at exponential phase with Qiagen Puregene Yeast / Bact kit B. PacBio sequencing libraries were prepared with NANOBIND CBBKIT RT PacBio. ST38_3 Nanopore sequencing library was prepared by using Native Barcoding Kit 24 V14 (ref SQK-NBD114.24) and sequencing was performed with flowcell R10.4.1 (ref FLO-MIN114) on MinION Mk1C device. Long-read PacBio sequences were assembled with hybridSPAdes v 3.15.5 for hybrid assembly of short and long reads42. Hybrid assembly of short 5 and long-read Nanopore sequences were performed by using Canu 2.243and Circlator 1.5.544.
[0174] For Illumina sequencing, DNA was extracted from stationary phase cultures by using the Qiagen Blood and Tissue DNeasy kit, libraries were prepared with the NEBNext Ultra II FS DNA Library Prep Kit and sequencing was performed with NovaSeq6000 or NextSeq500 sequencing platforms. Short-read Illumina sequences were assembled with SPAdes 3.15.545or 10 aligned to the reference sequences by using Breseq 0.35.746to identify SNPs, deletions, insertions and recombination events. IGV 2.11.947was used to visually confirm mutations events. Illumina- reads of the whole population of the evolved lineages collected at day 7, day 14, day 21 and day 28 were analysed for new junction evidence and coverage distribution by using Breseq 0.35.746with the -p option for pool sequencing, a polymorphism frequency cutoff 2.5% option with at 15 least 10 polymorphic reads. Plasmid coverage inside and outside the Tn6237 sequence was determined as the number of reads on three 15 kb-long regions, one in Tn6237, two outside, containing no IS, by using BAM files and the GRanges function of the GenomicAlignments package 1.22.1 in RStudio (R 3.6.3). The ratio was calculated as the ratio of the average of the number of reads mapping on the two regions outside Tn6237 to the number of reads mapping on 20 Tn6237. IS1 new junctions detected by Breseq were tested as potential Tn6237 integration sites by PCR using a primer located near the new IS1 insertion site and a primer in blaOXA-48 (Table 5). PCR products were Sanger-sequenced to confirm the insertion site. 11. Phylogenetic analysis of E. coli ST38 and sequence annotation
[0175] To contextualize the three ST38 strains used in this work we performed a phylogenetic 25 analysis using 1907 sequences retrieved from public databases (April 2020): 1248 assembled genome sequences from Enterobase (https: / / enterobase.warwick.ac.uk / ), 149 assembled genomes from the NCBI and 510 sequences retrieved as reads and assembled with SPAdes 3.12.045. QUAST 2.248was used to assess the assembly quality and contigs shorter than 500 bp werefiltered out for the phylogenetic analysis. A core genome alignment was generated with Parsnp 1.5.449, by using a finished genome sequence as reference. Maximum-Likelihood (ML) trees were generated with RAxML 8.2.1250using GTRGAMMA after removing regions of recombination with Gubbins51. The ST38 single locus variant ST963 strain CNRC6O47 was used as outgroup 5 to root the phylogenetic tree. Trees were visualized and annotated using ITOL52(https: / / itol.embl.de / ).
[0176] Genomes annotation was performed by using Prokka 1.14.553. Resistome and plasmidome were characterized by using ABRicate on the ResFinder db54(minimum coverage 60%, minimum identity 95%) and PlasmidFinder 2.1.155. 10 12. In silico characterization of ApsAB antiplasmid systems
[0177] Search for apsAB homologs were performed by using PSI-BLAST with three iterations on the recently introduced NCBI clustered nr database (https: / / blast.ncbi.nlm.nih.gov / ). This database is composed of representative sequences of clusters, that groups NCBI nr sequences sharing 90% identity and 90% length to other members of the cluster. Only sequences longer than 15 1100 a.a. residues (ApsA) and 500 a.a. residues (ApsB) were kept for PSI-BLAST iterations. CDS located downstream of apsA homologs were retrieved by using GCsnap 1.0.1756. Protein sequences were aligned by using MuscleW 3.8.31 (default options) under Jalview 2.11.357and a distance tree was created by Neighbor-Joining method using a BLOSUM62 matrix. Remote homology detection was also performed with the HHpred server 20 (https: / / toolkit.tuebingen.mpg.de / tools / hhpred) using PDB_mmCIF70_18_jun, SCOPe70_2_08, CATH_S40_v4.3 and UniProt-SwissProt-viral70_3_nov_2021 as target databases accessed in September 2023. Structural modelling was performed with Alphafold258implemented in Neurosnap (https: / / neurosnap.ai / ) and the resulting models were used as templates for similarity searches with Foldseek (https: / / search.foldseek.com / search). Structure annotation was performed 25 with Pymol 2.5.5 (The PyMOL Molecular Graphics System, Version 2.5.5 Schrödinger, LLC.). Alpha-fold or icn3d pdb models of 20 ApsB-like or DdmE-like proteins (Table 3) were recovered from Uniprot (https: / / www.uniprot.org / ) or ncbi (https: / / www.ncbi.nlm.nih.gov / Structure / icn3d / ) and aligned to ApsB structure model under Pymol 2.5.5 and the Root-mean-square deviation ofatomic positions (RMSD) was used as a proxy to estimate protein structure similarity. The genomic environment of apsAB homologs from 12 E. coli strains belonging to different STs was compared by using the web version of Clinker (https: / / cagecat.bioinformatics.nl / tools / chnker).13. Statistical tests
[0178] All graphs were generated with R (version 3.6.3) by using the R packages (tidyverse, forcats, ggplot2, ggpubr, rstatix, broom). Normality of data were assessed by using the Shapiro test. All the statistical analyses were performed with a pairwise two sample t.test and p-value were corrected following Benjamini-Hochberg correction (FDR).RESULTSExample 1. pOXA-48s induce a fitness cost and are unstable in ST38 E. coli
[0179] To determine whether the genetic background influences the persistence of pOXA-48 plasmids, three ST38 isolates belonging to different phylogenetic sublineages were chosen (Table 1). E. coli ST38 was used as a model because of its evolutionary history characterized by pervasive ARG integrations23 26in addition to its clinical relevance. Three related pOXA-48 plasmids (pOXA-48_l, pOXA-48_2 and pOXA-48_3) were transferred by conjugation from three clinical Klebsiella pneumoniae isolates, the most common pOXA-48 bearing species found in hospitals27. These pOXA-48 variants carried one or two IS7, n6237 structure being present only in pOXA-48_l (Fig. 1 A). The plasmid structure did not influence the meropenem minimum inhibitory concentration (MIC) of the transconjugant (TC) unlike the genetic background of the recipient strain with a higher MIC for ST38-1 TCs, 0.5 pg mF1, compared to 0.25 pg mF1and 0.38 pg mF1for ST38_2 and ST38_3 TCs respectively.
[0180] To assess the fitness cost induced by pOXA-48 the maximum growth rate of the TCs and isogenic plasmid-free strains were compared. The relative fitness was estimated as the ratio of doubling time (DT) of TC over plasmid-free strain. At least 15% DT increase was induced in lysogeny broth (LB) medium by pOXA-48 in ST38 1 (Fig. IB). In contrast, in this medium, the three plasmids induced a lower (<5%) fitness cost in ST38 2 and ST38 3 (Fig. IB). Fitness costwas dependent on the growth medium as in M9 glucose minimal medium it rose up to 12% for ST38_2 TCs, while decreasing to 7-11% in ST38_1 TCs.
[0181] As pOXA-48 plasmids are often costly to their host, their stability was quantified in the three ST38 strains, following ten-day serial passages of the TCs (ca. 80 generations) in LB in the absence of antibiotic. The three pOXA-48 plasmids were gradually lost and at day 10, more than 95% of the population in ST38 1 and 40-80% in ST38 2 and ST38 3 had lost pOXA-48, irrespective of the plasmid variant (Fig. 1C).Example 2. Twenty-eight days of in vitro evolution led to frequent W O\ 48 chromosomal integration in ST38 1
[0182] To determine whether the fitness of pOXA-48 TCs could be improved by plasmid-host coevolution, 28-day experimental evolutions of ST38_l / pOXA-48_l TCs were performed. Given the instability of pOXA-48, the growth medium was supplemented with subinhibitory concentration (0.1 pg ml’1) of meropenem every passage or every third passage. The ST38 1 plasmid-free strain was evolved as control. Five independent lineages were derived for each condition. Experimental evolution led to a rapid growth rate improvement for all TC lineages under both meropenem conditions (Fig. 2A). To characterize the populations, pool-sequencing were first performed after 7, 14, 21 and 28 days of evolution. It was observed a progressive decrease in read coverage of the pOXA-48_l DNA-region outside Tn6237 (Fig. 2B), suggesting the enrichment of bacteria having lost pOXA-48 but keeping / > / <2OXA-48.
[0183] To further characterize putative transposition events of Tn6237, meropenem-resistant colonies isolated at 28-day evolution for / > / <2OXA-48 and plasmid origin of replication (repA) were PCR-screened. / ? / aox \-48 repA' colonies were identified in all ST38_l / pOXA-48_l lineages. Two to three of these colonies per lineage were whole-genome-sequenced (WGS). Variant analysis revealed IS7 new junctions in the chromosome or in the ST38 1 IncFII plasmid. It was verified by PCR and Sanger sequencing that these new junctions corresponded to Tn6237 transposition. In total, 13 integrations at different positions in the chromosome and five in the IncFII plasmid were identified. Analysis of pool sequencing reads revealed 44 additional new junctions, considering a threshold of 5% of the reads. Twenty were chosen for PCR analysis52 among which 16 were confirmed as Tn6237 integration. Integrations were detected already at seven days of evolution and multiple integrations at multiple positions on the chromosome and the IncFII plasmid were selected in each lineage after 28 days.
[0184] To determine whether blaOXA-48 integration was associated with a fitness improvement, 5 three colonies were selected, one in the chromosome and two in the IncFII plasmid. The growth rate of the three clones was increased with a 10-15% reduction in DT compared to the original TC (Fig.2C) and the meropenem MIC decreased by 50% (0.25 vs 0.5 µg ml-1). Therefore, in ST38_1, blaOXA-48integration appears as a major mechanism to relieve the fitness cost associated with pOXA-48_1 carriage. 10 Example 3. blaOXA-48 integration depends on pOXA-48 structure and on recipient strain
[0185] To determine whether blaOXA-48integration during experimental evolution was only dependent on the presence of Tn6237, ST38_1 TCs carrying plasmids pOXA-48_1, pOXA-48_2 or pOXA-48_3 were evolved in LB with meropenem added every third passage. In contrast to pOXA-48_1, no fitness improvement throughout time was observed for pOXA-48_2, while 15 fitness of the population was slightly improved for pOXA-48_3 TCs at 28-day evolution (Fig. 2D). Bulk DNA sequencing of the whole populations at day 21 and at day 28 revealed after 28- day evolution only a relative decrease in read-coverage of the pOXA-48_3 region equivalent to the one lost in pOXA-48_1 (Fig. 9A and 9B) and a few new IS1 junctions. Two Tn6237 integrations in the chromosome were confirmed by PCR. Tn6237 transposition could have 20 occurred following homologous recombination between the two IS1999 copies reconstituting the Tn6237 structure (Fig.1A).
[0186] Similar experimental evolution (meropenem added every passage or every third passage in LB) of ST38_2 / pOXA48_1 and ST38_3 / pOXA48_1 TCs were performed. No evolution of the fitness was observed (Fig.2E) while PCR-screening of meropenem-resistant colonies at 28-day 25 did not identify any blaOXA-48+repA- colonies. This suggested that in the absence of a sufficient fitness cost, blaOXA-48integrations were not enriched enough to be detected. To test this hypothesis, a similar experiment was performed in minimal medium in which pOXA-48_1 induced a higher fitness cost to the ST38_2 TCs. A rapid improvement of growth rate was53 observed (Fig.2F). The inventors detected a few blaOXA-48+repA- colonies, associated with a Tn6237 chromosomal integration, in only two of the five evolved lineages by PCR screening. These colonies showed a full fitness recovery (Fig. 9C). Pool sequencing did not reveal any integration site at a frequency superior to 5%, nor a significant plasmid loss (Fig.9D). Therefore, 5 chromosomal integration and plasmid loss were not the main contributors to the fitness recovery observed in ST38_2 TCs, in contrast to ST38_1 TCs. Example 4. pOXA-48 plasmids are stabilized by inactivation or mutation of a novel antiplasmid system
[0187] In all evolved lineages, including those where Tn6237 transposition was selected, 10 bacteria still retaining the plasmid could be found after 28-day evolution. To identify other possible paths for plasmid-host coadaptation, WGS of three to five blaOXA-48+repA+colonies of each evolved lineage was performed. It was identified sporadic mutations occurring in the three ST38 evolved TCs and control plasmid-free strains including only nine in different loci of pOXA- 48. Mutations potentially decreasing susceptibility to carbapenems (in ompC or envZ for instance) 15 were encountered. Large chromosomal deletions ranging from 15 to 40 kb and encompassing mutS and rpoS were also observed in evolved ST38_1 lineages, carrying or not a pOXA-48. mutS loss led to a large number of mutations (n=13 to 38) likely resulting from an hypermutator phenotype.
[0188] On the other hand, in 62% (86 / 138) of evolved TCs, it was observed convergent 20 evolution with mutations (Table 2) in two adjacent genes, F3141 or F3140, located in a ST38_1 genomic island (Fig. 3A) or its ortholog in ST38_2. The most frequent mutations were IS1 insertions (n=67 with 37 different insertion sites). It was also detected insertions of other IS (n=3), non-synonymous (n=3, two different), non-sense (n=2), or frameshift (n=7, four different) mutations and complete or partial deletions (n=4) of these genes (Table 2). All (15 / 15) sequenced 25 blaOXA-48+repA+colonies from the M9-evolved ST38_2 pOXA-48_1 TCs carried mutations in orthologs of F3141 or F3140 and all (n=30) from the five LB-evolved lineages shared the same IS1 insertion in F3140 suggesting that it was present but not detectable in sequencing reads of the54 initial culture (Table 2). No mutation in F3141-F3140 orthologs was detected in ST38_3 strain evolved in LB.
[0189] In ST38_1, six out of the ten tested mutations led to a fitness improvement compared to the original TC, which was variable and always lower than following blaOXA-48 integration and 5 pOXA-48_1 loss (Fig.3B). In ST38_2 the six tested mutations in F3141-F3140 orthologs led, in M9, to a full fitness recovery (Fig.3B). This probably explains why they were selected at the expense of blaOXA-48 chromosomal integration in this strain. Three F3141 or F3140 mutated ST38_1 TCs and three mutated ST38_2 TCs were tested for plasmid stability by 10-day serial passages in the absence of meropenem. In all tested TCs an increase in plasmid persistence was 10 observed compared to the original TC (Fig. 3C). pOXA-48_1 was kept in more than 90% of ST38-2 mutated TCs. In ST38-1 the stability was more variable with 50 to 100% bacteria keeping pOXA-48_1.
[0190] Conjugation of the pOXA-48 plasmids to genetically modified derivatives of ST38_1 and ST38_3 deleted for F3141-F3140 confirmed the stabilization of the plasmid in the absence 15 of F3141-F3140 (Fig.3D). Of note, ST38_2 was not amenable to genetic manipulation under the conditions the inventors used. Furthermore, complementation in ST38_1∆F3141-F3140 strain with a pHV7-derivative plasmid expressing F3141-F3140 under the control of the arabinose inducible promoter led to more than 80% pOXA-48 loss at 48 hours following induction (Fig.3E). Induced expression of the operon with a non-sense mutation in either F3141 or in F3140 20 did not destabilize pOXA-48 confirming that the two proteins were needed for plasmid loss (Fig.3F).
[0191] To determine whether F3141-F3140 might be involved in a more global antiplasmid activity, the stability of five plasmids with different replication origins and copy numbers was compared. In addition to the IncL plasmid pOXA-48, F3141-F3140 deletion increased the 25 stability of p15A, ColE1 and pMB1-type plasmids, but had no effect on pSC101 and IncFII / IncFIB plasmid stability (Fig.4A). This confirmed that F3141-F3140 corresponds to a novel antiplasmid defence system that the inventors renamed apsAB for antiplasmid system AB.55
[0192] Interestingly, it was found that ApsAB also reduces the conjugation frequency of pOXA- 48 plasmids while comparing the transfer frequency with ST38_1 wild type or ST38_1 ∆apsAB as recipient strains (Fig.4B). This suggests that ApsAB interferes with foreign DNA acquisition. However, ApsAB did not affect the transformation frequency of p15A and ColE1 plasmids 5 (Fig.4C). As some antiplasmid systems are also involved in antiphage defence, the activity of ApsAB chromosomally expressed in MG1655 was tested against eight different phages from E. coli (lambda, T4, P1, 186cIts, CLB_P2, LF82_P8, T5 and T7)28and no antiphage activity was observed (Fig.5).
[0193] Finally, to determine whether ApsAB actively eliminates plasmids, the persistence of a 10 ColE1 plasmid over time following the induction of apsAB expression from a chromosomally integrated copy was quantified. Plasmid elimination was observed between 2h30 and 3h of arabinose induction and almost complete (>95%) by a six-hour induction, whether or not chloramphenicol was added at 2h30 to stop cell division. These results indicate that ApsAB actively eliminates the plasmid, likely by promoting its degradation (Fig.4D). 15 Example 5. ApsAB defence system is necessary for the selection of blaOXA-48 integration.
[0194] It was hypothesized that the elimination of pOXA-48 plasmids by ApsAB could contribute to the emergence of lineages with blaOXA-48inserted in the chromosome. To test this hypothesis, the apsAB operon was deleted in the original ST38_1 / pOXA-48_1 TC (Z103), which has the capacity to rapidly evolve towards Tn6237 integration. No significant difference in fitness 20 was detected between the deleted strain and the original TC (Fig. 6A). As expected, apsAB deletion led to pOXA-48_1 stabilization (Fig.6B). It was then performed experimental evolution in LB medium of five lineages of ST38_1Z103∆apsAB / pOXA-48_1. Contrary to the original ST38_1 / pOXA-48_1 (Fig.2A), no fitness recovery was detected after 28 days of evolution of ST38_1Z103∆apsAB / pOXA-48_1 (Fig.6C). No blaOXA-48+repA- colonies out of 120 tested 25 colonies at day 28 was detected by PCR screening. Similarly, pool sequencing of the whole population at days 7, 14, 21 and 28 showed no loss of pOXA-48_1 plasmid by analysing plasmid read-coverage and no Tn6237 transposition based on the PCR testing of the few new IS1 junctions56 (n=4) (Fig.6D). These results therefore show that plasmid destabilization through the activity of ApsAB is a key factor in the emergence of lineages with chromosomally integrated blaOXA-48. Example 6. ApsAB is the first characterized member of a broad family of Argonaute-like systems. 5
[0195] Sequence similarity search against the NCBI nr public database by BLASTP revealed positive matches of ApsA and ApsB only with proteins of unknown function. Nevertheless, using similarity search based on structure predictions, it was predicted in ApsA a central helicase domain with conserved residues characteristic of superfamily 2 helicase29and a C-terminal domain containing a PD-(D / E)XK- superfamily nuclease motif30(Fig.7A). Directed mutagenesis 10 showed that substitutions predicted to impair ATP hydrolysis (E533A) and helicase activity (K221A) completely abolished pOXA-48 plasmids destabilization while a mutation in the predicted nuclease active site (K1435A) had a partial effect (Fig.7B and Fig.10). On the other hand, structural modelling of ApsB revealed a loose structural similarity with prokaryotic Argonaute proteins (pAgos) acting as nucleic acid-guided endonucleases. However, ApsB lacked 15 typical PIWI and PAZ domains characteristic of Argonautes and was not identified among pAgos31(Fig.7C). Sequence alignment of ApsB homologs retrieved from the NCBI database identified a conserved Y[X]3K[X]nQG[X]nK motif (Fig.7Dand Table 3 ), reminiscent of the conserved residues Y[X]3K[X]nQ[X]nK characteristic of the MID domain of long pAgos that interacts with the 5’-end of guide nucleic acids32. The K413A substitution in this motif completely 20 abolished pOXA-48 destabilization, supporting the requirement of this motif for ApsAB antiplasmid activity (Fig.7B).
[0196] Association of an Argonaute-like protein and a protein with helicase and nuclease domains was reminiscent of the DdmDE plasmid defence system recently characterized in Vibrio cholerae13. No sequence similarity could be found between ApsB and DdmE. Their predicted 3D 25 structures were different (Root-Mean-Square Deviation of atomic positions (RMSD)=30.6) and DdmE did not contain the MID-like motif. However, searching for ApsA homologs by PSI- BLAST in databases unveiled a wide family of proteins ranging from ApsA-like to DdmD-like proteins. (Tables 4 and 6).
[0197] ApsAB-like systems were identified mainly among Enterobacterales but also in other gamma-proteobacteria and some beta-proteobacteria and cyanobacteria (Tables 3-4). In E. coli, complete or partially deleted apsAB homologs were located in at least three different families of genomic islands, some of them encoding other defence systems like Shango systems. The antiplasmid activity of two representatives of E. coli ApsAB-like systems was evaluated, showing92 / 92 % or 28 / 25 % protein sequence identity with ST38_1 ApsA / ApsB respectively. Both systems were found to destabilize a ColEl multicopy plasmid (Fig. 8), indicating that the anti plasmid activity is not restricted to ApsAB from ST38 strains.List of bacteriophagesName SourceEscherichia coli phage (Rousset et al., 2022)Escherichia coli phage T4 (Rousset et al., 2022)Escherichia coli phage Pl (Rousset et al., 2022)Escherichia coli phage 186Clts (Rousset et al., 2022)Escherichia coli phage CLB_P2 (Rousset et aL, 2022)Escherichia coli phage LF82_P8 (Rousset et al., 2022)Escherichia coli phage AL505_P2 (Rousset et al., 2022)Escherichia coli phage T5 (Rousset et al., 2022)Escherichia coli phage T7 (Rousset et al., 2022)Table 2: Main characteristics of the isolated bacteria sequenced after experimental evolutions , . . . . colonies with large deletions in ST38 1 . N° predictedNumber of Number of Colonies with Tn6237 Colonies with F3141- . . .. . . .. ~ Colonies with largeFvnhitinn evnsrimont including rpoS / mutS and leading to mutations per p Ineages sequenced colonies insertion sites$F3140 mutations2. . deletions in ST38 212. hypermutationR-%- colonyR aExperimental evolution 1ST38_1 / LB / No MEM 5 10 _ 2 1-30ST38_l / POXA-48_l / LB / MEM / day 5 25 11 5 7 1-38ST38_l / POXA-48_l / LB / MEM / 3 days 5 3016 715 1-34Experimental evolution 2ST38_1 / LB / No MEM 510- - - 0-2ST38_l / POXA-48_l / LB / MEM / 3 days 52812 13 7 0-24ST38_l / POXA-48_2 / LB / MEM / 3 days 515- 3 0-2ST38_l / POXA-48_3 / LB / MEM / 3 days 515- 11 1-3Experimental evolution 3ST38 2 / LB / No MEM 5 10 9 1-16ST38_2 / POXA-48_l / LB / MEM / day 5 15 15 12 0-2ST38_2 / POXA-48_l / LB / MEM / 3 days 5 19 19 17 0 15ST38 3 / LB / No MEM 5 12 22ST38_3 / POXA-48_l / LB / MEM / 3 days 5 15 0-2Experimental evolution 4ST38_2 / M9 / No MEM 5 10 0-2ST38_2 / POXA-48_l / M9 / MEM / 3 days 5 18 3 15 1-2Experimental evolution 5ST38 1 / LB / No MEM 5 10 5 1-41ST38_iaapsAB / POXA-48_l / LB / MEM / 3 days 5 15 8 1-38*: genotype of the isolated bacteria as deduced from the PCR detecting pOXA-48 repA and bl aoxA-48 ft : as determined after Illumina sequencing and Breseq analysis$: Tn6237 insertion site as suggested by the presence of a new juction by using breseq analysis of the sequence and as confirmed by PCR experiment and sequencing of the junction%: Large deletions including rpoS and mutS were recurrently observed in ST38-1 lineages after experimental evolution. They led to a hypermutation phenotype, characterized by a larger number of mutations than in other lineages &:Number of predicted mutations detected by breseq analysis excluding mutations in F3141-F3140, Tn6237 insertions and larges deletions@:the mutation in F3140-F3141 ortholog was identical for all sequenced colonies from every lineage after evolution of ST38_2 pOXA-48_l TC in LB, suggesting it derives from the enrichment during the experiment of a mutation already present in a minor fraction of the original TC bacteriaTable 3: Identification of apsA and apsB homologs by using PSI-BLAST analysis on NCBI clustered nr databaseTable 4 : Identification of apsA and apsB homologs by using PSI-BLAST analysis on NCBI clustered nr databaseTable 5: Plasmids used in this studyPlasmids Characteristics References CommentsCNR160E7-pOXA-48_l IncL , bla OXA-48 GCA_963924215CNR161Fl-pOXA-48_2 IncL , bla OXA-48 GCA_963924225CNR149J3-pOXA-48_3 IncL , bla OXA-48 GCA_963924235 pl5red pl5A, cat, araC ParaBAD -gam-bet-exo, sacB This studyP15A ori from pACY184, CmR and SacB from pMA76, araC, pBAD, gam, bet , exo from pKOBEG5pl5red Apra pl5A, apmR, araC Para BAD -gam-bet-exo, sacB This study ApraR cloned between Spe 1 and pvul of pl5Red and replacing CAT pl5Flip pl5A, apmR, araC, ParaBAD, flp, sacB This study FLP from pCP20 replacing gam, bet, exo in pl5Red pl5 Flip-Apra p!5A, cat, araC, ParaBAD, flp, sacB This study ApraR cloned between Spe 1 and pvu 1 of pl5 Flip and replacing CAT pACYt-apra pACY184Atet / ? Acot, pl5A ori, apmR This study pACYt-ZeoFRT pACY184AtetR, cat, pl5A ori, zeoR This study zeoR inserted between FRT sites at Hin dill and Bam HI sites.CNR146C9-pKPC IncFII / IncFIB , bia-KPC Collection pUC19-apra pMBl . apmR 1 pBbE8c ColEl, CAT, araC ParaBAD, mRFPl 2 pBbS8c pSClOl, CAT, araC ParaBAD, mRFPl 2 pBbE8c-Apra ColEl, apmR, araC ParaBAD, mRFPl This study Replacement of cat by apmR by cloning between Bsm Bl sites of vector pBbS8c-Apra pSClOl, apmR, araC ParaBAD, mRFPl This study Replacement of cat by apmR by cloning between Bsm Bl sites of vector pHV7 ori pl5A , ori fl, cat, araC ParaBAD 3 pHV7-empty ori pl5A , ori fl, apmR, araC ParaBAD This study Replacement of cat by apmR by cloning between Acc 1 and Neo 1 sites pHV7-F3141-F3140 ori pl5A , ori fl, apmR, araC ParaBAD, F3141-F3140 This study apsAB from ST38_1 cloned downstream pBAD of pHV7Apra by using a Bbs 1 cloning strategy pHV7-F3141[A633S; Y677*]-F3140 ori pl5A , ori fl, apmR, araC ParaBAD, F3141[A633S; Y677*]-F3140 This study pHV7-F3141 F3140[S158*] ori pl5A , ori fl, apmR, araC ParaBAD, F3141-F3140[S158*] This study pHV7-F3141[E533A]-F3140 ori pl5A , ori fl, apmR, araC ParaBAD, F3141[E533A]-F3140 This study pHV7-F3141[K1435A]-F3140 ori pl5A , ori fl, apmR, araC ParaBAD, F3141[K1435A]-F3140 This study pHV7-F3141[K221A]-F3140 ori pl5A , ori fl, apmR, araC ParaBAD, F3141[K221A]-F3140 This study pHV7-F3141-F3140[K413A] ori pl5A , ori fl, apmR, araC ParaBAD, F3141-F3140[K413A] This study pUCMiniTn7T-Lac pUC ori; AmpR ; GmR , lad , tac promoter, Tn7_R,L 4 miniTnzeoR (miniTnNotZeoBadRBK) R6K ori; ori T.-AmpR ; ZeoR , araC , ParaBAD, Tn7_R,L This study insertion of a Not / site in the MCS; zeoR replacing GmR by using BsrGl and Sacll sites; pBad+araC (from pHV7) replacing Lad and tac promoter by using Sac 1 and Nsi 1 sites, R6K ori+OriT from pTnS2 replacing pUC ori by using Pci 1 and Bsa 1 sites pTnS2 R6K ori, oriT, AmpR ,tnsA,B,C,D 4 miniTnopsAB R6K ori; ori T;AmpR; ZeoR, araC, ParaBAD, Tn7_R,L, apsAB This study apsAB from ST38_1 cloned between Notl and Spe / sites of miniTnNotZeoBadRGK miniTnops R6K ori; ori T;AmpR; ZeoR, araC , ParaBAD, Tn7_R,L, apsAB CNR36C9 This study apsAB from CNR36C9 cloned between Not / and Spe / sites of miniTnNotZeoBadRGK miniTnopsAB R6K ori; ori T;AmpR; ZeoR, araC, ParaBAD, Tn7_R,L, apsAB This study apsAB from CNR81D10 cloned between Not / and Spe / sites of miniTnNotZeoBadRGK l: Joussetet al.. Clin Infect Dis. 2018 Oct 15;67(9) : 1388-1394. doi: 10.1093 / cid / ciy293. PMID: 296883392: Lee et al. J Biol Eng. 2011 Sep 20;5:12. doi: 10.1186 / 1754-1611-5-12. PMID: 219334103: Voedts et al EMBO J 2021 Oct 1; 40(19) :el08126 doi: 10 15252 / embj 2021108126 Epub 2021 Aug 12 PMID: 343826984: Choi et al.P. Nat Methods. 2005 Jun;2(6) :443-8. doi: 10.1038 / nmeth765. PMID: 159089235: Chaveroche et al.. Nucleic Acids Res. 2000 Nov 15;28(22) :E97. doi: 10.1093 / nar / 28.22.e97. PMID: 110719516: Lennen et al. Nucleic Acids Res. 2016 Feb 29;44(4):e36. doi: 10.1093 / nar / gkvl090. Epub 2015 Oct 22. PMID: 26496947Table 6: Alpha-fold or icn3d pdb models of ApsB-like proteins recovered from databases (&) %Root-mean-square deviation of atomic positions (RMSD) was used as a proxy to estimate protein structure similarity between ApsB and ApsB-likeApsB-like protein Structure model& RMSD %SEQ ID NO: 298 A0A246NZ20_icn3d RMSD = 1.762 (4716 to 4716 atoms)SEQ ID NO: 300 A0A077PNJ9_icn3d RMSD = 1.342 (4223 to 4223 atoms)SEQ ID NO: 303 A0A109D9Al_icn3d RMSD = 0.775 (4009 to 4009 atoms)SEQ ID NO: 309 AF-A0A3P1SM97-F1 RMSD = 3.500 (3080 to 3080 atoms)SEQ ID NO: 312 A0A2E9T341_icn3d RMSD = 1.944 (3024 to 3024 atoms)SEQ ID NO: 319 F8H079_icn3d RMSD = 3.630 (3019 to 3019 atoms)SEQ ID NO: 333 A0A826VHB8_icn3d RMSD = 2.636 (2913 to 2913 atoms)REFERENCES1. Murray, C. J. L. et al. Global burden of bacterial antimicrobial resistance in 2019: a systematic analysis. The Lancet 399, 629-655 (2022)2. Rozwandowicz, M. et al. Plasmids carrying antimicrobial resistance genes in Enterobacteriaceae. J. Antimicrob. Chemother. 73, 1121-1137 (2018).3. Johnson, T. J. et al. In Vivo Transmission of an IncA / C Plasmid in Escherichia coli Depends on Tetracycline Concentration, and Acquisition of the Plasmid Results in a Variable Cost of Fitness. Appl. Environ. Microbiol. 81, 3561-3570 (2015).4. Starikova, I. et al. Fitness costs of various mobile genetic elements in Enterococcus faecium and Enterococcus faecalis. J. Antimicrob. Chemother. 68, 2755-2765 (2013).5. Bergstrom, C. T., Lipsitch, M. & Levin, B. R. Natural selection, infectious transfer and the existence conditions for bacterial plasmids. Genetics 155, 1505-1519 (2000).6. Brockhurst, M. A. & Harrison, E. Ecological and evolutionary solutions to the plasmid paradox. Trends Microbiol. 30, 534-543 (2022).7. San Millan, A., Heilbron, K. & MacLean, R. C. Positive epistasis between co-infecting plasmids promotes plasmid survival in bacterial populations. ISMEJ. 8, 601-612 (2014).8. Dahlberg, C. & Chao, L. Amelioration of the cost of conjugative plasmid carriage in Eschericha coli K12. Genetics 165, 1641-1649 (2003).9. Harrison, E., Guymer, D , Spiers, A. J., Paterson, S. & Brockhurst, M. A. Parallel Compensatory Evolution Stabilizes Plasmids across the Parasitism-Mutualism Continuum. Curr. Biol. 25, 2034-2039 (2015).10. Loftie-Eaton, W. et al. Compensatory mutations improve general permissiveness to antibiotic resistance plasmids. Nat. Ecol. Evol. 1, 1354-1363 (2017).11. Lopatkin, A. J. et al. Persistence and reversal of plasmid-mediated antibiotic resistance. Nat. Commun. 8, 1689 (2017).12. Boyle, T. A. & Hatoum-Aslan, A. Recurring and emerging themes in prokaryotic innate immunity. Curr. Opin. Microbiol. 73, 102324 (2023).13. Jaskolska, M., Adams, D. W. & Blokesch, M. Two defence systems eliminate plasmids from seventh pandemic Vibrio cholerae. Nature 604, 323-329 (2022).14. Mayo-Munoz, D., Pinilla-Redondo, R., Birkholz, N. & Fineran, P. C. A host of armor: Prokaryotic immune strategies against mobile genetic elements. Cell Rep. 42, 112672 (2023).15. Patino-Navarrete, R. et al. Specificities and Commonalities of Carbapenemase-Producing Escherichia coli Isolated in France from 2012 to 2015. mSystems 7, eOl 169-21 (2022).16. Pitout, J. D. D., Peirano, G., Kock, M. M., Strydom, K.-A. & Matsumura, Y. The Global Ascendency of OXA-48-Type Carbapenemases. Clin. Microbiol. Rev. 33, e00102-19 (2019).17. Wielders, C. C. H. et al. Epidemiology of carbapenem-resistant and carbapenemase- producing Enterobacterales in the Netherlands 2017-2019. Antimicrob. Resist. Infect. Control 11, 57 (2022).18. Poirel, L., Bonnin, R. A. & Nordmann, P. Genetic Features of the Widespread Plasmid Coding for the Carbapenemase OXA-48. Antimicrob. Agents Chemother. 56, 559-562 (2012).19. Alonso-del Valle, A. et al. Variability of plasmid fitness effects contributes to plasmid persistence in bacterial communities. Nat. Commun. 12, 2653 (2021).20. Hendrickx, A. P. A. et al. bla OXA-48-like genome architecture among carbapenemase- producing Escherichia coli and Klebsiella pneumoniae in the Netherlands. Microb. Genomics 7, (2021).21. Turton, J. F. et al. Clonal expansion of Escherichia coli ST38 carrying a chromosomally integrated OXA-48 carbapenemase gene. J. Med. Microbiol. 65, 538-546 (2016).22. Fonseca, E. L., Morgado, S. M., Caldart, R. V. & Vicente, A. C. Global Genomic Epidemiology of Escherichia coli (ExPEC) ST38 Lineage Revealed a Virulome Associated with Human Infections. Microorganisms 10, 2482 (2022).23. Emeraud, C. et al. Emergence and Polyclonal Dissemination of OXA-244-Producing Escherichia coli , France. Emerg. Infect. Dis. 27, 1206-1210 (2021).24. Falgenhauer, L. et al. Cross-border emergence of clonal lineages of ST38 Escherichia coli producing the OXA-48-like carbapenemase OXA-244 in Germany and Switzerland. Ini. J. Antimicrob. Agents 56, 106157 (2020).25. Guenther, S. et al. Chromosomally encoded ESBL genes in Escherichia coli of ST38 from Mongolian wild birds. J. Antimicrob. Chemother. 72, 1310-1313 (2017).26. Shawa, M. et al. Novel chromosomal insertions of ISEcpl-blaCTX-M-15 and diverse antimicrobial resistance genes in Zambian clinical isolates of Enterobacter cloacae and Escherichia coli. Antimicrob. Resist. Infect. Control 10, 79 (2021).27. Le6n-Sampedro, R. et al. Pervasive transmission of a carbapenem resistance plasmid in the gut microbiota of hospitalized patients. Nat. Microbiol. 6, 606-616 (2021).28. Rousset, F. et al. Phages and their satellites encode hotspots of antiviral systems. Cell Host Microbe 30, 740-753. e5 (2022).29. White, M. F. Structure, function and evolution of the XPD family of iron-sulfur-containing 5’— >3’ DNA helicases. Biochem. Soc. Trans. 37, 547-551 (2009).30. Knizewski, L., Kinch, L. N., Grishin, N. V., Rychlewski, L. & Ginalski, K. Realm of PD- (D / E)XK nuclease superfamily revisited: detection of novel families with modified transitive meta profile searches. BMC Struct. Biol. 1, 40 (2007).31. Lisitskaya, L., Aravin, A. A. & Kulbachinskiy, A. DNA interference and beyond: structure and functions of prokaryotic Argonaute proteins. Nat. Commun. 9, 5165 (2018).32. Miyoshi, T., Ito, K., Murakami, R. & Uchiumi, T. Structural basis for the recognition of guide RNA and target DNA heteroduplex by Argonaute. Nat. Commun. 7, 11846 (2016).33. Kuang, H., Yang, Y., Luo, H. & Lv, X. The impact of three carbapenems at a single-day dose on intestinal colonization resistance against carbapenem-resistant Klebsiella pneumoniae. mSphere 8, e00479-23 (2023).34. Dionisio, F., Zilhao, R. & Gama, J. A. Interactions between plasmids and other mobile genetic elements affect their transmission and persistence. Plasmid 102, 29-36 (2019).35. San Millan, A., Toll-Riera, M., Qi, Q. & MacLean, R. C. Interactions between horizontally acquired genes create a fitness cost in Pseudomonas aeruginosa. Nat. Commun. 6, 6845 (2015).36. Georjon, H. & Bernheim, A. The highly diverse antiphage defence systems of bacteria. Nat. Rev. Microbiol. 21, 686-700 (2023).37. Millman, A. et al. An expanded arsenal of immune systems that protect bacteria from phages. Cell Host Microbe 30, 1556-1569. e5 (2022).38. Song, X. etal. Catalytically inactive long prokaryotic Argonaute systems employ distinct effectors to confer immunity via abortive infection. Nat. Commun. 14, 6970 (2023).39. Jackson, S. A., Fellows, B. J. & Fineran, P. C. Complete Genome Sequences of the Escherichia coli Donor Strains STI 8 and MFDpir. Microbiol. Resour. Announc. 9, e01014-20 (2020).40. Choi, K.-H. et al. A Tn7-based broad-range bacterial cloning and expression system. Nat. Methods 2, 443-448 (2005).41. Datsenko, K. A. & Wanner, B. L. One-step inactivation of chromosomal genes in Escherichia coli K-12 using PCR products. Proc. Natl. Acad. Sci. 97, 6640-6645 (2000).42. Antipov, D., Korobeynikov, A., McLean, J. S. & Pevzner, P. A. HYBRID SPADES : an algorithm for hybrid assembly of short and long reads. Bioinformatics 32, 1009-1015 (2016).43. Koren, S. et al. Canu: scalable and accurate long-read assembly via adaptive k -mer weighting and repeat separation. Genome Res. 27, 722-736 (2017).44. Hunt, M. et al. Circlator: automated circularization of genome assemblies using long sequencing reads. Genome Biol. 16, 294 (2015).45. Prjibelski, A., Antipov, D., Meleshko, D., Lapidus, A. & Korobeynikov, A. Using SPAdes De Novo Assembler. Curr. Protoc. Bioinforma. 7Q, (2020).46. Deatherage, D. E. & Barrick, J. E. Identification of Mutations in Laboratory-Evolved Microbes from Next-Generation Sequencing Data Using breseq. in Engineering and Analyzing Multicellular Systems (eds. Sun, L. & Shou, W.) vol. 1151 165-188 (Springer New York, New York, NY, 2014).47. Robinson, J. T. etal. Integrative genomics viewer. Nat. Biotechnol. 29, 24-26 (2011).48. Gurevich, A., Saveliev, V., Vyahhi, N. & Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 29, 1072-1075 (2013).49. McKenna, A. et al. The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 20, 1297-1303 (2010).50. Stamatakis, A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30, 1312-1313 (2014).51. Croucher, N. J. et al. Rapid phylogenetic analysis of large samples of recombinant bacterial whole genome sequences using Gubbins. Nucleic Acids Res. 43, el5-el5 (2015).52. Letunic, I. & Bork, P. Interactive Tree Of Life (iTOL) v4: recent updates and new developments. Nucleic Acids Res. 47, W256-W259 (2019).53. Seemann, T. Prokka: rapid prokaryotic genome annotation. Bioinformatics 30, 2068-2069 (2014).54. Zankari, E. et al. Identification of acquired antimicrobial resistance genes. J. Antimicrob. Chemother. 67, 2640-2644 (2012).55. Carattoli, A. et al. In silico detection and typing of plasmids using PlasmidFinder and plasmid multilocus sequence typing. Antimicrob. Agents Chemother. 58, 3895-3903 (2014).56. Pereira, J. GCsnap: Interactive Snapshots for the Comparison of Protein-Coding Genomic Contexts. J. Mol. Biol. 433, 166943 (2021).57. Waterhouse, A. M., Procter, J. B., Martin, D. M. A., Clamp, M. & Barton, G. J. Jalview Version 2 — a multiple sequence alignment editor and analysis workbench. Bioinformatics 25, 1189-1191 (2009).58. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021).59. Sullivan, M. J., Petty, N. K. & Beatson, S. A. Easyfig: a genome comparison visualizer. Bioinformatics 27, 1009-1010 (2011).60. Florensa, A. F., Kaas, R. S., Clausen, P. T. L. C., Aytan-Aktug, D. & Aarestrup, F. M. ResFinder - an open online resource for identification of antimicrobial resistance genes in nextgeneration sequencing data and prediction of phenotypes from genotypes. Microb. Genomics 8, (2022).
Claims
CLAIMS1. An ApsAB system comprising an isolated ApsA protein and an isolated ApsB protein, wherein the ApsA protein is a protein having the sequence of SEQ ID NO: 1, a homolog thereof, or a variant thereof; and the ApsB protein is a protein having the sequence of SEQ ID NO: 2, a homolog thereof, or a variant thereof.
2. The ApsAB system of claim 1, wherein the ApsA protein comprises a SF2 helicase domain and a nuclease domain.
3. The ApsAB system of claim 2, wherein the helicase domain contains six helicase motifs having the sequences:T(G / A)XGK; DE(E / L / V / Q)(H / E)XXY; (P / A)EXX(L / T / V)XXX(L / T / V); SAT;(S / T)X(Y / F)X(S / G)AXXG(L / I / V)(N / D); and Q(A / T)(I / V / L)GRXER, where X is any amino acid and the amino acids between parentheses indicate alternatives.
4. The ApsAB system of claim 2 or claim 3, wherein the nuclease domain contains a nuclease motif having the sequence:(G / E)XXXE(X)2o-5o(Y / F)D(X)i2-i7(L / DV / M)DXKX(W / Y), where X is any amino acid; the amino acids between parentheses indicate alternatives; (X)i2-i7 and (X)2o-so represent a sequence of 12 to 17 and 20 to 50 amino acids, respectively, where each amino acid of the sequence is any amino acid.
5. The ApsAB system of any one of claims 1 to 4, wherein the ApsB protein contains a nucleic acid guide 5 ’end binding motif having the sequence:DXYXXXK(X)i2-i7Q(A / G) (X)35-9oE(L / I)XXK,where X is any amino acid; the amino acids between parentheses indicate alternatives; (X)i2-i7 and (X)3s-9o represent a sequence of 12 to 17 and 35 to 90 amino acids, respectively, where each amino acid of the sequence is any amino acid.
6. The ApsAB system of any one of claims 1 to 5, wherein, the ApsB protein has five B- sheets, ordered 32145, where B-strands 2 and 4 are anti-parallel to the other B-strands.
7. The ApsAB system of any one of claims 3 to 6, wherein the first helicase motif is TGFGK or TASGK; the second helicase motif is DELHEAY or DEEHEAY; the third helicase motif is PEVMLLRLL, PEAMLLRLL or PEIALLRVL; the fifth helicase motif is SSYKSAGTGLN or SHFQGAGTGLN; and / or the sixth helicase motif is QAVGRVER or QTIGRTER.
8. The ApsAB system of claim 7, wherein the six helicase motifs are chosen from: TGFGK, DELHEAY, PEVMLLRLL, SAT, SSYKSAGTGLN and QAVGRVER; TGFGK, DELHEAY, PEAMLLRLL, SAT, SSYKSAGTGLN and QAVGRVER; TASGK, DEEHEAY, PEIALLRVL, SAT, SHFQGAGTGLN and QTIGRTER9. The ApsAB system of any one of claims 4 to 8, wherein the nuclease motif is chosen from: GNVGE(X)3iFD(X)nIDVKRW and GNIGE(X)3iFD(X)nIDVKNW.
10. The ApsAB system of any one of claims 1 to 9, wherein the ApsA and / or ApsB protein has at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the protein of SEQ ID NO: 1 or SEQ ID NO:2.
11. The ApsAB system of any one of claims 1 to 10, wherein the ApsA and / or ApsB is a protein of any of SEQ ID NO: 3 to 6 or a variant thereof having at least 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with a protein of any of SEQ ID NO: 3 to 6.
12. The ApsAB system of any one of claims 1 to 11, wherein the ApsA and / or ApsB protein contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 amino acid changes from any of the proteins of SEQ ID NO: 1 to 6.
13. The ApsAB system of any one of claims 1 to 12, wherein the ApsA and ApsB protein has the sequence of any of SEQ ID NO: 1 to 6.
14. The ApsAB system of any one of claims 1 to 12, wherein the ApsA and / or ApsB is a protein of any of SEQ ID NO: 43 to 337.
15. The ApsAB system of any one of claims 1 to 14, wherein the ApsA protein has endonuclease activity.
16. The ApsAB system of any one of claims 1 to 15, wherein the ApsB protein binds a nucleic acid guide molecule.
17. The ApsAB system of claim 16, wherein the nucleic acid guide is single-stranded RNA, DNA or mixed RNA / DNA molecule.
18. The ApsAB system of claim 16 or 17, wherein the nucleic acid guide is small interfering DNA molecule.
19. The ApsAB system of claim 17 or 18, wherein the guide molecule is 5-end phosphorylated small interfering DNA molecule20. The ApsAB system of any one of claims 16 to 19, wherein the guide molecule is complementary to the sequence of a DNA target.
21. The ApsAB system of any one of claims 16 to 20, which further comprises the nucleic acid guide molecule.
22. The ApsAB system of any one of claim 21, wherein the nucleic acid guide molecule is bound to the ApsB protein.
23. The ApsAB system of any one of claims 1 to 22, which comprises a complex of the ApsA and ApsB proteins.
24. The ApsAB system of any one of claims 1 to 23, which has nucleic acid-guided endonuclease activity.
25. The ApsAB system of any one of claims 20 to 24, which cleaves a DNA target.
26. The ApsAB system of any one of claims 1 to 25, wherein the isolated ApsA or ApsB protein is labelled.
27. A composition comprising the ApsA and ApsB proteins of the ApsAB system of any one of claims 1 to 26.
28. An isolated AspA or ApsB protein of the ApsAB system of any one of claims 1 to 26.
29. A recombinant nucleic acid encoding the ApsA or ApsB protein of claim 28.
30. A recombinant nucleic acid encoding the ApsA and ApsB proteins of claim 28.
31. The recombinant nucleic acid of claim 29 or 30, having at least 80%, 90%, 93%, 95%, 96%, 97%, 98%, 99% or 100% identity with any of the sequences SEQ ID NO: 14 to 19.
32. A vector comprising the recombinant nucleic acid of any one of claims 29 to 31.
33. The vector of claim 32, wherein the vector is a mammalian expression vector.
34. The vector of claim 32, wherein the vector is a bacterial expression vector.
35. The vector of claim 32 or 34 having the sequence of any of SEQ ID NO: 7 to 13.
36. A cell comprising the vector of any one of claims 32 to 35.
37. A method of producing an ApsA or ApsB protein comprising introducing the expression vector of any one of claims 33 to 35 encoding an ApsA or ApsB protein into a host cell and expressing the ApsA or ApsB protein from the vector.
38. The method of claim 37, further comprising isolating the expressed protein from the host cell.
39. A method of cleaving a DNA target comprising:- providing an ApsB protein as defined in any one of claims 1, 5-6, 10-14, 16, 23 and 26; binding the ApsB protein to single stranded nucleic acid guide that is complementary to a dsDNA target to form an ApsB-guide complex;- contacting the ApsB-guide complex with the dsDNA target;- contacting the dsDNA target with an ApsA protein as defined in any one of claims 1-4, 7-15, 23 and 26; and- unwinding and cleaving the dsDNA target with the ApsA protein.
40. The method of claim 39, wherein the method is performed in vitro.
41. The method of claim 39 or 40, wherein the single stranded nucleic acid guide complementary to a dsDNA target is 5'-phosphorylated.
42. The method of any one of claims 39 to 41, wherein the method is performed in mammalian cells.
43. A kit comprising a nucleic acid guide as defined in any one of claims 17 to 20 and 22, an ApsB protein, and optionally an ApsA protein as defined in any one of claims 1 to 16, 22, 23 and 26.
44. The kit of claim 43, further comprising a fluorescent reporter nucleic acid, which has a fluorescent group and a quenching group.
Citation Information
Patent Citations
Recombinant immunoglobulin preparations, methods for their preparation, DNA sequences, expression vectors and recombinant host cells therefor
EP0125023A1
Process for the production of a chimera monoclonal antibody
EP0171496A2
Chimeric receptors by DNA splicing and expression
EP0173494A2
Mouse-human chimaeric immunoglobulin heavy chain, and chimaeric DNA encoding it
EP0184187A2
Transgenic non-human animals capable of producing heterologous antibodies
GB2272440A