Compositions and methods for control of specific bacterial populations

WO2025054631A3PCT designated stage expired Publication Date: 2025-06-12MASSACHUSETTS EYE & EAR INFARY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/046002
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-10
Filing Date
2024-09-10
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The rise of antibiotic-resistant bacteria, particularly vancomycin-resistant Enterococci (VRE), poses a significant threat to modern healthcare, as existing antibiotics are ineffective against these resistant strains.

Method used

Development of efagin, a phage-derived antimicrobial agent that specifically binds to bacterial cell wall polysaccharides, including enterococcal polysaccharide antigen (EPA), allowing for targeted killing of VRE and other enterococcal strains.

Benefits of technology

Efagins demonstrate targeted antibacterial activity, effectively killing or inhibiting the growth of VRE and other enterococcal strains, offering a novel approach to combat antibiotic-resistant infections.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention features E. faecalis phage-derived antimicrobial agent (efagins). The invention also features compositions and methods of using the efagins to treat bacterial infections.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPOSITIONS AND METHODS FOR CONTROL OF SPECIFIC BACTERIAL POPULATIONS

[0002] CROSS-REFERENCE TO RELATED APPLICATION

[0003] This application claims the benefit of U.S. Patent Application Serial No. 63 / 537,523, filed on September 10, 2023, the disclosure of which is incorporated by reference.

[0004] STATEMENT AS TO FEDERALLY FUNDED RESEARCH

[0005] This invention was made with government support under AI083214 and AI110818 awarded by the National Institute of Allergy and Infectious Diseases, National Institutes of Health. The government has certain rights in the invention.

[0006] BACKGROUND OF THE INVENTION

[0007] The rise in antibiotic-resistance across the globe has become a major threat to modern healthcare. As the number of newly approved antibiotics has dwindled over the past decade, the need for novel agents to combat drug-resistant microbes has greatly increased. Never has the need for new treatment options been more evident than at the present time, when the evolution of resistance in bacteria has outpaced modern drug development efforts.

[0008] Enterococci are a genus of gram-positive, round-shaped bacteria that commonly live in the gut, although they can cause infection anywhere in the body. The genus has a high amount of intrinsic resistance to some classes of antibiotics, but is generally sensitive to vancomycin. However, particularly virulent strains that are resistant even to vancomycin are an emerging and are a growing problem, particularly in healthcare facilities. Such vancomycin-resistant strains are referred to as VRE. The two main VRE species are vancomycin-resistant Enterococcus faecium (E. faecium) and vancomycin- resistant Enterococcus faecalis (E. faecalis). VRE can exist in the body, typically in the gastrointestinal tract, without causing a disease or other harmful effects. However, VRE sometimes cause local disease in the gastrointestinal tract and they can invade sites outside the gastrointestinal tract causing disease, for example, in the bloodstream, abdomen, or urinary tract. VRE in the bloodstream can be particularly problematic because, once in the bloodstream, VRE can cause sepsis, meningitis, pneumonia, or endocarditis.

[0009] There is accordingly an urgent need to develop new antimicrobial agents for treating bacterial infections.

[0010] SUMMARY OF THE INVENTION

[0011] In one aspect, the invention features an isolated efagin antibacterial particle, or a portion thereof, that specifically binds a bacterial cell wall polysaccharide.

[0012] In some embodiments, the efagin, or a portion thereof, specifically binds an enterococcal polysaccharide antigen (EPA). In some embodiments, the efagin, or a portion thereof, specifically binds an Enterococcus faecalis EPA.

[0013] In some embodiments, the efagin, or a portion thereof, specifically binds a non-enterococcus cell wall polysaccharide. In some embodiments, the efagin, or a portion thereof, includes a receptor binding protein.

[0014] In some embodiments, the receptor binding protein is a group 1 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 7, 16, and 25.

[0015] In some embodiments, the receptor binding protein is a group 2 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 34, 43, 52, and 61 .

[0016] In some embodiments, the receptor binding protein is a group 3 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 70, 79, and 88.

[0017] In some embodiments, the receptor binding protein specifically binds to the cell wall polysaccharide (for example, EPA).

[0018] In some embodiments, the efagin, or a portion thereof, includes a baseplate / distal tail protein, an endopeptidase tail, a receptor binding protein, a holin, and an amidase. In some embodiments, the efagin, or a portion thereof, further includes one or more of a tail completion protein, a major tail protein, a hypothetical protein, and a tail tape measure protein.

[0019] In some embodiments, the efagin, or a portion thereof, is a chimeric efagin, or a portion thereof.

[0020] In another aspect, the invention features a composition including an efagin, or a portion thereof, described herein.

[0021] In some embodiments, the composition includes group 1 and group 2 efagins, or a portion thereof.

[0022] In some embodiments, the composition includes group 1 and group 3 efagins, or a portion thereof.

[0023] In some embodiments, the composition includes group 2 and group 3 efagins, or a portion thereof.

[0024] In some embodiments, the composition includes group 1 , group 2, and group 3 efagins, or a portion thereof.

[0025] In some embodiments, the composition includes a chimeric efagin, or a portion thereof.

[0026] In some embodiments, the composition further includes an antibiotic.

[0027] In another aspect, the invention features an isolated bacterium including an efagin, or a portion thereof.

[0028] In some embodiments, the bacterium is a member of the genus Enterococcus.

[0029] In some embodiments, the bacterium is commensal.

[0030] In some embodiments, the bacterium is an opportunistic pathogen.

[0031] In another aspect, the invention features a method of treating a subject having a bacterial infection, including administering an effective amount of an efagin, or a portion thereof, described herein, a composition described herein, or an isolated bacterium described herein.

[0032] In some embodiments, the bacterial infection is a wound infection, a urinary tract infection, bacteremia, or infective endocarditis. In some embodiments, the bacterial infection is an enterococcal infection. In some embodiments, the enterococcal infection is an Enterococcus faecalis infection.

[0033] In some embodiments, the bacterial infection is a nosocomial infection.

[0034] In some embodiments, the efagin, or a portion thereof, the composition, or the isolated bacterium is administered to the skin, the upper respiratory tract, the oral cavity, or the vagina of the subject.

[0035] In another aspect, the invention features a nucleic acid molecule or set of nucleic acid molecules encoding an efagin, or a portion thereof, described herein.

[0036] In some embodiments, the nucleic acid molecule has sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to the sequence of any one of SEQ ID NOs: 91 -93.

[0037] In another aspect, the invention features an expression vector or set of expression vectors including a nucleic acid molecule or set of nucleic acid molecules described herein.

[0038] In another aspect, the invention features a host cell including (i) a nucleic acid molecule or set of nucleic acid molecules described herein or (ii) an expression vector or set of expression vectors described herein.

[0039] In some embodiments, the host cell is a bacterial cell, yeast cell, mammalian cell, insect cell, or plant cell. In some embodiments, the bacterial cell is a commensal species.

[0040] In another aspect, the invention features a method of producing an efagin, or a portion thereof, including:

[0041] (a) culturing a host cell described herein under conditions where the efagin, or a portion thereof, is expressed; and

[0042] (b) isolating the efagin, or a portion thereof, expressed in (a) from the host cell culture, thereby producing the efagin, or a portion thereof.

[0043] Advantageously, the compositions and methods described herein provide for antibacterial agents, termed efagins, that are specific for certain bacterial cell wall carbohydrates, such as enterococcal polysaccharide antigen (EPA). Moreover, variations in the receptor binding protein of the efagin correspond to variations in target cell wall carbohydrate (e.g., EPA) type. The target specificity of the efagin can advantageously be tuned by altering the sequence of the receptor binding protein, producing a chimeric efagin. Another advantage provided herein is that these antibacterial agents may be used to specifically kill E. faecalis target strains, as well as other enterococci and non-enterococci that share elements of the EPA structure in their cell walls.

[0044] Other features and advantages of the invention will be apparent from the following detailed description and figures, and from the claims.

[0045] Definitions

[0046] “Percent (%) sequence identity” with respect to a reference polypeptide sequence is defined as the percentage of amino acids in a candidate sequence that are identical to the amino acids in the reference polypeptide sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity. Alignment for purposes of determining percent amino acid sequence identity can be achieved in various ways that are within the capabilities of one of skill in the art, for instance, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. For example, percent sequence identity values may be generated using the sequence comparison computer program BLAST. As an illustration, the percent sequence identity of a given amino acid sequence, A, to, with, or against a given amino acid sequence, B, (which can alternatively be phrased as a given amino acid sequence, A that has a certain percent sequence identity to, with, or against a given amino acid sequence, B) is calculated as follows:

[0047] 100 multiplied by (the fraction X / Y) where X is the number of amino acids scored as identical matches by a sequence alignment program (e.g., BLAST) in that program’s alignment of A and B, and where Y is the total number of amino acids in B. It will be appreciated that where the length of amino acid sequence A is not equal to the length of amino acid sequence B, the percent sequence identity of A to B will not equal the percent sequence identity of B to A. An efagin described herein can have at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity to any of the sequences found in Tables 1 -10, for example, to one or more of a tail completion protein, a major tail protein, a hypothetical protein, a tail tape measure protein, a baseplate / distal tail protein, an endopeptidase tail, a receptor binding protein, a holin, and an amidase sequence found in Tables 1 -10.

[0048] As used herein, the term “specifically binds” indicates that the binding is selective and can be discriminated from unwanted or non-specific interactions. Molecules and binding partner pairs that “specifically bind” generally have a Kd of at least in the pM range, e.g., 1 -10 pM, 10-100 pM, 100-200 pM, 200-300 pM, 300-400 pM, 400-500 pM, 500-600 pM, 600-700 pM, 700-800 pM, 800-900 pM, or 900-1000 pM. An exemplary assay for measuring binding affinity is provided in Walter et al., J Virol. 2008 Mar; 82(5): 2265-2273.

[0049] As used herein, the term “efagin” refers to an E. faecalis phage-derived antimicrobial agent.

[0050] As used herein, the term “bacterial cell wall polysaccharide” refers to a polysaccharide that is covalently bound to peptidoglycan. Exemplary bacterial cell wall polysaccharides include capsular polysaccharides and wall teichoic acids.

[0051] As used herein, the terms “enterococcal polysaccharide antigen” and “EPA” are used interchangeably and refer to a cell wall polysaccharide in enterococci. EPA is a homolog of cell wall teichoic acids in other gram positive species.

[0052] An “isolated” efagin, or a portion thereof, refers to a preparation of the efagin, or a portion thereof, devoid of at least some of the other components that may also be present where the efagin or a similar substance naturally occurs or is initially prepared from. Thus, for example, an isolated efagin, or a portion thereof, may be prepared by using a purification technique to enrich it from a source mixture. In some embodiments, an isolated efagin is at least 50% by weight free from proteins and naturally-occurring molecules with which it is naturally associated. Preferably, the preparation is at least 70% (e.g., at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) by weight free from proteins and naturally- occurring molecules with which it is naturally associated. Purity can be measured by any appropriate method, for example, column chromatography or PEG precipitation. By “bacterial infection” is meant the invasion of a host by pathogenic bacteria. For example, the infection may include the excessive growth of bacteria that are normally present in or on the body of a human or growth of bacteria that are not normally present in or on a human. More generally, a bacterial infection can be any situation in which the presence of a bacterial population(s) is damaging to a host body. Thus, a subject has a bacterial infection when an excessive amount of a bacterial population is present in or on the person’s body, or when the presence of a bacterial population(s) is damaging the cells or other tissue of the subject.

[0053] As used herein, the terms “receptor binding protein” and “tail fiber protein” are used interchangeably and refer to a protein derived from a phage protein harbored by an enterococcal bacterium (for example, a strain of E. faecalis) that contributes to recognition and attachment of the phage to the surface of a target cell.

[0054] As used herein, the term “chimeric efagin” refers to an efagin that includes sequences from phage harbored by two or more enterococci bacteria (for example, strains of E. faecalis) or engineered to include a non-naturally occurring efagin. In a chimeric efagin, the target strain specificity is determined by the receptor binding protein.

[0055] As used herein, “treating” in reference to a disease or condition, refer to an approach for obtaining beneficial or desired results, e.g., clinical results. Beneficial or desired results can include, but are not limited to, reduction or elimination of a bacterial infection. The effect can be prophylactic in terms of completely or partially preventing a disease or infection or symptom thereof and / or can be therapeutic in terms of a partial or complete cure for a disease or infection and / or adverse effect attributable to the disease or infection. “Treatment,” as used herein, covers any treatment of a disease or condition or infection in an animal, particularly in a mammal (e.g., a human), and includes: (a) preventing the disease or infection from occurring in a subject which can be predisposed to the disease or infection but has not yet been diagnosed as having it; (b) inhibiting the disease or infection, e.g., arresting its development; and (c) relieving the disease or infection, e.g., reducing or eliminating a bacterial infection.

[0056] As used herein, the term “effective amount” refers to an amount of an efagin, or a portion thereof, a composition, or an isolated bacterium including an efagin, or a portion thereof, described herein, sufficient to effect the recited result, e.g., to kill a bacterium, lyse a bacterium, inhibit growth of a bacterium (e.g., by at least 20% (e.g., at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100%) when compared to a control), or treat a subject having a bacterial infection.

[0057] As used herein, “subject” refers to an animal, e.g., preferably a mammal, such as a human. Other mammals include, but are not limited to, farm animals (e.g., cows, pigs, horses, cattle, sheep, etc.), pets (e.g., cats and dogs), and primates. Other farm animals include poultry such as chickens, turkeys, and birds, as well as other commercially produced animals including fish, shrimp, and other types of animals potentially vulnerable to bacterial infection.

[0058] As used herein, the term “nosocomial infection” refers to a disease or infection occurring or originating in a hospital or other healthcare facility. Nosocomial infections are typically infections that are contracted in a hospital or other healthcare facility, and can be caused by infectious agents, e.g., bacteria, such as bacteria that are resistant to antibiotics. In certain aspects, a nosocomial infection is not present or incubating prior to the subject being admitted to the hospital or healthcare facility, which is acquired or contracted after the subject's admittance to the hospital or healthcare facility.

[0059] As used herein, the term “host cell” refers to a cell that includes the necessary cellular components needed to express one or more efagins, or a portion thereof, from their corresponding nucleic acids. The nucleic acids are typically included in nucleic acid vectors that can be introduced into the host cell by conventional techniques known in the art (transformation, transfection, electroporation, calcium phosphate precipitation, direct microinjection, etc.). A host cell may be a prokaryotic cell, e.g., a bacterial cell, a eukaryotic cell, e.g., a mammalian cell (e.g., a CHO cell or a HEK293 cell), a yeast cell, an insect cell, or a plant cell.

[0060] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood to one of ordinary skill in the art to which this disclosure belongs. For any term present in the art which is identical to any term expressly defined in this disclosure, the term's definition presented in this disclosure will control in all respects. Although methods and materials similar or equivalent to those described herein can be used in the practice of the disclosed methods and compositions, the exemplary methods and materials are described herein.

[0061] BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application with color drawings will be provided by the Office upon request and payment of the necessary fee.

[0063] FIG. 1 is a schematic representation of the genetic organization of the efagin region in the Enterococcus faecalis genome. The efagin genes (highlighted in blue) span a 14.6 kb region from EF1276 to EF1293. Gene function prediction was performed by combining information from Prokka, Pgam, Kegg TmHMM, and Phyre2. Notably, this region does not encode capsid proteins or proteins involved in phage genome excision and replication. Its genomic neighborhood (grey) is highly conserved. The receptor-binding protein, highlighted in dark blue, exhibits variations within the efagin cluster. These variations were used to classify different efagin types. The efagin element as a gene organization and gene content that suggests a distinct evolutionary origin from a yet to be identified phage ancestor. Efagins appear to exist in E. faecalis in at least 3 structural classes that vary at the C-terminus of EF1291 , which correlates with variation in target cell selection.

[0064] FIG. 2 is a representative Transmission Electron Microscopic (TEM) image of phage tail-like particles, efagins, produced by E. faecalis strain OG1 RF. Morphological features of efagins were visualized and measured using TEM and Fiji / lmageJ software.

[0065] FIGS. 3A-C are a series of representative images showing that efagins are responsible for cell killing and exhibit target strain-specific activity. A) Efagins from E. faecalis strain OG1 RF contained within a polyethylene glycol (PEG) precipitate showed activity against some E. faecalis strains seeded in soft agar (left; X98, FA2-2, ATCC 19433, DIV0223b), but not other strains (right; OG1 RF, B653, V583, and T3). An identical control PEG precipitate made from OG1 RF cultures in which the efagin operon was deleted exhibit no killing activity, proving that the killing activity is ascribable to the efagin itself. B) A set of isogenic deletions of this gene cluster in several E. faecalis strains. Each strain is an E. faecalis species that differs in its cell wall carbohydrate. Each parental strain and isogenic deletant were compared for the ability to inhibit the growth of other enterococcal strains. Efagins show activity against some target E. faecalis strains in soft agar (upper), but not others (lower). C) Efagins, unlike bona fide phages (e.g., M13), appear unable to self-replicate, infect adjacent cells, and form plaques. Instead, target cell killing appears to be a single hit event.

[0066] FIGS. 4A-C are a series of schematics showing grouping of efagin sequences by variation in the EF1291 C-terminal region. A) Efagin sequences grouped by variation in EF1291 C-terminal region (efagin positions 12,800 to 13,134 of 2,774 E. faecalis species), using hierarchical clustering. Variants (including SNPs and insertions / deletions) were identified by comparison arbitrarily to an E. faecalis V583 prototype. Efagin types are as follows: type A: B3751 ; type B: T9; type C: FA2-2; type D: COM1 ; type E: COM7; type F: B653; type G: 0G1 RF; type H: T1 ; and type I: VET195. B) Representative variation in EF1291 derived from comparison of 3,069 diverse E. faecalis genomes indicates the existence of 3 structural groups. Efagin variation clusters at the C-terminal encoding end of EF1291 which appears to be a unique variable region (efagin nucleotide positions 12,800 to 13,134 in the E. faecalis V583 genome). C) Comparative analysis of C-terminal variation in EF1291 derived from examination of 3,069 diverse E. faecalis genomes reveals the existence of three main structural groups that correlate with variation in target cell selectivity. Structural variation in efagin particles nearly exclusively clusters at the C-terminal end of EF1291 (nucleotide positions 12,800 to 13,134 in the E. faecalis V583 genome), which encodes a protein inferred to be a phage tail fiber protein. Conserved residues across all efagin types are underlined, while residues highlighted in bold are variations conserved only within each structural group.

[0067] FIGS. 5A-D are a series of schematics showing that the evolution of enterococcal polysaccharide antigen (EPA) and EF1291 are correlated. Identity was detected by EF1291 structural modeling. A) Alignment of EF1291 to Listeria gp15. Absent detectable primary sequence identity, Phyre2 was used to identify similarity with known protein structural motifs. A high confidence hit was found to the receptorbinding domain of Listeria phage PSA protein gp15, shown to recognize wall teichoic acid (WTA) and confer host target cell specificity. B) Correlated evolution of EF1291 and EPA. Since EPA is a form of WTA, the correlation between the presence of EPA genes and EF1291 SNPs was examined using a whole genome phylogeny of 237 distinct E. faecalis that spans the species' diversity. The correspondence between variation in EF1291 (blue) and in EPA genes (red) is shown, indicating that the evolution of these genes may be correlated. C) Predicted structure of EF1291 , highlighting residues correlated with EPA. D) Structure of Listeria gp15, highlighting residues known to be involved in binding to WTA.

[0068] FIGS. 6A-B are a series of representative schematics of the enterococci cell wall, highlighting A) the EPA and B) diagram of the EPA operon based on Guerardel et al. mBio. 2020; 11 (2): e00277-20. E. faecalis EPA gene cluster consists of a conserved core set of genes (grey) upstream of a group of variable genes (pink).

[0069] FIGS. 7A-B are a series of schematics showing that efagins target strains of heterologous EPA type and not self. A) Specificity of killing by representative efagins of four different SNP types against 13 different EPA types. B) Frequency of combinations of efagin types and EPA types present in E. faecalis representative genomes from this study. Strains grouped in the same EPA type share similarity in the variable genes of their EPA operons. FIG. 8 shows a spectrum of efagin activity among different E. faecalis EPA types. Representative enterococcal strains carrying specific efagin types are specific for killing enterococcal strains carrying different EPA types. Strains grouped in the same EPA type share similarity in the variable genes of their EPA operons.

[0070] DETAILED DESCRIPTION OF THE INVENTION

[0071] The disclosure provides novel bactericidal agents for treating bacterial infections, termed efagins {Enterococcus faecalis phage-derived antimicrobial agents). Efagins are anti-enterococcal entities that evolved from elements of phages, with antibacterial properties. The efagins are characterized by a stable protein complex derived ancestrally from phage tails. They may be purified and used systematically or locally to control discrete populations of bacteria in treating or preventing infection.

[0072] The present invention is based, in part, on the discovery by the present inventors that a gene cluster, previously annotated and publicly regarded as a defective phage in the genome of E. faecalis strain V583, is ubiquitous in all isolates of E. faecalis, with important functional variations, and that it encodes a new type of antibacterial factor of a class generally referred to as a phage tail-like bacteriocin. Moreover, the present inventors have found that this antibacterial factor specifically binds a cell wall carbohydrate termed enterococcal polysaccharide antigen (EPA), and that target EPA specificity is determined by variations in the C-terminus of an efagin tail fiber protein EF1291 . The inventors discovered that these agents may be used to specifically inhibit E. faecalis target strains, as well as other enterococci and non-enterococci that share elements of the EPA structure in their cell walls.

[0073] I. Efag ins

[0074] The disclosure provides agents for treating bacterial infections, termed efagins. Efagins represent a new type of anti-enterococcal entity that evolved from elements of phages, with antibacterial properties. They appear to be uniquely found in the species E. faecalis, and are encoded as part of the core chromosome of that species. Their antibacterial property, however, extends to other enterococcal species and non-enterococcal species.

[0075] The genes for efagin are clustered in the chromosome of E. faecalis and are highly conserved, indicating that it is under selection and is important to the biology of that species. However, the 3’ end of one of the genes associated with efagin expression is unusually polymorphic. This polymorphism was observed to co-vary with polymorphisms in another region of the E. faecalis chromosome that encodes the E. faecalis homolog of wall teichoic acid, termed “enterococcal polysaccharide antigen” or EPA.

[0076] The efagin encoded by one strain of E. faecalis is not effective against the producing strain, but is effective against other strains of E. faecalis with different EPA content. Structural modeling showed that the polymorphic efagin gene within was related to that encoding a tail fiber of a Listeria phage that bound the listeria cell wall polysaccharide. Moreover, the polymorphisms occurring within the efagin could be mapped to polysaccharide binding sites within the listeria phage tail fiber.

[0077] The finding that efagins target non-homologous strains indicates they play a natural role in defense. As the efagins specifically bind the cell wall polysaccharide of other E. faecalis strains, the defense is likely against ecological displacement by the other E. faecalis lineages. Moreover, the finding that efagin activity extends to some other enterococcal species indicates that this defense is broader than only against other E. faecalis strains. Other evolved phage-derived antibacterial factors have been described from a few other types of bacteria, but none are closely related to efagins, and so far, no other bacterial species has been identified where these elements are part of the core chromosome of that species.

[0078] 1. E fag in structure

[0079] In some embodiments, the efagin, or a portion thereof, comprises a phage-like structure. For example, the efagin, or a portion thereof, may comprise a morphological structure resembling at least a phage tail. The structure may be determined using, for example, transmission electron microscopy (TEM).

[0080] In some embodiments, the efagin, or a portion thereof, comprises a receptor binding protein. A receptor binding protein refers to a protein derived from phage that contributes to recognition and attachment of the phage to the surface of a target cell. A receptor binding protein may also be known as a tail fiber protein from derived from phage.

[0081] In some embodiments, the efagin, or a portion thereof, does not have a phage head. A phage head is a phage-like structure that contains a phage’s genetic material.

[0082] In some embodiments, efagins, or a portion thereof, are phage tail-like particles of between about 10 nm to about 1000 nm in contour length (e.g., between about 10 nm to about 50 nm, between about 50 nm to about 100 nm, between about 100 nm to about 150 nm, between about 150 nm to about 200 nm, between about 200 nm to about 250 nm, between about 250 nm to about 300 nm, between about 300 nm to about 350 nm, between about 350 nm to about 400 nm, between about 400 nm to about 450 nm, between about 450 nm to about 500 nm, between about 500 nm to about 550 nm, between about 550 nm to about 600 nm, between about 600 nm to about 650 nm, between about 650 nm to about 700 nm, between about 700 nm to about 750 nm, between about 750 nm to about 800 nm, between about 800 nm to about 850 nm, between about 850 nm to about 900 nm, between about 900 nm to about 950 nm, or between about 950 nm to about 1000 nm). In some embodiments, the efagins, or a portion thereof, are between about 100 nm to about 200 nm in length. In some embodiments, the efagins, or a portion thereof, are about 160 nm in length.

[0083] In some embodiments, the efagin, or a portion thereof, comprises a baseplate / distal tail protein, an endopeptidase tail, a receptor binding protein, a holin, and an amidase.

[0084] In some embodiments, the efagin, or a portion thereof, further comprises one or more of a tail completion protein, a major tail protein, a hypothetical protein, and a tail tape measure protein.

[0085] Tables 1 -10, below, provide exemplary amino acid sequences that may be included in an efagin described herein. In some embodiments, an efagin may include an amino acid sequence having sequence identity to any one of the sequences as described herein. Table 1. Exemplary efagin amino acid sequences from E. faecalis strain OG1RF

[0086] Table 2. Exemplary efagin amino acid sequences from E. faecalis strain T1

[0087] Table 3. Exemplary efagin amino acid sequences from E. faecalis strain B653

[0088] Table 4. Exemplary efagin amino acid sequences from E. faecalis strain FA2-2

[0089] Table 5. Exemplary efagin amino acid sequences from E. faecalis strain Com1

[0090] Table 6. Exemplary efagin amino acid sequences from E. faecalis strain E1Sol

[0091] Table 7. Exemplary efagin amino acid sequences from E. faecalis strain VET-195 Table 8. Exemplary efagin amino acid sequences from E. faecalis strain B3751

[0092] Table 9. Exemplary efagin amino acid sequences from E. faecalis strain T9

[0093] Table 10. Exemplary efagin amino acid sequences from E. faecalis strain Com7

[0094] 2. Efagin types

[0095] The entire efagin genetic unit is approximately 14.6 kbp in length. The present inventors have examined the genomes of over 3,000 E. faecalis strains and found almost as many SNP variants of efagins. The vast majority of these SNPs are inconsequential to activity. However, variants can be clustered into groups as described herein, based on co-occurring sequence variations.

[0096] The present inventors have identified 3 groups differing in conserved patterns of variation that also correlate with variation in target cell specificity.

[0097] Variations in the receptor binding protein contribute to the target strain specificity of the efagin, or a portion thereof. For example, in E. faecalis, the receptor binding protein is EF1291 (corresponding to nucleotide positions 12,800 to 13,134 in the E. faecalis V583 genome). The efagin genes in E. faecalis are highly conserved, except for the C-terminus of EF1291 . Efagin types are defined by the variable region in the C-terminus of EF1291 . Each group of efagin receptor binding proteins includes a common set of amino acid mutations that are only conserved within the group and that differ between other groups. At least 3 groups of efagin receptor binding proteins exist in different E. faecalis strain genomes. Alignment of the whole efagin cluster (14.6 Kb) of 12 strains (OG1 RF, T1 , B653, V583, FA2-2, Com1 , E1 Sol, VRT-195, B3751 , T9, Com7, T6) from 3 efagin groups using Mauve genome alignment viewer showed 13,91 1 (95.1 %) identical sites. Thus, there are fewer than 5% variations (-700 SNPS and indels) among these 12 E. faecalis representing the 3 efagin groups. Alignment results are provided in Table 1 1 , below.

[0098] Table 11. Alignment of the whole efagin cluster of 12 strains (OG1 RF, T1 , B653, V583, FA2-

[0099] 2, Com1, E1Sol, VRT-195, B3751, T9, Com7, and T6) from 3 efagin groups using Mauve genome alignment

[0100] Group 1

[0101] In some embodiments, the efagin receptor binding protein is a group 1 receptor binding protein. An efagin, or a portion thereof, having a group 1 receptor binding protein is a group 1 efagin, or a portion thereof.

[0102] Alignment of the whole efagin cluster (14.6Kb) of 4 strains of group 1 efagins (OG1 RF, T1 , B653, V583) using Mauve genome alignment viewer showed 14,278 (97.6%) identical sites. Alignment results are provided in Table 12, below.

[0103] Table 12. Alignment of the whole efagin cluster (14.6 Kb) of 4 strains of group 1 efagins (OG1RF, T1, B653, and V583) using Mauve genome alignment viewer

[0104] Group 1 receptor binding proteins have at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) sequence identity to a group 1 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 7, 16, and 25) . In some embodiments, Group 1 receptor binding proteins have at least 97% (e.g., 97.6%) sequence identity to a group 1 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 7, 16, and 25).

[0105] Exemplary group 1 receptor binding proteins include receptor binding proteins from E. faecalis strain OG1 RF, E. faecalis strain T1 , and E. faecalis strain B653.

[0106] Group 2

[0107] In some embodiments, the efagin receptor binding protein is a group 2 receptor binding protein. An efagin, or a portion thereof, having a group 2 receptor binding protein is a group 2 efagin, or a portion thereof.

[0108] Alignment of the whole efagin cluster (14.6 Kb) of 4 strains of Group 2 efagins (FA2-2, Com1 , E1 Sol, VET-195) using Mauve genome alignment viewer showed 14,290 (97.7%) identical sites. Alignment results are provided in Table 13, below.

[0109] Table 13. Alignment of the whole efagin cluster (14.6 Kb) of 4 strains of Group 2 efagins (FA2-2, Com1, E1Sol, and VET-195) using Mauve genome alignment viewer

[0110] Group 2 receptor binding proteins have at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) sequence identity to a group 2 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 34, 43, 52, and 61 ). In some embodiments, Group 2 receptor binding proteins have at least 97% (e.g., 97.7%) sequence identity to a group 2 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 34, 43, 52, and 61 ).

[0111] Exemplary group 2 receptor binding proteins include receptor binding proteins from E. faecalis strain FA2-2, E. faecalis strain Com1 , E. faecalis strain E1 Sol, and E. faecalis strain VET-195.

[0112] Group 3

[0113] In some embodiments, the efagin receptor binding protein is a group 3 receptor binding protein. An efagin, or a portion thereof, having a group 3 receptor binding protein is a group 3 efagin, or a portion thereof. Alignment of the whole efagin cluster (14.6Kb) of 4 strains of Group 3 efagins (B3751 , T9, T6, Com7) using Mauve genome alignment viewer showed 14,305 (97.8%) identical sites. Alignment results are provided in Table 14, below.

[0114] Table 14. Alignment of the whole efagin cluster (14.6Kb) of 4 strains of Group 3 efagins (B3751, T9, T6, Com7) using Mauve genome alignment viewer

[0115] Group 3 receptor binding proteins have at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) sequence identity to a group 3 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 70, 79, and 88). In some embodiments, Group 3 receptor binding proteins have at least 97% (e.g., 97.8%) sequence identity to a group 3 receptor binding protein sequence disclosed herein (e.g., any one of SEQ ID NOs: 70, 79, and 88).

[0116] Exemplary group 3 receptor binding proteins include receptor binding proteins from E. faecalis strain B3751 , E. faecalis strain T9, E. faecalis strain Com7, and E. faecalis strain T6.

[0117] In some embodiments, the efagin, or a portion thereof, is a chimeric efagin, or a portion thereof. Chimeric efagins include efagin sequences derived from two or more E. faecalis species. In a chimeric efagin, or a portion thereof, the target strain specificity is determined by the receptor binding protein.

[0118] The efagin groups described herein are the predominant and most commonly isolated groups in the strain collection and in publicly available genome sequences. However, the current efagin groups do not target all known variations in the EPA cell wall carbohydrate operon, and it is therefore likely that there are 2 or 3 more rarely occurring variant groups to be discovered and well defined.

[0119] 3. Efagin target strain specificity

[0120] The efagins, or portions thereof, described herein specifically bind a bacterial cell wall polysaccharide. Cell wall polysaccharides serve to cover the bacterium with a layer of glycans that is directly exposed to the environment. Exemplary bacterial cell wall polysaccharides include capsular polysaccharides and wall teichoic acids. In some embodiments, the receptor binding protein specifically binds to the cell wall polysaccharide (e.g., EPA). In some embodiments, the receptor binding protein specifically binds to the cell wall polysaccharide with a Kd of at least in the pM range, as described herein. In some embodiments, binding of an efagin to a cell wall polysaccharide of a bacterium results in lysis of the bacterium. In some embodiments, binding of an efagin to a cell wall polysaccharide of a bacterium results in killing of the bacterium. In some embodiments, binding of an efagin to a cell wall polysaccharide of a bacterium results in growth inhibition of the bacterium (i.e., bacteriostasis).

[0121] In some embodiments, the efagin, or a portion thereof, specifically binds to a wall teichoic acid. In some embodiments, the receptor binding protein specifically binds to a wall teichoic acid. Wall teichoic acids are anionic glycopolymers that are covalently attached to peptidoglycan of Gram-positive bacteria. In enterococci, wall teichoic acids are known as enterococcal polysaccharide antigen (EPA).

[0122] In some embodiments, the efagin, or a portion thereof, specifically binds an EPA. In some embodiments, the receptor binding protein specifically binds to an EPA. For example, the efagin, or a portion thereof, may specifically bind an EPA from an enterococcus species, such as Enterococcus alcedinis, Enterococcus alishanensis, Enterococcus aquimarinus, Enterococcus asini, Candidatus Enterococcus avicola, Enterococcus avium, Enterococcus bovis, Enterococcus bulliens, Enterococcus burkinafasonensis, Enterococcus caccae, Enterococcus camelliae, Enterococcus canintestini, Enterococcus canis, Enterococcus casseliflavus, Enterococcus cecorum, Enterococcus cloacae, Enterococcus coli, Enterococcus columbae, Enterococcus crotali, Enterococcus devriesei, Enterococcus diestrammenae, Enterococcus dispar, Enterococcus dongliensis, Enterococcus durans, Enterococcus entomosocium, Enterococcus eurekensis, Enterococcus faecalis, Enterococcus faecium, Enterococcus flavescens, Enterococcus florum, Enterococcus gallinarum, Enterococcus gilvus, Enterococcus haemoperoxidus, Enterococcus hawaiiensis, Enterococcus hermanniensis, Enterococcus hirae, Enterococcus hulanensis, Enterococcus innesii, Enterococcus italicus, Enterococcus lacertideformus, Enterococcus lactis, Enterococcus larvae, Enterococcus lemanii, Enterococcus malodoratus, Enterococcus massiliensis, Enterococcus mediterraneensis, Enterococcus montenegrensis, Enterococcus moraviensis, Enterococcus mundtii, Enterococcus nangangensis, Enterococcus olivae, Enterococcus pallens, Enterococcus pernyi, Enterococcus phoeniculicola, Enterococcus pingfangensis, Enterococcus plantarum, Enterococcus porcinus, Enterococcus pseudoavium, Enterococcus quebecensis, Enterococcus raffinosus, Enterococcus ratti, Enterococcus rattus, Enterococcus rivorum, Enterococcus rotai, Enterococcus saccharolyticus, Enterococcus saccharominimus, Enterococcus saigonensis, Enterococcus sanguinicola, Enterococcus seriolicida, Enterococcus silesiacus, Enterococcus solitarius, Enterococcus songbeiensis, Enterococcus spodopteracolus, Candidatus Enterococcus stercoravium, Candidatus Enterococcus stercoripullorum, Enterococcus sulfureus, Enterococcus termitis, Enterococcus thailandicus, Enterococcus timonensis, Enterococcus ureasiticus, Enterococcus ureilyticus, Enterococcus viikkiensis, Enterococcus villorum, Enterococcus wangshanyuanii, Enterococcus xiangfangensis, Enterococcus xinjiangensis, or any enterococcal bacterium described herein. In some embodiments, the efagin, or a portion thereof, specifically binds an Enterococcus faecalis EPA. In some embodiments, the receptor binding protein specifically binds an Enterococcus faecalis EPA.

[0123] In some embodiments, the efagin, or a portion thereof, is a group 1 efagin, or a portion thereof, and specifically targets an E. faecalis strain having the EPA type of, or having an EPA type that is similar to, E. faecalis strain RMC1 , E. faecalis strain FA2-2, E. faecalis strain JH2-2, E. faecalis strain DIV2445a, E. faecalis strain T6, E. faecalis strain E1 Sol, E. faecalis strain Com7, E. faecalis strain Com1 , E. faecalis strain 39-5, E. faecalis strain ATCC 19433, E. faecalis strain DIV0213b, E. faecalis strain E99, E. faecalis strain T21 , E. faecalis strain X98, or E. faecalis strain B653.

[0124] In some embodiments, the efagin, or a portion thereof, is a group 2 efagin, or a portion thereof, and specifically targets an E. faecalis strain having the EPA type of, or having an EPA type that is similar to, E. faecalis strain BFS026, E. faecalis strain BFS024, E. faecalis strain T14, E. faecalis strain T17, E. faecalis strain DIV2445a, E. faecalis strain RMC5, E. faecalis strain ATCC35038, E. faecalis strain UAA769, E. faecalis strain T1 , E. faecalis strain T13, E. faecalis strain Com6, E. faecalis strain OG1 RF, E. faecalis strain T9, E. faecalis strain BFS016A, E. faecalis strain DIV0205e, E. faecalis strain BFS107 or E. faecalis strain BFS020.

[0125] In some embodiments, the efagin, or a portion thereof, is a group 3 efagin, or a portion thereof, and specifically targets an E. faecalis strain having the EPA type of, or having an EPA type that is similar to, E. faecalis strain DIV2445a, E. faecalis strain T6, E. faecalis strain E1 Sol, E. faecalis strain Com7, E. faecalis strain T17, E. faecalis strain RM4679, E. faecalis strain BFS026 (AZ122), E. faecalis strain BFS024, E. faecalis strain T14, E. faecalis strain DIV1217, E. faecalis strain Fly2, E. faecalis strain V583, E. faecalis strain MMH594, E. faecalis strain X98, or E. faecalis strain B653.

[0126] Efagins within a group can specifically target enterococcal strains with similar EPA types. EPA operons possess a common set of core genes followed by a cluster of variable genes (see Palmer et al., mBio. 2012 Mar 1 ;3(1 ):e00318-11 ). EPA types are classified by commonalities and differences occurring in the variable genes present in a given strain. As is described in Palmer et al., variable genes in the EPA operon occur between orthologs of epaR (EF2177 in V583) and EF2165. These variable genes encode predicted glycosyltransferases and other proteins with likely roles in extracellular polysaccharide production.

[0127] In some embodiments, the efagin, or a portion thereof, specifically binds a non-enterococcus cell wall polysaccharide, such as a cell wall polysaccharide from any other bacterium possessing an accessible carbohydrate ligand to which an efagin binds. For example, the efagin, or a portion thereof, may specifically bind a cell wall polysaccharide, of a species or strains of the genus Streptococcus, Staphylococcus, Listeria, Clostridium and others. Examples of species include Streptococcus pyogenes, Staphylococcus aureus, Listeria monocytogenes, and Clostridium difficile.

[0128] II. Compositions

[0129] The invention features compositions that include the efagins, or a portion thereof, described herein. In some embodiments, the composition comprises group 1 and group 2 efagins, or a portion thereof. In some embodiments, the composition comprises group 1 and group 3 efagins, or a portion thereof. In some embodiments, the composition comprises group 2 and group 3 efagins, or a portion thereof. In some embodiments, the composition comprises group 1 , group 2, and group 3 efagins, or a portion thereof.

[0130] In some embodiments, the composition comprises a chimeric efagin, or a portion thereof.

[0131] In some embodiments, the composition further comprises an antibiotic. An antibiotic is an agent that has the capacity to inhibit the growth of and / or to kill, bacteria. Antibiotics typically used to treat enterococcal infections include one or more of penicillin, ampicillin, amoxicillin, piperacillin, tazobactam, vancomycin, ceftriaxone, imipenem, meropenem, gentamicin, streptomycin, linezolid, doxycycline, rifampicin, daptomycin, linezolid, tedizolid, tigecycline, eravacycline, and omadacycline. Other antibiotics useful for treating enterococcal infections include, but are not limited to, aminoglycosides (e.g., amikacin, gentamicin, kanamycin, neomycin, netilmicin, streptomycin, tobramycin, paromomycin and the like), ansamycins (e.g., geldanamycin, herbimycin and the like), carbacephem (e.g., loracarbef), carbapenems (e.g., ertapenem, doripenem, imipenem / cilastatin, meropenem and the like), cephalosporins (e.g., first generation (e.g., cefadroxil, cefazolin, cefalotin, cefalexin and the like), second generation (e.g., cefaclor, cefamandole, cefoxitin, cefprozil, cefuroxime and the like), third generation (e.g., cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone and the like), fourth generation (e.g., cefepime and the like) and fifth generation (e.g., ceftobiprole and the like), glycopeptides (e.g., teicoplanin, vancomycin and the like), macrolides (e.g., azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, troleandomycin, telithromycin, spectinomycin and the like), monobatams (e.g., aztreonam and the like), penicillins (e.g., amoxicillin, ampicillin, azlocillin, carbenicillin, cloxacillin, dicloxacillin, flucloxacillin, mezlocillin, meticillin, nafcillin, oxacillin, penicillin, piperacillin, ticacillin and the like), polypeptides (e.g., bacitracin, colistin, polymyxin B and the like), quinolones (e.g., ciprofloxacin, enoxacin, gatifloxacin, levofloxacin, lomefloxacin, moxifloxacin, norfloxacin, ofloxacin, trovafloxacin and the like), sulfonamides (e.g., mafenide, prontosil, sulfacetamide, sulfamethizole, sulfanilamide, sulfasalazine, sulfisoxazole, trimethoprim, trimethoprim-sulfamethoxazole and the like), tetracyclines (e.g., demeclocycline, doxycycline, minocycline, oxytetracycline, tetracycline and the like), and others (e.g., arsphenamine, chloramphenicol, clindamycin, lincomycin, ethambutol, fosfomycin, fusidic acid, furazolidone, isoniazid, linezolid, metronidazole, mupirocin, nitrofurantoin, platensimycin, pyrazinamide, quinupristin / dalfopristin, rifampin, and tinidazol). In some embodiments, the antibiotic is specific for enterococci.

[0132] In some embodiments, the composition is a pharmaceutical composition. In some embodiments, a pharmaceutical composition of the invention including an efagin, or a portion thereof, described herein may be used in combination with other agents (e.g., therapeutic biologies and / or small molecules (e.g., antibiotics)) or compositions in a therapy. In addition to a therapeutically effective amount of the efagin, or a portion thereof, the pharmaceutical composition may include one or more pharmaceutically acceptable carriers or excipients (Flemington’s Pharmaceutical Sciences 16th edition, Osol, A. Ed. (1980)), which can be formulated by methods known to those skilled in the art. In some embodiments, a pharmaceutical composition of the invention includes a nucleic acid molecule (DNA or RNA, e.g., mRNA) encoding a polypeptide described herein, or a vector containing such a nucleic acid molecule.

[0133] Acceptable carriers and excipients in the pharmaceutical compositions are nontoxic to recipients at the dosages and concentrations employed. Acceptable carriers and excipients may include buffers such as phosphate, citrate, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives (such as octadecyldimethylbenzyl ammonium chloride; hexamethonium chloride; benzalkonium chloride; benzethonium chloride; phenol, butyl or benzyl alcohol; alkyl parabens such as methyl or propyl paraben; catechol; resorcinol; cyclohexanol; 3-pentanol; and m-cresol); low molecular weight (less than about 10 residues) polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, or lysine; monosaccharides, disaccharides, and other carbohydrates including glucose, mannose, or dextrins; chelating agents such as EDTA; sugars such as sucrose, mannitol, trehalose or sorbitol; salt-forming counter-ions such as sodium; metal complexes (e.g. Zn-protein complexes); and / or non-ionic surfactants such as polyethylene glycol (PEG). Pharmaceutical compositions of the invention can be administered parenterally in the form of an injectable formulation. Pharmaceutical compositions for injection can be formulated using a sterile solution or any pharmaceutically acceptable liquid as a vehicle. Pharmaceutically acceptable vehicles include, but are not limited to, sterile water, physiological saline, and cell culture media (e.g., Dulbecco’s Modified Eagle Medium (DMEM), a-Modified Eagles Medium (a-MEM), F-12 medium). Formulation methods are known in the art, see e.g., Banga (ed.) Therapeutic Peptides and Proteins: Formulation, Processing and Delivery Systems (3rd ed.) Taylor & Francis Group, CRC Press (2015).

[0134] The pharmaceutical compositions may be prepared in microcapsules, such as hydroxylmethylcellulose or gelatin-microcapsule and poly-(methylmethacrylate) microcapsule. The pharmaceutical compositions of the invention may also be prepared in other drug delivery systems such as liposomes, albumin microspheres, microemulsions, nanoparticles, and nanocapsules. Such techniques are described in Remington: The Science and Practice of Pharmacy 22nd edition (2012). The pharmaceutical compositions may also be prepared as a sustained-release formulation. Suitable examples of sustained-release preparations include semipermeable matrices of solid hydrophobic polymers containing the polypeptides described herein. Examples of sustained release matrices include polyesters, hydrogels, polylactides, copolymers of L-glutamic acid and y ethyl-L- glutamate, non-degradable ethylene-vinyl acetate, degradable lactic acid-glycolic acid copolymers such as LUPRON DEPOTTM, and poly-D-(-)-3-hydroxybutyric acid. Some sustained-release formulations enable release of molecules over a few months, e.g., one to six months, while other formulations release pharmaceutical compositions of the invention for shorter time periods, e.g., days to weeks. The pharmaceutical composition may be formed in a unit dose form as needed.

[0135] The formulations to be used for in vivo administration are generally sterile. Sterility may be readily accomplished, e.g., by filtration through sterile filtration membranes.

[0136] III. Isolated bacterium including an efagin, or a portion thereof

[0137] The invention features an isolated bacterium comprising an efagin, or a portion thereof. Advantageously, an isolated bacterium comprising an efagin, or a portion thereof, may function as a live cell (e.g., probiotic) therapy. The isolated bacterium may also be genetically engineered to enhance or improve probiotic properties (e.g., an engineered probiotic).

[0138] In some embodiments, the bacterium is a member of the genus Enterococcus. Exemplary members of the genus Enterococcus include Enterococcus alcedinis, Enterococcus alishanensis, Enterococcus aquimarinus, Enterococcus asini, Candidatus Enterococcus avicola, Enterococcus avium, Enterococcus bovis, Enterococcus bulliens, Enterococcus burkinafasonensis, Enterococcus caccae, Enterococcus camelliae, Enterococcus canintestini, Enterococcus canis, Enterococcus casseliflavus, Enterococcus cecorum, Enterococcus cloacae, Enterococcus coli, Enterococcus columbae, Enterococcus crotali, Enterococcus devriesei, Enterococcus diestrammenae, Enterococcus dispar, Enterococcus dongliensis, Enterococcus durans, Enterococcus entomosocium, Enterococcus eurekensis, Enterococcus faecalis, Enterococcus faecium, Enterococcus flavescens, Enterococcus florum, Enterococcus gallinarum, Enterococcus gilvus, Enterococcus haemoperoxidus, Enterococcus hawaiiensis, Enterococcus hermanniensis, Enterococcus hirae, Enterococcus hulanensis, Enterococcus innesii, Enterococcus italicus, Enterococcus lacertideformus, Enterococcus lactis, Enterococcus larvae, Enterococcus lemanii, Enterococcus malodoratus, Enterococcus massiliensis, Enterococcus mediterraneensis, Enterococcus montenegrensis, Enterococcus moraviensis, Enterococcus mundtii, Enterococcus nangangensis, Enterococcus olivae, Enterococcus pallens, Enterococcus pernyi, Enterococcus phoeniculicola, Enterococcus pingfangensis, Enterococcus plantarum, Enterococcus porcinus, Enterococcus pseudoavium, Enterococcus quebecensis, Enterococcus raffinosus, Enterococcus ratti, Enterococcus rattus, Enterococcus rivorum, Enterococcus rotai, Enterococcus saccharolyticus, Enterococcus saccharominimus, Enterococcus saigonensis, Enterococcus sanguinicola, Enterococcus seriolicida, Enterococcus silesiacus, Enterococcus solitarius, Enterococcus songbeiensis, Enterococcus spodopteracolus, Candidatus Enterococcus stercoravium, Candidatus Enterococcus stercoripullorum, Enterococcus sulfureus, Enterococcus termitis, Enterococcus thailandicus, Enterococcus timonensis, Enterococcus ureasiticus, Enterococcus ureilyticus, Enterococcus viikkiensis, Enterococcus villorum, Enterococcus wangshanyuanii, Enterococcus xiangfangensis, or Enterococcus xinjiangensis.

[0139] In some embodiments, the bacterium is commensal. For example, the bacterium may be any commensal species that naturally occurs in the healthy human microbiome. In some embodiments, the bacterium is a commensal member of the genus Enterococcus. Commensal members of the genus Enterococcus are described in Lebreton et al. Enterococcus Diversity, Origins in Nature, and Gut Colonization. 2014 Feb 2. In: Gilmore et al., editors. Enterococci: From Commensals to Leading Causes of Drug Resistant Infection. Boston: Massachusetts Eye and Ear Infirmary; 2014. Commensal bacteria can advantageously colonize areas of the human body including the skin, the upper respiratory tract, the oral cavity, and the vagina. Such bacteria may be administered to treat a bacterial infection.

[0140] In some embodiments, the bacterium is an opportunistic pathogen.

[0141] IV. Methods of treatment

[0142] The efagins, or a portion thereof, described herein specifically bind to target bacterial strains and kill bacteria, lyse bacteria, or inhibit growth of bacteria. Consequently, the efagins, or a portion thereof, described herein can be used to treat bacterial infections. Accordingly, the disclosure provides for a method of treating a subject having a bacterial infection, comprising administering an effective amount of an efagin, or a portion thereof, described herein, a composition described herein, or an isolated bacterium including an efagin, or a portion thereof, described herein.

[0143] In some embodiments, the bacterial infection is a wound infection, a urinary tract infection, bacteremia, or infective endocarditis.

[0144] In some embodiments, the bacterial infection is an enterococcal infection, as described in

[0145] Lebreton et al. infra. For example, E. faecalis and E. faecium cause a variety of infections, including endocarditis, urinary tract infections, prostatitis, intra-abdominal infection, cellulitis, skin, soft tissue, and wound infections, as well as concurrent bacteremia. In some embodiments, the enterococcal infection is an Enterococcus faecalis infection.

[0146] In some embodiments, the subject has a nosocomial infection, which a disease or infection occurring or originating in a hospital or other healthcare facility. Nosocomial infections are typically infections that are contracted in a hospital or other healthcare facility, and can be caused by infectious agents, e.g., bacteria that are resistant to antibiotics. In certain aspects, a nosocomial infection is not present or incubating prior to the subject being admitted to the hospital or healthcare facility, which is acquired or contracted after the subject's admittance to the hospital or healthcare facility.

[0147] For example, enterococcal species are prominent nosocomial pathogens because they are normal flora in the human gastrointestinal tract. Antimicrobial resistance allows for their survival in an environment with heavy antimicrobial usage, and they contaminate the hospital environment and survive for prolonged periods of time.

[0148] Subjects that may be treated according to the methods of the disclosure are animals, e.g., preferably a mammal, such as a human. Other mammals include, but are not limited to, farm animals (e.g., cows, pigs, horses, cattle, sheep, etc.), pets (e.g., cats and dogs), and primates. Other farm animals include poultry such as chickens, turkeys, and birds, as well as other commercially produced animals including fish, shrimp, and other types of animals potentially vulnerable to bacterial infection.

[0149] An efagin, or a portion thereof, or a composition described herein may be administered to the subject by intravenous, parenteral, intramuscular, intranasal, subcutaneous, percutaneous, topical, intraperitoneal, inhalation, systemic, and oral administration.

[0150] An isolated bacterium including an efagin, or a portion thereof, described herein may be administered to the subject by oral, topical, inhalation, nasal, vaginal, or rectal administration.

[0151] In some embodiments, an efagin, or a portion thereof, a composition, or an isolated bacterium including an efagin, or a portion thereof, described herein is administered to the skin, the upper respiratory tract, the oral cavity, or the vagina.

[0152] The efagin, or a portion thereof, the composition, or the isolated bacterium may be administered in an effective amount, which is an amount of the efagin, or a portion thereof, the composition, or the isolated bacterium sufficient to effect the recited result, e.g., to kill a bacterium, lyse a bacterium, inhibit growth of a bacterium, or treat a subject having a bacterial infection.

[0153] Dosages of an efagin, or a portion thereof, or of a composition described herein may range from sub-milligram quantities for treating localized infections, to oral or rectal administration of milligram to gram quantities to depopulate select microbes from the digestive tract. Dosages of isolated bacteria including an efagin, or a portion thereof, described herein may range, for example, from about 105cells to about 1015cells.

[0154] V. Nucleic acids, vectors, and host cells

[0155] The efagins, or a portion thereof, described herein can be produced from a host cell. A host cell refers to a vehicle that includes the necessary cellular components, e.g., organelles, needed to express the polypeptides described herein from their corresponding nucleic acid or set of nucleic acids. In some embodiments, the nucleic acid or set of nucleic acids encoding the efagin, or a portion thereof, is endogenous to the host cell. In some embodiments, the nucleic acid or set of nucleic acids encoding the efagin, or a portion thereof, is heterologous to the host cell.

[0156] In some embodiments, the nucleic acid or set of nucleic acids may be included in an expression vector or set of expression vectors that can be introduced into the host cell by conventional techniques known in the art (e.g., transformation, transfection, electroporation, calcium phosphate precipitation, direct microinjection, infection, or the like). The choice of nucleic acid vectors depends in part on the host cells to be used. Generally, preferred host cells are of either prokaryotic (e.g., bacterial) or eukaryotic (e.g., mammalian) origin.

[0157] The efagins, or a portion thereof, described herein are isolated, which is a preparation of the efagin, or a portion thereof, devoid of at least some of the other components that may also be present where the efagin, or a portion thereof, or a similar substance naturally occurs or is initially prepared from. Thus, for example, an isolated efagin, or a portion thereof, may be prepared by using a purification technique to enrich it from a source mixture. Accordingly, in some embodiments, the disclosure provides for methods of producing an efagin, or a portion thereof, comprising (a) culturing a host cell or set of host cells comprising a nucleic acid molecule or set of nucleic acid molecules or expression vector or set of expression vectors encoding an efagin, or a portion thereof, under conditions where the efagin, or a portion thereof, is expressed; and (b) isolating the efagin, or a portion thereof, expressed in (a) from the host cell culture, thereby producing the efagin, or a portion thereof.

[0158] 1. Nucleic acids

[0159] The efagins, or a portion thereof, of the disclosure may be provided as a nucleic acid molecule or set of nucleic acid molecules encoding the efagin.

[0160] The efagin, or a portion thereof, may be encoded by a nucleic acid molecule or set of nucleic acid molecules having a nucleotide sequence that is at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100%) identical to any one of SEQ ID NOs: 91 -93. In some embodiments, the efagin, or a portion thereof, may be encoded by a nucleic acid molecule or set of nucleic acid molecules having a nucleotide sequence that is at least 95% identical to any one of SEQ ID NOs: 91 -93.

[0161] An exemplary nucleic acid encoding a group 1 efagin is provided in SEQ ID NO: 91 , below, and is available under GenBank assembly no. GCA 000172575.2.

[0162] Group 1 (OG1 RF_efagin GCA_000172575.2_ASM17257v2_efagin) (SEQ ID NO: 91 ): TTAAGTTAACCCGGTATTTAATACCATCACAGAATAATCCCTAAATTCAAAAATTCGTCCTTTATAAAG ATAGGTGGCGCCATACTTTTGTCGATAGTACGTAATCACATTTTTTAATGTTTCAACATCGATTTCCAA AAATTCAGCACAAGAATAATGGTTACTTAAGCCAGCTTCAGAGCAACGAATTAAATCGTCCAATGTAA CTAATTGCTCTAAGGCCACATTGCGGGCTTTCAACTCTTGCTTCCGATTTTCTGTACAATTTTGATTTA AAATAGTGCCAACGGATGTTTGATAATGCCCGTATTCCTCGGCTAAAATGTTCTTCTTTTGACGGGTA CTCAATGTTTTTTCAATATAAATTTTGCCATTTCGATACAACCCATAACAACCTGTTTGATTGTATAAAT CAATTTCTAAAACCGTCACATCCTTTTGAATGGAACTGACGAGTTTTTCATAATCATTCACTTTACCAT TCACCTGCCATTACTTAATAAAAAGAGTAGATGGGAAGTTAATCCTCTTTTTTGTCAGAAGAAATCGAT TGTTGATATTTGGCATCAATTTCATCCAAGTAATCGTGAATTTTCTCGATTTCTTCTTTTGAAAATATCT

[0163] TTTCTGGATCCCCAGCGTGGGCAGCCAGTGTCATTTCGTCACGGCGAGGGAATGAGAGAATTTCTG

[0164] CTTGCGTTTGTTGTTCGTGTAATTGTTGTTCGGCAAATTGATAAACGATGGCTTGTCGTTGGGGCTCT

[0165] AATTTTTTATAAATGCGGTCGATAGTTGAGGCATTTTTTTCCATGGGAACCTCTTGTCCCATCAACCA

[0166] GGCCTCATTAATGTCTAATGCATCTGCAATGCGATAAACTTTGTCTTGTTTTGCTTCGTAACGACCAG

[0167] CTAGCCAATCGCTGATCGAAGATTTACCGATGCCAGTTTTTTTCGCTAAATCACTAGGTTTGATGTTC

[0168] TTGGCTGTTAAAGCCTCTTTTAAACGGACAGCAAAAATGTTCATATTTCCGAACCTCCTATTGCCTCA

[0169] AAGTATACTATACCTGTTTAGAGAAGGCAAGTAACAATCGCAATTAAATTTAGCAGTTCATAAAACCG

[0170] AACTTATGAGTTGACAGTGAAAATATTTGATGCTAAGATAGAGCCATTCAAATGTTCGGTTAACTGAA

[0171] CTATTTTTTTGAACATTTTGTTCGGAAAACCGTATTTTTGTAGAAAGGGTGAAGAAATGAGTCGGAATT

[0172] ATAAAAAAACGTTGTCGGATATGTTACTTTTAGCAATTATTTTATTAATAAGCAGTGTCTCAATAAAAAT

[0173] TGGGGCCATCGTGATTGGTATGATTGGCCTCATGGAATTACTAACAGAGTAACAATAATTTAGTCAAA

[0174] AGGAGAGAAGCGGATTGGCAGAACGACGCATGTTTGCAAAAACGATTATTGATAGTGATGCGTTTTT

[0175] AGATATGCCCTTATCAAGTCAGGCCCTGTATTTTCATTTAGCGATGCGTGCCGATGATGATGGATTTA

[0176] TCAATAATCCCAAAAAATTGCAGCGAATGGTCGGTTGTGGGGAAGACGATCTAAAATTGCTAATGGTT

[0177] AAAAAATTTATTCTAGTATTTGAAAGCGGTGTGATCGTTATCAAACATTGGAAAATTCATAATTATATTC

[0178] GCAGTGATCGTTACAAAGCAACCTTGTATCAAGAAGAGAAAAATCAGATTGTTGAAAAAAATAGCAAA

[0179] GCTTATACGTTTAAAGCAGAATCGTCTGTCGGTGGTCAACCAGCTGACTACCAACGGTTACCACAGG

[0180] AAAGCATAGTCCAGTCTAAGTTAGGTCAGAGTCAAGGTAGTAGTTCAGAAAACGATTGTTTAAAGATG

[0181] ATTTATTATTTTTATGAGGAAAACGGCTTTGGTACACTGGCCTCAAAAACAAGCCAAGATTTTAAGTAT

[0182] TGGTTGCAAGATTTTATACAAAAAGGGGCTAGCCAAGAGGAAGCATGCCAATTAATCTTGCATGCTTT

[0183] AGGAATTGCCGTCGATCGAAATAAACGGAATTACGGCTATGTAAATGCTATTTTGAAAAGTTGGGAG

[0184] CAACAAAATTATTTATCCGTACATGAAGTTCTGGTAAATGATAAAAAACAAGTGTTGGAGCATGCGCC

[0185] GCAAATGACAGAAGAATATCAAGAGTTAGGTTTTTAAAGAAAGGAGGAAATCAGTATGCATGCGACA

[0186] GATCAAACTTTTCAAATACTATTGAGTCAATTGTTAGAAAAAGTTGAAGACCGTTGTCCTGAATGTGG

[0187] CAGTGAACAATATGTTTGGCAACAAAAAAATAAAGATGGCACAGAACGTTGTGCCCCAACTTGTTGGT

[0188] CGTGTGGGTATAAAATGCTAAAAAAACATGAACAACAAGCCAATCAACAACGTTCTCAAGAGAGTTTT

[0189] ATGGCACGTACACAAAAATTTTTTCATCAAGGGTCCTTAATTGCTGATGATGCGCTACGGCAATGTCG

[0190] TTTAACCAATTACCAAACCACTGAATTAGAAACAAGACAAGCAAAAGAACGGGCCTTAGCAGCAGTTT

[0191] CAGCGATTGTTGAAGGAAAGCCAATCCACGTTATTTTTTCAGGGAAACCTGGTGTCGGTAAAAGTCA

[0192] TTTGGCTATCAGTATTTTAGTTGAAGTCTTAGAACGCTCTGCATATCAAAAGTATTGTTTATTTGTCAG

[0193] CTACTCTGAGTTATTAGAAAAACTAAAAATGTCCATGAATGAATCGGCCAAAAGCCAAGCAAAGGCTC

[0194] AAGCGTATATTACTAGAATGAAAAAAGCAGACGTTTTGGTTTTAGATGATTTAGGTGCTGAATTAGGA

[0195] ATTAAAAATAAAGTTAGTACGGATTTTAATAATGACATCTTAAACCGAATTTTAGAAGCTAGACAGAAT

[0196] AAAGCAACTATTTTTACTACTAATTTTTCTGGAAAACAACTGGTGGAGGCCTATGGAACACGCATTATT

[0197] TCTCGTCTAATGAAGCACGCCAGTGGCTATGTTTTCCAATATAAAGACACAACAGACAAACGAATGAG

[0198] GAGTGTGAAATAAATATGTTAACAATTATTATTGGGTTTATCTTTTGGACAATGACACTAATGTTAGGT

[0199] TATCTAATTGGTGAAAGAGAAGGCCGTAAACATGAGTAATTTAACAAAACGTAAAAAAGATTTATTTGA

[0200] AATGAAAAGCGTTGTATTTAAAGATATTTCAAAGCAACAAAGCGAAAAAGCACAAAAAAGAAAACGAC

[0201] TCTTACAACTAATGAATCAATATCCCGATTGGGCAAGTCAAAAAAATAAACTTATTATGCAGGAAATTC AAGAATTAGGACAAGCAATCGGTAATTGGTCGATGGATCAATCAAGACCCATCCAATCCATCAAGGC

[0202] CGCATCGTTTACAAAAAGCGAGTATCTCTATTTAATTTGGCTCGGTTATTCAGATGAAGCGATTCGTC

[0203] ACGGCTTAGACATGTCGAAAGAGTGTTATTTTATTTATCGATTAACACTTTTAAATAATAAAAGTAAAG

[0204] GAGATTAACCAATGCGTACGTCAACATTTAATTATATCAAAGATATTTTAGCAGACTTTTATAAAACAG

[0205] AAGAGTATATCCGTCAACGGGAAGAAGAATTACGGCACCCTTATCAAGAAGCAGATTTAAATGCTGG

[0206] TATTAGAGGACAAGGACTTCACTCTGTAGTGACCGAACGAATGGCGATTACGATAGCTATGGATCGT

[0207] CGTCTGTGGAACTTAGAGAGAAATCGAGACATTATCAAAAATTGTTTAGCCGAAGCGGATGAACAAA

[0208] CGCGCGTGATTATTGAAGAACTATATATGAAAAAACGGCCCTCTTTAACATTAATTGGACTTGCCCAG

[0209] CAATTATTTATTAGTAAAAGCCAAGCCTATAAATTAAGAAATCATTTCTTTGAAGCGGTGGCGGATGAA

[0210] CTAGGGATGTAAACATGGAAAAAGCGTGGAATTTTTTCAGGTGTCAACATGGTAAATTAATAGTGTCG

[0211] AAAGAGATAGATAAACGTGAGGCAACCAAAAAAATGAAGACACGGAATTCTATGATTTTGACTGCTTT

[0212] CTTGTGTCAGCTATGAAGGAGCAGAAAATGCCGGCTACTTTCAAGATCCTTCATTTTGACTAGAAGAG

[0213] AGCCAATTTGTTAACCAATCCTGAATTTTTTGAATGGAAAGGTGGCGCTAAAAATGAATGAAGCGGAA

[0214] CAAGAGTTATATGAAGCCCTTGTTACAATCTGCCAGACGTCAGGATTTTTGTTGCTAGAGGAACTGCC

[0215] GACAGATTTACCAGATCAGCCATTTGTTTATTTAGGTGATAGTAAAGAATTACCTAAGCCAACTAAATC

[0216] AGCTATTTTGGGAGAAATTGAATTAATAATGCATGTTTATGGTGCGTTATCTGAACGACAACAAATTTC

[0217] TACAATTAAAGGAACGATTTTACGGCAGGCAACCAGTAACTTAAAACGAACGGCTCATTTTAATTGGG

[0218] GTATCAAACATCAAGAAGTCAAAGCACAAATGGTAAAAGATACCAAACAAATGAAAAAAACAATTTGG

[0219] CATGCTGTATTACCATTACACATGCAATTTTACTAGGAGGAATTATCAATGGGAGAAGTTATGCAAGG

[0220] AAAAGACCGTATTTTATTAGTTCGTCGCTTGGATGAAGCAGCGACAAAGAAAGCAATGAAACCCTTAT

[0221] TTCAAATTGAACATGAATGGGAATTCTCACGTGAATCGAGCGGTACGCAAACAAAAGATGGCGTTGC

[0222] GAATGCTGTATCTGGTTTAGAAGTTACGTTATCGTTAAGCGGTTTAGCCTCTCGAGATGATGAAAATT

[0223] TATACATGAAAGACGCAGTCGAAGATGGCATCTTAATGGAATTTTGGGATGTTGATTTAAAAGGTGAA

[0224] AAAAATGCGGAAGGTAAATATCCAGCAATTTATGCCCAAGGTTATGTAAATTCATGGAGTTTACCAGC

[0225] CAATGTAGAAGAATTAGTAGAAATCGAAACAGAAGCCTCTATTAATGGCAAGCCACAAGATGGCTTTG

[0226] CAACAATAGAAGCAGATATTATTGCGGAAGCACAATATGCGTTCCAAGATACCGTTCCAGAGAAAGC

[0227] ACCACAACCTGGCGAATAATCAAAAAGTGTTGAATTTTAGGAGGATAAAAAATGAATTTAGAGATTAA

[0228] CGGAAAAACAATTGAAGTGAAATTTACGATTGGCGCGATTCGCGAATTAGATAAACGTTACCAAATTG

[0229] AAAATGGCGCTGCCAAATTCGGCATGGGCATCAGTTCAGCAATGATTTATTTACGCCAATACAATCCA

[0230] GTAATCTTAGTTGACATCATGGAAGCTTTACAAAGTGGGCAATTAAAACTAGGTAAGTCGGAAATTGA

[0231] AGCATGGTTAATGACCCAAGATGTCAAAAAACTTTCAGATGATTTGCTTAAAGAAATGGGAAAGCAAC

[0232] CTCTTACAAAACCAATGATCGATCAGTTCAGCAAAGAAGCGAAGAAAGCAGAAGCGCAAGCGACCAA

[0233] CTAATTAAAACGAGCGATGACGTGTATCACGACATCGCTCTTTCTGCTTTTCGCTACTTAGGCTGTCA

[0234] TTCATTTGAAGAAGTGGATCGGATGACCATGTCTGAATTTGAATTACGAATGATTGCTTTTAATTTAGC

[0235] AGAAGTAGATGAAGAGCGGAAAAGGCACGAGCTTGCCTACTTAAATGTTAAAGCGCAAGCGACGAA

[0236] CAAAAAAGGAAAACCCGTTTTTGAAAGCTTTAAAAGTTTTTATGATTATGAAAAACGAGTTGCTGAAGT

[0237] TCTGGCAGCTAACCAGCCACAACGAACGAAATTAAATGAGCGGAAAAAAACGCAACTTGCCACTGTG

[0238] GCAGAGCGTCTACGCCGCTATCGAGAAGGGAGGAGAGTAGATGGAGAATGACAAAGAAAAAACGCC

[0239] GTTGTCGGAGGCAAAGAAAAGCCTTGCAGGCGTCCAACAAGCATTAAAAAGTATGAGCGGTGAGTAT

[0240] GCCTTATTAAGTGGATATTTAGGGAAAATTAGTGCGGGTGTCAATCAGTCAGCCACGGTCATGAACA CATTTAAAACCGTCATGCAACAATCTGGAGAAACAGTGAAAAAAACAGGAGACGAAACAGCAAAGGC

[0241] AGCAGATCAAATGAACACAGAGTTAACAGATTCTGCTGAACAAGCCGGTGAAGCAGCTGAAAAAGCG

[0242] GGGAAAGAAACCTCTGATGGCTTTACTAATGCACAAAATAATATGCTGAGCTTTGGGACGGCCATGA

[0243] CTAGTGCCGTTTCCTTACCTATGCTGAACGTTTTAAAAACAGCTATGGGCGTCGGTGCTGGGGTCAG

[0244] TGGCGAATTTCAAGGAATGCAAGGACTGATTATGGCCAGTGCAGGAGGGATTTCTGATTCATTGCAA

[0245] GGCGAGTTGCAAGGGGCATTGACTCAGATGAATCAATCATTTGAAGCGGCGGCACAAGTGATTCAAA

[0246] GCGTGATGGCTCCAGGAATGGAAATTTTGGTTCAAGTGGTTATCACAGTCGTCAAAGGCATTACAGC

[0247] TTTGGTGAATTTATTTATCAAATTACCAAAACCCGTCCAAGTTTTTATTGTTGCCATTATGGGCATTTTA

[0248] GCCGCCATTGGGCCCATGTTGATTATGGTAACGATGGCTCAGCAAAAATTTCAACAGTTTAGTGCTG

[0249] GTTTGGCTCTTGTACAAGGAAACATTGGGAAGTTAGGTGGTGGCTTATCAAAACTAAGTGCTAGTTTT

[0250] AGTGCCTTAGGTGGAGGACCATTAATTTTAATTGTAGCAGCCGTTTTAGCAGCGGTAGCAGCGTTTA

[0251] TTTATTTCTATAAAACCAATGAAACATTTAGAAATAGTATCAATAGCTTAGCTAGTGCCATTCAAGGAG

[0252] CTGTTTCAGCGGCGTTTGGCAAATTGGTAGGATTGCTACAACAGATCCAGCCGGCCTTTCAGCAAGT

[0253] AATGGCAGTTTTTAAACAATTTTTTGCAGTAGGCTTAGAGAAAATGGCGACTATTTTTTCAACAATTGG

[0254] TCGTGTGCTAGCAGGCGTTTTTGCCAGCGGCTTGCAATTAGGTAGTAACTTATTAGGGCAATTCGGT

[0255] GGCACCTTTGACAAAGCTGGTTTAGCGGTTGGTCTTTTGGTAAAAGTTCTGACAAAGGTTGCACTGG

[0256] CTGCATTAGGAATTTCTGGGCCGTTTGGTCTAATTATTTCCTTGATTGTTTCATTCGTGACGGCCTGG

[0257] ATGAAAACCGGTGATTTGAGTGCGGGTGGTATTACCCAAGTCTTTGATAATTTAGGTAACACGATTAC

[0258] ATCTGTTACAACAATGCTTGCGGCTAATCTACCGAAAGTTATACAACTTTTTACAACAGTCTTAACCAG

[0259] TATTCTCGGGAAAATAACAGAAGCTATTCCAAGCATCGTCACCGCGTTATCTAGTTTAATTACGTTAAT

[0260] TGTTGGTGCGATCGTTGCCAATTTGCCAGTCTTAATTGAAGCGTCAACACAAATTATTACTACGTTGA

[0261] TTCAGGGGATTACAACAGTCTTACCAATGTTGATAGAAGTTGGTTTGAGCTTATTAATGACTTTAGTTA

[0262] ATGCGATTGTCACCGCCTTGCCAACAATTACAACTGCAGCGATTAATATCATCACTACATTAGTGACA

[0263] GCTTTTGTCACAGCGTTACCAATGCTAGTTACAGCAGGTGTTTCAATTATCACAGCCTTAGTCAATGC

[0264] ATTTGTTACGATGTTACCGTTGATTTTGACTGCTGGTTTACAAATTTTGATGGCATTAATCACTGGGAT

[0265] TATGACGATTTTACCTCAGTTAATTCAATCAGCGCTGACGATTATTCTAGCGTTAGTGACAGCGTTGA

[0266] TAGGTGCCTTACCACAGATTATCAGCGCAGGTGTCAAATTGTTAATGGCGTTAATTCAAGGAATTATT

[0267] TCGATTTTACCAACCTTAGTTGCGGCAGCTATTACCTTAATTTTGACATTGGTAAATGCCTTAATTGGT

[0268] GCCTTGCCACAAATCATCAGCGCAGGCGTCAAATTGCTAATGGCTTTGATCCAAGGGATTATTTCAAT

[0269] TTTACCGCAACTGGTTACTGCAGCAATTACGCTAATTACCGCTTTAATGGGTGCGTTAATCAATGCGT

[0270] TGCCACAGTTGTTAAGTGCTGGGATTCAACTGATTCAAGCCTTAATTAATGGTGTACTCAGTCTATTG

[0271] GGTGCCTTGCTGTCCGCAGCAGGAACATTAATCTCACAAATGATCACGAAGATTGGTTCTTATTTTGG

[0272] TCAACTGTTAGCTTCGGGCGGACAGTTAGTTGAAAATATCAAAAATGGGGTTACCAATGCAGCCGAT

[0273] CAGGTAAAAACTGCCATTGGTTCTGTAATTGAAGGTGCTTGGCAAGCAATCCAAGGTTGGTTTTCAAA

[0274] ATTCACCGATGCCGGTGCGAATATTGTCGGTATGATTGCTGATGGAATTACAGGCGCAATTGGAAAA

[0275] GCCAAAGAAGCAATCGATGGGGTCGTCAGTAAAATTCGTAACTTTTTACCATTTTCACCAGCAAAAGA

[0276] AGGTCCCTTATCTGATTTGCATAAATTGAATTTCGGCGGCACGATTGCCACGGGGATTTATGCAGGC

[0277] GAAACAGCCGTTAGTAGAGCAATGGCTTCTATTTTAGATTTACCGCTGTTAAATGATTTTGCCTTGGA

[0278] CTTAGCTGGTCGAGGAAACTTCACGGCAACGATTGACCATCGTTTAGAAAATGATGCATACAATCGA

[0279] CCATTATTTGTGACAGTAGAGTCAACGTTAGATGGAAAAGTTGTCGCAGCAACTACGGCGCCTTATTT AGCAACAGAGTTACAACGACAACAAGTGAAACAAAATAACCGCTTAGGAAGGAGAGGATAACATGTA

[0280] TAAATTTGTTGATACCAATCAAGCAACTCATTCAACGCCTCTTCCTTCAGAAGCGTTGAATTTTAACGG

[0281] CCAATTTTTAGAAAAAGTCATCCCTGGCTATCAAACATTATCAGTTTCAGGACGAGAATTAGTTCCAA

[0282] GCGAAATTGAAAGCTATCAATTAGGGATTCGTGATGGTAAACGTCACGTTTATGCGCGAATTCCAGA

[0283] ACGAGAATTAACAGTCAAATATCGCCTTTCAGCTGTGAATAATGAAGCATTTCGAGATGCATTTAATC

[0284] ATTTAAACGTTGCTTTGTTTACGGAAAAAGACGTTTCTATTTGGTTTAACGATGAACCGGAAATGCTGT

[0285] GGTTTGGCAGTAAGTCTTCAGTGAGTGATGTACCCGAAGGTGTTAACCAAGTAACAGGCACCTTTAC

[0286] TTTATTGCTTTCTGATCCGTATAAATACACACGGAGTGATGCGACTAGTGTGATGTGGGGTTCGCCAA

[0287] CCATTACATTTCAAGCGAATTACTTAATGGGGAATACAGGCTCAGGTGCAGTTGATTTTCCAATTTTA

[0288] ATTGAAGGCGGGGCTTATTGGGGATCAACCATGATTACCTTTCAAAATCGGGCTTACACGATGGGGG

[0289] ATTTAGGCAAAGAAGTTCGGCCAATTGAAATTTATCCTACGGTTGAAGGATTAAAAGTCAAACCGACC

[0290] ATTATTTTAACAGGAACCGGACGTGGTGTTTGGATTAAAACACGGAACGATACAATTAACTTAGGAGA

[0291] CTTTGATCGTTCGGAAATTATTATCGATACTGAAAATTTTTATCTGACAAAAAATGGTGCACCGATGAT

[0292] TCGACCAATGAACGATTTTCATCTATATCCCAATGAACCGCTGTATATTCAAGCCAAAGATAGCGACT

[0293] TCCGCTTGACGATTCGCTATCCTAACCGATTTGTGTAGGAGGGTGATTAAATGTTAATGGCGCTGGA

[0294] TTTGAAAAGAACATATACGGCAATCTTGGATAATGCCTATCAAGTCAGTTATGAAAAAATAGAGAACA

[0295] AAATTGGGAGTTTAGATTTTACCATGCCACTAGATGATCCTAAAAATGAATTTATTGAAGAAATGCAAT

[0296] GGGTGGAACTGACCGACAATGAGAATGAATATATTGGTTTATATCGCGTGATGCCAACCACAATTAA

[0297] GAAAGATGCGAACAATAATCAAATTCACTACTCTGCCACAGAAGCATTATGTACCTTAGGCGATACTG

[0298] TCCTATTTGGTTGTCACGAAATTAAAAACAAAACAACGAAAGAGGCCATTCAATTTCTATTAAATAAAC

[0299] AAAAAACAAAGCATTGGGTCCTAAAAAAATGTGATTTTTCAAGGAAATTAACCTATAAATGGGAGAAT

[0300] GAAAATGGGCTAGTCGAGCCTTTATTTAGCATCCCAGCCGATTTCGAAGAGGAATATCTTTGGCAAT

[0301] GGAATACAGAGGTCTATCCTTTTGAACTTTCATTAGTCAAACCGCCAACAGAACCAGTTGCGCGAATT

[0302] CAAGAAGGTTACAACATGCAAGGATTTGAAATAGAACGTAATCCCAAGATGCTAATCAATCGGATTTA

[0303] TCCATTAGGTTCAGGCGAAGGTGTTAACAAAGTCAATATTCGCTCGGTCAATCAAGGGGTTCCGTAT

[0304] TTAGAGAACAAGGCCGCAATTGACCGCTATGGTTTATTAGAGTCAATTTGGGTGGAACAGCGTTTTTC

[0305] TGATCCCAAGGCATTAAAGGAAAATGCTTTGCGAATGTTAGAAGAATGGACCAAACCACAAGTTTCTT

[0306] GGGTAGTGACTGCAGCTGATTTAATTAAATTAACAGATCAACCTTTGGCAATCGATCGTTTGCGGTTG

[0307] GGTACGGTTATCATGATTAATACGAATGAGTTTGGGAGTGTCAACCTTCGTATCAAAAAAGAAAGTAA

[0308] AAAAGATGTCTTTGGTGCCCCCCAAGACATTCAGCTAGAGTTAGGAAACCTGCAAGAAACAATTCATA

[0309] GTACCATGACAGCTTTCAGTCGGAAACAAGAGATTAGCGAAACTTACGCACAAGGGGCGACGACAC

[0310] TTTTAAATCGTTCAATACAAAGAGAACTTAGCAAGACACAGCCAGTGGAGCTGAATTTATACTTTGAC

[0311] GAGGACATTCTTTATGTAAACACCGCAGAATTAACGTTCAAGGCAACTGCTAAAGGACCTTCGCATTC

[0312] TGTAACGAATATTGATTTGGTAGTGGATGGCAAAAAATTACCCCAACTATCATTGCAACAACAACGGC

[0313] TAAACATTTTGAATTATTTACGAAAAACAACAGATGGAAAAATCGAACGCGGCAATCACACGCTTCAA

[0314] TTTTTCTCTCATCAGCCACTATGGTTGGATGCTTCGGTCATCTGTCGTGTGTATATTCAATCCCAATTG

[0315] GGTGGCCAGTTTTAATAAAATAATGAAAACTAGAGGAGTGTGACGAAATGTCAGTAGAACATATTGAA

[0316] GAATTAGATACCCTGAATCAAGGTCGCCTTAAAATCAATGCAATCTTGGATCAGTCGAATGCATCAGC

[0317] TGAGAAAGTTGATGCTTACCAAGTCCAGTTAACGAATGGAATTTCTGAAGCGAAAAACATCGCAGAT

[0318] GAAGCTGGCAAAGAAGCCGTACAAATTGCCACCGATGCAGGCAATCAAGCAAATGAAACAGCCAAC CAAGCGATGAACAATGCCAAAACAGCAATCACGATTGCAGGAAATGCAGTTTCAACGGCAAATAATA ATAAACAAGAATTTGATACTTTGCGAAATGATTTCGATCAATTAGTAGCAGAAGCAGGTGATAGTAAT CCAGAAATTGTCCAAGCACGCACAGATACACAAGGCATCAAACAAGCTACTCTGGCGAATCGTCTTC AAATTGATTTGAATGACCGTATGACAAAAGCAGACGGTATTTCTTTATTGGCTAAGCCAACTACTGTC AAAATGAAGTTAGACTTTAACGGTAAAACGGCCGGCAATACAGCCACCAATGCAAACAGTTATTACAC TGATTTTACGGCAAAAATTCTTAAGAAGCCAACAGACGTTTGGGAGGAAGTTTCCCAAGCGGACTAC AATAAAATGGCCAGCCGTGATGATGAGGGTGTGAAAACAGGTTCCACCCAAAGCGGTGTGATTCCG CAACAATTAGCGGCCTTCAATCTCGTTGAAGCCGCTAAAAAATTAATTCCACAAATGTTTGAAACATT CACAACTGACGAGGCGGTGGCATTTATTCGCCAGAACGTTCAATTTTTTACGATTAATCAACGTGTGA AAGCCGCTGCGCCCAATAATCAAACGATTAAAATCGCTACGTATTTACCAACTACGGATAATTGGGTA ACTCAAATCCAAGAATCAGCAAAAGAGTTTGGCGATTTTTCAATTCAAATCAATGATCAGAATTTTATC ACAGATGAAGGTTTCATTTATTTAATGAGCTATACAGATTCATCGAATGGGGTAACGCCAGCTAGCTT AGAAGTTGATTACGTGGGGCTTCATATTGGTCTGTCTGTTGATGCCCAAGCGGTTTTAGCGAAGAGT GGTTTTGTTCAAGCAGAGCAACTCAAGACCCATGTGGAAAATCAGGATAATCCGCACCAAGTAACCG CTGAACAAGTGGGGCTAGGCAATGTAGAAAATTATGGCTTCGCATCAGACAGCGAAGCAGTCGCGG GAACTTTAACGAGTAAATATATGCACCCGAAAAACGTTGCGGAAGCGATTAAAGGTCAAGCTGTGAC ACAAACAGGTGATCAAGAGATTGCTGGGGTGAAGAATTTTGTAACTATGCCAACCGTCAATGGTGTG CCTTTTGAATCTTCTAAAATGGCCATTTATGAAGCTAGTGGAGTCGGTGAAGTCGGGGCAAAATATCA GGCGGCCTTTAATAAGGATAATATGAAATTTGTATTAATTAGGGTAGGAAATCGTGTCGATGCATTCG TAAGATGTAATTTGAGTGATCCGACGAAATTGAATAATAATTTGGTTAAAGTGTTTACTGTTCCAACAG GATATACATTATCGACGAAGATTACAAAGGGAATATGGAATTTGGCGTTAACTGCTATGCAATATACC

[0319] TTCCCTCAACCGAATTGTGCAGGTTTATATGAGATGGGAAATCAAGGAATTCTTTTTGGTGCTAACCG TGCTG G AAATATTTACCTAC AAG G AAGTTG GTAC ACG G ACG ATCCGTTTCC AAC AAAATAAC AAG CTA GTTTAG G AGTCTTTTTATG G AACGTTATCTC AAC AC AATAATAATG CTTTTAAG C ATTTTCG GTGG G AT TGTCGTACGTTTATTAGGTGGATTAGATCAATTGTTGGATGTCTTCCTCTTTTTAATTATTGTCGATTT CATCACAGGTTGGATTAAGGCAATCGCCACAAAAGAATTGTCCAGTCGGATTGGTATGCTCGGAATT GCGAAAAAAGTGACGATGTTATTTGTGGTTGCCGTAGCGGTTCGTGTTGAAAAAGTTGTGGGGAACA ATTTGCCAATTCGGGAAATGGTTCTGATTTTTTACATTGCGAACGAAGGACTTTCTTTTTTTGAAAACA TTGCGACCTTTATTCCTATGCCGAAAAAGTTAAAAGAGTTATTTATTCAGTTAAAAAATAAAGATGATT AAGTAGAAGTGGTCGGGACAAACGTAGAACTTTCGGCTGATTGCCGAAGAAATTACTTCTGTCCCGC CATTTATCTGCAGGTTTAAGCCGTGGAAGGGAAGTTATTTTGACTTTCCTTTCATGGCTTTTTTAAGAA AGGAGCATGCTATGTTTAAAAAATTAATGATTCAACTTGCTTTAGTGATTGGCTTAAGTTTAACGATTC CGATGACGGCTTGCGCTTACACTATCGAAGCGGATCCAATCAACTTTACTTATTTTCCCGGCTCTGCA AGCAATGAATTAATTGTTTTACATGAATCTGGAAACGAGCGGAACCTAGGACCACACAGTTTAGACAA TGAAGTGGCCTATATGAAACGAAATTGGTCAAATGCTTATGTCTCATATTTTGTCGGATCTGGTGGAC GAGTGAAACAATTAGCTCCTGCTGGCCAAATTCAATATGGCGCAGGTTCTTTAGCTAATCAAAAAGCC TATGCGCAAATCGAATTGGCTCGAACGAATAATGCGGCGACGTTTAAAAAAGATTATGCTGCCTATGT TAATTTGGCCCGTGATTTGGCTCAGAACATTGGTGCTGATTTTTCGCTAGACGATGGAACAGGTTATG GAATAGTCACTCATGATTGGATTACAAAAAATTGGTGGGGAGATCATACAGATCCTTATGGTTATTTA GCGCGTTGGGGGATTAGTAAAGCACAGTTGGCACAAGATTTACAAACGGGCGTTTCTGAAACAGGT GAGACTGTCATTATTCAGCCAGGTAAACCTAATGCGCCAAAATATCAAGTAGGACAAGCAATTCGTTT

[0320] CACTTCAATCTATCCAACACCAGATGCTTTAATCAATGAACATCTATCAGCAGAGGCACTTTGGACAC

[0321] AAGTAGGAACAATTACAGCGAAATTACCCGACCGACAAAACCTTTACCGTATTGAAAATAGCGGACAT

[0322] TTGTTAGGTTATGTGAACGACGGCGACATTGCTGAACTTTGGCGCCCGCAAACGAAGAAATCATTTC

[0323] TAATTGGTGTGGACGAAGGTATTGTTTTAAGAGCGGGCCAACCTAGCCTGTCAGCCCCAATTTATGG

[0324] TATTTGGCCTAAAAATACTCGCTTTTATTACGATGCGTTTTATATTGCAGATGGGTATGTTTTTATTGG

[0325] TGGGACAGATACGTCAGGCGCGAGAATTTATTTGCCAATCGGACCAAACGATGGCAACGCACAGAA TACATGGGGATCATTTACTAGCTAA

[0326] An exemplary nucleic acid encoding a group 2 efagin is provided in SEQ ID NO: 92, below, and is available under GenBank assembly no. GCA_000395175.1 .

[0327] Group 2 (Com1_efagin GCA_000395175.1_Ente_faec_Com1_V1_efagin) (SEQ ID NO: 92):

[0328] TTAAGTTAACCCAGTATTTAATACCATCACAGAATAATCCCTAAATTCAAAAATCCGTCCTTTATAAAG

[0329] ATAAGTGGCGCCATACTTTTGTCGATAGTACGTAATCACATTTTTTAATGTTTCAACATCGATTTCCAA

[0330] AAATTCAGCACAAGAATAATGGTTGCTTAAGCCAGCTTCAGAGCAACGAATTAAATCGTCTAATGTAA

[0331] CTAATTGCTCTAAGGCCACATTACGGGCTTTCAACTCTTGCTTCCGATTTTCTGTACAATTTTGATTTA

[0332] AAATAGTGCCAACGGATGTTTGATAATGCCCGTATTCCTCGGCTAAAATGTTCTTCTTTTGACGGGTA

[0333] CTCAATGTTTTTTCAATATAAATTTTGCCATTTCGATACAACCCATAACAACCTGTTTGATTGTATAAAT

[0334] CAATTTCTAAAACCGTCACATCCTTTTGAATGGAACTGACGAGTTTTTCATAATCATTCACTTTACCAT

[0335] TCACCTGCCATTACTTAATAAAAAGAGTAGATGGGAAGTTAATCCTCTTTTTTGTCAGAAGAAATCGAT

[0336] TGTTGATATTTGGCATCAATTTCATCCAAGTAATCGTGAATTTTCTCGATTTCTTCTTTTGAAAATATCT

[0337] TTTCTGGATCCCCAGCGTGGGCAGCCAGTGTCATTTCGTCACGGCGAGGGAATGAGAGAATTTCTG

[0338] CTTGCGTTTGTTGTTCGTGTAATTGTTGTTCGGCAAATTGATAAACGATGGCTTGTCGTTGGGGTTCT

[0339] AATTTTTTATAAATGCGGTCGATAGTTGAGGCATTTTTTTCCATGGGAACCTCTTGTCCCATCAACCA

[0340] GGCCTCATTAATATCTAATGCATCTGCAATGCGATAAACTTTGTCTTGTTTTGCTTCGTAACGACCAG

[0341] CTAGCCAATCGCTGATCGAAGATTTACCGATGCCAGTTTTTTTCGCTAATTCACTAGGTTTGATGTTC

[0342] TTGGCTGTTAAAGCCTCTTTTAAACGGATAGCAAAAATGTTCATATTTCCGAACCTCCTATTGCCTCAA

[0343] AGTATACTATACCTGTTTAGAGAAGGCAAGTAACAATCGCAATTAAATTTAGCAGTTCATAAAACCGA

[0344] ACTTATGAGTTGACAGTGAAAATATTTGATGCTAAGATAGAGCCATTCAAATGTTCGGTTAACTGAAC

[0345] TATTTTTTTGAACACTTTGTTCGGAAAACCGTATTTTGGTGGAAAGGGTGAAGAAATGAGTCGGAATT

[0346] ATAAAAAAACGTTGTCGGATATGTTACTTTTAGCAATTATTTTATTAATAAGCAGTGTCTCAATAAAAAT

[0347] TGGGGCCATCGTGATTGGTATGATTGGCCTCATGGAATTACTAACAGAGTAACAATAATTTAGTCAAA

[0348] AGGAGAGAAGCGGATTGGCAGAACGACGCATGTTTGCAAAAACGATTATTGATAGTGATGCGTTTTT

[0349] AGATATGCCCTTATCAAGTCAGGCCCTGTATTTTCATTTAGCGATGCGTGCCGATGATGATGGATTTA

[0350] TCAATAATCCCAAAAAATTGCAGCGAATGGTTGGTTGTGGGGAAGACGATCTAAAATTGCTAATGGTT

[0351] AAAAAATTTATTCTAGTATTTGAAAGTGGTGTGATCGTTATCAAACATTGGAAAATTCATAATTATATTC

[0352] GCAGTGATCGTTACAAACCAACCTTGTATCAAGAAGAGAAAAATCAGATTGTTGAAAAAAATAGCAAA

[0353] GCTTATACGTTTAAAGCAGAATCGTCTGTCAGTGGTCAACCAGCTGACTACCAACGGTTACCACAGG

[0354] AAAGCATAGTCCAGTCTAAGTTAGGTCAGAGTCAAGGTAGTAGTTCAGAAAACGATTGTTTAAAGACG

[0355] ATTTATCATTTTTATGAGGAAAACGGCTTTGGTACACTGGCCTCAAAAACAAGCCAAGATTTTAAGTAT TGGTTGCAAGATTTTATACAAAAAGGGGCTAGCCAAGAGGAAGCATGCCAATTAATCTTGCATGCTTT

[0356] AGGAATTGCCGTCGATCGAAATAAACGGAATTACGGCTATGTAAATGCTATTTTGAAAAGTTGGGAG

[0357] CAACAAAATTATTTATCCGTACATGAAGTTCTGGTAAATGATAAAAAACAAGTGTCGGGGCATGCGCC

[0358] GCAAATGACAGAAGAATATCAAGAGTTAGGTTTTTAAAGAAAGGAGGAAATCAGTATGCATGCGACA

[0359] GATCAAACTTTTCAAATACTATTGAGTCAATTGTTAGAAAAAGTTGAAGACCGTTGTCCTGAATGTGG

[0360] CAGTGAACAATATGTTTGGCAACAAAAAAATAAAGATGGCACAGAACGTTGTGCCCCAACTTGTTGGT

[0361] CGTGTGGGTATAAAATGCTAAAAAAACATGAACAACAAGCCAATCAACAACGTTCTCAAGAGAGTTTT

[0362] ATGGCACGTACACAAAAATTTTTTCATCAAGGGTCCTTAATTGCTGATGATGCGCTACGGCAATGTCG

[0363] TTTAACCAATTACCAAACCACTGAATTAGAAACAAGACAAGCAAAAGAACGGGCCTTAGCAGCAGTTT

[0364] CAGCGATTATTGAAGGAAAGCCAATCCACGTTATTTTTTCAGGGAAACCTGGTGTCGGTAAAAGTCAT

[0365] TTGGCTATCAGTATTTTAGTTGAAGTCTTAGAACGCTCTGCATATCAAAAGTATTGTTTATTTGTCAGC

[0366] TACTCTGAGTTATTAGAAAAACTAAAAATGTCCATGAATGAGTCGACCAAAAGCCAAGCAAAGGCTCA

[0367] AGCGTATATTACTAGAATGAAAAAAGCAGACGTTTTGGTCTTAGATGATTTAGGTGCTGAATTAGGAA

[0368] TTAAAAATAAAGTTAGTACGGATTTTAATAATGACATCTTAAATCGAATTTTAGAAGCTAGACAGAATA

[0369] AAGCAACTATTTTTACTACTAATTTTTCTGGAAAACAACTGGTGGAGGCCTATGGAACACGCATTATTT

[0370] CTCGTCTAATGAAGCACGCCAGTGGCTATGTTTTCCAATATAAAGACACAACAGACAAACGAATGAG

[0371] GAGTGTGAAATAAATATGTTAACAATTATTATTGGGTTTATCTTTTGGACAATGACACTAATGTTAGGT

[0372] TATCTAATTGGTGAAAGAGAAGGCCGTAAACATGAGTAATTTAACAAAACGTAAAAAAGATTTATTTGA

[0373] AATGAAAAGCGTTGTATTTAAAGATATTTCAAAGCAACAAAGCGAAAAAGCACAAAAAAGAAAACGAC

[0374] TCTTACAACTAATGAATCAATATCCCGATTGGGCAAGTCAAAAAAATAAACTTATTATGCAGGAAATTC

[0375] AAGAATTAGGACAAGCAATCGGTAATTGGTCGATGGATCAATCAAGACCCATCCAATCCATTAAGGC

[0376] CGCAACGTTTACAAAAAGCGAGTATCTCTATTTAATTTGGCTCGGTTATTCAGATGAAGCGATTCGTC

[0377] ACGGCTTAGAAATGTCGAAAGAGTGTTATTTTATTTATCGATTAACACTTTTAAATGAATAAAAGTAAA

[0378] GGAGATTAACCAATGCGTACGTCAACATTTAATTATATCAAAGATATTTTAGCAGACTTTTATAAAACA

[0379] GATGAGTATATCCGGCAACGGGAAGAAGAATTACGGCACCCTTATCAAGAAGCAGATTTAAATGCTG

[0380] GTATTAGAGGACAAGGACTTCACTCTGTAGTGACCGAACGAATGGCGATTACGATAGCTATGGATCG

[0381] TCGTCTGTGGAACTTAGAGAGAAATCGAGACATTATCAAAAATTGTTTAGCTGAAGCGGATGAACAAA

[0382] CGCGCGTGATTATTGAAGAACTATATATGAAAAAACGGCCCTCTTTAACATTAATTGGACTTGCCCAG

[0383] CAATTATTTATTAGTAAAAGCCAAGCCTATAAATTAAGAAATCATTTCTTTGAAGCGGTGGCGGATGAA

[0384] CTAGGGATGTAAACATGGAAAAAACGTGGAATTTTTTCAGGTGTCAACATGGTAAATTAATAGTGTCG

[0385] AAAGAGATAGATAAACGTGAGGCAACCAAAAAAATGAAGACACGGAATTCTATAATTTTGACTGCTTT

[0386] CTTGTGTCAGCTATGAAGGAGCAGAAAATGCCGGCTACTTTCAAGATCCTTCATTTTGACTAGAAGAG

[0387] AGCCAATTTGTTAACCAATCCTGAATTTTTTGAATGGAAAGGTGGCGCTAAAAATGAATGAAGCGGAA

[0388] CAAGAGTTATATGAAGCCCTTGTTGCAATCTGCCAGACGTCAGGATTTTTGTTGCTAGAGGAACTGC

[0389] CGACAGATTTACCAGATCAGCCATTTGTTTACTTAGGTGATAGTAAAGAATTACCTAAGCCAACTAAA

[0390] TCAGCTATTTTGGGTGAAATTGAATTAATAATGCATGTTTATGGTGCGTTATCTGAACGACAACAAATT

[0391] TCTACAATTAAAGGAACGATTTTACGGCAGGCAACCAGTAACTTAAAACGAACGGCTCATTTTAATTG

[0392] GGGTATCAAACATCAAGAAGTCAAAGCACAAATGGTAAAAGATACCAAACAAATGAAAAAAACAATTT

[0393] GGCATGCTGTACTACCATTACACATGCAATTTTACTAGGAGGAATTATCAATGGGAGAAGTTATGCAA

[0394] GGAAAAGACCGTATTTTATTAGTTCGTCGCTTGGATGAAGCAGCGACAAAGAAAGCAATGAAACCTTT ATTTCAAATTGAACATGAATGGGAATTCTCACGTGAATCGAGCGGTACGCAAACAAAAGATGGCGTT

[0395] GCGAATGCTGTTTCTGGTTTAGAAGTTACGTTATCGTTAAGCGGTTTAGCCTCTCGAAATGATGAAAA

[0396] TTTATACATGAAAGACGCAGTTGAAGATGGCATCTTAATGGAATTTTGGGATGTTGATTTAAAAGGTG

[0397] AAAAAAATGCGGAAGGTAAATATCCAGCAATTTATGCCCAAGGTTATGTAAATTCATGGAGTTTACCA

[0398] GCCAATGTAGAAGAATTAGTAGAAATCGAAACAGAAGCCTCTATTAATGGCAAACCACAAGATGGTTT

[0399] TGCAACAGTAGAAGCAGATATTATTGCAGAAGCACAATATGCGTTCCAAGATACCGTTCCAGATAAAG

[0400] CACCACAACCTGGCGAATAATCAAAAAGTGTTGAATTTTAGGAGGATAAAAAATGAATTTAGAGATTA

[0401] ACGGAAAAACAATTGAAGTGAAATTTACGATTGGCGCGATTCGCGAATTAGATAAACGTTACCAAATT G AAAATG GCG CTG CC AAATTCG GC ATG GG C ATC AGTTC AG C AATG ATTTATTTACG CC AATAC AATC CAGTAATCTTAGTTGACATCATGGAAGCTTTACAAAGTGGGCAATTAAAAATAGGTAAGTCGGAAATT

[0402] GAAGCATGGTTAATGACCCAAGATGTCAAAAAACTTTCAGATGATTTGCTTAAAGAAATGGGAAAGCA

[0403] ACCTCTTACAAAACCAATGATCGATCAGTTCAGCAAAGAAGCGAAGAAAGCAGAAGCGCAAGCGACC

[0404] AACTAATTAAAACGAGCGATGACGTGTATCACGACATCGCTCTTTCTGCTTTTCGCTACTTGGGCTGT

[0405] CGTTCATTTGAAGAAGTGGATCGGATGACCATGTCTGAATTTGAATTACGAATGATTGCTTTTAATTTA

[0406] GCAGAAGTAGATGAAGAGCGGAAAAGGCACGAGCTTGCCTACTTAAATGTTAAAGCGCAAGCGACA

[0407] AACAAAAAAGGAAAACCCGTTTTTGAAAACTTTAAAAGTTTTTATGATTATGAAAAACGAGTTGCTGAA

[0408] GTTCTGGCAGCTAACCAGCCACAACGAACGAAATTAAATGAGCGGAAAAAAACGCAACTTGCCACTG

[0409] TGGCAGAGCGTCTACGCCGCTATCGAGAAGGGAGGAGAGTAGATGGAGAATGACAAAGAAAAAACG

[0410] CCGTTATCGGAGGCAAAGAAAAGCCTTGCAGGCGTCCAACAAGCATTAAAAAGTATGAGCGGTGAG

[0411] TATGCCTTATTAAGTGGATATTTAGGGAAAATTAGTGCGGGTGTCAATCAGTCAGCCACGGTCATGAA

[0412] CACATTTAAAACCGTCATGCAACAATCTGGAGAAACAGTGAAAAAAACAGGAGACGAAACAGCAAAG

[0413] GCAGCAGATCAAATGAACACAGCGTTAACAGATTCTGCTGAACAAGCCGGTGAAGCAGCTAAAAAAG

[0414] CGGGGAAAGAAACCTCTGATGGCTTTACTAATGCACAAAATAATATGCTGAGCTTTGGGACGGCCAT

[0415] GACTAGTGCCGTTTCCTTACCTATGCTGAACGTTTTAAAAACAGCTATGGGCGTCGGTGCTGGGGTC

[0416] AGTGGCGAATTTCAAGGAATGCAAGGACTGATTATGGCTAGTGCAGGAGGGATTTCTGATTCATTGC

[0417] AAGGCGAGTTGCAAGGGGCATTGACTCAGATGAATCAATCATTTGAAGCGGCGGCACAAGTGATTCA

[0418] AAGCGTGATGGCTCCAGGAATGGAAATTTTGGTTCAAGTGGTTATCACAGTCGTCAAAGGCATTACA

[0419] GCTTTGGTGAATTTATTTATCAAATTACCAAAACCCGTCCAAGTTTTTATTGTTGCCATTATGGGCATT

[0420] TTAGCCGCCATTGGGCCCATGCTGATTATGGTAACGATGGCTCAGCAAAAATTTCAACAGTTTAGTG

[0421] ATGGTTTGGTTCTTGTAAAAGGAAACATTGGGAAGTTAGGTGGTGGCTTATCAAAACTAAGTGCTAGT

[0422] TTTAGTGCCTTAAGTGGAGGACCGTTAATTTTAATTGTAGCAGCCGTTTTAGCAGCGGTAGCAGCGTT

[0423] TATTTATTTCTATAAAACCAATGAAACATTTAGAAATAGTATCAATAGCTTAGCTAGTGCCATTCAAGG

[0424] AGCTGTTTCAGCGGCGTTTGGCAAATTGGTAGGATTGCTACAACAGATCCAGCCGGCCTTTCAGCAA

[0425] GTAATGGCAGTTTTTAAACAATTTTTTGCAGTAGGCTTAGAGAAAATGGCGACTATTTTTTCAACAATT

[0426] GGTCGTGTGCTAACAGGCGTTTTTGCCAGCGGTTTGCAATTAGGTAGTAACTTATTAGGGCAATTTG

[0427] GTGGCACCTTTGACAAAGCTGGTTTAGCGGTTGGTCTTTTGGTAAAAGTTCTGACAAAAGTTGCACT

[0428] GGCTGCATTAGGAATTTCTGGGCCGTTTGGTCTAATTATTTCCTTGATTGTTTCATTCGTGACGGCCT

[0429] GGATGAAAACCGGTGATTTGAGTGCGGGTGGTATTACCCAAGTCTTTGATAATTTAGGTAACACGATT

[0430] ACATCTGTTACAACAATGTTGGCAGCTAATCTACCGAAAGTTATACAACTTTTTACAACCGTCTTAACC

[0431] AGTATTCTCGGGAAAATAACAGAAGCTATTCCAAGCATCGTAACCGCGTTATCTAGTTTAATTACGTT AATTGTTGGTGCGATCGTTGCCAATTTGCCAGTCTTAATTGAAGCGGCAACGCAAATTATTACTACGT

[0432] TGATTCAGGGGATTACAACAGTCTTACCAATGTTGATAGAAGTTGGTTTGAGCTTATTAATGACTTTA

[0433] GTTAATGCGATTGTCACCGCCTTGCCAACAATTACAACTGCAGCGATTAATATCATCACTACATTAGT

[0434] GACAGCTTTTGTCACAGCGTTACCAATGCTAGTTACAGCAGGTGTTTCAATTATCACAGCCTTAGTCA

[0435] ATGCATTTGTTACTATGTTACCGTTGATTTTGACTGCTGGTTTACAAATTTTGATGGCATTAATCACTG

[0436] GGATTATGACGATTTTACCTCAGTTAATTCAATCAGCGCTGACGATTATTCTAGCGTTAGTGACTGCG

[0437] TTGGTAGGTGCCTTGCCACAGATTATCAGCGCAGGTGTCAAATTGTTAATGGCGTTAATTCAAGGAAT

[0438] TATTTCGATTTTACCAACCTTAGTTGCGGCAGCTATTACCTTAATTTTGACATTGGTAAATGCCTTAAT

[0439] TGGTGCCTTGCCACAAATCATCAGCGCAGGCGTCAAATTGCTAATGGCTTTGATCCAAGGGATTATT

[0440] TCAATTTTACCGCAACTGGTTACTGCAGCAATTACGCTAATTACCGCTTTAATGGGTGCGTTAATCAA

[0441] TGCGTTGCCACAGTTGTTAAGTGCTGGGATTCAACTGATTCAAGCCTTAATTAATGGTGTACTCAGTC

[0442] TATTGGGTGCCTTGCTGTCCGCAGCAGGAACATTAATCTCACAAATGATCACGAAGATTGGTTCTTAT

[0443] TTTGGTCAACTGTTAGCTTCGGGCGGACAGTTAGTTGAAAATATCAAAAATGGGGTTACCAATGCAG

[0444] CCGATCAGGTAAAAAATGCCATTGGTTCTGTAATTGAAGGTGCTTGGCAAGCAATCCAAGGTTGGTT

[0445] TTCAAAATTCACCGATGCCGGTGCGAATATTGTCGGTATGATTGCTGATGGAATTACAGGCGCAATT

[0446] GGAAAAGCCAAAGAAGCAATCGATGGGGTCGTCAGTAAAATTCGTAACTTTTTACCATTTTCACCAGC

[0447] AAAAGAAGGTCCCTTATCTGATTTGCATAAATTGAATTTCGGCGGCACGATTGCCACGGGGATTTATG

[0448] CAGGCGAAACAGCCGTTAGTAGAGCAATGGCTTCTATTTTAGACTTACCGCTGTTAAATGATTTTGCC

[0449] TTGGACTTAGCTGGTCGAGGAAACTTCACAGCAACGATTGACCATCGTTTAGAAAATGATGCATACAA

[0450] TCGACCATTATTTGTGACAGTAGAGTCAACGTTAGATGGGAAAGTTGTCGCAGCAACTACGGCGCCT

[0451] TATTTAGCAACAGAGTTACAACGACAACAAGTGAAACAAAATAACCGCTTAGGAAGGAGAGGATAAC

[0452] ATGTATAAATTTGTTGATACCAATCAAGCAACTCATTCAACGCCTCTTCCTTCAGAAGCGTTGAATTTT

[0453] AACGGTCAATTTTTAGAAAAAGTCATCCCTGGCTATCAAACATTATCAGTTTCAGGACGAGAATTAGT

[0454] TCCAAGCGAAATTGAAAGCTATCAATTAGGGATTCGTGATGGCAAACGCCACGTTTATGCGCGGATT

[0455] CCAGAACGAGAATTAACAGTCAAATATCGACTTTCAGCTGTGAATAATGAAGCGTTTCGAGATGCATT

[0456] TAATCATTTAAATGTTGCTTTGTTTACGGAAAAAGACGTTTCTATTTGGTTTAACGATGAACCGGAAAT

[0457] GCTGTGGTTTGGTAGTAAGTCTTCAGTGAGTGATGTACCCGAAGGCGTTAACCAAGTAACAGGCACT

[0458] TTTACTTTATTGCTTTCTGATCCGTATAAATACACACGGAGTGATGCGACTAGTGTGATGTGGGGTTC

[0459] GCCAACCATTACATTTCAAGCGAATTACTTAATGGGGAATACAGGCTCAGGTGCAGTTGATTTTCCAA

[0460] TTTTAATTGAAGGTGGAGCCTATTGGGGATCAACCATGATTACCTTTCAAAATCGGGCTTACACGATG

[0461] GGGGATTTAGGCAAAGAAGTTCGGCCAATTGAAATTTATCCTACGGTCGAAGGGTTAAAAGTCAAAC

[0462] CGACCATTATTTTAGCAGGAACCGGACGTGGTGTTTGGATTAAAACAAGGAACGATACAATTAACTTA

[0463] GGTGACTTTGATCGTTCGGAAATTATCATTGATACTGAAAATTTTTATCTGACAAAAAATGGTGCACC

[0464] GATGATTCGACCAATGAACGATTTTTATCTATATCCCAATGAACCGCTGTATATTCAAGCCAAAGATA

[0465] GCGACTTCCGCTTGACGATTCGCTATCCTAACCGATTTGTGTAGGAGGGTGATTAAATGTTAATGGC

[0466] GCTGGATTTGAAAAGAACATATACGGCAATTTTGGACAATGCCTATCAAGTCAGTTATGAAAAAATAG

[0467] AGAACAAAATTGGCAGTTTAGATTTTACCATGCCACTAGATGATCCTAAAAATGAATTTATTGAAGAAA

[0468] TGCAATGGGTGGAACTGACCGACAATGAGAATGAATATATTGGTTTATATCGCGTGATGCCAACCAC

[0469] AATTAAGAAAGATGCGAACAATAATCAAATTCACTACTCTGCCACAGAAGCATTATGTACCTTAGGCG

[0470] ATACTGTCCTATTTGGTTGTCACGAAATTAAAAACAAAACAACGAAAGAGGCCATTCAATTTCTATTGA ATAAACAAAAAACAAAGCATTGGGTCCTGAAAAAATGTGATTTTTCAAGGAAATTAACCTATAAATGG

[0471] GAGAATGAAAACGGGCTAGTCGAGCCTTTATTTAGCATCCCAGCCGATTTCGAAGAGGAATATCTTT

[0472] GGCAATGGAATACAGAGGTCTATCCTTTTGAACTTTCATTAGTCAAACCACCAACAGAACCAGTTGCG

[0473] CGAATTCAAGAAGGTTACAATATGCAAGGATTCGAAATAGAACGTAATCCCAAGATGCTAATCAATCG

[0474] GATTTATCCATTAGGTTCAGGCGAAGGTGTTAACAAAGTCAATATTCGCTCGGTCAATCAAGGGGTTC

[0475] CGTATTTAGAGAACAAGGCCGCAATTGACCGCTATGGTTTATTGGAGTCAATTTGGGTGGAACAGCG

[0476] TTTTTCTGATCCCAAGGCATTAAAGGAAAATGCTTTGCGAATGTTAGAAGAATGGACCAAACCACAAG

[0477] TTTCTTGGGTAGTGACTGCAGCTGATTTAATTAAATTAACAGATCAACCTTTGGCAATCGATCGTTTG

[0478] CGGTTGGGTACGGTTATCATGATTAATACGAATGAGTTTGGGAGTGTCAACCTTCGTATCAAAAAAGA

[0479] AAGTAAAAAAGATGTCTTTGGTGCCCCCCAAGACATTCAGCTAGAGTTGGGAAATCTGCAAGAAACA

[0480] ATTCATAGTACCATGACAGCTTTCAGTCGGAAACAAGAGATTAGCGAAACTTACGCACAAGGGGCGA

[0481] CGACACTTTTAAATCGTTCAATACAAGGAGAACTTAGCAAGACACAGCCAGTGGAGCTGAATTTATAC

[0482] TTTGACGAGGACATTCTTTATATAAACACCGCAGAATTAACGTTCAAGGCAACTGCTAAAGGACCTTC

[0483] GCATTCTGTAACGAACATTGATTTGGTAGTGGATGGCAAAAAATTACCCCAACTATCATTGCAACAAC

[0484] AACGGCTAAACATTTTGAGTTATTTACGAAAAACAACAGATGGAAAAATCGAACGCGGCAATCACACG

[0485] CTTCAATTTTTCTCTCATCAGCCACTATGGTTGGATGCTTCGGTCATCTGTCGTGTGTATATTCAATCC

[0486] CAATTGGGTGGCCAGTTTTAATAAAATAATGAAAACTAGAGGAGTGTGACGAAATGTCAGTAGAACAT

[0487] ATTGAAGAATTAGATACCCTGAATCAAGGTCGCCTTAAAATCAATGCAATCTTGGATCAGTCGAATGC

[0488] ATCAGCTGAGAAAGTAGATGCTTACCAAGTCCAGTTAACGAATGGAATTTCTGAAGCGAAAAACATC

[0489] GCAGATGAAGCTGGCAAAGAAGCCGTACAAATTGCCACCGATGCAGGCAATCAAGCAAATGAAACA

[0490] GCCAACCAAGCGATGAACAATGCCAAAACAGCCATCACGATTGCAGGAAATGCAGTTTCAACGGCAA

[0491] ATAATAATAAACAAGAATTTGATACTTTGCGAAATGATTTCGATCAATTAGTAGCAGAAGCAGGTGATA

[0492] GTAATCCAGAAATTGTCCAAGCACGCACAGATACACAAGGCATCAAACAAGCTACTCTAGCGAATCG

[0493] TCTTCAAATTGATTTGAATGACCGCATGACAAAAGCAGATGGCATTTCTTTATTGGCTAAGCCAACTA

[0494] CTGTCAAAATGAAGTTAGACTTTAACGGTAAAACGGCCGGCAATACAGCCACCAATGCAAACAGTTA

[0495] TTCCACTGATTTTACGGCTAAAATTCTTAAAAAGCCAACAGACGTTTGGGAGGAAGTTTCCCAAGCGG

[0496] ACTACAATAAAATGGCCAGCCGTGATGATGAGGGTGTGAAAACAGGTTCCACCCAAAGTGGTGTAAT

[0497] TCCACAACAGTTAGCGGCCTTCAATCTCGTTGAAGCCGCAAAAAAATTAATTCCACAAATGTTTGAAA

[0498] CATTCACAACTGACGAGGCGGTGGCATTTGTTCGCCAGAACGTTCAATTTTTTACGATTAATCAACGT

[0499] GTGAAAGCCGCTGCGCCCAATAATCAAACGATTAAAATCGCTGCGTATTTACCAACTACGGATAATTG

[0500] GGTAACTCAAATCCAAGAATCAGCAAAAGAGTTCAGCGATTTTTCAATTCAAATCAATGATCAGAATTT

[0501] TATCACAGATGAAGGGTTCATTTATTTAATGAGCTATACAGATTCATCGAATGGGGTGACGCCAGCCA

[0502] GCGTAGAAGTTGATTACGTGGGACTTCATATTGGTCTGTCTGTTGATGCCCAAGCGGTTTTAGCGAA

[0503] GAGTGGTTTTGTTCAGGCAGAGCAACTCAATACCCATGTGGAAAATCAGGATAATCCACACCAAGTA

[0504] ACCGCTGAACAAGTTGGGCTGGGCAATGTAGAAAATTATGGCTTCGCATCAGACAGCGAAGCAGTC

[0505] GCGGGAACTTTAACGAGTAAATATATGCACCCGAAAAACGTTGCGGAAGCGATTAAAGGTCAAGCTG

[0506] TGACACAAACAGGTGACCAAGAGATTGCTGGGATGAAGAATTTTGTAACTATGCCAACCGTCAATGG

[0507] TGTGCCTTTGGAATCCTCTAAAATGGCCATTTATGAAGCTAGTGGAGTCGGTGAAGTCGAGGCAAAA

[0508] TATCAGGCGGCCTTTAATAAAGTGAATATGAAATTTGTATTAATCAGAGTTGGAAATCGTGTCGATGC

[0509] ATTTGTAAGATGTAATTTGAGTGATCCAACGAAATTGAACAGTCAGGTGATCAAAGTTTTTAATGTTCC CACAGGTTATAGTATAAATTCTGTTCTCAAAAGAAATATTTGGAATATTCCTTTAACGGCAGTTCAATA

[0510] CAACTTTCCTCAACCAATTTGTACAGCCTTGTACGAAATAGATAATAAGGGAATTATATTCTGTTCTAA

[0511] CCGTGCTGGAAATATTTACCTCCAAGGAAGTTGGTACACAGACGATCCGTTTCCAACAAAATAACAA

[0512] GCTAGTTTAGGAGTCTTTTTATGGAACGTTATCTCAACACAATAACAATGCTTTTAAGCATTTTCGGTG

[0513] GGATTGTCGTACGTTTATTAGGTGGATTAGATCAATTGTTGGATGTCTTCCTCTTTTTAATTATTGTCG

[0514] ATTTCATCACAGGTTGGATTAAGGCAATCGCCACAAAAGAATTGTCCAGTCGGATTGGTATGCTCGG

[0515] AATTGCGAAAAAAGTTACGATGTTATTTGTGGTTGCCGTAGCGGTTCGTGTTGAAAAAGTTGTGGGG

[0516] AACAATTTGCCAATTCGGGAAATGGTTCTGATTTTTTACATTGCGAACGAAGGACTTTCTTTTTTTGAA

[0517] AATATTGCGACCTTTATTCCTATGCCGAAAAAGTTAAAAGAGTTATTTATTCAGTTAAAAAATAAAGAT

[0518] GATTAAGTAGAAGTGGTCGGGACAAATGTAGAACTTTCGGCTGATTGCCGAAGAAATTACTTCTGTC

[0519] CCGCCATTTATCTGCAGGTTTAAGCCGTGGAAGGGAAGTTGTTTTGACTTTCCTTTCATGGCTTTTTT

[0520] AAGAAAGGAGTATGCTATGTTTAAAAAATTAATGATTCAACTTGCTTTAGTGATTGGCTTAAGTTTAAC

[0521] GATTCCGATGACGGCTTGCGCTTACACCATCGAAGCGGATCCAATCAACTTTACTTATTTTCCAGGCT

[0522] CTGCAAGCAATGAATTAATTGTTTTACATGAATCAGGAAACGAGCGGAACCTAGGACCACACAGTTTA

[0523] GACAATGAAGTGGCCTATATGAAACGAAATTGGTCAAATGCTTATGTCTCATATTTTGTCGGATCTGG

[0524] TGGACGAGTGAAACAATTAGCTCCTGCTGGTCAAATTCAATATGGCGCAGGTTCTTTAGCTAATCAAA

[0525] AAGCCTATGCGCAAATCGAATTGGCTCGAACGAATAATGCGGCGACGTTTAAAAAAGATTATGCTGC

[0526] CTATGTTAATTTGGCCCGTGATTTGGCTCAGAACATTGGTGCTGATTTTTCGCTGGACGATGGAACA

[0527] GGTTATGGCATAGTCACTCATGATTGGATTACAAAAAATTGGTGGGGAGATCATACAGATCCTTATGG

[0528] TTATTTAGCGCGTTGGGGGATTAGTAAAGCGCAGTTGGCACAAGATTTACAAACGGGCGTTTCTGAA

[0529] ACAGGTGAGACTGTCATTGTTCAGCCAGGTAAACCTAATGCACCAAAATATCAAGTAGGACAAGCAA

[0530] TTCGTTTCACTTCAATCTATACAACACCAGATGCTTTAATCAATGAACATCTATCAGCAGAGGCACTTT

[0531] GTACACAGGTAGGAACAATTACAGCGAAATTACCCGACCGACAAAACCTTTACCGTATTGAAAATAG

[0532] CGGACATTTGTTAGGTTATGTGAACGACGGCGACATTGCTGAACTTTGGCGCCCGCAAACGAAGAAA

[0533] TCATTTCTAATTGGTGTGGACGAAGGTATTGTTTTAAGAGCAGGACAACCTAGTTTGTCAGCACCTAT

[0534] TTATGGTATCTGGCCCAAAAATACTCGCTTTTATTACGATGCGTTTTATATTGCAGATGGGTATGTTTT

[0535] TATTGGTGGAACAGATACGACAGGCGCGAGAATTTATTTGCCAATCGGACCAAACGATGGCAACGCA CAGAATACATGGGGATCATTTACTAGCTAA

[0536] An exemplary nucleic acid encoding a group 3 efagin is provided in SEQ ID NO: 93, below, and is available under GenBank assembly: GCA_000393035.1 .

[0537] Group 3 (T9_efagin GCA_000393035.1_Ente_faec_T9_V1_efagin) (SEQ ID NO: 93)

[0538] TTAAGTTAACCCAGTATTTAATACCATCACAGAATAATCCCTAAATTCAAAAATCCGTCCTTTATAAAG

[0539] ATAAGTGGCGCCATACTTTTGTCGATAGTACGTAATCACATTTTTTAATGTTTCAACATCGATTTCCAA

[0540] AAATTCAGCACAAGAATAATGGTTGCTTAAGCCAGCTTCAGAGCAACGAATTAAATCGTCTAATGTAA

[0541] CTAATTGCTCTAAGGCCACATTGCGAGCTTTCAACTCTTGCTTCCGATTTTCTGTACAATTTTGATTTA

[0542] AAATAGTGCCAACGGATGTTTGATAATGCCCGTATTCCTCGGCTAAAATGTTCTTCTTTTGACGGGTA

[0543] CTCAATGTTTTTTCAATATAAATTTTGCCATTTCGATACAACCCATAACAACCTGTTTGATTGTATAAAT

[0544] CAATTTCTAAAACCGTCACATCCTTTTGAATGGAACTGACGAGTTTTTCATAATCATTCACTTTACCAT

[0545] TCACCTGCCATTACTTAATAAAAAGAGTAGATGGGAAGTTAATCCTCTTTTTTGTCAGAAGAAATCGAT TGTTGATATTTGGCATCAATTTCATCCAAGTAATCGTGAATTTTCTCGATTTCTTCTTTTGAAAATATCT

[0546] TTTCTGGATCCCCAGCGTGGGCAGCCAGTGTCATTTCGTCACGGCGAGGGAATGAGAGAATTTCTG

[0547] CTTGCGTTTGTTGTTCGTGTAATTGTTGTTCGGCAAATTGATAAACGATGGCTTGTCGTTGGGGTTCT

[0548] AATTTTTTATAAATGCGGTCGATAGTTGAGGCATTTTTTTCCATGGGAACTTCTTGTCCCATCAACCAG

[0549] GCCTCATTAATGTCTAATGCATCTGCAATGCGATAAACTTTGTCTTGTTTTGCTTCGTAACGACCAGC

[0550] TAGCCAATCGCTGATTGAAGATTTACCGATGCCAGTTTTTTTCGCTAAATCACTAGGTTTGATGTTCTT

[0551] GGCTGTTAAAGCCTCTTTTAAACGGACAGCAAAAATGTTCATATTTCCGAACCTCCTATTGCCTCAAA

[0552] GTATACTATACCTGTTTAGAGAAGGCAAGTAACAATCGCAATTAAAATCACTGGTTCATAAAACCGAA

[0553] CTTATGAGTTGACAGTGAAAATATTTGATGCTAAGATAGAGCCATTCGAATGTTCGGTTAACTGAACT

[0554] ATTTTTTTGAATACTTTGTTCGGAAAACCGTATTTTTGTAGAAAGGGTGAAGAAATGAGTCGGAATTAT

[0555] AAAAAAACGTTGTCGGATATGTTACTTTTAGCAATTATTTTATTAATAAGCAGTGTCTCAATAAAAATTG

[0556] GAGCCATCGTGATTGGTATGATTGGCCTCATGGAATTACTAACAGAATAACAATAATTTAGTCAAAAG

[0557] GAGAGAAGCGGATTGGCAGAACGACGCATGTTTGCAAAAACGATTATTGATAGTGATGCGTTTTTAG

[0558] ATATGCCCTTATCAAGTCAGGCCCTGTATTTTCATTTAGCGATGCGTGCCGATGATGATGGATTTATC

[0559] AATAATCCCAAAAAATTGCAGCGAATGGTCGGTTGTGGGGAAGACGATCTAAAATTGTTGATGGTTAA

[0560] AAAATTTATTCTAGTATTTGAAAGTGGTGTGATCGTTATCAAACATTGGAAAATTCATAATTATATTCGC

[0561] AGTGATCGTTACAAACCAACCTTGTATCAAGAAGAGAAAAATCAGATTGTTGAAAAAAATAGCAAAGC

[0562] TTATACGTTTAAAGTAGAATCGTCTGTCAGTGGTCAACCAGCTGACTACCAACGGTTACCACAGGAAA

[0563] GCATAGTCCAGTCTAAGTTAGGTCAGAGTCAAGGTAGTAGTTCAGAAAACGATTGTTTAAAGATGATT

[0564] TATCATTTTTATGAGGAAAACGGCTTTGGTACACTGGCCTCAAAAACAAGCCAAGATTTTAAGTATTG

[0565] GTTGCAAGATTTTATACAAAAAGGGGCTAGCCAAGAGGAAGCATGCCAATTAATCTTGCATGCTTTAG

[0566] GAATTGCCGTCGATCGAAATAAACGGAATTACGGCTATGTAAATGCTATTTTGAAAAGTTGGGAGCAA

[0567] CAAAATTATTTATCCGTACATGAAGTTCTGGTAAATGATAAAAAACAAGTGTCGGAGCATGCGCCGCA

[0568] AATGACAGAAGAATATCAAGAGTTAGGTTTTTAAAGAAAGGAGGAAATCAGTATGCATGCGACAGATC

[0569] AAACTTTTCAAATACTATTGAGTCAATTGTTAGAAAAAGTTGAAGACCGTTGTCCTGAATGTGGCAGT

[0570] GAACAATATGTTTGGCAACAAAAAAATAAAGATGGCACAGAACGTTGTGCCCCAACTTGTTGGTCGT

[0571] GTGGGTATAAAATGCTAAAAAAACATGAACAAGAAGCCACTCAACAACGTTCTCAAGAGAGTTTTATG

[0572] GCACGTACACAAAAATTTTTTCATCAAGGGGCCTTAATTGCTGATGATGCGCTACGGCAATGTCGTTT

[0573] AACCAATTACCAAACCACTGAATTAGAAACAAGACAAGCAAAAGAACGGGCCTTAGCAGCAGTTTCA

[0574] GCGATTGTTGAAGAAAAGCCAATTCACGTTATTTTTTCAGGGAAACCTGGTGTCGGTAAAAGTCATTT

[0575] GGCTATCAGTATTTTAGTTGAAGTCTTAGAACGCTCTGCATATCAAAAGTATTGTTTATTTGTCAGCTA

[0576] CTCTGAGTTATTAGAAAAACTAAAAATGTCCATGAATGAATCGGCCAAAAGCCAAGCAAAGGCTCAAG

[0577] CGTATATTACTAGAATGAAAAAAGCAGACGTTTTGGTCTTAGATGATTTAGGTGCTGAATTAGGAATT

[0578] AAAAATAAAGTTAGTACGGATTTTAACAATGACATCTTAAACCGAATTTTAGAAGCTAGACAGAATAAA

[0579] GCAACTATTTTTACTACTAATTTTTCTGGAAAACAACTGGTGGAGGCCTATGGAACACGTATTATTTCT

[0580] CGTCTAATGAAGCACGCCAGTGGCTATGTTTTCCAATATAAAGACACAACAGACAAGCGAATGAGGA

[0581] GTGTGAAATAAATATGTTAACAATTATTATTGGGTTTATCTTTTGGACAATGACACTGATGTTAGGTTA

[0582] TCTAATTGGTGAAAGAGAAGGCCGTAAACATGAGTAATTTAACAAAATGTAAAAAAGATTTATTTGAAA

[0583] TGAAAAGCGTTGTATTTAAAGATATTTCAAAGCAACAAAGCGAAAAAGCACAAAAAAGAAAACGACTC

[0584] TTACAACTAATGAATCAATATCCCGATTGGGCAAGTCAAAAAAATAAACTTATTATGCAGGAAATTCAA GAATTAGGACAAGCAATCGGTAATTGGTCGATGGATCAATCAAGACCCATCCAATCCATCAAGGCCG CATCGTTTACAAAAAGCGAGTATCTCTATTTAATTTGGCTCGGTTATTCAGATGAAGCGATTCGTCAC GGCTTAGACATGTCGAAAGAGTGTTATTTTATTTATCGATTAACACTTTTAAATGAATAAAAGTAAAGG AGATTAACCAATGCGTACGTCAACATTTAATTATATCAAAGATATTTTAGCAGACTTTTATAAAACAGA TGAGTATATCCGGCAACGGGAAGAAGAATTACGGCACCCTTATCAAGAAGCAGATTTAAATGCTGGT ATTAGAGGACAAGGACTTCACTCTGTAGTGACCGAACGAATGGCGATTACGATAGCTATGGATCGTC GTCTGTGGAACTTAGAGAGAAATCGAGACATTATCAAAAATTGTTTAGCTGAAGCGGATGAACAAAC GCGCGTGATTATTGAAGAACTATATATGAAAAAACGGCCCTCTTTAACATTAATTGGACTTGCCCAGC AATTATTTATTAGTAAAAGCCAAGCCTATAAATTAAGAAATCATTTCTTTGAAGCGGTGGCGGATGAAC TAGGCATGTAAACATGGAAAAAGCGTGGAATTTTTTCAGGTGTCAACATGGTAAATTAATAGTGTCGA AAGAGATAGATAAACGTGAGGCAACCAAAAAAATGAAGACACGGAATTCTATGATTTTGACTGCTTTC TTGTGTCAGCTATGAAGGAGCAGAAAATGCCGGCTACTTTCAAGATCCTTCATTTTGACTAGAAGAGA GCCAATTTGTTAACCAATCCTGAATTTTTTGAATGGAAAGGTGGCGCTAAAAATGAATGAAGCGGAAC AAGAGTTATATGAAGCCCTTGTTGCAATCTGCCAGACGTCAGGATTTTTGTTGCTAGAGGAACTGCC GACAGATTTACCAGATCAGCCATTTGTTTACTTAGGTGATAGTAAAGAATTACCTAAGCCAACTAAAT

[0585] CAGCTATTTTGGGTGAAATTGAATTAATAATGCATGTTTATGGTGCGTTATCTGAACGACAACAAATTT CTACAATTAAAGGAACGATTTTACGGCAGGCAACCAGTAACTTAAAACGAACGGCTCATTTTAATTGG GGTATCAAACATCAAGAAGTCAAAGCACAAATGGTAAAAGATACCAAACAAATGAAAAAAACAATTTG GCATGCTGTACTACCATTACACATGCAATTTTACTAGGAGGAATTATCAATGGGAGAAGTTATGCAAG GAAAAGACCGTATTTTATTAGTTCGTCGCTTGGATGAAGCAGCGACAAAGAAAGCAATGAAACCCTTA TTTCAAATTGAACATGAATGGGAATTCTCACGTGAATCGAGCGGTACGCAAACAAAAGATGGCGTTG CGAATGCTGTATCTGGTTTAGAAGTTACGTTATCGTTAAGCGGTTTAGCCTCTCGAGATGATGAAAAT TTATACATGAAAGACGCAGTCGAAGATGGCATCTTAATGGAATTTTGGGATGTTGATTTAAAAGGTGA AAAAAATGCGGAAGGTAAATATCCAGCAATTTATGCCCAAGGTTATGTAAATTCATGGAGTTTACCAG CCAATGTAGAAGAATTAGTAGAAATTGAAACAGAAGCCTCTATTAATGGCAAACCACAAGATGGTTTT GCAACAGTAGAAGCAGATATTATTGCAGAAGCACAATATGCGTTCCAAGATACCGTTCCAGATAAAG CACCACAACCTGGCGAATAATCAAAAAGTGTTGAATTTTAGGAGGATAAAAAATGAATTTAGAGATTA ACGGAAAAACAATTGAAGTGAAATTTACGATTGGCGCGATTCGCGAATTAGATAAACGTTACCAAATT

[0586] G AAAATG GCG CTG CC AAATTCG GC ATG GG C ATC AGTTC AG C AATG ATTTATTTACG CC AATAC AATC CAGTAATCTTAGTTGACATCATGGAAGCTTTACAAAGTGGGCAATTAAAAATAGGTAAGTCGGAAATT GAAGCATGGTTAATGACCCAAGATGTCAAAAAACTTTCAGATGATTTGCTTAAAGAAATGGGAAAGCA ACCTCTTACAAAACCAATGATCGATCAGTTCAGCAAAGAAGCGAAGAAAGCAGAAGCGCAAGCGACC AACTAATTAAAACG AG CG ATG ACGTGTATC ACG AC ATCG CTCTTTCTG CTTTTCG CTACTTAG GCTGT CATTCATTTGAAGAAGTGGATCGGATGACCATGTCTGAATTTGAATTACGAATGATTGCTTTTAATTTA GCAGAAGTAGATGAAGAGCGGAAAAGGCACGAGCTTGCCTACTTAAATGTTAAAGCGCAAGCGACA AACAAAAAAGGAAAACCCGTTTTTGAAAGCTTTAAAAGTTTTTATGATTATGAAAAACGAGTTGCTGAA GTTCTGGCAGCTAACCAGCCACAACGAACAAAATTAAATGAGCGGAAAAAAACGCAACTTGCCACTG TGGCAGAGCGTCTACGCCGCTATCGAGAAGGGAGGAGAGTAGATGGAGAATGACAAAGAAAAAACG CCGTTATCGGAGGCAAAGAAAAGCCTTGCAGGCGTCCAACAAGCATTAAAAAGTATGAGCGGTGAG TATGCCTTATTAAGTGGATATTTAGGGAAAATTAGTGCGGGTGTCAATCAGTCAGCCACGGTCATGAA CACATTTAAAACCGTCATGCAACAATCTGGAGAAACAGTGAAAAAAACAGGAGACGAAACAGCAAAG

[0587] GCAGCAGATCAAATGAACACAGCGTTAACAGATTCTGCTGAACAAGCCGGTGAAGCAGCTAAAAAAG

[0588] CGGGGAAAGAAACCTCTGATGGCTTTACTAATGCACAAAATAATATGCTGAGCTTTGGGACGGCCAT

[0589] GACTAGTGCCGTTTCCTTACCTATGCTGAATGTTTTAAAAACAGTTATGGGCGTCGGTGCTGGGGTC

[0590] AGTGGCGAATTTCAAGGAATGCAAGGACTGATTATGGCCAGTGCAGGAGGGATTTCTGATTCATTGC

[0591] AAGGCGAGTTGCAAGGGGCATTGACTCAGATGAATCAATCATTTGAAGCGGCGGCACAAGTGATTCA

[0592] AAGCGTGATGGCTCCAGGAATGGAAATTTTGGTTCAAGTGGTTATCACAGTCGTCAAAGGCATTACA

[0593] GCTTTGGTGAATTTATTTATCAAATTACCAAAACCCGTCCAAGTTTTTATTGTTGCCATTATGGGCATT

[0594] TTAGCCGCCATTGGGCCCATGTTGATTATGGTAACGATGGCTCAGCAAAAATTTCAACAGTTTAGTGC

[0595] TGGTTTGGCTCTTGTACAAGGAAACATTGGGAAGTTAGGTGGTGGCTTATCAAAACTAAGTGCTAGTT

[0596] TTAGTGCCTTAGGTGGAGGACCATTAATTTTAATTGTAGCAGCCGTTTTAGCAGCGGTAGCAGCGTTT

[0597] ATTTATTTCTATAAAACCAATGAAACATTTAGAAATAGTATCAATAGCTTAGCTAGTGCCATTCAAGGA

[0598] GCTGTTTCAGCGGCGTTTGGCAAATTGGTAGGACTGCTACAACAGATCCAGCCGGCCTTTCAACAAA

[0599] TAATGGCAGTTTTTAAACAATTTTTTGCAGTAGGCTTAGAGAAAATGGCGACTATTTTTTCAACAATTG

[0600] GTCGTGTGCTAGCAGGCGTTTTTGCCAGCGGTTTGCAATTAGGTAGTAACTTATTAGGGCAATTTGG

[0601] TGGCACCTTTGACAAAGCTGGTTTAGCGGTTGGTCTTTTGGTAAAAGTTCTGACAAAGGTTGCACTG

[0602] GCTGCATTAGGAATTTCTGGGCCGTTTGGTCTAATTATTTCCTTGATTGTTTCATTCGTGACGGCCTG

[0603] GATGAAAACCGGTGATTTGAGTGCGGGTGGTATTACCCAAGTCTTTGATAATTTAGGTAACACGATTA

[0604] CATCGGTTACAACAATGCTGGCAGCTAATCTACCGAAAGTTATACAACTTTTTACAACAGTCTTAACC

[0605] AGTATTCTCGGGAAAATAACAGAAGCTATTCCAAGCATCGTAACCGCGTTATCTAGTTTAATTACGTT

[0606] AATTGTTGGTGCGATCGTTGCCAATTTGCCAGTCTTAATTGAAGCGTCAACACAAATTATTACTACGT

[0607] TGATTCAGGGGATTACAACAGTCTTACCAATGTTGATAGAAGTTGGTTTGAGCTTATTAATGACTTTA

[0608] GTTAATGCGATTGTCACCGCCTTGCCAACAATTACAACTGCAGCGATTAATATCATCACTACATTAGT

[0609] GACAGCTTTTGTCACAGCGTTACCAATGCTAGTTACAGCAGGTGTTTCAATTATCACAGCCTTAGTCA

[0610] ATGCATTTGTTACGATGTTACCGTTGATTTTGACTGCTGGTTTACAAATTTTGATGGCATTAATCACTG

[0611] GGATTATGACGATTTTACCTCAGTTAATTCAATCAGCGCTGACGATTATTCTAGCGTTAGTGACAGCG

[0612] TTGATAGGTGCCTTACCACAGATTATCAGCGCAGGTGTCAAATTGTTAATGGCGTTAATTCAAGGAAT

[0613] TATTTCGATTTTACCAACCTTAGTTGCGGCAGCTATTAACTTAATTTTGACATTGGTAAATGCCTTAAT

[0614] TGGTGCCTTGCCACAAATCATCAGCGCAGGCGTCAAATTGCTAATGGCTTTGATCCAAGGGATTATT

[0615] TCAATTTTACCGCAACTGGTTACTGCAGCAATTACGCTAATTACCGCTTTAATGGGTGCGTTAATCAA

[0616] TGCGTTGCCACAGTTGTTAAGTGCTGGGATTCAACTGATTCAAGCCTTAATTAATGGTGTACTCAGTC

[0617] TATTGGGTGCCTTGCTGTCCGCAGCAGGAACATTAATCTCACAAATGATCACGAAGATTGGTTCTTAT

[0618] TTTGGTCAACTGTTAGCTTCGGGCGGACAGTTAGTTGAAAATATCAAAAATGGGGTTACCAATGCAG

[0619] CCGATCAGGTAAAAACTGCCATTGGTTCTGTAATTGAAGGTGCTTGGCAAGCAATCCAAGGTTGGTT

[0620] TTCAAAATTCACCGATGCCGGTGCGAATATTGTCGGTATGATTGCTGATGGAATTACAGGCGCAATT

[0621] GGAAAAGCCAAAGAAGCAATCGATGGGGTCGTCAGTAAAATTCGTAACTTTTTACCATTTTCACCAGC

[0622] AAAAGAAGGTCCCTTATCTGATTTGCATAAATTGAATTTCGGCGGCACGATTGCCACGGGGATTTATG

[0623] CAGGCGAAACAGCCGTTAGTAGAGCAATGGCTTCTATTTTAGATTTACCGCTGTTAAATGATTTTGCC

[0624] TTGGACTTAGCTGGTCGAGGAAACTTCACGGCAACGATTGACCATCGTTTAGAAAATGATGCATACA

[0625] ATCGACCATTATTTGTGACAGTAGAGTCAACGTTAGATGGAAAAGTTGTCGCAGCAACTACGGCGCC TTATTTAGCAACAGAGTTACAACGACAACAAGTGAAACAAAATAACCGCTTAGGAAGGAGAGGATAA

[0626] CATGTATAAATTTGTTGATACCAATCAAGCAACTCATTCAACGCCTCTTCCTTCAGAAGCGTTGAATTT

[0627] TAACGGCCAATTTTTAGAAAAAGTCATCCCTGGTTATCAAACATTATCAGTTTCAGGACGAGAATTAG

[0628] TTCCAAGCGAAATTGAAAGCTATCAATTAGGGATTCGTGATGGCAAACGTCACGTTTATGCGCGAATT

[0629] CCAGAACGAGAATTAACAGTCAAATATCGCCTTTCAGCTGTGAATAATGAAGCATTTCGAGATGCATT

[0630] TAATCATTTAAACGTTGCTTTGTTTACGGAAAAAGACGTTTCTATTTGGTTTAACGATGAACCGGAAAT

[0631] GCTGTGGTTTGGCAGTAAGTCTTCAGTGAGTGATGTACCCGAAGGTGTTAACCAAGTAACAGGCACC

[0632] TTTACTTTATTGCTTTCTGATCCGTATAAATACACACGGAGTGATGCGACTAGTGTGATGTGGGGTTC

[0633] GCCAACCATTACATTTCAAGCGAATTACTTAATGGGGAATACAGGCTCAGGTGCAGTTGATTTTCCAA

[0634] TTTTAATTGAAGGCGGGGCTTATTGGGGATCAACCATGATTACCTTTCAAAATCGGGCCTACACGAT

[0635] GGGGGATTTAGGCAAAGAAGTTCGGCCAATTGAAATTTATCCTACGGTCGAAGGATTAAAAGTCAAA

[0636] CCGACCATTATTTTAACAGGAACCGGACGTGGTGTTTGGATTAAAACACGGAACGATACAATTAACTT

[0637] AGGAGACTTTGATCGTTCGGAAATTATCATTGATACTGAAAATTTTTATCTGACAAAAAATGGTGCACC

[0638] GATGATTCGACCAATGAACGATTTTTATCTATATCCCAATGAACCGCTGTATATTCAAGCCAAAGATA

[0639] GCGACTTCCGCTTGACGATTCGCTATCCTAACCGATTTGTGTAGGAGGGTGATTAAATGTTAATGGC

[0640] GCTGGATTTGAAAAGAACATATACGGCAATTTTGGACAATGCCTATCAAGTCAGTTATGAAAAAATAG

[0641] AGAACAAAATTGGCAGTTTAGATTTTACCATGCCACTAGATGATCCTAAAAATGAATTTATTGCAGAAA

[0642] TGCAATGGGTGGAACTAACCGACAATGAGAATGAATATATTGGTTTATATCGCGTGATGCCAACCAC

[0643] AATTAAGAAAGATGCGAACAATAATCAAATTCACTACTCTGCCACAGAAGCATTATGTACCTTAGGCG

[0644] ATACTGTCCTATTTGGTTGTCACGAAATTAAAAACAAAACAACGAAAGAGGCCATTCAATTTCTATTGA

[0645] ATAAACAAAAAACAAAGCATTGGGTCCTAAAAAAATGTGATTTTTCAAGGAAATTAACCTATAAATGGG

[0646] AGAATGAAAACGGGCTAGTCGAGCCTTTATTTAGCATCCCAGCCGATTTCGAAGAGGAATATCTTTG

[0647] GCAATGGAATACAGAGGTCTATCCTTTTGAACTTTCATTAGTCAAACCGCCAACAGAACCAGTTGCGC

[0648] GAATTCAAGAAGGTTACAACATGCAAGGATTTGAAATAGAACGTAATCCCAAGATGCTAATCAATCGG

[0649] ATTTATCCATTAGGTTCAGGCGAAGGTGTTAACAAAGTCAATATTCGCTCGGTCAATCAAGGGGTTCC

[0650] GTATTTAGAGAACAAGGCCGCAATTGACCGCTATGGTTTATTGGAGTCAATTTGGGTGGAACAGCGT

[0651] TTTTCTGATCCCAAGGCATTAAAGGAAAATGCTTTGCGAATGTTAGAAGAATGGACCAAGCCACAAGT

[0652] TTCTTGGGTAGTGACTGCAGCTGATTTAATTAAATTAACAGATCAACCTTTGGCAATCGATCGTTTGC

[0653] GGTTGGGCACGGTTATCATGATTAATACGAATGAATTTGGGAGTGTCAACCTTCGTATTAAAAAAGAA

[0654] AGCAAAAAAGATGTCTTTGGTGCCCCCCAAGACATTCAGCTAGAGTTGGGAAACCTGCAAGAAACAA

[0655] TTCATAGTACCATGACAGCTTTCAGTCGGAAACAAGAGATTAACGAAACTTACGCACAAGGGGCGAC

[0656] GACACTTTTAAATCGTTCAATACAAGTAGAACTTAACAAGACACAGCCAGTGGAGCTAAATTTATACTT

[0657] TGACGAGGACATTCTTTATGTAAACACCGCAGAATTAACGTTCAAGGCAACTGCTAAAGGACCTTCG

[0658] CATTCTGTAACGAACATTGATTTGGTAGTGGATGGTAAAAAATTACCCCAACTATCATTGCAACAACA

[0659] ACGGCTAAACATTTTGAGTTATTTACGAAAAACAACAGATGGAAAAATCGAACGCGGCAATCACACG

[0660] CTTCAATTTTTCTCTCATCAGCCACTATGGTTGGATGCTTCGGTCATCTGTCGTGTGTATATTCAATCC

[0661] CAATTGGGTGGCCAGTTTTAATAAAATAATGAAAACTAGAGGAGTGTGACGAAATGTCAGTAGAACAT

[0662] ATTGAAGAATTAGATACCCTGAATCAAGGTCGCCTTAAAATCAATGCAATCTTGGATCAGTCGAATGC

[0663] ATCAGCTGAGAAAGTAGATGCTTACCAAGTCCAGTTAACGAATGGAATTTCTGAAGCGAAAAACATC

[0664] GCAGATGAAGCTGGCAAAGAAGCCGTACAAATTGCCACTGATGCAGGCAATCAAGCAAATGAAACA GCCAACCAAGCGATGAACAATGCCAAAACAGCCATCACGATTGCAGGAAATGCAGTTTCAACGGCAA

[0665] ATAATAATAAACAAGAATTTGATACTTTACGAAATGATTTTGATCAATTAGTAGCAGAAGCGGGTGATA

[0666] GTAATCCAGAAATTGTCCAAGCACGCACAGATACACAAGGCATCAAACAAGCTACCTTAGCGAATCG

[0667] TCTTCAAATTGATTTGAATGACCGTATGACAAAAGCAGACGGTATTTCTTTATTGGCTAAGCCAACCA

[0668] CTGTCAAATTGAAGTTAGACTTTAACGGTAAAACGGCCGGCAATACAGCCACCAATGCAAACAGTTAT

[0669] TCCACTGATTTTACGGCTAAAATTCTTAAGAAGCCAACAGACGTTTGGGAGGAAGTTTCCCAAGCGG

[0670] ACTACAATAAAATGGCCAGCCGTGATGATGAGGGCGTGAAAACAGGTTCCACCCAAAGCGGTGTGA

[0671] TTCCGCAACAATTAGCGGCCTTCAATCTCGTTGAAGCCGCTAAAAAATTAATTCCACAAATGTTTGAA

[0672] ACAGTCACAACTGACGAGGCGGTGGCATTTATTCGCCAGAACGTTCAATTTTTTACGATTAATCAACG

[0673] TGTGAAAGCCGCTGCGCCCAATAATCAAACGATTAAAATCGCTACGTATTTACCAACTACGGATAATT

[0674] GGGTAACTCAAATCCAAGAATCAGCAAAAGAGTTCAGCGATTTTTCAATTCAAATCAATGATCAGAAT

[0675] TTTATCACAGATGAAGGTTTCATTTATTTAATGAGCTATACAGATTCATCGAATGGGGTAACGCCAGC

[0676] TAGCTTAGAAGTTGATTACGTGGGGCTTCATATTGGTCTGTCTGTTGATGCCCAAGCGGTTTTAGCG

[0677] AAGAGTGGTTTTGTTCAAGCAGAGCAACTCAATACCCATGTGGAAAATCAGGATAATCCGCACCAAG

[0678] TAACCGCTGAACAAGTGGGGCTAGGCAATGTAGAAAATTATGGCTTCGCATCAGACAGCGAAGCAGT

[0679] CGCGGGAACTTTAACGAGTAAATATATGCACCCGAAAAACGTTGCGGAAGCGATTAAAGGTCAAGCT

[0680] GTGACACAAACAGGTGATCAAGAGATTGCTGGGGTGAAGAATTTTGTAACTATGCCAACCGTCAATG

[0681] GTGTGCCTTTGGAATCTTCTAGAATGGCCATTTATGAAGCTAGTGGAGTCGGTGAAGTCGAGGCAAA

[0682] ATATCAGGCGGCCTTTAATAAGGATAATATGAAATTTGTATTAATTAGGGTGGGAAATCGTGTCGATG

[0683] CATTTGTAAGATGTAATTTGAGTGATCCAACGAAATTGAATAACCACATGCCTAAAGTATTTAACATAC

[0684] CAACTGGGTACAAAATGTCCTCAAAAATAAGTGCTAGTGTCTGGAATATTCCGCTCTCAGTTGCACCT

[0685] TACGTTTTTCCGTATCCTAATTGTAATGCACTATATGAAATTGGAAATCAAGGGATAATTTTTGCCTCA

[0686] AGCAGAGCCGGAAATGTTTACCTCCAAGGAAGTTGGTACACGGACGATCCGTTTCCGACAAAATAAC

[0687] AAGCTAGTTTAGGAGACTTTTTATGGAACGTTATCTCAACACAATAACAATGCTTTTAAGCATTTTCGG

[0688] TGGGATTGTCGTACGTTTATTAGGCGGATTAGATCAATTGTTGGATGTCTTCCTCTTTTTAATTATTGT

[0689] CGATTTCATCACAGGTTGGATTAAGGCAATCGCCACAAAAGAATTGTCCAGTCGGATTGGTATGCTC

[0690] GGAATTGCGAAAAAAGTGACGATGTTATTTGTGGTTGCTGTAGCGGTTCGTGTTGAAAAAGTTGTGG

[0691] GGAACAATTTGCCAATTCGGGAAATGGTTCTGATTTTTTACATTGCGAACGAAGGACTTTCTTTTTTTG

[0692] AAAACATTGCGACCTTTATTCCTATGCCGAAAAAGTTAAAAGAGTTATTTATTCAGTTAAAAAATAAAG

[0693] ATGATTAAGTAGAAGTGGTCGGGACAAACGTAGAACTTTCGGCTGATTGCCGAAGAAATTACTTCTG

[0694] TCCCGCCATTTATCTGCAGGTTTAAGCCATGAAAGGGAAGTTATTTTGACTTTCCTTTCATGGCTTTTT

[0695] TTAGAAAGGAGTATGCTATGTTTAAAAAATTAATGATTCAACTTGCTTTAGTGATTGGCTTAAGTTTAA

[0696] CGATTCCGATGACGGCTTGCGCTTACACCATCGAAGCGGATCCAATCAACTTTACTTATTTTACAGGC

[0697] TCTTCAAGCAATGAATTAATTGTTTTACATGAATCAGGAAACGAGCGGAACCTAGGACCACACAGTTT

[0698] AGACAATGAAGTGGCCTATATGAAACGAAATTGGTCAAATGCTTATGTCTCATATTTTGTCGGATCTG

[0699] GTGGACGAGTGAAACAATTAGCTCCTGCTGGCCAAATTCAATATGGCGCAGGTTCTTTAGCTAATCA

[0700] AAAAGCCTATGCGCAAATCGAATTGGCTCGCACAAATAATGCGGCGACGTTTAAAAAAGATTATGCT

[0701] GCCTATGTTAATTTGGCCCGTGATTTGGCTCAGAACATTGGTGCTGATTTTTCGCTGGACGATGGAA

[0702] CAGGTTATGGAATAGTCACTCATGATTGGATTACAAAAAATTGGTGGGGAGATCATACAGATCCTTAT

[0703] GGTTATTTAGCGCGTTGGGGGATTAGTAAAGCGCAGTTGGCACAAGATTTACAAACGGGCGTTTCTG AAACAGGTGAGACTGTCATTATTCAGCCTGGTAAACCTAATGCGCCAAAATATCAAGTAGGACAAGC AATTCGTTTCACTTCAATCTATCCAACACCAGATGCTTTAATCAATGAACATCTATCAGCAGAGGCACT TTGGACACAAGTAGGAACAATTACAGCGAAATTACCCGACCGACAAAACCTTTACCGTATTGAAAATA GCGGACATTTGTTAGGTTATGTGAACGACGGCGACATTGCTGAACTTTGGCGCCCGCAAAAGAAGAA ATCATTTCTAATTAGTGTGAACGAAGGCATTGTTTTAAGAGCAGGACAACCTAGCCTGTCAGCCCCAA TTTATGGTATTTGGCCTAAAAATACTCGCTTTTATTACGATGCGTTTTATATTGCAGATGGGTATGTTT TTATTGGTGGAACAGATACAACAGGCGCGAGAATTTATTTGCCAATCGGACCAAACGATGGCAACGC ACAGAATACATGGGGATCATTTACTAGCTAA

[0704] A nucleic acid molecule or set of nucleic acid molecules encoding the amino acid sequence of a polypeptide described herein may be prepared by a variety of methods known in the art. These methods include, but are not limited to, oligonucleotide-mediated (or site-directed) mutagenesis, PCR mutagenesis, ligation, and overlap extension PCR. A nucleic acid molecule encoding a polypeptide described herein may be obtained using standard techniques, e.g., gene synthesis. Alternatively, a nucleic acid molecule encoding an efagin, or a portion thereof, may be mutated to include specific amino acid substitutions using standard techniques in the art, e.g., QuikChange™ mutagenesis. Nucleic acid molecules can be synthesized using a nucleotide synthesizer or PCR techniques.

[0705] 2. Vectors

[0706] In some embodiments, the nucleic acid or set of nucleic acids encoding an efagin, or a portion thereof, is heterologous to the host cell. A nucleic acid sequence encoding a polypeptide described herein may be inserted into a vector or set of vectors capable of replicating and expressing the nucleic acid molecule in prokaryotic or eukaryotic host cells. Many vectors are available in the art and can be used for the purpose of the invention. Each vector may include various components that may be adjusted and optimized for compatibility with the particular host cell. For example, the vector components may include, but are not limited to, an origin of replication, a selection marker gene, a promoter, a ribosome binding site, a signal sequence, the nucleic acid sequence encoding protein of interest, and a transcription termination sequence.

[0707] In some embodiments, the nucleic acid or set of nucleic acids encoding an efagin, or a portion thereof, is endogenous to the host cell (e.g., the host cell is an enterococcus species). For example, the nucleic acid or set of nucleic acids encoding an efagin, or a portion thereof, may be found in the genome of the host cell. In another example, the nucleic acid or set of nucleic acids encoding an efagin or a portion thereof is endogenous to the host cell and is inserted into a vector or set of vectors in which the nucleic acid or set of nucleic acids are operably linked to a heterologous regulatory expression unit (e.g., a heterologous promoter).

[0708] 3. Host cells

[0709] In some embodiments, bacterial cells may be used as host cells for the invention. For example, host cell expression systems could include any harmless commensal species that naturally occurs in the healthy human microbiome, such as a live cell (e.g., engineered probiotic) therapy.

[0710] In some embodiments, E. coli cells may be used as host cells for the invention. Examples of E. coli strains include, but are not limited to, E. coli 294 (ATCC®31 ,446), E. coli A 1776 (ATCC®31 ,537, E. coli BL21 (DE3) (ATCC® BAA- 1025), and E. coli RV308 (ATCC®31 ,608).

[0711] In some embodiments, mammalian cells may be used as host cells for the invention. Examples of mammalian cell types include, but are not limited to, human embryonic kidney (HEK) (e.g., HEK293, HEK 293F), Chinese hamster ovary (CHO), HeLa, COS, PC3, Vero, MC3T3, NSO, Sp2 / 0, VERY, BHK, MDCK, W138, BT483, Hs578T, HTB2, BT20, T47D, NSO, CRL7O3O, and HsS78Bst cells.

[0712] In addition, plant cells and insect cells may also be used as host cells for the invention.

[0713] Different host cells have characteristic and specific mechanisms for the posttranslational processing and modification of protein products (e.g., glycosylation). Appropriate cell lines or host systems may be chosen to ensure the correct modification and processing of the polypeptide expressed. The above-described expression vectors may be introduced into appropriate host cells using conventional techniques in the art, e.g., transformation, transfection, electroporation, calcium phosphate precipitation, and direct microinjection. Once the vectors are introduced into host cells for protein production, host cells are cultured in conventional nutrient media modified as appropriate for inducing promoters, selecting transformants, or amplifying the genes encoding the desired sequences. Methods for expression of therapeutic proteins are known in the art, see, for example, Paulina Baibas, Argelia Lorence (eds.) Recombinant Gene Expression: Reviews and Protocols (Methods in Molecular Biology), Humana Press; 2nd ed. 2004 and Vladimir Voynov and Justin A. Caravella (eds.) Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology) Humana Press; 2nd ed. 2012.

[0714] 4. Protein production, recovery, and purification

[0715] Host cells used to produce the polypeptides described herein may be grown in media known in the art and suitable for culturing of the selected host cells. Examples of suitable media for bacterial host cells include Luria broth (LB) plus necessary supplements, such as a selection agent, e.g., ampicillin. Examples of suitable media for mammalian host cells include Minimal Essential Medium (MEM), Dulbecco’s Modified Eagle’s Medium (DMEM), Expi293™ Expression Medium, DMEM with supplemented fetal bovine serum (FBS), and RPMI-1640. Host cells are cultured at suitable temperatures, such as from about 20 °C to about 39 °C, e.g., from 25 °C to about 37 °C, preferably 37 °C, and CO2 levels, such as 5 to 10%. The pH of the medium is generally from about 6.8 to 7.4, e.g., 7.0, depending mainly on the host organism. If an inducible promoter is used in the expression vector of the invention, protein expression is induced under conditions suitable for the activation of the promoter.

[0716] In some embodiments, depending on the expression vector and the host cells used, the expressed protein may be secreted from the host cells (e.g., mammalian host cells) into the cell culture media. Protein recovery may involve filtering the cell culture media to remove cell debris. The proteins may be further purified. A polypeptide described herein may be purified by any method known in the art of protein purification, for example, by chromatography (e.g., ion exchange, affinity, and size-exclusion column chromatography), centrifugation, differential solubility, or by any other standard technique for the purification of proteins.

[0717] In other embodiments, host cells may be disrupted, e.g., by osmotic shock, sonication, or lysis, to recover the expressed protein. Once the cells are disrupted, cell debris may be removed by centrifugation or filtration. In some instances, a polypeptide can be conjugated to marker sequences, such as a peptide to facilitate purification. An example of a marker amino acid sequence is a hexa-histidine peptide (His- tag), which binds to nickel-functionalized agarose affinity column with micromolar affinity. Other peptide tags useful for purification include, but are not limited to, the hemagglutinin “HA” tag, which corresponds to an epitope derived from influenza hemagglutinin protein (Wilson et al., Cell 37:767, 1984).

[0718] Alternatively, the polypeptides described herein can be produced by the cells of a subject (e.g., a human), e.g., in the context of gene therapy, by administering a vector (such as a viral vector (e.g., a retroviral vector, adenoviral vector, poxviral vector (e.g., vaccinia viral vector, such as Modified Vaccinia Ankara (MVA)), adeno-associated viral vector, and alphaviral vector)) containing a nucleic acid molecule encoding the polypeptide. The vector, once inside a cell of the subject (e.g., by transformation, transfection, electroporation, calcium phosphate precipitation, direct microinjection, infection, etc.) will promote expression of the polypeptide, which is then secreted from the cell. If treatment of a disease or disorder is the desired outcome, no further action may be required. If collection of the protein is desired, blood may be collected from the subject and the protein purified from the blood by methods known in the art.

[0719] 5. Efagin isolation from enterococcal strains

[0720] Efagins may be purified from any enterococcal strain as described herein, for example, an enterococcal strain grown according to standard methods. Typically, efagin-expressing enterococcal cultures are grown in a nutritional medium such as Brain Heart Infusion, and the bacteria are removed by centrifugation. The supernatant culture liquid is then filtered through 0.22 pm or 0.45 pm porosity filters to remove any remaining bacteria. The efagins then can be precipitated from the culture fluid by the addition of polyethylene glycol and NaCI, and the efagins then collected by centrifugation. They may be further purified by ion exchange chromatography taking advantage of the low isoelectric point of the major protein of the tail shaft, and then by size exclusion chromatography taking advantage of their large size compared to individual proteins and other solutes.

[0721] Exemplary enterococcus strains from which an efagin may be purified include, but are not limited to, Enterococcus alcedinis, Enterococcus alishanensis, Enterococcus aquimarinus, Enterococcus asini, Candidatus Enterococcus avicola, Enterococcus avium, Enterococcus bovis, Enterococcus bulliens, Enterococcus burkinafasonensis, Enterococcus caccae, Enterococcus camelliae, Enterococcus canintestini, Enterococcus canis, Enterococcus casseliflavus, Enterococcus cecorum, Enterococcus cloacae, Enterococcus coli, Enterococcus columbae, Enterococcus crotali, Enterococcus devriesei, Enterococcus diestrammenae, Enterococcus dispar, Enterococcus dongliensis, Enterococcus durans, Enterococcus entomosocium, Enterococcus eurekensis, Enterococcus faecalis, Enterococcus faecium, Enterococcus flavescens, Enterococcus florum, Enterococcus gallinarum, Enterococcus gilvus, Enterococcus haemoperoxidus, Enterococcus hawaiiensis, Enterococcus hermanniensis, Enterococcus hirae, Enterococcus hulanensis, Enterococcus innesii, Enterococcus italicus, Enterococcus lacertideformus, Enterococcus lactis, Enterococcus larvae, Enterococcus lemanii, Enterococcus malodoratus, Enterococcus massiliensis, Enterococcus mediterraneensis, Enterococcus montenegrensis, Enterococcus moraviensis, Enterococcus mundtii, Enterococcus nangangensis, Enterococcus olivae, Enterococcus pallens, Enterococcus pernyi, Enterococcus phoeniculicola, Enterococcus pingfangensis, Enterococcus plantarum, Enterococcus porcinus, Enterococcus pseudoavium, Enterococcus quebecensis, Enterococcus raffinosus, Enterococcus ratti, Enterococcus rattus, Enterococcus rivorum, Enterococcus rotai, Enterococcus saccharolyticus, Enterococcus saccharominimus, Enterococcus saigonensis, Enterococcus sanguinicola, Enterococcus seriolicida, Enterococcus silesiacus, Enterococcus solitarius, Enterococcus songbeiensis, Enterococcus spodopteracolus, Candidatus Enterococcus stercoravium, Candidatus Enterococcus stercoripullorum, Enterococcus sulfureus, Enterococcus termitis, Enterococcus thailandicus, Enterococcus timonensis, Enterococcus ureasiticus, Enterococcus ureilyticus, Enterococcus viikkiensis, Enterococcus villorum, Enterococcus wangshanyuanii, Enterococcus xiangfangensis, or Enterococcus xinjiangensis. In some embodiments, the efagin may be isolated from an Enterococcus faecalis strain.

[0722] The following examples are provided to illustrate, not limit the invention.

[0723] EXAMPLES

[0724] Example 1. Genetic organization of the efagin cluster

[0725] A gene cluster was identified in the first complete Enterococcus genome sequence in 2003, that of Enterococcus faecalis strain V583, which was annotated as “Phage 2”. Subsequent studies by others showed that this “phage” appeared to be non-functional, likely owing to its small size and potentially missing functions. However, the present inventors observed that this gene cluster was highly conserved, with well-maintained reading frames, in every E. faecalis genome sequenced since. This implied that, although it may not encode an inducible phage, there was an important function associated with this gene cluster that was central to the biology of this species.

[0726] This Phage 2 cluster, which the present inventors have termed “efagin”, is found in all E. faecalis, but no other species, and it always occurs in a fixed position in the chromosome. A schematic representation of the genetic organization of the efagin region in the E. faecalis genome is shown in FIG. 1 . Gene function prediction was performed by combining information from Prokka, Pgam, Kegg TmHMM, and Phyre2. The efagin genes (highlighted in blue) span a 14.6 kb region from EF1276 to EF1293. Notably, this region does not encode capsid proteins or proteins involved in phage genome excision and replication. Its genomic neighborhood (grey) is highly conserved. The receptor-binding protein, highlighted in dark blue, exhibits variations within the efagin cluster.

[0727] In summary, the genes for efagin are clustered in the chromosome of E. faecalis and are highly conserved, indicating that it is under selection and is important to the biology of that species. Among all enterococcal genomes examined, efagins appear to be unique to the species E. faecalis. Within E. faecalis they are encoded as part of its core genome and are in every strain. Example 2. Efagin activity is specific to certain E. faecalis target strains

[0728] By Transmission Electron Microscopy (TEM), efagins appear to be phage tail-like particles of 160 nm in contour length that lack a phage head and capsid structure (FIG. 2).

[0729] To demonstrate the role of efagin in target cell killing, an isogenic E. faecalis strain OG1 RF was constructed in which the efagin gene cluster was deleted. The parental strain (OG1 RF) and isogenic deletant were compared for the ability to inhibit the growth of other E. faecalis strains. The results shown in FIGS. 3A-B demonstrate the role of the efagin in target cell killing.

[0730] For example, in FIG. 3A, efagins from E. faecalis strain OG1 RF contained within a polyethylene glycol (PEG) precipitate showed activity against some E. faecalis strains seeded in soft agar (left; X98, FA2-2, ATCC 19433, DIV0223b), but not other strains (right; OG1 RF, B653, V583, and T3). An identical control PEG precipitate made from OG1 RF cultures in which the efagin operon was deleted exhibit no killing activity, proving that the killing activity is ascribable to the efagin itself.

[0731] As another example, in FIG. 3B, each strain is an E. faecalis species that differs in its cell wall carbohydrate. Cell wall carbohydrates vary within a species, likely because they are recognized by phages. As a result, there are a number (approximately 20) different cell wall carbohydrate synthesis operons in different E. faecalis strains. This makes the efagin somewhat strain specific as opposed to species specific. As multidrug resistant hospital strains are lineages that often differ in this operon from commensal strains, this specificity could be of therapeutic value by sparing comparatively harmless commensal E. faecalis. This is likely to be true for other enterococcal species as well (see, for example, Palmer et al., mBio. 2012 Mar 1 ;3(1):e00318-11 ).

[0732] The results in FIGS. 3A-B further demonstrate that the efagin encoded by one strain of E. faecalis is not effective against the producing strain itself, but is effective against other strains of E. faecalis. Their antibacterial property targets other strains of E. faecalis, but not the producing strain (see, e.g., test of OG1 RF efagin on an OG1 RF lawn in FIG. 3A). Unlike the bona fide phage M13, diluted efagins do not self replicate and yield plaques (FIG. 3C).

[0733] In summary, efagins are responsible for target cell killing, and have strain-specific activity.

[0734] The Examples below show that the strain-specific activity of efagin is possible because of specific structural variations in efagins, along with variations in the operons that encode its binding target, enterococcal polysaccharide antigen (EPA). Efagin activity also extends to other enterococcal species and non-enterococcal species possessing the efagin target within their cell wall.

[0735] Example 3. Structural classes of efagin exist in different E. faecalis strain genomes

[0736] Insights into the mechanism of efagin-mediated bacterial killing were derived from examination of polymorphisms within the efagin operons in different strains. The genes for efagin are clustered in the chromosome of E. faecalis and are very highly conserved, indicating that it is under selection and is important to E. faecalis biology (FIG. 1 ). However, the 3’ end of one of the genes in the operon, gen EF1291 (shown as “receptor binding protein” in FIG. 1 ), shows a high level of polymorphism. Closer examination of the predicted translation product of this gene showed a distinct pattern of variation indicating three different protein structural classes (FIGS. 4A-C). Grouping of efagin sequences by variation in the EF1291 C-terminal region (efagin positions 12,800 to 13,134 of 2,774 E. faecalis), using hierarchical clustering, is shown in FIG. 4A. Variants (including SNPs and insertions / deletions) were identified by comparison arbitrarily to an E. faecalis V583 prototype. Efagin types A-l are as follows: type A: B3751 ; type B: T9; type C: FA2-2; type D: COM1 ; type E: COM7; type F: B653; type G: 0G1 RF; type H: T1 ; and type I: VET195. Efagin types A-l fall into structurally related groups as shown in FIG. 4B. Efagins within the same structural groups show similar, overlapping or identical target cell specificities. Efagin types A-l show a non-comprehensive list of the amino acid variations that occur in the receptor binding domain. They can be grouped as shown in FIG. 4B based on similar patterns of variation occurring near the C-terminus, and this grouping appears to coincide with target cell specificity determination. Group 1 efagin includes efagin type F (B653), type G (OG1 RF), and type H (T1 ); group 2 efagin includes efagin type C (FA2-2), type D (COM1 ), and type I (VET195); and group 3 efagin includes efagin type A (B3751 ), type B (T9), and type E (COM7).

[0737] The present inventors have found that there are many variations that result from mutations leading to amino acid changes. Many of these changes appear not to affect structure or function. Although they may possess non-consequential variations elsewhere, there was a discernable pattern to mutations resulting in common amino acids in the C-terminal half of the EF1291 protein that could be clustered into groups. These changes do have functional consequences that result in altering the specificity for target cells that have different variations of EPA in their cell walls.

[0738] A comparative analysis of C-terminal variation in EF1291 derived from examination of 3,069 diverse E. faecalis genomes is shown in FIGS. 4B-C, revealing the existence of three main structural groups that correlate with variation in target cell selectivity. Structural variation in efagin particles nearly exclusively clusters at the C-terminal end of EF1291 (nucleotide positions 12,800 to 13,134 in the E. faecalis V583 genome), which encodes a protein inferred to be a phage tail fiber protein.

[0739] In summary, the efagin genes are highly conserved except the C-terminus of EF1291 , and efagin types are defined by this variable region. At least 3 structural classes of efagin exist in different E. faecalis strain genomes.

[0740] The Examples below show that these genetic variations within EF1291 , typified by that encoded within group representative strains E. faecalis OG1 RF, Com1 and T9 (FIG. 4C), are directly linked to the variability in efagin target cell spectrum of activity, and structural modeling indicates that this results from binding alternative cell wall carbohydrates.

[0741] Example 4. The evolution of EPA and EF1291 are correlated

[0742] This unique pattern of variation in the C-terminus of EF1291 , a protein inferred informatically to be a potential tail fiber protein, was hypothesized to specify a variable binding target on the cell surface of E. faecalis and possibly other species targeted for killing.

[0743] Absent detectable primary sequence identity, Phyre2 was used to identify similarity with known protein structural motifs. A high confidence hit was found to the receptor-binding domain of Listeria phage PSA protein gp15, shown to recognize wall teichoic acid (WTA) and confer host target cell specificity. Alignment results of EF1291 to Listeria gp15 are shown in FIG. 5A. 26% of the length of EF1291 matched with gp15 (C-terminal end of the gene). Since the enterococcal polysaccharide antigen (EPA) is a form of WTA, the correlation between the presence of EPA genes and EF1291 SNPs was examined using a whole genome phylogeny of 237 distinct E. faecalis that spans the species' diversity. Results are shown in FIG. 5B. Correspondence between variation in EF1291 (blue) and in EPA genes (red), indicates that the evolution of these genes may be correlated.

[0744] The predicted structure of EF1291 and the structure of Listeria gp15 are shown in FIGS. 5C and 5D, respectively. Variable regions of EF1291 (highlighted in FIG. 5C) correspond to WTA binding sites in the head domain of Listeria gp15 (FIG. 5D). The presence / absence of four EPA genes correlates with EF1291 SNP patterns.

[0745] The E. faecalis EPA gene cluster is shown in FIGS. 6A-B, consisting of a conserved core set of genes (grey) upstream of a group of variable genes (pink). A representative schema of the enterococci cell wall is provided, highlighting the EPA (FIG. 6A) and diagram of the EPA operon based on Guerardel et al. mBio. 2020; 11 (2): e00277-20 (FIG. 6B).

[0746] In summary, the predicted structure of the C-terminal third of EF1291 is similar to Listeria phage gp15, which binds wall teichoic acid and determines the phage's host specificity.

[0747] Example 5. Classes differ in specific variations in the C-terminus of one efagin protein (EF1291) that determines the target cell specificity

[0748] As is described in Examples 3 and 4, above, the pattern of translated EF1291 C-terminal polymorphism was observed to co-vary with polymorphisms in another region of the E. faecalis chromosome that encodes the E. faecalis homolog of cell wall teichoic acid - in enterococci, termed “enterococcal polysaccharide antigen” or EPA. Structural modeling also showed that the polymorphic efagin gene was distantly related to a tail fiber of a Listeria phage that binds a listeria cell wall polysaccharide, providing additional support for a model where efagin binds to a cell surface carbohydrate in targeting susceptible cells.

[0749] Patterns of efagin killing of strains of various EPA types is shown in FIG. 7A, demonstrating that efagins targets strains of heterologous EPA type and not self. The frequency of combinations of Efagin types and EPA types present in 237 distinct E. faecalis genomes from this study is shown in FIG. 7B. Each EPA type has a similar EPA operon structure and is inferred to have a cell wall carbohydrate of similar composition. An EPA type refers to an EPA operon that has a structure that differs from other structures by the presence or absence of a gene. The EPA operons of E. faecalis typically possess a conserved core of genes followed by a group of variable genes. EPA types are grouped by similarities in the variable gene content. Variations tested are represented by the various strains indicated, with strains grouped together in the same box being those that share similarity in the variable genes of their EPA operons.

[0750] To further confirm that the polymorphic C-terminus of a putative efagin tail fiber protein EF1291 specified the type of EPA carbohydrate of different E. faecalis strains targeted by efagin, this domain from the efagin gene EF1291 of OG1 RF (efagin Group 1 ; FIGS. 4B-C) was swapped with that of strain E. faecal is T9 (efagin Group 3; FIGS. 4B-C). As predicted, the resulting chimeric form, now of an OG1 RF efagin with the strain T9 targeting domain of EF1291 exhibited the targeting pattern of the T9 donor of the variable domain and, as predicted, targeted different E. faecalis EPA types (FIG. 8).

[0751] The target of efagin was further confirmed to be EPA by selecting for spontaneous efagin resistant mutant strains (of strains of several different EPA types). Genome sequencing of resistant mutants showed that they possessed specific mutations, including base changes and deletions, that alter the function of genes in the respective EPA operons of the resistant mutants.

[0752] These results demonstrate that efagin cell killing agents target a cell wall polysaccharide termed the extracellular polysaccharide antigen (EPA) which is the homolog of cell wall teichoic acids in other gram positive species. The structural classes of efagin differ in specific variations in the C-terminus of one efagin protein (EF1291 ) that determines the target cell specificity.

[0753] The fact that efagins target non-homologous strains suggests they play a natural role in defense. That they target the cell wall polysaccharide of other E. faecalis strains suggests that they defend from or possibly promote ecological displacement by competing E. faecalis lineages. That the activity extends beyond E. faecalis io other enterococcal species suggests that this ecological competition may be broader than only between other E. faecalis strains. This discovery provides insights into novel mechanisms of bacterial defense and killing.

[0754] In summary, these results provide a new EPA-targeting antimicrobial agent encoded within the genomes of all E. faecalis strains examined, which was initially misidentified as a defective phage in the original E. faecalis genome annotation. The activity correlates highly with EPA type of target cell. The presence of efagins in the chromosome may explain in part the ability of E. faecalis to be a successful generalist that allows it to establish itself in new hosts including hospitalized patients. This element likely affects ecological succession and strain displacement, and may also serve as a phage abortive infection mechanism. Given that the antimicrobial activity is strongly related to the EPA type of the target cell, these entities have therapeutic value for the selective elimination of specific pathogenic strains of enterococci based on wall teichoic acid / EPA type, sparing commensal strains, and as agents that target other pathogenic bacteria possessing similar cell wall carbohydrates as well.

[0755] OTHER EMBODIMENTS

[0756] While the invention has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations of the invention following, in general, the principles of the invention and including such departures from the present disclosure come within known or customary practice within the art to which the invention pertains and may be applied to the essential features hereinbefore set forth.

[0757] All publications, patents, and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0758] Some embodiments of the technology described herein can be defined according to any of the following numbered embodiments:

[0759] E1 . An isolated efagin antibacterial particle, or a portion thereof, that specifically binds a bacterial cell wall polysaccharide. E2. The efagin, or a portion thereof, of E1 , wherein the efagin, or a portion thereof, specifically binds an enterococcal polysaccharide antigen (EPA).

[0760] E3. The efagin, or a portion thereof, of E2, wherein the efagin, or a portion thereof, specifically binds an Enterococcus faecalis EPA.

[0761] E4. The efagin, or a portion thereof, of E1 , wherein the efagin, or a portion thereof, specifically binds a non-enterococcus cell wall polysaccharide.

[0762] E5. The efagin, or a portion thereof, of any one of E1 -E4, wherein the efagin, or a portion thereof, comprises a receptor binding protein.

[0763] E6. The efagin, or a portion thereof, of E5, wherein the receptor binding protein is a group 1 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 7, 16, and 25.

[0764] E7. The efagin, or a portion thereof, of E5, wherein the receptor binding protein is a group 2 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 34, 43, 52, and 61.

[0765] E8. The efagin, or a portion thereof, of E5, wherein the receptor binding protein is a group 3 receptor binding protein having sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to any one of SEQ ID NOs: 70, 79, and 88.

[0766] E9. The efagin, or a portion thereof, of any one of E5-E8, wherein the receptor binding protein specifically binds to the cell wall polysaccharide (for example, EPA).

[0767] E10. The efagin, or a portion thereof, of any one of E5-E9, wherein the efagin, or a portion thereof, comprises a baseplate / distal tail protein, an endopeptidase tail, a receptor binding protein, a holin, and an amidase.

[0768] E11 . The efagin, or a portion thereof, of E10, further comprising one or more of a tail completion protein, a major tail protein, a hypothetical protein, and a tail tape measure protein.

[0769] E12. The efagin, or a portion thereof, of any one of E1 -E11 , wherein the efagin, or a portion thereof, is a chimeric efagin, or a portion thereof.

[0770] E13. A composition comprising the efagin, or a portion thereof, of any one of E1 -E12.

[0771] E14. The composition of E13, comprising group 1 and group 2 efagins, or a portion thereof.

[0772] E15. The composition of E13, comprising group 1 and group 3 efagins, or a portion thereof.

[0773] E16. The composition of E13, comprising group 2 and group 3 efagins, or a portion thereof.

[0774] E17. The composition of E13, comprising group 1 , group 2, and group 3 efagins, or a portion thereof.

[0775] E18. The composition of any one of E13-E17, wherein the composition comprises a chimeric efagin, or a portion thereof.

[0776] E19. The composition of any one of E13-E18, further comprising an antibiotic.

[0777] E20. An isolated bacterium comprising an efagin, or a portion thereof.

[0778] E21 . The isolated bacterium of E20, wherein the bacterium is a member of the genus Enterococcus.

[0779] E22. The isolated bacterium of E20 or E21 , wherein the bacterium is commensal.

[0780] E23. The isolated bacterium of E20 or E21 , wherein the bacterium is an opportunistic pathogen. E24. A method of treating a subject having a bacterial infection, comprising administering an effective amount of the efagin, or a portion thereof, of any one of E1 -E12, the composition of any one of E13-E19, or the isolated bacterium of any one of E20-E23.

[0781] E25. The method of E24, wherein the bacterial infection is a wound infection, a urinary tract infection, bacteremia, or infective endocarditis.

[0782] E26. The method of E24 or E25, wherein the bacterial infection is an enterococcal infection.

[0783] E27. The method of E26, wherein the enterococcal infection is an Enterococcus faecalis infection.

[0784] E28. The method of any one of E24-E27, wherein the bacterial infection is a nosocomial infection.

[0785] E29. The method of any one of E24-E28, wherein the efagin, or a portion thereof, the composition, or the isolated bacterium is administered to the skin, the upper respiratory tract, the oral cavity, or the vagina of the subject.

[0786] E30. A nucleic acid molecule or set of nucleic acid molecules encoding the efagin, or a portion thereof, of any one of E1 -E12.

[0787] E31 . The nucleic acid molecule or set of nucleic acid molecules of E30, wherein the nucleic acid molecule has sequence identity (e.g., at least 75% (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity) to the sequence of any one of SEQ ID NOs: 91 -93.

[0788] E32. An expression vector or set of expression vectors comprising the nucleic acid molecule or set of nucleic acid molecules of E30 or E31 .

[0789] E33. A host cell comprising (i) the nucleic acid molecule or set of nucleic acid molecules of E30 or E31 or (ii) the expression vector or set of expression vectors of E32.

[0790] E34. The host cell of E33, wherein the host cell is a bacterial cell, yeast cell, mammalian cell, insect cell, or plant cell.

[0791] E35. The host cell of E34, wherein the bacterial cell is a commensal species.

[0792] E36. A method of producing an efagin, or a portion thereof, comprising:

[0793] (a) culturing the host cell of any one of E33-E35 under conditions where the efagin, or a portion thereof, is expressed; and

[0794] (b) isolating the efagin, or a portion thereof, expressed in (a) from the host cell culture, thereby producing the efagin, or a portion thereof.

[0795] Other embodiments are within the following claims.

Claims

CLAIMS1 . An isolated efagin antibacterial particle, or a portion thereof, that specifically binds a bacterial cell wall polysaccharide.

2. The efagin, or a portion thereof, of claim 1 , wherein the efagin, or a portion thereof, specifically binds an enterococcal polysaccharide antigen (EPA).

3. The efagin, or a portion thereof, of claim 2, wherein the efagin, or a portion thereof, specifically binds an Enterococcus faecalis EPA.

4. The efagin, or a portion thereof, of claim 1 , wherein the efagin, or a portion thereof, specifically binds a non-enterococcus cell wall polysaccharide.

5. The efagin, or a portion thereof, of any one of claims 1 -4, wherein the efagin, or a portion thereof, comprises a receptor binding protein.

6. The efagin, or a portion thereof, of claim 5, wherein the receptor binding protein is a group 1 receptor binding protein having sequence identity to any one of SEQ ID NOs: 7, 16, and 25.

7. The efagin, or a portion thereof, of claim 5, wherein the receptor binding protein is a group 2 receptor binding protein having sequence identity to any one of SEQ ID NOs: 34, 43, 52, and 61 .

8. The efagin, or a portion thereof, of claim 5, wherein the receptor binding protein is a group 3 receptor binding protein having sequence identity to any one of SEQ ID NOs: 70, 79, and 88.

9. The efagin, or a portion thereof, of claim 5, wherein the receptor binding protein specifically binds to the cell wall polysaccharide.

10. The efagin, or a portion thereof, of claim 5, wherein the efagin, or a portion thereof, comprises a baseplate / distal tail protein, an endopeptidase tail, a receptor binding protein, a holin, and an amidase.11 . The efagin, or a portion thereof, of claim 10, further comprising one or more of a tail completion protein, a major tail protein, a hypothetical protein, and a tail tape measure protein.

12. The efagin, or a portion thereof, of claim 1 , wherein the efagin, or a portion thereof, is a chimeric efagin, or a portion thereof.

13. A composition comprising the efagin, or a portion thereof, of claim 1 .

14. The composition of claim 13, comprising group 1 and group 2 efagins, or a portion thereof.

15. The composition of claim 13, comprising group 1 and group 3 efagins, or a portion thereof.

16. The composition of claim 13, comprising group 2 and group 3 efagins, or a portion thereof.

17. The composition of claim 13, comprising group 1 , group 2, and group 3 efagins, or a portion thereof.

18. The composition of claim 13, wherein the composition comprises a chimeric efagin, or a portion thereof.

19. The composition of claim 13, further comprising an antibiotic.

20. An isolated bacterium comprising an efagin, or a portion thereof.21 . The isolated bacterium of claim 20, wherein the bacterium is a member of the genus Enterococcus.

22. The isolated bacterium of claim 20, wherein the bacterium is commensal.

23. The isolated bacterium of claim 20, wherein the bacterium is an opportunistic pathogen.

24. A method of treating a subject having a bacterial infection, comprising administering an effective amount of the efagin, or a portion thereof, of claim 1 , the composition of claim 13, or the isolated bacterium of claim 20.

25. The method of claim 24, wherein the bacterial infection is a wound infection, a urinary tract infection, bacteremia, or infective endocarditis.

26. The method of claim 24, wherein the bacterial infection is an enterococcal infection.

27. The method of claim 26, wherein the enterococcal infection is an Enterococcus faecalis infection.

28. The method of claim 24, wherein the bacterial infection is a nosocomial infection.

29. The method of claim 24, wherein the efagin, or a portion thereof, the composition, or the isolated bacterium is administered to the skin, the upper respiratory tract, the oral cavity, or the vagina of the subject.

30. A nucleic acid molecule or set of nucleic acid molecules encoding the efagin, or a portion thereof, of claim 1 .31 . The nucleic acid molecule or set of nucleic acid molecules of claim 30, wherein the nucleic acid molecule has sequence identity to the sequence of any one of SEQ ID NOs: 91 -93.

32. An expression vector or set of expression vectors comprising the nucleic acid molecule or set of nucleic acid molecules of claim 30.

33. A host cell comprising (i) the nucleic acid molecule or set of nucleic acid molecules of claim 30 or (ii) the expression vector or set of expression vectors of claim 32.

34. The host cell of claim 33, wherein the host cell is a bacterial cell, yeast cell, mammalian cell, insect cell, or plant cell.

35. The host cell of claim 34, wherein the bacterial cell is a commensal species.

36. A method of producing an efagin, or a portion thereof, comprising:(a) culturing the host cell of claim 33 under conditions where the efagin, or a portion thereof, is expressed; and(b) isolating the efagin, or a portion thereof, expressed in (a) from the host cell culture, thereby producing the efagin, or a portion thereof.

Citation Information

Patent Citations

  • Enterocins and methods of using the same

    US20210309703A1