Phylogenetic marker of erwiniaceae and use thereof
The use of concatenated protein sequences from the shikimate pathway and D-arabitol 4-dehydrogenase markers resolves the phylogenetic challenges of Erwiniaceae, enabling accurate classification and phylogenetic analysis of Erwiniaceae bacteria.
Patent Information
- Application Number
- PCT/KR2025/004560
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-04
- Publication Date
- 2025-10-16
AI Technical Summary
The phylogeny of Erwiniaceae remains unresolved due to high evolutionary rates of molecular markers and variation in phylogenetic trees, making it difficult to classify and distinguish genera and species within the family.
A marker composition using the concatenated sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC, along with D-arabitol 4-dehydrogenase, is employed for phylogenetic classification and taxonomic phylogeny of Erwiniaceae bacteria.
This approach provides a reliable method for classifying Erwiniaceae bacteria, resolving phylogenetic ambiguities and distinguishing genera and species accurately, as demonstrated by the ME phylogenetic tree and phenotypic characteristics.
Smart Images

Figure KR2025004560_16102025_PF_FP_ABST
Abstract
Description
Phylogenetic markers of the Erwiniaceae and their uses
[0001] The present invention relates to phylogenetic markers of the Erwiniaceae family and uses thereof.
[0002] Erwiniaceae is a new family separated from Enterobacteriaceae, consisting of Gram-negative, rod-shaped, non-spore-forming microorganisms (International Journal of Systematic and Evolutionary Microbiology66, 5575-5599 (2016)). Since Erwinia amylovora was first proposed as a distinct taxon, nine genera have been named: Erwinia, Tatumella, Pantoea, Buchnera, Wigglesworthia, Phaseolibacter, Mixta, Winslowiella, and Duffyella. Members of Erwiniaceae have formed symbiotic or pathogenic relationships with plants, insects, humans, and animals in the natural environment and on the International Space Station. Most species have genome sizes ranging from 2.75 to 5.88 Mb, with a GC content of 44.4% to 58%. B. Genetic variations in intracellular symbiont strains such as Candidatus aphidicola (0.64 Mb and 25.2% GC content), W. glossinidia (0.71 Mb and 23.8% GC content), Candidatus Pantoea carbekii (1.15 Mb and 30.6% GC content), Candidatus Erwinia haradaeae (1.09 Mb and 30.6% GC content), and E. dacicola (2.70 Mb and 52.8% GC content) suggest the diversity of Erwiniaceae species.
[0003] Members of the Erwiniaceae family are agriculturally and clinically important. Despite pathogenic variation, researchers have explored symbiotic relationships between pests and diseases to continuously manage agriculture and human health through biological control. For example, P. agglomerans has been considered a promising biological control agent for various bacterial and fungal plant diseases (plant and clinical strains. BMC Microbiology9, 1-18 (2009)), and the low-molecular-weight lipopolysaccharide (IP-PA1) produced by this strain has a potent analgesic effect and is effective in treating various human and animal disorders (Annals of Agricultural and Environmental Medicine23 (2016)). A rapidly growing number of Pantoea isolates harboring diverse genes for industrial and agricultural applications in diverse hosts and environments, including soil and water, are producing antibiotics such as pantocins, herbicolins, microcins, and phenazines, some of which can be used to treat the fire blight pathogen, E. Effective against amylovora (Genome Announcements1, e00904-00913 (2013)).
[0004] Recent genetic approaches have provided new insights into the genetic diversity and phylogenetic analysis of Erwiniaceae, particularly members of the genera Erwinia, Mixta, Pantoea, and Tatumella, using a large number of previously uncharacterized or unknown genes. Numerous genome-based studies have investigated the evolution of Erwiniaceae, which possess diverse ecological, pathophysiological, and genetic properties. However, the phylogeny of all genera within Erwiniaceae remains unresolved due to the high evolutionary rate of molecular markers and the variation in their phylogenetic trees.
[0005] The purpose of the present invention is to provide a marker composition for phylogenetic classification of Erwiniaceae.
[0006] In addition, the purpose of the present invention is to provide a method for determining the taxonomic phylogeny of Erwiniaceae bacteria.
[0007] In order to solve the above problem, the present invention provides a marker composition for phylogenetic classification of Erwiniaceae, including a preparation for detecting the connecting sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC.
[0008] In addition, the present invention provides a method for determining the taxonomic phylogeny of Erwiniaceae bacteria.
[0009] In the present invention, in order to overcome the problem of difficulty in classifying the Erwiniaceae through conventional gene or protein sequence analysis or in forming a phylogenetic tree showing evolutionary relationships, the common metabolic pathways of the type strain species included in all genera belonging to the Erwiniaceae and the closely related Enterobacteriaceae and Pectobacteriaceae species were analyzed, and it was confirmed that the protein sequence analysis connecting the five AroBEKAC protein sequences involved in the shikimate pathway connected to the quinone metabolite synthesis of the electron transport chain involved in aromatic amino acids and respiration and the zeaxanthin pigment synthesis pathway can be used as a new marker to classify at least the genera and species of Erwiniaceae bacteria and to explain the evolutionary relationships.
[0010] Figure 1 is a diagram comparing the colony colors of the strains:
[0011] 1: PD-1;
[0012] 2:M. tenebrionisKCTC 72449 T ;
[0013] 3:E. rhaponticiKACC 22740 T ;
[0014] 4:P. agglomeransKACC 15275 T ;
[0015] 5:P. ananatisKACC 22739 T ; and
[0016] 6:P. stewartiiKACC 22737 T .
[0017] Figure 2 is a diagram showing the UPGMA phylogenetic tree of the 16S rRNA gene of the PD-1 strain:
[0018] Dotted line a: 98.7% similarity threshold for species; and
[0019] Dotted line b: 95% similarity threshold for the genus.
[0020] Figure 3 is a diagram showing the UPGMA phylogenetic tree of Erwiniaceae species based on the concatenated nucleotide sequences of the atpD, gyrB, infB, and rpoB genes.
[0021] Figure 4 is a diagram showing an ANI (Average nucleotide index) matrix constructed through pairwise comparison after analyzing the genomes of the PD-1 strain and the reference strain of the Erwiniaceae species.
[0022] Figure 5 is a diagram showing a neighbor-joining tree constructed using ANI-derived distance measurements between the genomes of the PD-1 strain and a reference strain of the Erwiniaceae species.
[0023] Figure 6 is a diagram showing an average amino acid (AAI) matrix constructed through pairwise comparison of protein sequence sets of a PD-1 strain and a reference strain of the Erwiniaceae species.
[0024] Figure 7 is a diagram showing a neighbor-joining tree constructed using AAI-derived distance measurements of the PD-1 strain and the reference strain of the Erwiniaceae species.
[0025] Figure 8 is a diagram showing the AAI-based Neighbor-joining tree of the PD-1 strain and the reference strain of the Erwiniaceae species and the mutation of chemotaxonomic marker candidates:
[0026] Half-filled and white boxes: partial or complete deletions in the genome;
[0027] P: plasmid-encoded crt operon included in yellow box;
[0028] Aro: chorismate synthase AroBEKAC (3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC);
[0029] Isp: octaprenyl diphosphate synthase IspAB ((2E,6E)-farnesyl diphosphate synthase IspA and octaprenyl diphosphate synthase IspB);
[0030] Ubi: ubiquinone 합성 효소 UbiABCDEFGHIJKX(4-hydroxybenzoate octaprenyltransferase UbiA, ubiquinone biosynthesis regulatory protein kinase UbiB, chorismate lyase UbiC, 4-hydroxy-3-polyprenylbenzoate decarboxylase UbiD, ubiquinone / menaquinone biosynthesis C-methylase UbiE, 3-demethoxyubiquinol 3-hydroxylase UbiF, 3-demethylubiquinol 3-O-methyltransferase UbiG, 2-octaprenyl-6-methoxyphenyl hydroxylase UbiH, 2-octaprenylphenol hydroxylase UbiI, accessory factor UbiJ, accessory factor UbiK, anaerobic accessory factor UbiT, ubiquinone biosynthesis protein UbiU, ubiquinone biosynthesis protein UbiV 및 UbiX family flavin prenyltransferase);
[0031] Men: menaquinone synthase MenABCDEHI (1,4-dihydroxy-2-naphthoate polyprenyltransferase MenA, 1,4-dihydroxy-2-naphthoyl-CoA synthase MenB, O-succinylbenzoate synthase MenC, 2-succinyl-5-enolpyruvyl-6-hydroxy-3-cyclohexene-1-carboxylic-acid synthase MenD, O-succinylbenzoate--CoA ligase MenE, 2-succinyl-6-hydroxy-2,4-cyclohexadiene-1-carboxylate synthase MenH and 1,4-dihydroxy-2-naphthoyl-CoA hydrolase MenI); and
[0032] Crt: cartenoid synthase CrtE-(Idi)-Crt CrtZ).
[0033] Figure 9 is a diagram showing the ME phylogenetic tree constructed based on the concatenated sequences of AroBEKAC proteins:
[0034] Parentheses: Number of amino acid residues (aa) in the concatenated AroBEKAC sequence.
[0035] Figure 10 is a diagram showing the results of detection, purification, and quantification of major respiratory quinones derived from Escherichia coli K-12 MG1655.
[0036] Figure 11 is a diagram showing the analysis of respiratory quinones under aerobic and anaerobic culture conditions of the PD-1 strain and the reference strain of the Erwiniaceae species.
[0037] Figure 12 is a diagram analyzing respiratory quinones under aerobic and anaerobic culture conditions of Enterobacteriaceae and Erwiniaceae strains.
[0038] Figure 13 is a diagram confirming the sugar metabolism characteristics of the PD-1 strain:
[0039] A: Growth curve of PD-1 strain in M9 minimal medium containing D-glucose, D-xylitol, D-xylose, D-arabitol, D-arabinose and L-arabinose; and
[0040] B: ME phylogenetic tree of D-arabitol dehydrogenase (DalD) in PD-1 strains and closely related strains.
[0041] Hereinafter, the present invention will be described in detail with reference to the attached drawings and embodiments thereof. However, the following embodiments are provided as examples of the present invention. If a detailed description of a technology or configuration well known to those skilled in the art is judged to unnecessarily obscure the gist of the present invention, such detailed description may be omitted, and the present invention is not limited thereby. The present invention is capable of various modifications and applications within the scope of the following claims and equivalents interpreted therefrom.
[0042] Additionally, the terminology used in this specification is intended to appropriately express preferred embodiments of the present invention, and may vary depending on the intent of the user or operator, or the customs of the field to which the present invention pertains. Therefore, the definitions of these terms should be determined based on the contents throughout this specification. Throughout this specification, when a part is said to "include" a certain component, unless specifically stated otherwise, this does not mean that other components are excluded, but rather that other components may be included.
[0043] Unless otherwise defined, all technical terms used in this invention have the same meaning as commonly understood by those skilled in the art. While preferred methods and samples are described herein, similar or equivalent methods are also included within the scope of the present invention. The contents of all publications cited as references herein are incorporated herein by reference.
[0044] In one aspect, the present invention relates to a marker composition for phylogenetic classification of Erwiniaceae, comprising an agent that detects concatenated sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA and chorismate synthase AroC.
[0045] In one embodiment, the composition may include an agent that detects a linked sequence of a base sequence encoding 3-dehydroquinate synthase AroB, a base sequence encoding shikimate dehydrogenase AroE, a base sequence encoding shikimate kinase AroK, a base sequence encoding 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and a base sequence encoding chorismate synthase AroC.
[0046] In one embodiment, the 3-dehydroquinate synthase AroB can comprise the amino acid sequence of SEQ ID NO: 1, the shikimate dehydrogenase AroE can comprise the amino acid sequence of SEQ ID NO: 2 or 3, the shikimate kinase AroK can comprise the amino acid sequence of SEQ ID NO: 4 or 5, the 3-phosphoshikimate 1-carboxyvinyltransferase AroA can comprise the amino acid sequence of SEQ ID NO: 6, and the chorismate synthase AroC can comprise the amino acid sequence of SEQ ID NO: 7.
[0047] In one embodiment, the linking sequence may be a protein sequence or a base sequence encoding the protein sequence sequentially including or linked to five protein sequences consisting of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA and chorismate synthase AroC, and may include a base sequence encoding the amino acid sequence of SEQ ID NO: 8 to which WP_173632419.1, WP_173632496.1, WP_173632418.1, WP_173633610.1 and WP_173634887.1 are linked.
[0048] In one embodiment, the composition may further comprise an agent that detects a base sequence encoding D-arabitol 4-dehydrogenases.
[0049] In one embodiment, the D-arabitol 4-dehydratase may comprise the amino acid sequence of SEQ ID NO: 9 or 10.
[0050] In one embodiment, a protein sequence or a base sequence encoding the protein sequence concatenated with the five protein sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC is a sequence specifically conserved in the Erwiniaceae family, and thus can be used as a taxonomic marker for Erwiniaceae bacteria.
[0051] In one embodiment, the agent for detecting a base sequence encoding a gene or protein may be a primer pair, a probe, or a primer pair and a probe that specifically recognize a nucleic acid sequence of the gene, a nucleic acid sequence complementary to the nucleic acid sequence, a fragment of the nucleic acid sequence and the complementary sequence, and the detection thereof may be performed by one or more methods selected from the group consisting of polymerase chain reaction, real-time RT-PCR, reverse transcription polymerase chain reaction, competitive RT-PCR, nuclease protection assay (RNase, S1 nuclease assay), in situ hybridization, nucleic acid microarray, northern blot, DNA chip, multiplex PCR, or ddPCR.
[0052] In one embodiment, the agent for detecting a protein may be a peptide, antibody or immunologically active fragment thereof, aptamer, avidity multimer or peptidomimetics that specifically binds to the full-length or fragment thereof of the protein, and the detection may be performed by one or more methods selected from the group consisting of Western blot, enzyme linked immunosorbent assay (ELISA), immunofluorescence, radioimmunoassay (RIA), radioimmunodiffusion, immunoelectrophoresis, tissue immunostaining, immunoprecipitation assay, complement fixation assay, FACS, mass spectrometry or protein microarray.
[0053] The term “detection” or “measurement” used in the present invention means analyzing the presence or absence of a detected or measured object and whether there is a change.
[0054] The term "primer" as used herein refers to a short nucleic acid sequence having a short free 3-terminal hydroxyl group, which can form base pairs with a complementary template and serves as a starting point for copying the template strand. The primer can initiate DNA synthesis in the presence of a reagent for polymerization (i.e., DNA polymerase or reverse transcriptase) and four different nucleoside triphosphates in an appropriate buffer and temperature.
[0055] In the present invention, the term "probe" refers to a nucleic acid fragment such as RNA or DNA, ranging from a few bases to several hundred bases in length, capable of forming a specific binding with a base sequence encoding a target gene or protein, and is labeled so as to enable detection of the presence or absence of a base sequence encoding a specific gene or protein. The probe can be produced in the form of an oligonucleotide probe, a single-stranded DNA probe, a double-stranded DNA probe, an RNA probe, etc.
[0056] The primers or probes of the present invention can be chemically synthesized using the phosphoramidite solid support method or other well-known methods. These nucleic acid sequences can also be modified using many means known in the art. Non-limiting examples of such modifications include methylation, capping, substitution with one or more homologs of a natural nucleotide, and modifications between nucleotides, such as modification with uncharged linkers (e.g., methyl phosphonate, phosphotriester, phosphoramidate, carbamate, etc.) or charged linkers (e.g., phosphorothioate, phosphorodithioate, etc.).
[0057] As used herein, the term "antibody" is a term known in the art and refers to a specific protein molecule directed against an antigenic site. For the purposes of the present invention, the antibody refers to an antibody that specifically binds to the marker protein of the present invention, and the method for producing the antibody can be produced using a well-known method. This also includes a partial peptide that can be produced from the protein. The form of the antibody of the present invention is not particularly limited, and a polyclonal antibody, a monoclonal antibody, or a part thereof that has antigen binding property is also included in the antibody of the present invention, and all immunoglobulin antibodies are included. Furthermore, the antibody of the present invention also includes special antibodies such as humanized antibodies.
[0058] In one aspect, the present invention relates to a method for determining the taxonomic phylogeny of Erwiniaceae bacteria, comprising the steps of: 1) analyzing the linking sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC in a candidate bacterium; and 2) determining that the candidate bacterium is a bacterium of the Erwiniaceae family if the candidate bacterium includes a base sequence encoding the linking sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC of a control bacterium of the Erwiniaceae family.
[0059] In one embodiment, the linking sequence may comprise the amino acid sequence of SEQ ID NO: 8.
[0060] In one embodiment, after step 2), a step of determining that the candidate bacterium is Paramixta manurensis of the Erwiniaceae family may be additionally included if the candidate bacterium includes a base sequence encoding D-arabitol 4-dehydratase.
[0061] The present invention is described in more detail through the following examples. However, the following examples are intended only to concretize the content of the present invention and are not intended to limit the present invention.
[0062] Example 1. Isolation of a novel strain of PD-1
[0063] Compost waste of mushroom (P. eryngii) culture substrate (oak sawdust 70-80% (w / w) and wheat / rice 20-30% (w / w)) from Hwasun-gun, Jeollanam-do, Republic of Korea (34.99939 N, 126.91040 E) was collected in a 50 mL Falcon tube, and 1 g of wet weight was homogenized in 100 mL of sterilized water at 150 rpm for 24 h at 25°C. Centrifugation was performed at 4,200 × g for 2 min, and 200 μL of the supernatant was streaked on LB (Luria Bertani) agar plates (BD, Sparks MD, USA) and cultured at 37°C for 48 h. At this time, E. coli K-12 MG1655, S. Typhimurium ATCC 14028, and E. pyrinus KCTC 2590, E. teleogrylli SCU-B244, E. rhapontici KACC 22740 (Fig. 1, 3), M. tenebrionis KCTC 72449 (Fig. 1, 2), Pantoea agglomerans KACC 15275 (Fig. 1, 4), P. ananatis KACC 22739 (Fig. 1, 5), and P. stewartii KACC 22737 (Fig. 1, 6) were used. Afterwards, a novel strain PD-1 was isolated, which formed non-adhesive, round, convex yellow colonies (lighter yellow than the color of the control strain Pantoea species) with a diameter of 1.5-2.5 mm on an R2A agar plate (Fig. 1, 1). The novel strain PD-1 isolated in this way was cultured in BD TS (tryptic soy) liquid medium and R2A medium to prepare an 80% (v / v) glycerol stock and stored at -80°C.
[0064] Example 2. Identification of the species and genus of PD-1 using conventional methods.
[0065] 2-1. 16S rRNA analysis
[0066] To identify the novel strain PD-1 isolated in Example 1, 16S rRNA gene sequence analysis was performed. Specifically, the PD-1 genome was isolated using the MG Genomic DNA Purification Kit (MGmed, Korea), and the 16S rRNA portion was amplified by PCR using universal primers (27F: 5′-AGAGTTTGATCMTGGCTCAG-3′ and 1492R: 5′-TACGGYTACCTTGTTACGACTT-3′), and analyzed by Sanger sequencing. The 16S rRNA sequence thus derived (GenBank accession number MN197605) was used to calculate 16S rRNA gene pairwise similarity using the EzBioCloud database and tools.
[0067] As a result, the 16S rRNA sequence of PD-1 showed the highest similarity (98.13%) with that of P. vagans LMG 241999 (EF688012), and the UPGMA phylogenetic tree showed that PD-1 formed a separate branch among the clades of E. typographi, P. flectans, and Rosenbergiella species clusters (Fig. 2). However, the large genera Erwinia, Pantoea, and Mixta could not form a cluster with the strains of that type, and the endosymbiotic B. aphidicola and W. glossinidia formed independent clades distant from other outgroup Brenneri species belonging to the Pectobacteriaceae family.
[0068] 2-2. MLSA
[0069] To identify the species and genus of the novel Erwiniaceae strain PD-1 and to improve the phylogeny of Erwiniaceae, multilocus sequence analysis (MLSA) was performed using concatenated sequences of housekeeping genes including atpD, gyrB, infB, and rpoB genes.
[0070] In the MLSA phylogenetic tree, the PD-1 strain was distinguished from the Erwinia-Winslowiella lineage, which was clearly distinguished from the Mixta-Pantoea-Duffyella lineage (Fig. 3). In addition, Fig. 3 shows the differences between the PD-1 strain isolated from mushroom compost and [Pantoea]beijingensisJZB2120001 (= LMG27579) isolated from the fruiting body of P. eryngii. T ) showed a close correlation between them.
[0071] As described above, conventional MLSA showed a high similarity between the type strains and the concatenated sequences, and thus placed the species more accurately than 16S rRNA sequence analysis. However, it was still difficult to distinguish the genus Winslowiella from the genus Erwinia, and the phylogenetic tree constructed based on 16S RNA or MLSA data had the problem of making it difficult to distinguish Erwiniaceae strains at the species and genus levels. Therefore, in order to confirm the phylogeny of the newly isolated PD-1 strain in the present invention, a method other than the conventional method was required.
[0072] Example 3. Sequence analysis of PD-1
[0073] 3-1. Whole Genome Sequencing (WGS) Analysis
[0074] For whole-genome DNA sequencing of PD-1 using the PacBio RS II system (Pacific Biosciences, Menlo Park, CA, USA), high-quality DNA of ≥40 kb was prepared using AMPure PB magnetic beads (Beckman Coulter Inc., Brea, CA, USA). Genomic DNA was quantified using a NanoDrop spectrophotometer and a Qubit fluorometer, and 200 ng / μL of DNA extract was electrophoresed on a field-inversion gel to confirm its quality. A 10-μL DNA library was prepared using the PacBio DNA Template Prep Kit 1.0, and SMRTbell template was annealed using the PacBio DNA / Polymerase Binding Kit P6. C4 chemistry and 240-min movies were performed using the PacBio DNA Sequencing Kit 4.0 and SMRT Cell 8M, and whole genome sequencing of the PD-1 strain was performed using Illumina HiSeq. The raw data were assembled using HGAP3 (Hierarchical Genome Assembly Process), and the corresponding annotation was performed using the RAST (Rapid Annotation using Subsystem Technology) server.
[0075] Whole genome sequencing (WGS) performed on PacBio RSII and Illumina sequencing platforms assembled two contigs from 4,617,009 bp of the circular chromosome (GenBank accession number CP054212.1) and 88,619 bp of the circular plasmid (GenBank accession number CP054213.1). The genome assembly of the PD-1 strain was confirmed to include 4,525 genes containing 4,358 protein-coding sequences using the NCBI Prokaryotic Genome Annotation Pipeline (PGAP). In addition, the whole genome sequence of the PD-1 strain was registered with NCBI under GenBank accession numbers (chromosome: CP054212 and plasmid: CP054213).
[0076] 3-2. DDH (DNA-DNA hybridization) and ANI (average nucleotide identity) analysis
[0077] In silico DNA-DNA hybridization (DDH) and average nucleotide identity (ANI) were used to obtain estimates of the overall similarity between the genome sequences of the PD-1 strain and the Erwiniaceae reference strain. The DDH and ANI values between the entire genomes of the PD-1 strain and the Erwiniaceae reference strain were found to be lower than the standard cutoff values for species identification (DDH: 70% and ANI: 95%) (Figs. 4 and 5 and Table 1). In addition, in the PD-1 strain, approximately 20% of the genome sequence differed from the sequence of the nearest relative (NR) [Pantoea] beijingensis JZB2120001 (Table 1).
[0078] Through this, ANI data confirmed that the PD-1 strain is a new species, but its genus could not be defined through conventional pairwise genome comparison.
[0079] [Correction pursuant to Rule 91 04.06.2025]
[0080]
[0081] * Notes: The ANI-AAI matrix was reconstructed using the ANI and AAI values between the PD-1 strain and the reference strains of Erwiniaceae selected from Figures 4 to 7. Yellow boxes: species clustered at a threshold of 80%.
[0082] Example 4. Reclassification of PD-1 species and genus
[0083] To determine the taxonomic position of the PD-1 strain in the Erwiniaceae, the average amino acid index (AAI) between the PD-1 strain and the protein database of the Erwiniaceae reference strain was determined (Fig. 6). The AAI-based neighbor-joining tree showed a cutoff value of 68% for genus delineation in Erwiniaceae, indicating that the genus is monophyletic (Fig. 7). Based on the analysis tree, the PD-1 strain was identified as a novel species in a new genus between the Winslowella and Mixta clusters of Erwiniaceae, and was named Paramixta manurensis (gen. nov., sp. nov.). It was deposited with the Korea Research Institute of Bioscience and Biotechnology on April 28, 2023, and assigned the accession number "KCTC13848BP."
[0084] Example 5. Marker derivation for phylogenetic classification of Erwiniaceae
[0085] 5-1. Derivation of chemical taxonomic markers for Erwiniaceae species
[0086] To reconstruct the phylogeny of Erwiniaceae, an AAI tree was constructed based on AAI data, and the presence of genetic mutations in enzymes of the ubiquinone (Q) and menaquinone (MK) biosynthetic enzymes, chorismate (Aro), octaprenyl diphosphate (Isp), ubiquinone (Ubi), menaquinone (Men), and cartenoid (Crt) biosynthetic pathways (Table 3), which are candidates for chemotaxonomic markers of Erwiniaceae species, was examined in the strains.
[0087] [Correction pursuant to Rule 91 04.06.2025]
[0088] As a result, five proteins of the shikimate pathway that biosynthesize chorismate, which are commonly conserved without mutation in Erwiniacea: 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC (hereinafter referred to as AroBEKAC), were derived as a set of molecular markers for phylogenetic classification of Erwiniaceae (Fig. 8).
[0089] 5-2. ME phylogenetic tree and evolutionary distance analysis
[0090] A minimum evolution (ME) phylogenetic tree was constructed based on the concatenated sequence of AroBEKAC proteins in PD-1 strains. Specifically, the optimal phylogenetic tree was found to be greater than 90% of 1000 replicate trees in a bootstrap test, and was constructed at a scale of units of the number of substituted amino acids per site calculated using the JTT matrix-based method and the Close-Neighbor-Interchange algorithm at search level 1 in MEGA X. A total of 1587 positions were present in the final dataset, and the number of amino acid residues (aa) in the concatenated AroBEKAC sequence is indicated in parentheses.
[0091] The branching patterns of the ME phylogenetic tree constructed based on the concatenated sequence of the AroBEKAC protein in the PD-1 strain were similar to those shown in the AAI tree of Example 5-1, and the phylogenetic tree using the AroBEKAC protein showed a wider evolutionary distance at a rate of 0.175 amino acid substitutions per site than the MLSA distance of 0.05 nucleotide substitutions per site for the evolutionary distance classified at the genus level using the reference strain of Erwiniaceae (Fig. 9). The phylogenetic tree constructed in this way was consistent with the results of previous studies, thereby confirming that AroBEKAC is a reliable and effective molecular marker for taxonomic and phylogenetic analysis of Erwiniaceae strains.
[0092] Example 6. Phenotypic Characterization of PD-1 Strains
[0093] 6-1. Analysis of respiratory quinone compounds
[0094] To analyze the phenotypic characteristics of PD-1 strains, respiratory quinone compounds, which are chemotaxonomic markers of Erwiniaceae strains, were analyzed by HPLC-UV / mass analysis. At this time, the differences between the PD-1 strain and four Enterobacteriaceae strains (E. coli K-12 MG1655, S. Typhimurium ATCC 14028, Enterobacter pyrinus KCTC 2590T, and Entomohabitans teleogrylli SCU-B244T) and five Erwiniaceae strains (E. rhapontici KACC 22740T, M. tenebrionis KCTC 72449T, P. agglomerans KACC 15275T, P. ananatis KACC 22739T, and P. stewartii KACC 22737T) were compared, and the molar ratio of quinone and menaquinone compounds in the solvent-extracted samples was determined using a standard curve generated for each purified compound. Specifically, the PD-1 strain and reference strains were cultured in TS medium under aerobic or anaerobic conditions until the optical density at 600 nm became 1.0. Anaerobic conditions were maintained in a glove box filled with a 5% H2 / 10% CO2 / 85% N2 mixture, and the culture tubes were tightly sealed and cultured at 37°C for 24 h with shaking at 200 rpm for E. coli and S. Typhimurium, and at 30°C for PD-1 and other strains. After culture, the cells were centrifuged at 4,000 rpm (3,515 × g) for 10 min and the cell pellets were washed with deionized water. After centrifugation at 4,000 rpm for 10 minutes, the wet weight of the collected cells was measured, and 9 times the volume of methanol-petroleum ether (1:1, v / v) was added to extract respiratory quinones.After vortexing and sonication for 10 min, the petroleum ether supernatant was transferred to a new tube, dried in vacuo, and dissolved in 200 μL of 100% ethanol. The HPLC column was an Inertsil ODS-3V (4.6 × 150 mm, 5-μm particle size, GL Sciences, Tokyo, Japan), heated at 53°C, and anhydrous MeOH containing 0.5% formic acid was pumped at a flow rate of 1 mL / min. An internal standard of 100 μg / mL ubiquinone 10 (Q10) was included in the sample analysis, and ESI-MS analysis was operated in positive mode. Mass spectrometry conditions were curtain gas 30 psi, source temperature 500 °C, spray voltage 5500 V, and ion source gas 50 psi. In addition, E. coliMG1655 cells grown aerobically in 1 L TS medium were purified, and the molar concentrations of the isolated quinone compounds (Q8, DMK8, and MK8) were determined by chromatographic separation on a ZORBAX ODS column (9.4 × 250 mm) and a Waters Nova-Pak C18 column (3.9 × 150 mm) using methanol as the eluent (Fig. 10). Ubiquinone 10 (Q10) was used as an exogenous standard at this time.
[0095]
[0096] As a result, the three major quinones detected were ubiquinone (ubiquinone 8, Q8), demethylmenaquinone (demethylmenaquinone 8, DMK8), and menaquinone (menaquinone 8, MK8). The quinone composition of the PD-1 strain showed that the levels of DMK8 and MK8 decreased compared to Q8 during the transition from aerobic to anaerobic conditions (Figs. 11 and 12). These quinone profiles under aerobic and anaerobic conditions were similar to those of M. tenebrionis KCTC 72449. T and E. rhaponticiKCTC 22740 T , but showed a negative correlation with the quinone profiles of four Enterobacteriaceae strains (Table 4). The high Q8 levels of PD-1 strains suggest that PD-1 strains may play an important role in respiration, using oxygen and nitrogen as electron acceptors under aerobic and anaerobic conditions, respectively, like E. coli. As shown in Fig. 8, some strains of the genera Pantoea and Duffyella lost menaquinones, whereas some reference strains of the genera Mixta and Rosenbergiella contained menaquinone and carotenoids biosynthetic genes and operons, and these strains were found to acquire or lose menaquinones at various rates when acquiring the plasmid-associated car operon or vice versa. In other words, a high level of genetic variation was found among members of Erwiniaceae using respiratory quinone and carotenoids markers.
[0097] 6-2. FAME Analysis
[0098] To distinguish the PD-1 strain from its closest type, FAME (Fatty acid methyl ester) analysis and API test were performed and compared with data from previously known Erwiniaceae strains. Specifically, FAME samples from cells grown in TS medium were prepared according to the method described in the International Journal of Systematic and Evolutionary Microbiology43, 162-173 (1993), and analyzed using a gas chromatography (GC)-mass spectrometer (Shimadzu GC-17A) equipped with a Supelco SP-2560 capillary GC column. The analysis conditions were as follows: initial temperature of 100°C, maintained for 5 min, then increased to 240°C at 3.5°C / min, and maintained for 30 min. In addition, FAME components were identified using a FAME standard mixture (Sigma-Aldrich, cat# 1269119).
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107] As a result, the FAME composition of the PD-1 strain was similar to several strains of the genera Mixta and Winslowiella (Pearson correlation coefficient > 0.9), but different from selected reference species of the genera Erwinia and Pantoea (Table 5). In addition, the API 20E test results of the PD-1 strain showed that it was significantly higher than those of closely related strains, such as M. theicola QC88-366. T and [Pantoea]beijingensisJZB2120001 showed high similarity (similarity index > 0.6) with the results of the PD-1 strain (Table 6). All selected reference strains of species closely related to the PD-1 strain produced D-glucose, D-mannitol, melibiose and L-arabinose. When the numerical data obtained from the API 20E and 50CHB / E tests were examined in other Erwinia members, the characteristic of the PD-1 strain to produce acids from glycerol and erythritol was shown to be a distinguishing feature from its closest strain, [Pantoea]beijingensisJZB2120001, and other reference strains of the genus Erwinia (Tables 7 and 8). Compared with the reference strains of the genera Winslowiella and Erwinia, which are closely related to the PD-1 strain, the PD-1 strain was found to be positive for acid production from glycerol and erythritol and negative for esculin hydrolysis (Table 8).
[0108] Based on these phylogenetic and phenotypic characteristics, PD-1 T Strain (KCTC13848BP T ) was proposed as the reference strain for Paramixta maurensisgen. nov., sp. nov.
[0109] 6-3. Analysis of sugar metabolism characteristics
[0110] The glucose metabolism characteristics of the PD-1 strain were investigated, and it was found that it could utilize D-arabitol and glucose as sole carbon and energy sources, but did not grow in M9 minimal medium containing D-xylitol, D-xylose, D-arabinose, and L-arabinose (Fig. 13A). Genome analysis revealed that the PD-1 strain contained two homologous genes encoding D-arabitol 4-dehydrogenases (gene: DalD) (WP_173634172.1 and WP_173633231.1), which are classified as mannitol dehydrogenase family proteins. The genome of the PD-1 strain contains a set of genes encoding D-arabitol 4-dehydrogenases, D-xylulokinase XylB (synonym, AtlK), and ribulose-phosphate 3-epimerase for D-arabitol catabolism via D-xylulose-5-phosphate and D-ribulose-5-phosphate in the pentose phosphate pathway (Table 9). The PD-1 strain has a complete set of genes for glycolysis via the Embden-Meyerhof-Parnas pathway, the pentose phosphate pathway, the TCA cycle, and other pathways involving D / L-fucose, D-galactose, D-mannose, and D-mannitol utilization, excluding L-rhamnose. In addition, the ME phylogenetic tree also showed that the two DalD proteins of the PD-1 strain evolved from different dalD genes in the Erwiniaceae strains, and the different dalD genes on the chromosome showed a gene arrangement between the two clades, and the PD-1 strain and P. alliiLMG 24248 TExcept for the single dalD' gene, most other clades, including T. citrea and T. morbirosei, shared the dalD-xylB operon structure (Fig. 13B). In this context, there is a difference in that the PD-1 strain and the reference strain of the genus Winslowiella containing the dalD-xylB gene can utilize D-arabitol, whereas the [Pantoea] beijingensis JZB2120001 strain lacking the dalD gene cannot (Table 9). Since such genetic variation in DalD has not been determined in many bacteria, including the genera Duffyella, Mixta, Phaseolibacter, Wigglesworthia, and Buchnera, it can be used as an indicator to distinguish species and genera in Erwiniaceae strains. This genetic finding is consistent with the results of substrate utilization tests that characteristically utilized various polyols, including glycerol, erythritol, and D-arabitol.
[0111]
[0112]
Claims
A marker composition for phylogenetic classification of Erwiniaceae, comprising an agent for detecting concatenated sequences of 1.3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC.
2. A marker composition for phylogenetic classification of the Erwiniaceae, comprising an agent for detecting a linked sequence of a base sequence encoding 3-dehydroquinate synthase AroB, a base sequence encoding shikimate dehydrogenase AroE, a base sequence encoding shikimate kinase AroK, a base sequence encoding 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and a base sequence encoding chorismate synthase AroC, according to claim 1.
3. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the 3-dehydroquinate synthase AroB comprises an amino acid sequence of sequence number 1 in the first paragraph.
4. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the shikimate dehydrogenase AroE comprises an amino acid sequence of sequence number 2 or 3 in the first paragraph.
5. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the shikimate kinase AroK comprises an amino acid sequence of SEQ ID NO: 4 or 5 in the first paragraph.
6. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the 3-phosphoshikimate 1-carboxyvinyltransferase AroA comprises an amino acid sequence of sequence number 6 in the first paragraph.
7. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the chorismate synthase AroC comprises an amino acid sequence of sequence number 7 in claim 1.
8. A marker composition for phylogenetic classification of the Erwiniaceae, comprising a preparation for detecting an amino acid sequence of sequence number 8 or a base sequence encoding the same, according to claim 1.
9. A marker composition for phylogenetic classification of the Erwiniaceae family, further comprising a preparation for detecting a base sequence encoding D-arabitol 4-dehydrogenases in accordance with claim 1.
10. A marker composition for phylogenetic classification of the Erwiniaceae, wherein the D-arabitol 4-dehydratase comprises an amino acid sequence of SEQ ID NO: 9 or 10 in paragraph 9. 11.1) Analyzing the linked sequences of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA and chorismate synthase AroC in candidate bacteria; and 2) A method for determining the taxonomic phylogeny of Erwiniaceae bacteria, comprising a step of determining that the candidate bacteria is a bacterium of the Erwiniaceae family if the candidate bacteria includes a base sequence encoding a connecting sequence of 3-dehydroquinate synthase AroB, shikimate dehydrogenase AroE, shikimate kinase AroK, 3-phosphoshikimate 1-carboxyvinyltransferase AroA, and chorismate synthase AroC of a control Erwiniaceae bacterium.
12. A method for determining the taxonomic phylogeny of Erwiniaceae bacteria in claim 11, wherein the connecting sequence comprises an amino acid sequence of sequence number 8.
13. A method for determining the taxonomic phylogeny of bacteria of the Erwiniaceae family, further comprising a step of determining that the candidate bacteria is Paramixta manurensis of the Erwiniaceae family if the candidate bacteria includes a base sequence encoding D-arabitol 4-dehydratase after step 2).
Citation Information
Patent Citations
Nucleic acid molecules for detecting bacteria and phylogenetic units of bacteria
US7202027B1
KR20220168831A