Genome identification, quantification and visualization method for drug-resistant pathogenic bacteria in environmental sample

By using metagenomic assembly technology, the problem of identifying and quantifying drug-resistant pathogens in environmental samples has been solved, enabling precise identification and risk assessment at the genomic scale, and improving the accuracy and visualization of environmental antibiotic resistance risk assessment.

CN122024852APending Publication Date: 2026-05-12EAST CHINA NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EAST CHINA NORMAL UNIV
Filing Date
2026-02-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently identify and quantify drug-resistant pathogens in environmental samples, particularly in complex microbial communities where the host association of antibiotic resistance genes is insufficient. They also lack the ability to conduct genome-scale risk assessment and genetic transmission risk assessment, and lack an integrated analytical workflow.

Method used

Using metagenomic assembly as the core analysis unit, this study employs metagenomic high-throughput sequencing, assembly, binning, quality assessment, open reading frame prediction, database alignment, quantitative analysis, and visualization techniques to identify, quantify, and visualize the genetic structure of drug-resistant pathogens in environmental samples.

Benefits of technology

It enables the accurate identification and quantification of drug-resistant pathogens in environmental samples without relying on microbial isolation and culture, improving the accuracy and interpretability of risk assessment, enhancing the visualization of drug resistance genetic structures, and demonstrating good versatility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024852A_ABST
    Figure CN122024852A_ABST
Patent Text Reader

Abstract

The invention provides a method for identifying, quantifying and visualizing genomes of drug-resistant pathogenic bacteria in environmental samples, which is suitable for identifying dangerous drug-resistant pathogenic bacteria in various environments. The method comprises the following steps: (1) based on a metagenome sequencing big data set, obtaining a high-quality bacterial genome by using genome quality control, assembly and binning technologies; (2) identifying a drug-resistant bacterium genome carrying an antibiotic drug-resistant gene, further annotating a virulence factor, and comparing with an authoritative pathogenic bacterium directory to realize accurate identification of drug-resistant pathogenic bacteria; and (3) mapping the reference genome of the drug-resistant pathogenic bacteria through the metagenome so as to quantify the abundance of the drug-resistant pathogenic bacteria, and further carrying out visual analysis on the drug-resistant pathogenic bacteria mediated drug-resistant genes, related movable genetic elements and other gene neighborhood structures. The method does not depend on isolated culture, realizes full-spectrum analysis of drug-resistant pathogenic bacteria in environmental samples, and provides important technical support for accurate and scientific evaluation of drug resistance risks of environmental bacteria antibiotics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of environmental microbiome and antibiotic resistance risk assessment technology, specifically involving an integrated analysis method based on metagenomic-assembled genomes (MAG) for the identification, abundance quantification, and genome structure visualization of drug-resistant pathogens in environmental samples. This method is applicable to various environmental samples, including water, air, and soil. By correlating antibiotic resistance, pathogenicity, and genetic structure information at the genomic scale, it enables systematic analysis and risk assessment of high-risk drug-resistant pathogens in the environment. Background Technology

[0002] The continued spread of antibiotic resistance (AR) has become a significant threat to global public health. Environmental systems, including river bodies, wastewater treatment plants, farmland soils, and drinking water distribution networks, are widely considered important media for the generation, accumulation, and spread of antibiotic resistance genes (ARGs). In environmental media closely related to human health, the microbial community is complex and diverse in origin. Some microorganisms may simultaneously carry antibiotic resistance genes and pathogenic factors, posing a potential threat to public health once they enter the human body. Therefore, the accurate identification, quantification, and transmission risk assessment of high-risk antibiotic-resistant pathogens in environmental samples are key scientific and technological issues in environmental resistance research and public health surveillance.

[0003] Currently, research on drug-resistant pathogens mainly relies on traditional isolation and culture techniques or short-sequence metagenomic data analysis, both of which suffer from low throughput and poor accuracy, making it difficult to meet the high-throughput and accurate screening needs for high-risk drug-resistant pathogens in complex bacterial communities in environmental samples.

[0004] 1) Insufficient association between ARGs and hosts: Annotating and counting the abundance of ARGs based on sequencing reads or short contigs makes it difficult to reliably associate ARGs with specific microbial hosts at the genome scale. This makes it impossible to identify the "specific microbial individual carrying the drug-resistant gene," thus limiting the accurate identification of drug-resistant risk vectors. 2) Lack of a systematic judgment mechanism for screening drug-resistant pathogens: Focusing solely on the presence or abundance of ARGs makes it difficult to effectively integrate ARG host information with authoritative pathogen lists. It fails to distinguish whether drug-resistant genes originate from low-risk environmental background bacteria or high-risk human pathogens, leading to significant uncertainty in environmental drug resistance risk assessment results. 3) Lack of genome-scale risk quantification methods: Using ARG gene abundance as an evaluation indicator lacks relative abundance quantification methods for the high-risk unit of "pathogen genomes carrying drug-resistant genes," making it difficult to reflect the true exposure level of drug-resistant pathogens in environmental samples. 4) Insufficient capacity for genetic transmission risk assessment: There is a lack of systematic analysis of the genome structure of drug-resistant pathogens, especially the lack of characterization of adjacent mobile genetic elements (MGEs) of ARGs, making it difficult to determine the transmission potential and risk hotspots of drug-resistant genes through horizontal gene transfer mechanisms. 5) Lack of standardized, integrated workflow: A unified methodology has not yet been established that can simultaneously complete genome reconstruction, joint assessment of drug resistance and pathogenicity, quantification of pathogenicity, and visualization of genetic structure. This limits the versatility of this technology in different environmental samples and its practical application in public health surveillance.

[0005] Therefore, there is an urgent need for an analysis method for drug-resistant pathogens in the environment with the genome as the core analytical unit. By simultaneously integrating information on antibiotic resistance, pathogenicity, and genetic structure at the genomic scale, this method can achieve accurate identification, effective quantification, and visual assessment of potential transmission risks of drug-resistant pathogens in environmental samples. Summary of the Invention

[0006] The purpose of this invention is to overcome the technical shortcomings of existing environmental drug resistance research, such as unclear association between drug resistance genes and pathogenic bacteria hosts, lack of genomic-scale quantitative means for drug-resistant pathogens, and difficulty in intuitively assessing the risk of drug resistance genetic transmission, which rely on isolation and culture or gene fragment analysis methods. The invention proposes an integrated analysis method for the identification, quantification, and visualization of genomic structure of drug-resistant pathogens in environmental samples, with metagenomic assembly of the genome as the core analysis unit.

[0007] The method of this invention can systematically correlate antibiotic resistance, pathogenicity and genetic structure information at the same genome scale without relying on microbial isolation and culture, so as to achieve accurate identification, abundance quantification and potential transmission risk analysis of drug-resistant pathogens in environmental samples, thereby improving the accuracy, reliability and interpretability of environmental antibiotic resistance risk assessment results.

[0008] The specific technical solution for achieving the objective of this invention is as follows:

[0009] A method for genome identification, quantification, and visualization of drug-resistant pathogens in environmental samples, comprising the following steps:

[0010] Step 1: Obtain total microbial DNA from environmental samples, perform metagenomic high-throughput sequencing, and obtain high-quality metagenomic data (Clean Reads) after sequence quality control and removal of host sequences;

[0011] Step 2: Use assembly software to assemble the Clean Reads obtained in Step 1 to obtain contigs. Then, use various binning algorithms to cluster the contigs and reconstruct metagenomic assembly genomes (MAGs).

[0012] Step 3: Perform quality assessment and screening on the MAGs obtained in Step 2 to obtain a high-quality set of MAGs for subsequent analysis;

[0013] Step 4: Perform open reading frame (ORF) prediction on the high-quality MAGs screened in Step 3, and align the predicted protein sequences to the Antibiotic Resistance Genes (ARG) database to identify MAGs carrying antibiotic resistance genes and determine them as MAGs that host antibiotic resistance genes.

[0014] Step 5: In the drug resistance gene host MAGs identified in Step 4, the predicted protein sequences are further aligned to the Virulence Factor Genes (VFG) database to screen MAGs that carry both antibiotic resistance genes and virulence factor genes as potential drug-resistant pathogen genomes.

[0015] Step 6: Compare the genomes of the potential drug-resistant pathogens obtained in Step 5 with the authoritative list of pathogens from the World Health Organization (WHO) to confirm the genomes of drug-resistant pathogens in the environmental samples.

[0016] Step 7: Use quantitative analysis software to map the Clean Reads obtained in Step 1 onto the genomes of the drug-resistant pathogens identified in Step 6, and calculate their abundance in environmental samples;

[0017] Step 8: For the drug-resistant pathogen genome in Step 7, extract the genomic regions containing antibiotic resistance genes and / or virulence factor genes, annotate the surrounding mobile genetic elements (MGEs), and generate the corresponding gene neighborhood structure visualization map.

[0018] Furthermore, step 3 specifically includes:

[0019] Step 3-1: Using CheckM software, assess the completeness and contamination of each MAG based on a lineage-specific marker gene set.

[0020] Step 3-2: Select MAGs with a completeness greater than a preset threshold and a contamination level less than a preset threshold as high-quality MAGs for subsequent analysis. The completeness is greater than 70% or 90%, and the contamination level is less than 10%.

[0021] Furthermore, step 4 specifically includes:

[0022] Step 4-1: Use Prodigal to predict open reading frames on the filtered MAGs;

[0023] Step 4-2: Align the predicted protein sequence with specialized antibiotic resistance gene databases such as SARG, CARD, or ResFinder using BLAST or DIAMOND tools;

[0024] Step 4-3: Set the filtering threshold as follows: e value less than or equal to 1×10 -10 The sequence similarity is greater than or equal to 80%, and the coverage is greater than or equal to 75%. Based on this comparison and screening results, MAG carrying antibiotic resistance genes was identified and determined as the host MAG of the resistance gene.

[0025] Furthermore, step 5 specifically includes:

[0026] Step 5-1: The protein sequence of the drug resistance gene host MAG determined in Step 4 is compared with the virulence factor gene database VFDB using BLAST or DIAMOND tools;

[0027] Step 5-2: Set the filtering threshold as follows: e value less than or equal to 1×10 -10 MAGs carrying both antibiotic resistance genes and virulence factor genes were screened for sequence similarity ≥ 70% and coverage ≥ 70% as potential drug-resistant pathogen genomes.

[0028] Furthermore, the authoritative list of pathogens includes the list of key or critical pathogens published by the World Health Organization (WHO).

[0029] Furthermore, step 7 specifically includes:

[0030] Step 7-1: Using CoverM or Salmon quantitative analysis software, map the high-quality metagenomic data Clean Reads obtained in Step 1 onto the genome of the target drug-resistant pathogen MAG identified in Step 6;

[0031] Step 7-2: Calculate the relative abundance of each target MAG and normalize the result to the number of mapped sequences per kilobase per million (RPKM) value. The RPKM calculation formula is as follows:

[0032] RPKM

[0033] In the formula: Mapped Reads is the number of sequences mapped to the target MAG; Total Mapped Reads is the total number of mappings to all MAGs; Length is the genome length of the target MAG (in bases); The RPKM value characterizes the relative abundance level of the target drug-resistant pathogen in the environmental microbial community and can be used to quantitatively analyze the environmental distribution of drug-resistant pathogens.

[0034] In the technical solution provided by this invention, the preferred platform for high-throughput sequencing is Illumina NovaSeq. FastP software is used for quality control and filtering of the raw sequencing data. For the assembled and screened metagenomic genome, Prodigal software is used to predict open reading frames. In the functional gene annotation stage, the conventional functions of genes annotated in the NR database are compared using BLAST or DIAMOND tools, and antibiotic resistance genes and virulence factor genes are accurately identified by comparing with specialized databases such as SARG and VFDB. In the final visualization stage, gene neighborhood maps are drawn using the gggenes package based on the R language environment, and automated scripts are configured to standardize the analysis workflow and generate maps in batches.

[0035] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects:

[0036] (1) Breaking through the dependence on traditional pure culture methods, this invention enables the systematic identification of unculturable or difficult-to-culture drug-resistant pathogens. Traditional research on drug-resistant pathogens mainly relies on pure microbial culture and drug susceptibility testing. These methods suffer from stringent culture conditions, long cycles, low throughput, and the inability to cover a large number of unculturable or difficult-to-culture microorganisms in the environment, leading to significant bias in the identification results of environmental drug-resistant pathogens. This invention is based on metagenomic assembly and genome analysis, enabling the full-spectrum identification of drug-resistant pathogens in environmental samples without relying on microbial isolation and culture. This significantly expands the scope of monitorable targets and improves the comprehensiveness of environmental drug resistance risk assessment.

[0037] (2) Achieving precise association between antibiotic resistance genes and host microorganisms at the genomic scale avoids the problem of "separate analysis of genes and hosts" in traditional metagenomic analysis. Existing metagenomic analyses mostly stop at the level of sequencing reads or spliced ​​fragments, making it difficult to determine the true host origin of antibiotic resistance genes and answer the key question of "which specific microorganisms carry resistance genes". This invention uses metagenomic assembly of the genome as the smallest unit of analysis, annotating antibiotic resistance genes at the scale of complete or near-complete genomes, fundamentally achieving a reliable correspondence between resistance genes and their host microorganisms, and providing a solid foundation for the precise identification of resistance risk vectors.

[0038] (3) By comprehensively evaluating drug resistance and pathogenicity information at the genomic level, the specificity and reliability of drug-resistant pathogen identification are improved. Traditional methods usually only focus on the presence or abundance of antibiotic resistance genes, making it difficult to distinguish whether the resistance genes originate from environmental background bacteria or risky human pathogens, and easily overestimating the risk of environmental drug resistance. This invention integrates information on antibiotic resistance genes, virulence factor genes, and authoritative pathogen lists at the same metagenomic assembly genomic scale to perform multiple determinations of drug-resistant pathogens, effectively reducing the risk of false positive identification and improving the scientificity and reliability of the screening results for environmental drug-resistant pathogens.

[0039] (4) For the first time, the relative abundance of drug-resistant pathogens is quantified at the genome level, improving the accuracy of environmental exposure risk assessment. Existing technologies mostly use the abundance of a single antibiotic resistance gene as a risk assessment indicator, which is difficult to reflect the true exposure level of "individual pathogens carrying drug-resistant genes" in the environment. This invention uses the assembled genome of confirmed drug-resistant pathogen metagenomics as the quantitative object, and through sequencing sequence mapping and normalization calculation, it realizes the genome-level relative abundance quantification of drug-resistant pathogens in environmental samples, thus elevating environmental drug resistance risk assessment from the "gene level" to the "pathogen level".

[0040] (5) To achieve visualization of drug resistance genetic structures and enhance the ability to intuitively interpret the risk of horizontal gene transfer. Traditional metagenomic analysis procedures usually lack systematic analysis of the genetic environment surrounding drug resistance genes, making it difficult to determine the potential transmission capacity of drug resistance genes. This invention, through joint annotation and structural visualization of antibiotic resistance genes, virulence factor genes and their adjacent mobile genetic elements in the genome of drug-resistant pathogens, can intuitively display the characteristics of drug resistance genetic structures and potential hotspots for horizontal gene transfer, providing direct evidence for assessing the risk of drug resistance transmission.

[0041] (6) A standardized and integrated analytical process is constructed, which has good versatility and scalability. Under a unified technical framework, this invention integrates key steps such as metagenomic assembly, genome quality control, joint determination of drug resistance and pathogenicity, quantification of pathogenic bacteria and visualization of genetic structure, forming a standardized and integrated analytical method that is applicable to a variety of environmental sample types and is easy to promote and apply in environmental drug resistance monitoring, public health risk assessment and regulatory practice.

[0042] In summary, this invention can effectively improve the systematicness, accuracy, and visualization of the identification and risk assessment of drug-resistant pathogens in environmental samples, providing a reliable technical means for environmental public health risk management and drug-resistant pollution prevention and control, and has significant practical application value. Attached Figure Description

[0043] Figure 1 is a schematic diagram of the overall process of the present invention;

[0044] Figure 2 is a schematic diagram showing the abundance changes of target drug-resistant pathogens and the specific types of drug-resistant genes they carry in the wastewater treatment plant (influent, activated sludge and effluent) and the receiving river system (upstream, near the discharge outlet and downstream).

[0045] Figure 3 is a visualization map of the gene neighborhood structure constructed for the genomes of two drug-resistant pathogens in an embodiment of the present invention. Detailed Implementation

[0046] The embodiments of the present invention are described in conjunction with the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments.

[0047] See Figure 1 The method for genomic identification, quantification, and visualization of drug-resistant pathogens in environmental samples according to the present invention specifically includes:

[0048] Step 1: Environmental Sample Collection and Microbial DNA Acquisition

[0049] Representative environmental samples were collected, including but not limited to the receiving river system of the wastewater treatment plant. After pretreatment to remove large particulate impurities, the samples were enriched with microbial biomass by filtration or centrifugation, and total microbial DNA was extracted. The concentration and quality of the obtained DNA were measured to meet the requirements of subsequent metagenomic sequencing analysis.

[0050] Step 2: Metagenomic high-throughput sequencing and data preprocessing

[0051] Metagenomic high-throughput sequencing was performed on the total microbial DNA obtained in step 1 to obtain raw sequencing data. The raw sequencing data underwent quality control and filtering to remove low-quality sequences, adapter sequences, and host-derived sequences, resulting in a high-quality metagenomic sequencing dataset.

[0052] Step 3: Metagenome Assembly and Binning

[0053] Based on the high-quality metagenomic sequencing data obtained in step 2, de novo assembly was performed to obtain contiguous sequences; further, based on sequence coverage and nucleotide composition characteristics, the contiguous sequences were binned and reconstructed to obtain multiple metagenomic assembled genomes.

[0054] Step 4: Metagenomic Assembly, Genome Quality Assessment and Screening

[0055] The quality of the metagenomic assembled genomes obtained in step 3 was assessed, and high-quality metagenomic assembled genomes were selected based on genome integrity and contamination indicators, which were then used as the analysis objects for subsequent functional annotation and host association analysis.

[0056] Step 5: Annotation of drug resistance genes and their host associations

[0057] Using the high-quality metagenomic assembly genome selected in step 4 as the analysis unit, open reading frames were predicted, and the predicted protein sequences were compared with the antibiotic resistance gene database to identify the types and subtypes of antibiotic resistance genes carried in the metagenomic assembly genome, thereby determining the host of antibiotic resistance genes at the genomic scale.

[0058] Step 6: Virulence factor annotation and its host association

[0059] The protein sequences of the metagenomic assembly genomes of the drug-resistant gene hosts identified in step 5 are further compared with the virulence factor database to identify the virulence factor genes they carry. Based on the coexistence of antibiotic resistance genes and virulence factor genes in the same metagenomic assembly genome, metagenomic assembly genomes that simultaneously carry antibiotic resistance genes and virulence factor genes are screened as potential drug-resistant pathogen genomes.

[0060] Step 7: Screening and confirmation of pathogens

[0061] The potential drug-resistant pathogens obtained in step 6 were compared with the WHO authoritative list of pathogens to confirm the genomes of drug-resistant pathogens in the environmental samples.

[0062] Step 8: Quantitative analysis of the abundance of drug-resistant pathogens

[0063] Using quantitative analysis software (such as CoverM or Salmon), the high-quality metagenomic data Clean Reads obtained in step 2 were mapped onto the genomes of the drug-resistant pathogens MAG identified in step 7. The relative abundance of each target MAG in the sample was calculated. The relative abundance was normalized to the number of mapped sequences per kilobase per million (RPKM) value, calculated using the following formula:

[0064] RPKM

[0065] Where: Mapped Reads is the number of sequences mapped to the target MAG; Total Mapped Reads is the total number of mappings to all MAGs; Length is the genome length of the target MAG (in bases); the RPKM value characterizes the relative abundance level of the target drug-resistant pathogen in the environmental microbial community.

[0066] Step 9: Visualization of the genome structure of drug-resistant pathogens

[0067] For the genomes of drug-resistant pathogens that have been quantified and pose potential public health risks, genomic regions containing antibiotic resistance genes and / or virulence factor genes are extracted. Functional genes within these regions and their adjacent mobile genetic elements are annotated, and a visual map of gene neighborhood structure is generated to intuitively display the characteristics of drug-resistant genetic structure and the potential risk of horizontal gene transfer.

[0068] Example

[0069] This embodiment is used to verify the applicability, stability and operability of the proposed method for genome-level identification, quantification and structural visualization of drug-resistant pathogens in environmental samples in a typical wastewater treatment plant and its receiving river water environment system.

[0070] 1) Sample collection and pretreatment

[0071] This embodiment selects a representative wastewater treatment plant-receiving river system within a megacity as the research object. The selected research system covers multiple wastewater treatment plants (WWTPs) employing different treatment processes and their corresponding receiving rivers, which can systematically reflect the impact of wastewater treatment processes and discharge behavior on the distribution characteristics of drug-resistant pathogens in the environment.

[0072] Environmental samples collected included influent, secondary sedimentation tank sludge, and effluent samples from multiple wastewater treatment plants. Simultaneously, upstream, adjacent, and downstream sampling points were established in the receiving rivers corresponding to the discharge outlets of each wastewater treatment plant, forming a spatially continuous river sample sequence to assess the variation patterns of drug-resistant pathogens within wastewater treatment units and the emission impact gradient (Table 1). Sample collection was strictly conducted in accordance with relevant industry standards. Liquid samples were pretreated to remove large suspended solids before being enriched with microbial biomass through filtration or centrifugation; sludge samples were homogenized and directly used for microbial DNA extraction. All samples were immediately stored at 4 °C after collection and pretreated within 24 hours to minimize the impact of changes in microbial community structure during sample storage on subsequent metagenomic analysis results.

[0073] It should be understood that those skilled in the art can make reasonable adjustments to the number of sewage treatment plants, the location of sampling points, and the sampling volume according to the specific research area, the scale of sewage treatment, and the type of river, without affecting the implementation and technical effect of the method of the present invention.

[0074] Table 1 Sampling Points and Sampling Volume Settings

[0075]

[0076] 2) Extraction and metagenomic analysis of total microbial DNA

[0077] After sample processing, total microbial DNA was extracted using the FastDNA SPIN Kit for Soil. The extracted DNA concentration was detected using a Qubit 4.0 nucleic acid / protein quantitative fluorometer and a 1 dsDNA HS (high sensitivity) analysis kit (Thermo Fisher Scientific, USA), and DNA purity was determined using a NanoDrop Spectrophotometer (Thermo Fisher, USA) according to protocol to ensure that the DNA purity, integrity, and concentration met the requirements for metagenomic sequencing. The DNA fragments were then homogenized to approximately 400 bp using a Covaris ultrasonic homogenizer, and sequencing libraries were constructed using a commercial library construction kit (such as the NEXTFLEX Rapid DNA-Seq Kit). The libraries were then subjected to paired-end sequencing using the Illumina NovaSeq 6000 platform to obtain high-coverage, high-quality metagenomic sequences. Those skilled in the art can adjust the sequencing depth and library strategy according to project needs without affecting the effectiveness of this method.

[0078] 3) Reconstructing the genome of drug-resistant bacteria

[0079] Metagenomic sequencing yielded 24 samples and approximately 240 Gb of raw reads. To ensure data quality, reads were preprocessed using FastP, including removing adapter sequences, discarding reads with an average quality score below Q30, filtering fragments shorter than 150 bp, and removing reads containing N bases to obtain high-quality clean reads. The quality-controlled clean reads were then assembled using MEGAHIT software via de novo assembly, generating contigs through a multi-k-mer iterative strategy. Subsequently, the binning results from MetaBAT2 and CONCOCT were integrated using MetaWRAP's binning module, achieving high-precision merging and purification using the bin_refinement module. The reassemble_bins module was then used to reassemble the initially obtained MAGs to improve genome continuity. All obtained MAGs were annotated using GTDB-Tk for phylogenetic classification, and CheckM was used to assess the integrity and contamination of each MAG. In this embodiment, the screening thresholds were set to integrity ≥ 70% and contamination ≤ 10% to ensure that the MAGs used for drug resistance and pathogenicity analysis have reliable genomic structures.

[0080] 4) Identification of pathogenic and drug-resistant bacteria

[0081] For all high-quality MAGs, Prodigal software was used to predict their open reading frames (ORFs), and BLASTP was used to align the ORF sequences to a structured SARG database to identify and annotate antibiotic resistance genes (ARGs). The alignment parameter was set to e-value ≤ 1×10⁻⁶. -10 The sequence similarity is ≥ 80% and the coverage is ≥ 70% to ensure the accuracy and conservation of ARG annotation. MAGs carrying ARGs are marked as drug-resistant bacterial genomes. Simultaneously, MAG ORFs are compared with virulence factor databases (such as the VFDB core library) using BLAST or DIAMOND, with screening thresholds set at sequence similarity ≥ 70% and coverage ≥ 70% to identify potential VFG-carrying MAGs. MAGs carrying both ARGs and VFGs are further compared with the WHO's list of key pathogens (Table 2) to screen for drug-resistant pathogenic MAGs, forming the key identification output of the method of this invention.

[0082] Table 2: WHO List of Priority Bacterial Pathogens (Updated 2024)

[0083]

[0084] * The family Enterobacterales includes: Klebsiella pneumoniae, Escherichia coli, Enterobacter spp., Serratia spp., Proteus spp., Citrobacter spp., and Morganella spp.

[0085] The family Enterobacterales includes: Klebsiella pneumoniae, Escherichia coli, and Enterobacter spp.

[0086] 5) Quantitative analysis of the abundance of drug-resistant pathogens

[0087] For the metagenomically assembled genomes (MAGs) of drug-resistant pathogens obtained through the aforementioned screening steps, this embodiment further employs CoverM (v0.6.1) software to quantitatively analyze their relative abundance in different environmental samples. Specifically, high-quality metagenomic data (Clean Reads) obtained from each environmental sample are mapped to the genomic sequences of the corresponding drug-resistant pathogen MAGs, and the relative abundance of each target MAG is calculated in CoverM's "relative_abundance" mode, using RPKM as the abundance unit. This method enables quantitative distribution analysis of drug-resistant pathogens at the genomic scale in different types of environmental samples. In this embodiment, drug-resistant pathogen MAGs of *Acinetobacter johnsonii*, *Acinetobacter indicus*, *Escherichia coli*, *Pseudomonas aeruginosa*, and *Vibrio cholerae* were identified. Among them, certain MAGs (such as MAG5 belonging to Pseudomonas aeruginosa and MAG140 belonging to Escherichia coli) correspond to pathogens of key concern to the World Health Organization (WHO), and the types and numbers of antibiotic resistance genes they carry are significantly higher than those of other identified MAGs.

[0088] The abundance quantification results obtained based on the method of this invention show that different drug-resistant pathogens exhibit significantly differentiated distribution characteristics in the influent, treatment units, effluent of wastewater treatment plants, as well as in the upstream, adjacent areas of discharge outlets, and downstream sections of receiving rivers. This quantitatively reflects the presence level and changing trend of drug-resistant pathogens in the wastewater treatment process and the emission impact gradient. Figure 2 This paper presents the abundance variations of target drug-resistant pathogens and the specific types of drug-resistant genes they carry in wastewater treatment plant (influent, activated sludge, and effluent) and receiving river systems (upstream, near the discharge outlet, and downstream). The results show that the method of this invention can achieve stable and reproducible quantitative analysis of drug-resistant pathogens at the genomic scale in complex wastewater treatment plant-receiving river environments. It should be noted that the specific pathogen species and their abundance distribution involved in this embodiment are only used to illustrate the technical effects of the method of this invention, and the invention is not limited to these specific species or their abundance ranges.

[0089] 6) Visualization of the genome structure of drug-resistant pathogens

[0090] To visually demonstrate the gene arrangement and structural characteristics of antibiotic resistance genes (ARGs), virulence factor genes (VFGs), and their surrounding mobile genetic elements (MGEs) in the metagenomic assembly genome (MAG) of drug-resistant pathogens, this embodiment constructs a gene neighborhood structure visualization map for two drug-resistant pathogen genomes to show the distribution of resistance genes and their neighboring genes. Figure 3 Specifically, key genomic regions (contigs or scaffolds) containing ARGs and / or VFGs are extracted from confirmed drug-resistant pathogens (MAGs), and gene annotation information within these regions is integrated, including antibiotic resistance gene types, virulence factor gene types, adjacent open reading frames (ORFs), inserted sequences (IS elements), integrases, transposases, and other functional elements. Based on this, the gggenes package (based on the R language environment) is used, combined with a custom automated script, to visualize the gene neighborhood structure of the aforementioned genomic regions, generating a genome structure map of the drug-resistant pathogens. This example highlights the genome structure examples of two drug-resistant pathogens, MAG5 (belonging to *Pseudomonas aeruginosa*) and MAG140 (belonging to *Escherichia coli*). The generated maps clearly show the relative positions of antibiotic resistance genes and mobile genetic elements, gene arrangement directions, local gene cluster structural features, and potential horizontal gene transfer hotspots.

[0091] The visualization process described above does not require manual drawing of each image; it can be batch-generated using automated scripts with preset parameters, making it suitable for high-throughput analysis of the genome structures of drug-resistant pathogens in large-scale environmental samples. The visualized maps obtained in this embodiment are only used to illustrate the operability and applicability of the method of the present invention in the analysis of the genome structures of drug-resistant pathogens. It should be understood that different environmental sample sources, sequencing depths, and microbial community compositions may lead to differences in gene arrangement structures and map morphology, but these differences do not affect the implementation of the method of the present invention or its technical effects.

[0092] In summary, this embodiment verifies the effectiveness and stability of the method of the present invention in accurately identifying, quantifying the abundance of, and visualizing the genome structure of drug-resistant pathogens in complex aquatic environments such as wastewater treatment plants and their receiving rivers. The method of the present invention can systematically output the genome structure information of drug-resistant pathogens within a unified analytical framework, providing reliable technical support for analyzing the distribution characteristics, potential transmission mechanisms, and risk assessment of drug-resistant genes in wastewater treatment and environmental discharge processes. It has good versatility and application value.

Claims

1. A method for genomic identification, quantification, and visualization of drug-resistant pathogens in environmental samples, characterized in that, Includes the following steps: Step 1: Obtain total microbial DNA from environmental samples, perform metagenomic high-throughput sequencing, and obtain high-quality metagenomic data (Clean Reads) after sequence quality control and removal of host sequences; Step 2: Use assembly software to assemble the Clean Reads obtained in Step 1 to obtain contigs. Then, use various binning algorithms to cluster the contigs and reconstruct the metagenomic assembly genome (MAG). Step 3: Perform quality assessment and screening on the MAGs obtained in Step 2 to obtain a high-quality set of MAGs; Step 4: Perform open reading frame (ORF) prediction on the high-quality MAGs screened in Step 3, and align the predicted protein sequences to the antibiotic resistance gene ARG database to identify MAGs carrying antibiotic resistance genes and determine them as MAGs that host antibiotic resistance genes. Step 5: In the drug resistance gene host MAG identified in Step 4, the predicted protein sequence is further aligned to the virulence factor gene VFG database to screen MAGs that carry both antibiotic resistance genes and virulence factor genes as potential drug-resistant pathogen genomes. Step 6: Compare the genomes of potential drug-resistant pathogens obtained in Step 5 with the authoritative list of pathogens from the World Health Organization (WHO) to confirm the genomes of drug-resistant pathogens in the environmental samples. Step 7: Use quantitative analysis software to map the Clean Reads obtained in Step 1 onto the genomes of the drug-resistant pathogens identified in Step 6, and calculate their abundance in environmental samples; Step 8: For the drug-resistant pathogen genome in Step 7, extract the genomic regions containing antibiotic resistance genes and / or virulence factor genes, annotate the surrounding mobile genetic elements (MGEs), and generate the corresponding gene neighborhood structure visualization map.

2. The method according to claim 1, characterized in that, Step 3 specifically includes: Step 3-1: Using CheckM software based on a lineage-specific marker gene set, assess the integrity and contamination of each MAG; Step 3-2: Select MAGs with a completeness greater than a preset threshold and a contamination level less than a preset threshold as high-quality MAGs, wherein the completeness is greater than 70% or 90% and the contamination level is less than 10%.

3. The method according to claim 1, characterized in that, Step 4 specifically includes: Step 4-1: Use Prodigal to predict open reading frames on the filtered MAGs; Step 4-2: Align the predicted protein sequence with specialized antibiotic resistance gene databases such as SARG, CARD, or ResFinder using BLAST or DIAMOND tools. Step 4-3: Set the filtering threshold as follows: e value less than or equal to 1×10 -10 The sequence similarity is greater than or equal to 80%, and the coverage is greater than or equal to 75%. Based on this comparison and screening results, MAG carrying antibiotic resistance genes was identified and determined as the host MAG of the resistance gene.

4. The method according to claim 1, characterized in that, Step 5 specifically includes: Step 5-1: The protein sequence of the drug resistance gene host MAG determined in Step 4 is compared with the virulence factor gene database VFDB using BLAST or DIAMOND tools; Step 5-2: Set the filtering threshold as follows: e value less than or equal to 1×10 -10 The sequence similarity is ≥ 70%, the coverage is ≥ 70%, and MAGs carrying both antibiotic resistance genes and virulence factor genes are selected as potential drug-resistant pathogen genomes.

5. The method according to claim 1, characterized in that, The authoritative list of pathogens includes the list of key or critical pathogens published by the World Health Organization (WHO).

6. The method according to claim 1, characterized in that, Step 7 specifically includes: Step 7-1: Using CoverM or Salmon quantitative analysis software, map the high-quality metagenomic data Clean Reads obtained in Step 1 onto the genome of the target drug-resistant pathogen MAG identified in Step 6; Step 7-2: Calculate the relative abundance of each target MAG and normalize the result to the number of mapped sequences per kilobase per million, i.e., the RPKM value. The RPKM calculation formula is as follows: RPKM In the formula: Mapped Reads is the number of sequences mapped to the target MAG; Total Mapped Reads is the total number of mappings for all MAGs; Length is the genome length of the target MAG; The RPKM value characterizes the relative abundance level of the target drug-resistant pathogen in the environmental microbial community and can be used to quantitatively analyze the environmental distribution of drug-resistant pathogens.