Tracing analysis and multi-habitat homology evaluation method of antibiotic drug resistance gene in environmental sample
By combining metagenomic sequencing and machine learning algorithms, the problem of identifying the association between antibiotic resistance gene transmission in multiple habitats has been solved. This has enabled the accurate identification of the transmission characteristics and sources of antibiotic resistance genes in different habitats, providing a scientific basis for risk assessment and management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies are insufficient to accurately identify and track the transmission links and potential sources of antibiotic resistance genes in various ecological environments. Furthermore, existing source tracing analysis methods lack the integration of gene-level characteristics, making it difficult to reflect the true alternation relationships and potential source relationships between different habitats.
We employ metagenomic sequencing big data sets and machine learning algorithms to construct a method for tracing the origins of antibiotic resistance genes in multiple habitats. Through abundance matrix construction, the FEAST machine learning model, metagenomic assembly and binning technology, we reconstruct high-quality bacterial genomes. Combined with phylogenetic analysis, we identify and quantify the transmission characteristics of antibiotic resistance genes in different habitats.
It has enabled the precise identification and quantitative characterization of antibiotic resistance genes in multiple habitats, revealed the potential transmission links and homology between different habitats, and provided a scientific basis for the source identification and risk assessment of antibiotic resistance pollution.
Smart Images

Figure CN121983121A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of metagenomic bioinformatics analysis and environmental antibiotic resistance monitoring technology. Specifically, it relates to an analytical method based on metagenomic sequencing big data sets and machine learning algorithms to conduct source tracing analysis and gene homology assessment of antibiotic resistance genes in various ecological environments. This method enables the analysis of cross-habitat transmission relationships of antibiotic resistance genes between different ecological environments and the identification of potential sources, providing technical support for source control, transmission risk assessment, and environmental management of antibiotic resistance pollution. Background Technology
[0002] The extensive use and long-term emissions of antibiotics have led to the proliferation and spread of antibiotic-resistant bacteria (ARBs) and antibiotic resistance genes (ARGs) in the environment, making antibiotic resistance a serious global public health problem. ARGs have been widely detected in various ecological environments, including soil, water bodies, sediments, wastewater treatment systems, and human-related environments. The continuous accumulation and spread of ARGs in the environment may not only alter the structure and ecological function of microbial communities but also enter human-inhabited environments through food chains or the biosphere, where they can be carried by bacteria and microorganisms, posing a potential threat to public health and ecological security. Therefore, systematically identifying the cross-habitat transmission characteristics of antibiotic resistance genes in different ecological environments, the connectivity of resistance genes between different habitats, and potential sources of pollution are crucial issues that urgently need to be addressed in current research on environmental antibiotic resistance and pollution control.
[0003] With the development of high-throughput sequencing technology and metagenomics, detecting and annotating antibiotic resistance genes in the environment based on metagenomic big data sets has become a common research method. Existing studies have statistically analyzed the types and abundance of antibiotic resistance genes in different samples or habitats by comparing resistance gene databases, thereby revealing the compositional differences and abundance trends of resistance genes in different environments. However, this type of analysis mainly focuses on the abundance differences of resistance genes in a single habitat or between different samples, and it is difficult to further answer whether there are potential transmission links between different ecological environments, and the role played by different habitats in the formation of resistance pollution. To explore the transmission relationship of resistance genes between different environments, some studies have introduced co-occurrence analysis or network analysis methods to associate resistance genes with potential hosts in different samples or environments. These methods reveal the association characteristics between resistance gene hosts and environmental factors to a certain extent, but their analysis results usually rely on statistical correlation, making it difficult to distinguish between true alternation relationships and accidental co-occurrence phenomena between different habitats. In addition, correlation analysis often cannot directly reflect the potential source relationship of resistance genes between different ecological environments, making it difficult to provide direct evidence for the source identification of resistance pollution. On the other hand, regarding the issue of source analysis of microbial communities or functional genes, existing studies have proposed using source tracing models to estimate the source composition of microorganisms or functional genes in the target environment. Such methods can, to a certain extent, quantitatively analyze the contribution ratio of different sources to the target environment, providing new technical approaches for understanding the material or information inputs in complex environmental systems. However, existing source tracing analyses mostly focus on the microbial community level or only estimate source contributions based on abundance information, lacking an analytical framework that combines the specific sequence characteristics of drug resistance genes and host information, making it difficult to further reveal the intrinsic biological mechanisms of drug resistance gene transmission between different ecological environments.
[0004] Therefore, there is an urgent need for a tool that can accurately and rapidly identify the pollution transmission characteristics of antibiotic resistance genes in various environmental media. This invention, based on metagenomic big data sets and machine learning models, can systematically identify antibiotic resistance genes in a variety of different ecological environments. Furthermore, by integrating source contribution analysis and gene-level association characteristics, it reveals the connectivity and homology of resistance genes, providing crucial technical support for the source identification, transmission mechanism interpretation, and risk assessment of antibiotic resistance pollution. Summary of the Invention
[0005] The purpose of this invention is to provide a method for tracing the origins and assessing the connectivity of antibiotic resistance genes in multiple habitats based on metagenomic sequencing big data sets and machine learning algorithms. This method integrates abundance information of resistance genes, source contribution analysis, and gene-level phylogenetic homology analysis to identify potential transmission links of antibiotic resistance genes in different ecological environments, providing a reliable analytical method for identifying the source of antibiotic resistance pollution, recognizing transmission risks, and managing the environment.
[0006] This invention uses environmental samples from different ecological environments and media as research objects. It employs metagenomic sequencing and bioinformatics analysis methods to identify and quantitatively characterize thousands of antibiotic resistance genes across multiple habitats, constructing an abundance matrix for antibiotic resistance genes in various habitats. Based on this, the FEAST microbial source tracing model, based on the machine learning expectation-maximization algorithm, is introduced to analyze the potential sources of resistance gene composition in target habitats, quantifying the contribution rate of different potential source habitats to the target habitat and screening for potential connected habitats with source attributes. Furthermore, by combining metagenomic assembly and binning technologies, high-quality bacterial genomes from connected habitats are reconstructed, and the antibiotic resistance genes they carry and their microbial hosts are annotated and identified. Based on the types of resistance genes commonly carried by bacterial genomes in different habitats, their nucleotide sequences are further extracted and phylogenetic analysis is performed to compare the homology characteristics of resistance genes across different habitats, thereby elucidating the potential connectivity and cross-habitat transmission characteristics of antibiotic resistance genes in multiple habitats.
[0007] Specific implementation schemes for achieving the objectives of this invention:
[0008] A method for tracing and analyzing the origins of antibiotic resistance genes in environmental samples and assessing their homology across multiple habitats includes the following steps:
[0009] Step 1: Obtain environmental microbial metagenomic data for different habitats, specifically including:
[0010] 1-1: Collect environmental samples from different habitats;
[0011] 1-2: Extract microbial genomic DNA from environmental samples, perform metagenomic sequencing, and obtain raw metagenomic data from different habitats;
[0012] 1-3: Quality control was performed on the raw metagenomic data to filter out low-quality sequences with a base quality of less than 30 and a sequence length of less than 150 bp, resulting in high-quality metagenomic datasets of each environmental sample after quality control.
[0013] Step 2: Construct a multivariate abundance matrix of antibiotic resistance genes in different habitats: Based on the high-quality metagenomic datasets obtained in Steps 1-3, the ARGs-OAP tool is used to align with the structured antibiotic resistance gene database SARG v3.0 to accurately identify and quantify resistance genes in different habitats, and obtain the abundance matrix of diverse resistance genes in different habitats.
[0014] Step 3: Construct the FEAST pollution source model for microbial source tracing using machine learning algorithms to identify and quantify the contribution rate of "potential source habitats" to the composition of drug-resistant genes in the "target habitat": Based on the abundance matrix of diverse drug-resistant genes in different habitats obtained in Step 2, define one habitat as the target habitat and screen potential source habitats; based on the abundance matrix of drug-resistant genes in potential source habitats, construct the FEAST source tracing model using the machine learning expectation-maximization algorithm, and verify and optimize the model's prediction accuracy and generalization ability through the cross-residue-one-out iteration method and the simulated dataset method; use the model to predict the "source-sink" relationship between habitats and calculate the source contribution rate of each source habitat to the target habitat;
[0015] Step 4: Select habitats for connectivity and gene homology analysis: Based on the predicted relationship and contribution rate of the source-sink habitats obtained in Step 3, select habitats with a contribution rate >1% and perform connectivity and / or gene homology analysis with the target habitat. Define the selected habitats as potential connectivity habitats.
[0016] Step 5: Extract high-quality metagenomic datasets of potential connected habitats: From the high-quality metagenomic datasets obtained in Steps 1-3, filter and extract high-quality metagenomic datasets of potential connected habitats as described in Step 4.
[0017] Step 6: Reconstructing high-quality bacterial genomes from potentially connected habitats: Based on the high-quality metagenomic datasets from the potentially connected habitats extracted in Step 5, the MetaWRAP genome binning analysis platform was called via conda in the Linux terminal. The cat tool was used to merge the paired-end sequences of the high-quality metagenomic datasets of each sample into one file. The Megahit genome assembly tool was used to assemble the merged short genome sequences. The Metabat2, Concoct, and Maxbin algorithms were used to bin the assembled long sequence contig files. The binning results were purified and filtered using BIN_REFINEMENT, and finally, high-quality bacterial genomes with integrity ≥ 50% and contamination ≤ 10% were retained.
[0018] Step 7: Annotate antibiotic resistance genes and species on the high-quality bacterial genomes reconstructed in potentially connected habitats: For the high-quality bacterial genomes reconstructed in Step 6, use Prodigal to predict the protein sequences of open reading frames (ORFs) in the bacterial genome. Use the BLASTP tool to compare the ORF sequences with the structured antibiotic resistance gene database SARGv3.0, selecting sequences with similarity ≥ 80%, coverage ≥ 75%, and e-value ≤ 10. -7 The standard identification of drug resistance gene sequences was used to obtain high-quality bacterial genomes carrying drug resistance genes, which were defined as ARG-MAGs; at the same time, the GTDB_tk tool was used to annotate the species of ARG-MAGs to obtain the microbial host information and classification results of ARG-MAGs in potential connected habitats.
[0019] Step 8: Construct a phylogenetic tree of antibiotic resistance genes in potentially connected habitats: Based on the microbial hosts of ARG-MAGs in potentially connected habitats from Step 7, extract nucleotide sequences of the same class of antibiotic resistance genes and perform nucleotide sequence alignment using the clustalW tool; based on the alignment results, construct a phylogenetic tree of antibiotic resistance genes in potentially connected habitats using the Fasttree tool; extract nucleotide sequences of the same class of antibiotic resistance genes from the reconstructed high-quality bacterial genomes.
[0020] Step 9: Clarify the cross-habitat transmission characteristics of antibiotic resistance genes: Based on the phylogenetic tree of antibiotic resistance genes in potential connected habitats in Step 8, compare and analyze the homology characteristics of resistance genes between potential source habitats and target habitats based on clustering relationships, branching structures, and differences in evolutionary branch lengths, construct a multi-media transmission map of resistance genes and their inter-host transmission network; comprehensively evaluate the cross-habitat transmission characteristics of antibiotic resistance genes in potential connected habitats.
[0021] Furthermore, the different habitats mentioned in step 1-1 include soil, water bodies, sediments, sewage treatment systems, biological environments, and human activity environments.
[0022] Furthermore, step 2, which describes the precise identification and quantification of antibiotic resistance genes in different habitats, specifically includes: importing the quality-controlled sample metagenomic dataset file into the Linux terminal, calling the ARGs-OAP tool via conda, and performing precise annotation, classification, and quantitative analysis of antibiotic resistance genes in two steps: rapid screening and precise classification. The output results include quantitative results at three different levels: type, subtype, and gene. The gene level is selected as the abundance file for quantifying antibiotic resistance genes in multiple habitats.
[0023] Furthermore, step 3, which involves constructing the FEAST plastic source model using the expectation-maximization algorithm in machine learning, specifically includes: preparing input files for source tracing analysis, including an abundance file of antibiotic resistance genes and a classification information file. The abundance file is formatted as follows: each column represents a sample, and each row represents the category of antibiotic resistance genes. The classification information file defines the target habitat as the sink and other habitats as the source. The abundance file and classification information file are imported into R, and after installing the software packages required for the FEAST model, the FEAST plastic source model is constructed based on the expectation-maximization algorithm.
[0024] Furthermore, step 8, which involves extracting nucleotide sequences of antibiotic resistance genes of the same category, specifically includes: screening for common antibiotic resistance gene categories carried by ARG-MAGs; and using the seqkit tool to extract nucleotide sequences of the same category of antibiotic resistance genes from high-quality genomes of potentially connected habitats based on the open reading frame (ORF) sequence number and genome sequence number in the BLASTP alignment results.
[0025] Beneficial Effects: This invention provides a method for tracing the origins of antibiotic resistance genes in environmental samples and assessing their multi-habitat homology based on metagenomic bioinformatics analysis. By integrating antibiotic resistance gene abundance analysis, source tracing analysis, and phylogenetic analysis, it achieves a systematic analysis of the transmission characteristics of resistance genes across different habitats and the potential source habitat associations based on gene homology. This method is not dependent on a single environment or a single data type, and can be applied to various ecological environments and complex sample types, exhibiting strong versatility and adaptability. By screening connected habitats with potential source attributes through source tracing results, and further combining metagenomic assembly binning and microbial host annotation, it reveals the transmission associations of antibiotic resistance genes between different habitats from the perspectives of gene homology and bacterial-microbial host, providing a new analytical approach for understanding the cross-habitat transmission characteristics of antibiotic resistance pollution. This invention can comprehensively reflect the homology associations of resistance genes across different habitats. Moreover, the analytical process of this invention is based on high-throughput sequencing data, can be implemented on a Linux platform, has a clear operation process, and has strong reproducibility of results. It can provide a scientific basis and technical support for the identification of antibiotic resistance pollution sources, risk assessment, and the formulation of prevention and control strategies. Attached Figure Description
[0026] Figure 1 The results of the FEAST model source tracing analysis include a schematic diagram showing the proportion of potential source contributions of each habitat and unknown sources to marine habitats.
[0027] Figure 2 This is a schematic diagram of a phylogenetic tree of the TEM-1 and floR antibiotic resistance genes carried by Escherichia coli in potentially connected habitats. Detailed Implementation
[0028] The embodiments of the present invention are described in conjunction with the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments.
[0029] The present invention provides a method for tracing and analyzing the origins of antibiotic resistance genes in environmental samples and assessing their multi-habitat homology, specifically including:
[0030] 1) Obtain environmental microbial metagenomic data from samples in different ecological environments;
[0031] Specifically, the following steps are included:
[0032] 1-1) Samples were collected according to the sampling specifications for different habitats and environmental media (soil, water, sediment, sewage treatment systems, organism-related environments and human-related environments);
[0033] 1-2) Total DNA was extracted from samples from different habitats. After enrichment and processing according to the characteristics of samples from different media, DNA was extracted using commercial DNA kits.
[0034] 1-3) Take 1.5 μg of each DNA sample for high-throughput metagenomic sequencing. Ensure that the sequencing depth is sufficient and that the amount of metagenomic data after quality control of each sample is more than 10 Gb.
[0035] 1-4) Import the raw metagenomic sequencing data into the Linux terminal environment. After entering the corresponding software runtime environment, call fastp to perform quality control on the paired-end sequencing data. Set the parameters to "-q 30 -l 150" to filter low-quality sequences with a base quality < 30 and a length < 150 bp, and obtain clean reads for subsequent metagenomic analysis.
[0036] 2) Based on the clean reads alignment of structured antibiotic resistance gene data SARG v3.0 in different habitats, construct the gene-level ARGs abundance matrix;
[0037] Specifically, the following steps are included:
[0038] 2-1) The clean reads of different habitat metagenomic datasets obtained in steps 1-4 are used as input data and uniformly organized into the analysis environment for subsequent antibiotic resistance gene annotation and quantitative analysis;
[0039] 2-2) The ARGs-OAP annotation tool was invoked via conda, and this software tool was used to perform a two-stage comparative analysis of clean reads from different habitats. First, a rapid screening stage was used to perform a preliminary comparison of clean reads to remove sequences unrelated to antibiotic resistance genes. Subsequently, a precise comparison stage was used to perform high-precision comparison and quantitative analysis of the screened sequences with the structured antibiotic resistance gene database SARG v3.0 to obtain annotation information and abundance matrices of antibiotic resistance genes in samples from different habitats.
[0040] 2-3) Based on the comparison and quantification results, the antibiotic resistance gene annotation information is summarized according to different classification levels, including type, subtype, and gene levels. In the preferred method, the quantitative results of specific resistance gene gene levels are selected to construct an antibiotic resistance gene abundance matrix for samples from different habitats. The abundance matrix is arranged with samples as columns and antibiotic resistance genes as rows, and the matrix elements are the abundance values of the corresponding resistance genes in the samples, which are used for subsequent multi-habitat antibiotic resistance gene source tracing analysis and multi-habitat connectivity analysis.
[0041] 3) Based on the abundance matrix of ARGs in different habitats, the FEAST source tracing model for microorganisms was constructed using the expectation-maximization algorithm to conduct source tracing analysis on the drug resistance gene composition of the target habitat, and source habitats and target habitats with certain source attributes (contribution rate >1%) were selected as potential connected habitats.
[0042] Specifically, the following steps are included:
[0043] 3-1) Construct the input data file and classification information file for source tracing analysis using the FEAST model. The input data file is an abundance matrix of antibiotic resistance genes in samples from different habitats, with antibiotic resistance genes as rows and samples as columns, and the matrix elements representing the relative abundance of the corresponding antibiotic resistance gene in the sample. The classification information file is used to identify the habitat type to which different samples belong, where the target habitat sample is defined as the sink, and other potentially connected habitat samples are defined as the source.
[0044] 3-2) Import the antibiotic resistance gene abundance matrix file and classification information file into the R language runtime environment, install and load the software toolkit and its dependencies required for the FEAST model, and initialize the FEAST source tracing analysis process. Based on the classification information file and abundance file, call the expectation-maximization machine learning algorithm in the FEAST model to perform microbial source tracing analysis on samples from different habitats. Set the algorithm parameters ("different_sources_flag = 1, outfile = "demo"") to distinguish different source habitats and output the source analysis results of the antibiotic resistance gene composition in the target habitat. The analysis results include the contribution rate of different source habitats to the antibiotic resistance gene composition of the target habitat and the proportion of unattributed sources;
[0045] 3-3) Analyze the microbial source tracing results of the FEAST model. Based on the contribution rate of different habitats to the composition of antibiotic resistance genes in the target habitat, and in conjunction with the research objectives, set a preset contribution rate threshold. Select habitats with a contribution rate of at least >1% as potential connected habitats together with the target habitat. Select the clean reads corresponding to the potential connected habitats as input data for subsequent multi-habitat antibiotic resistance gene homology / connectivity analysis.
[0046] 4) Reconstruct high-quality bacterial genomes from potential connected habitats based on corresponding metagenomic datasets;
[0047] Specifically, the following steps are included:
[0048] 4-1) In the Linux terminal, place the clean reads of the selected high-quality metagenomic dataset with potential connectivity into the same file directory;
[0049] 4-2) Use the cat tool to merge all the forward sequence files into a single forward sequence file. Similarly, merge all the reverse sequences into a single reverse sequence file.
[0050] 4-3) In the Linux terminal, enter the MetaWRAP analysis platform via conda, call the Megahit tool, and use the merged total forward and reverse sequence files as input files to splice and assemble short sequences;
[0051] 4-4) In the MetaWRAP analysis platform, the assembled contig long sequence file and the clean reads used for assembly are used as input files. The binning command is called with the parameters set to "--metabat2 --concoct--maxbin". Three different algorithms are used to bin the long sequence contigs from different habitats.
[0052] 4-5) For the genomes obtained from binning, the bin_refinement tool was used in the MetaWRAP analysis platform to purify the genomes. The parameters were set to "-c 50 -x 10" to require integrity ≥ 50% and contamination ≤ 10%. Genomes from different habitats were purified and filtered to reconstruct high-quality bacterial genomes (MAGs) in potentially connected habitats.
[0053] 5) Species annotation of MAGs in the reconstructed potential connected habitats to obtain their microbial host information; at the same time, annotation of ARGs carried on MAGs is completed by comparing with the SARG v3.0 structured antibiotic resistance gene database.
[0054] Specifically, the following steps are included:
[0055] 5-1) In the Linux terminal, place all the obtained MAGs into the same directory;
[0056] 5-2) Use conda to call the GTDB_tk species annotation tool, with the directory where the MAGs are placed as the input file path, and set the parameter "--extension.fa" to set the output directory. After the species annotation is completed, view the summarized bacterial classification results file in the output directory to obtain the bacterial microbial host information of the MAGs.
[0057] 5-3) Annotate MAGs for different habitats using the structured antibiotic resistance gene reference database SARG v3.0. Use conda to call the Prodigal tool to predict the ORF of MAGs, using BLAST with "e-value ≤ 10". -7 "The predicted ORF protein sequences are compared with the SARG v3.0 database, and the similarity is set to ≥ 80% and the coverage is ≥ 75% to identify MAGs carrying ARGs (ARGs-Carrying MAGs, ACGs).
[0058] 6) Extract nucleotide sequences of antibiotic resistance genes of the same class from ACGs in potentially connected habitats, and construct a phylogenetic tree of antibiotic resistance genes in potentially connected habitats after nucleotide sequence alignment, so as to analyze the cross-habitat transmission characteristics and potential connectivity relationships of antibiotic resistance genes in potentially connected habitats.
[0059] Specifically, the following steps are included:
[0060] 6-1) Based on the results and process files of the BLAST comparison with the SARG v3.0 database in step 5-3, find the ORF protein sequence ID encoding ARGs and the genome ID of the MAGs containing the ORF sequence;
[0061] 6-2) Based on the ORF sequence ID and genome ID, use the seqkit tool to extract the nucleotide sequences of the same class of drug resistance genes in potential connected habitat ACGs;
[0062] 6-3) Use the clustalW tool to perform gene nucleotide sequence alignment on the extracted drug resistance gene sequences in potentially connected habitats. By statistically analyzing the overall consistency ratio and conserved site distribution characteristics among drug resistance gene sequences of the same class of antibiotics in different habitats, the sequence similarity level is quantitatively evaluated.
[0063] 6-4) Based on the nucleotide sequence alignment results, a phylogenetic tree is constructed using Fasttree via conda to analyze the nucleotide sequence alignment results of antibiotic resistance genes in multiple habitats, thereby obtaining the phylogenetic structure of resistance genes in multiple habitats. Based on a comprehensive analysis of multi-dimensional quantitative information such as clustering relationships, branching structures, differences in evolutionary branch distances, clustering stability, and branch support of the same type of resistance genes in different habitats, the homology of resistance genes in multiple habitats is assessed.
[0064] Example
[0065] An example of the application of a method for tracing the origins of antibiotic resistance genes in environmental samples and assessing their homology across multiple habitats.
[0066] In this embodiment, the metagenomic data used are derived from publicly available databases and datasets obtained from previous studies. The data includes raw metagenomic sequencing datasets from marine aquatic habitats and several non-marine habitats. The raw metagenomic sequencing data from each habitat sample underwent standardized quality control and processing before being used as input data for this embodiment.
[0067] In this embodiment, the analysis process is divided into a microbial source tracing and analysis module and a gene homology analysis module, which respectively verify the feasibility and technical effectiveness of the method of the present invention in the study of antibiotic resistance in samples involving multiple habitats from the perspectives of source analysis of drug resistance gene pollution and multi-habitat connectivity.
[0068] The specific steps are as follows:
[0069] 1) Preparation and preprocessing of multi-habitat metagenomic datasets
[0070] Metagenomic datasets from different habitats, including ocean, soil, natural water bodies, human excrement, and livestock and poultry excrement, were selected as the subjects of analysis. All metagenomic datasets used were from publicly available resource repositories or databases.
[0071] After importing metagenomic datasets from different habitats into the Linux terminal, fastp was used to perform quality control on the metagenomic datasets. The parameters were set to "-q 30 -l 150" to filter sequences with a quality score below 30 and a sequence length below 150 bp, thus obtaining clean reads from other habitats such as ocean and soil.
[0072] 2) Construct an abundance matrix of antibiotic resistance genes.
[0073] In the Linux terminal, the ARGs-OAP annotation tool is invoked via conda to annotate and quantify antibiotic resistance gene ARGs. Clean reads are placed in the same directory, which is used as the input path. Two-step filtering and comparison commands are enabled: `args_oap stage_one -i input -o output -f fa -t 8` and `args_oap stage_two -i output -t 8`. The `normalized_cell.gene.txt` file in the execution directory contains the preliminary quantitative results at the gene level of the resistance gene. Based on the annotation and quantitative results, the antibiotic resistance gene annotation information is summarized according to different habitat types, including marine, soil, human feces, and livestock manure habitats. The habitat classification information and gene quantitative abundance files are then compiled to construct an antibiotic resistance gene abundance matrix for different habitats. The abundance matrix is composed of samples as columns and antibiotic resistance genes as rows. The matrix elements are the relative abundance values of the corresponding resistance genes in the samples. The abundance unit is copy / cell. It is used for subsequent source tracing analysis of antibiotic resistance gene composition microorganisms and multi-habitat gene homology analysis.
[0074] 3) Microbial source tracing analysis of drug resistance genes in multiple habitats
[0075] Based on an abundance matrix integrating multi-habitat classification information and antibiotic resistance gene quantification results, a source tracing model FEAST for microorganisms is constructed using the expectation-maximization algorithm to analyze the source composition of antibiotic resistance genes in target habitats. First, input data files and classification information files for microbial source tracing analysis in the FEAST model are constructed. The input data file is an abundance matrix of antibiotic resistance genes in different habitats, with antibiotic resistance genes as rows and samples as columns. Matrix elements represent the relative abundance of the corresponding antibiotic resistance gene in the sample, with the unit of relative abundance being copy / cell. The classification information file is used to identify the habitat type of samples from different habitats, where marine habitats are defined as sinks (target habitats) and potential source habitats are defined as sources. Table 1 shows the classification information files and antibiotic resistance gene abundance files for nine different habitats: ocean, sewage treatment plants, human feces, cow dung, natural water bodies, soil, landfills, pig manure, and chicken manure. Import the antibiotic resistance gene abundance file and classification information file into the R language runtime environment, install and load the software packages and dependencies required for the FEAST model, initialize the FEAST model microbial source tracing analysis process, use the model to predict the "source-sink" relationship between habitats, and calculate the source contribution rate of each source habitat to the target habitat. Figure 1 The study presented information on the contribution rates of wastewater treatment plants, human excrement, cow dung, natural water bodies, soil, landfills, pig manure, and chicken manure habitats to the development of antimicrobial resistance genes in marine habitats, as well as the proportion of unattributed sources. From this data, natural water bodies, landfills, wastewater treatment plants, soil, human excrement, chicken manure, and pig manure were selected as potential source habitats.
[0076] Table 1. Classification information and abundance matrix files for nine different habitats.
[0077]
[0078] 2. Based on the phylogenetic relationships of antibiotic resistance genes and the analysis of gene homology across multiple habitats, the cross-habitat transmission of antibiotic resistance genes in different habitats was revealed. The steps are as follows:
[0079] 1) Screening metagenomic datasets of potential source habitats
[0080] Based on the microbial source tracing analysis results of the FEAST model, potential source habitats (contribution rate > 1%) were screened. The target habitat and potential source habitats were combined for subsequent connectivity analysis. The selected habitats were marine, soil, human feces, and livestock feces, defined as potential connectivity habitats, representing habitat types strongly associated with the drug resistance gene composition of marine habitats.
[0081] 2) Reconstructing high-quality bacterial genomes carrying ARGs in potentially connected habitats
[0082] Based on clean reads from the selected metagenomic dataset of connected habitats, Megahit was used to splice and assemble short-read sequences to obtain contigs of continuous assembled long sequences. Using the assembled contig files and corresponding clean read coverage information as input, the assembly results were binned in the MetaWRAP analysis platform. During binning, multiple binning algorithms, including MetaBAT2, CONCOCT, and MaxBin, were simultaneously invoked to cluster the contigs based on sequence composition features and coverage information, thereby obtaining preliminary MAGs. The bin_refinement module in MetaWRAP was used to assess and screen the binning results, removing low-quality genomes based on genome integrity and contamination indicators. Integrity thresholds were set ≥ 50%, and contamination thresholds ≤ 10%, reconstructing high-quality MAGs from potential connected habitats.
[0083] Furthermore, Prodigal was used to predict the ORFs of the reconstructed high-quality MAGs to obtain potential coding gene sequences in the genome. Subsequently, the predicted ORF protein sequences were compared with the structured antibiotic resistance gene reference database SARG v3.0 using the BLASTP tool, with criteria of "similarity ≥ 80%, coverage ≥ 75%, and e-value ≤ 10". -7 This process identifies MAGs carrying at least one ARG as metagenomic assembly genomes carrying antibiotic resistance genes (ARGs-Carrying MAGs, ACGs). Simultaneously, species classification annotation is performed on these ACGs. The GTDB-Tk tool is used to perform phylogenetic prediction on these genomes to obtain their corresponding bacterial classification information, thereby clarifying the microbial hosts of antibiotic resistance genes in different habitats.
[0084] 3) Construct a phylogenetic tree of antibiotic resistance genes of the same class in different habitats.
[0085] After obtaining ACGs from multiple habitats, the seqkit tool was used to process the antibiotic resistance genome files. Based on the ORF sequence ID and genome ID in the annotation information of the antibiotic resistance genes, the corresponding antibiotic resistance gene nucleotide sequences were extracted in batches, generating a target antibiotic resistance gene nucleotide sequence file containing sequences from different habitats. The clustalW software was used to perform multiple sequence alignment on the extracted antibiotic resistance gene nucleotide sequences to align homologous regions between nucleotide sequences from different habitats. clustalW uses a progressive alignment strategy to systematically compare the conservation and differences between sequences, thereby obtaining standardized sequence alignment result files, providing a foundation for subsequent phylogenetic analysis. The Fasttree software was used to analyze the multiple sequence alignment results, and a phylogenetic tree was quickly constructed based on the approximate maximum likelihood algorithm. Based on the proportion of nucleotide sequence similarity, the proportion of conserved sites, the coverage and structural integrity of antibiotic resistance genes, as well as the clustering relationships, evolutionary branch structure, cluster stability, branch support, and differences in evolutionary branch length of antibiotic resistance genes in different habitats, a comprehensive analysis of multi-dimensional quantitative information is used to determine the homology of antibiotic resistance genes in different habitats, providing a basis for revealing the cross-habitat diffusion characteristics of antibiotic resistance genes. Figure 2 The phylogenetic trees of the two antibiotic resistance genes TEM-1 and floR carried by Escherichia coli in marine, soil, human feces and livestock feces habitats are shown. Marine habitats are marked in red. Enlarged partial images show the habitats that are connected to the resistance genes in the marine habitats. These habitats are characterized by being adjacent to the marine habitats and belonging to the same phylogenetic cluster.
[0086] In summary, this invention, based on metagenomic datasets from different habitats, utilizes screening of antibiotic resistance genes, gene and species annotation, high-quality bacterial genome reconstruction, and nucleotide sequence alignment and phylogenetic analysis of the same resistance gene sequence in different habitats. This allows for the study of the cross-habitat distribution characteristics of antibiotic resistance genes from a gene homology perspective, revealing potential connectivity and gene homology between different habitats at the antibiotic resistance gene level. This comprehensive analysis of the spread characteristics of antibiotic resistance genes across multiple habitats provides technical support for homology assessment and cross-habitat diffusion research of antibiotic resistance genes in multiple habitats. Furthermore, it offers a reliable data foundation and analytical approach for subsequent resistance risk assessment, pollution source tracing, and the formulation of environmental resistance control strategies.
Claims
1. A method for tracing and analyzing the origins of antibiotic resistance genes in environmental samples and assessing their homology across multiple habitats, characterized in that, Includes the following steps: Step 1: Obtain environmental microbial metagenomic data for different habitats, specifically including: 1-1: Collect environmental samples from different habitats; 1-2: Extract microbial genomic DNA from environmental samples, perform metagenomic sequencing, and obtain raw metagenomic data from different habitats; 1-3: Quality control was performed on the raw metagenomic data to filter out low-quality sequences with a base quality of less than 30 and a sequence length of less than 150 bp, resulting in high-quality metagenomic datasets of each environmental sample after quality control. Step 2: Construct a multivariate abundance matrix of antibiotic resistance genes in different habitats: Based on the high-quality metagenomic datasets obtained in Steps 1-3, the ARGs-OAP tool is used to align with the structured antibiotic resistance gene database SARG v3.0 to accurately identify and quantify resistance genes in different habitats, and obtain the abundance matrix of diverse resistance genes in different habitats. Step 3: Construct the FEAST pollution source model for microbial source tracing using machine learning algorithms to identify and quantify the source contribution rate of potential source habitats to the composition of drug-resistant genes in the target habitat: Based on the abundance matrix of diverse drug-resistant genes in different habitats obtained in Step 2, define one habitat as the target habitat and screen potential source habitats; based on the drug-resistant gene abundance matrix in potential source habitats, construct the FEAST source tracing model using the machine learning expectation-maximization algorithm, and verify and optimize the model's prediction accuracy and generalization ability through the cross-residue-one-out iteration method and the simulated dataset method; use the model to predict the source-sink relationship between habitats and calculate the source contribution rate of each source habitat to the target habitat; Step 4: Select habitats for connectivity and gene homology analysis: Based on the source-sink habitat prediction relationship and its contribution rate obtained in Step 3, select habitats with a contribution rate >1% and perform connectivity and / or gene homology analysis with the target habitat. Define the selected habitats as potential connectivity habitats. Step 5: Extract high-quality metagenomic datasets of potential connected habitats: From the high-quality metagenomic datasets obtained in Steps 1-3, filter and extract high-quality metagenomic datasets of potential connected habitats as described in Step 4. Step 6: Reconstructing high-quality bacterial genomes from potentially connected habitats: Based on the high-quality metagenomic datasets from the potentially connected habitats extracted in Step 5, the MetaWRAP genome binning analysis platform was called via conda in the Linux terminal. The cat tool was used to merge the paired-end sequences of the high-quality metagenomic datasets of each sample into one file. The Megahit genome assembly tool was used to assemble the merged short genome sequences. The Metabat2, Concoct, and Maxbin algorithms were used to bin the assembled long sequence contig files. The binning results were purified and filtered using BIN_REFINEMENT, and finally, high-quality bacterial genomes with integrity ≥ 50% and contamination ≤ 10% were retained. Step 7: Annotate antibiotic resistance genes and species on the high-quality bacterial genomes reconstructed in potentially connected habitats: For the high-quality bacterial genomes reconstructed in Step 6, use Prodigal to predict the protein sequences of open reading frames (ORFs) in the bacterial genome. Use the BLASTP tool to compare the ORF sequences with the structured antibiotic resistance gene database SARGv3.0, selecting sequences with similarity ≥ 80%, coverage ≥ 75%, and e-value ≤ 10. -7 The standard identification of drug resistance gene sequences was used to obtain high-quality bacterial genomes carrying drug resistance genes, which were defined as ARG-MAGs; at the same time, the GTDB_tk tool was used to annotate the species of ARG-MAGs to obtain the microbial host information and classification results of ARG-MAGs in potential connected habitats. Step 8: Construct a phylogenetic tree of antibiotic resistance genes in potentially connected habitats: Based on the microbial hosts of ARG-MAGs in potentially connected habitats from Step 7, extract nucleotide sequences of the same class of antibiotic resistance genes and perform nucleotide sequence alignment using the clustalW tool; based on the alignment results, construct a phylogenetic tree of antibiotic resistance genes in potentially connected habitats using the Fasttree tool; extract nucleotide sequences of the same class of antibiotic resistance genes from the reconstructed high-quality bacterial genomes. Step 9: Clarify the cross-habitat transmission characteristics of antibiotic resistance genes: Based on the phylogenetic tree of antibiotic resistance genes in potential connected habitats in Step 8, compare and analyze the homology characteristics of resistance genes between potential source habitats and target habitats based on clustering relationships, branching structures, and differences in evolutionary branch lengths, construct a multi-media transmission map of resistance genes and their inter-host transmission network; comprehensively evaluate the cross-habitat transmission characteristics of antibiotic resistance genes in potential connected habitats.
2. The method according to claim 1, characterized in that, The different habitats mentioned in step 1-1 include soil, water bodies, sediments, sewage treatment systems, biological environments, and human activity environments.
3. The method according to claim 1, characterized in that, Step 2, which describes the precise identification and quantification of antibiotic resistance genes in different habitats, specifically includes: importing the quality-controlled sample metagenomic dataset file into the Linux terminal, calling the ARGs-OAP tool via conda, and performing precise annotation, classification, and quantification analysis of antibiotic resistance genes in two steps: rapid screening and precise classification. The output results include quantification results at three different levels: type, subtype, and gene. The gene level is selected as the abundance file for quantifying antibiotic resistance genes in multiple habitats.
4. The method according to claim 1, characterized in that, Step 3, which describes the construction of the FEAST plastic source model using the expectation-maximization algorithm in machine learning, specifically includes: preparing input files for source tracing analysis, including an abundance file of antibiotic resistance genes and a classification information file. The abundance file is formatted as follows: each column represents a sample, and each row represents the category of antibiotic resistance genes. The classification information file defines the target habitat as the sink and other habitats as the source. The abundance file and classification information file are imported into R, and after installing the software packages required for the FEAST model, the FEAST plastic source model is constructed based on the expectation-maximization algorithm.
5. The method according to claim 1, characterized in that, Step 8, which involves extracting nucleotide sequences of antibiotic resistance genes of the same class, specifically includes: screening for common antibiotic resistance gene classes carried by ARG-MAGs; and using the seqkit tool to extract nucleotide sequences of the same class of antibiotic resistance genes from high-quality genomes of potentially connected habitats based on the open reading frame (ORF) sequence number and genome sequence number in the BLASTP alignment results.
Citation Information
Patent Citations
Pollutant tracing method based on environmental microbiome database
CN116682497A
Environmental drug resistance group analysis method based on metagenome
CN117174165A
Resistance gene detection method and system based on metagenomics
CN120412707A
Rescue system using Portable LED Lantern and
KR102464679B1
Computational exploration of the global microbiome for antibiotic discovery
WO2025050118A1