Machine learning based method for assessing the risk of bacterial antibiotic resistance in environmental samples
By identifying high-risk antibiotic resistance genes in environmental samples and combining them with machine learning models, this approach addresses the issues of limited risk identification dimensions and insufficient analysis of driving factors in existing technologies. It enables multi-dimensional correlation characterization and interpretability assessment of antibiotic resistance risks in environmental samples, thereby improving the accuracy and quantifiability of risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-01-29
- Publication Date
- 2026-07-21
Smart Images

Figure CN121983139B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of environmental engineering, bioinformatics, and environmental health risk assessment, and particularly to a method for comprehensively assessing the antibiotic resistance risk of bacteria in environmental samples using metagenomics analysis combined with machine learning models. This method is applicable to the identification of antibiotic resistance risk, analysis of driving factors, and quantitative characterization of risk in environmental media samples such as air, water, or soil. Background Technology
[0002] Antimicrobial resistance (AMR) has evolved from a simple clinical treatment challenge into a typical problem of novel environmental pollutants and environmental health risks. Environmental stresses such as antibiotics and their metabolites, disinfectants, and heavy metals can exert continuous selective pressure in wastewater treatment plants, livestock emissions, hospital and pharmaceutical industry emissions, urban runoff, and contaminated soil / sediment, promoting the enrichment and spread of antibiotic resistance genes (ARGs) in environmental microbial communities. ARGs in the environment can spread between different microorganisms through horizontal gene transfer mediated by mobile genetic elements, and under certain conditions, they can associate with pathogenic hosts, thus posing a potential risk to human health. Simultaneously, host bacteria of antibiotic resistance genes in environmental media may enter human exposure pathways via aerosols, water bodies, the food chain, and soil contact, posing serious health threats to humans. Therefore, conducting antibiotic resistance risk assessments on environmental samples requires not only focusing on the presence and abundance of resistance genes, but also comprehensively considering their mobility, host pathogenicity, and the effects of environmental influencing factors to achieve accurate risk identification and control.
[0003] In existing practices of environmental antibiotic resistance risk assessment, common problems include: (1) the risk identification dimension is too singular: some methods mainly rely on the detection and abundance of ARGs for characterization, making it difficult to consider key dimensions that are more directly related to population infection and health risks, such as "mobility" and "pathogenic host mediation"; (2) insufficient ability to analyze driving factors: environmental physicochemical indicators, meteorological and hydrological parameters and socioeconomic parameters may have an important impact on resistance risk factors, but traditional statistics or experience judgments are often difficult to handle multi-source heterogeneous factors and nonlinear relationships, and it is also difficult to give a stable ranking of influencing factors; (3) lack of interpretable and referable comprehensive quantitative framework, risk standardization methods and technical processes that include the above factors.
[0004] Therefore, there is an urgent need for a comprehensive risk assessment technology based on metagenomic big data sets and machine learning big models to identify high-risk antibiotic resistance genes, interpret key environmental driving factors, and ensure the comparability and interpretability of environmental samples. Summary of the Invention
[0005] The purpose of this invention is to provide a machine learning-based method for assessing the risk of bacterial antibiotic resistance in environmental samples, overcoming the limitations of existing technologies that rely solely on the presence of bacterial resistance genes to biasedly assess environmental resistance risk. This invention incorporates multidimensional risk attributes such as the mobility of resistance genes and host pathogenicity, and uses a machine learning model to achieve nonlinear impact analysis driven by multiple environmental factors and a comprehensive quantitative assessment of the antibiotic resistance risk in samples.
[0006] The specific technical solution for achieving the objective of this invention is as follows:
[0007] A machine learning-based method for assessing the risk of antibiotic resistance in bacteria in environmental samples includes the following steps:
[0008] Step 1: Obtain second-generation metagenomic sample data from environmental samples, and perform quality control to obtain a metagenomic dataset;
[0009] Step 2: Identify high-risk antibiotic resistance genes in environmental samples, including:
[0010] 2-1: The metagenomic dataset obtained in step 1 is assembled using the assembly module of the metaWRAP software to obtain the contiguous sequences of each sample, and the open reading frames are predicted using the prodigal software to obtain the open reading frame sequences.
[0011] 2-2: Align the open reading frame sequences obtained in step 2-1 with the SARG antibiotic resistance gene database to obtain contig sequences carrying antibiotic resistance genes;
[0012] 2-3: Align the open reading frame sequences obtained in step 2-1 with the mobileOG-db mobile genetic element database to obtain contiguous sequences carrying mobile genetic elements. Define antibiotic resistance genes and mobile genetic elements as having mobility when they are in the same contiguous group and the distance between them is less than 5 kb. Based on the definition, obtain a list of mobile antibiotic resistance genes.
[0013] 2-4: The contiguous sequences carrying antibiotic resistance genes obtained in step 2-2 are compared with the GTDB database for species annotation. The host of antibiotic resistance genes in the publicly available list of 1,005 clinically relevant species is defined as the pathogen. Based on the definition, contiguous sequences and antibiotic resistance gene lists whose hosts are pathogens are obtained.
[0014] 2-5: Define antibiotic resistance genes that are both mobile and host pathogens as high-risk antibiotic resistance genes. Based on the definition, the intersection of the list of mobile antibiotic resistance genes obtained in step 2-3 and the list of antibiotic resistance genes that host pathogens obtained in step 2-4 is recorded to obtain the types of high-risk antibiotic resistance genes in the environmental sample.
[0015] Step 3: Use the ARGs-OAP workflow to identify and quantify antibiotic resistance genes in the metagenomic dataset obtained in Step 1. Parameter setting: e value ≤ 10 -7 Similarity ≥ 80%, coverage ≥ 75%, minimum matching length ≥ 25 amino acids, and based on the types of high-risk antibiotic resistance genes obtained in step 2, the types and relative abundance of high-risk antibiotic resistance genes in the environmental samples are obtained.
[0016] Step 4: Obtain the environmental parameter index data corresponding to the environmental sample described in Step 1;
[0017] Step 5: Based on the types and relative abundance of high-risk antibiotic resistance genes obtained in Step 3 and the environmental parameter index data obtained in Step 4, construct the input dataset for building the machine learning model, and perform missing value processing and normalization on the dataset to obtain the processed dataset.
[0018] Step 6: Divide the processed dataset obtained in Step 5 into a training set and a test set, with the training set comprising 80% and the test set comprising 20%. Train at least two machine learning models based on the training set data, and obtain candidate optimal models through cross-validation, hyperparameter optimization, and feature selection. Analyze the test set data using the candidate optimal models and calculate the coefficient of determination R of the candidate optimal models. 2 Compared with the mean squared error (MSE), select the one with the highest R. 2 The model with the lowest value and the lowest MSE value is taken as the optimal model;
[0019] Step 7: Perform SHAP interpretability analysis on the optimal model obtained in Step 6, output the ranking results of environmental impact factors of high-risk antibiotic resistance genes, and screen the top three environmental impact factors in the ranking results.
[0020] Step 8: Using the three factors of antibiotic resistance, mobility, and host pathogenicity, establish a resistance risk assessment model with the following formula. Use a self-built Python script to count the number of contiguous sequences mentioned in Step 2, and calculate the antibiotic resistance risk value of bacteria in the environmental samples according to the formula:
[0021]
[0022] In the formula: Q Struct This indicates the antibiotic resistance risk value of the bacteria; N Contig N represents the total number of contiguous group sequences in the sample; ARG N represents the number of contiguous sequences containing only antibiotic resistance genes; ARG,MGE This indicates the number of contiguous sequences containing both antibiotic resistance genes and mobile genetic elements, specifically the physical co-localization of antibiotic resistance genes and mobile genetic elements within the same contiguous group and less than 5 kb apart; N ARG,MGE,PAT N represents the number of contiguous sequences in a host that is a pathogen and simultaneously carries antibiotic resistance genes and mobile genetic elements; ARG,PAT This indicates the number of contiguous sequences that are pathogens in the host and also carry antibiotic resistance genes.
[0023] Step 9: Using the antibiotic resistance risk value of bacteria in the environmental sample obtained in Step 8 and the top three environmental impact factors obtained in Step 7, obtain the comprehensive antibiotic resistance risk assessment result that can be quantified and explained in the environmental sample.
[0024] Furthermore, the environmental sample mentioned in step 1 includes one of the following media: air, water, soil, and sludge.
[0025] Furthermore, the environmental parameter index data corresponding to the environmental sample mentioned in step 4 includes the sample's physicochemical indicators, the meteorological and hydrological parameters of the sample's location, and the socio-economic and public service indicators of the sample's location. The sample's physicochemical indicators include one or more of temperature, pH value, salinity, dissolved oxygen, total nitrogen, total phosphorus, chemical oxygen demand, and biochemical oxygen demand. The meteorological and hydrological parameters of the sample's location include one or more of rainfall, air pressure, wind speed, particulate matter concentration, runoff, water depth, and hydraulic residence time. The socio-economic and public service indicators of the sample's location include one or more of GDP, population density, labor force, immigration rate, employment-to-population ratio, trade rate, aquaculture production, fertilizer consumption, current per capita health expenditure, number of hospital beds, and antibiotic usage.
[0026] Furthermore, the machine learning model described in step 6 is selected from at least two of the following: random forest model, extreme random tree model, extreme gradient boosting model, and lightweight gradient boosting machine model.
[0027] Beneficial Effects: This invention, based on metagenomic data from environmental samples, assesses antibiotic resistance genes using the criteria of "mobility and pathogen carriage," expanding antibiotic resistance risk identification from a single detection level to a multi-dimensional health risk association characterization. Furthermore, it integrates environmental influencing factors such as sample physicochemical indicators, meteorological and hydrological data, and socioeconomic factors, utilizing machine learning models and SHAP interpretability analysis to rank the importance of driving factors, providing a basis for risk management. A comprehensive risk assessment model for resistance groups is established to calculate the antibiotic resistance risk value of environmental samples. This enables an interpretable and quantifiable comprehensive antibiotic resistance risk assessment among different samples, improving the comparability of results and the directionality of risk management. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method for assessing the risk of antibiotic resistance in bacteria in environmental samples according to the present invention;
[0029] Figure 2 This is a schematic diagram showing the ranking results of environmental factors affecting high-risk antibiotic resistance genes in this invention.
[0030] Figure 3 This is a schematic diagram illustrating the risk assessment of bacterial antibiotic resistance in samples from different aquatic environments according to the present invention. Detailed Implementation
[0031] The embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments.
[0032] See Figure 1 The method for assessing the risk of bacterial antibiotic resistance in environmental samples according to the present invention includes:
[0033] 1) Obtain second-generation metagenomic sample data from environmental samples, and perform quality control to obtain a metagenomic dataset;
[0034] Step 1), specifically includes:
[0035] 1-1) Second-generation metagenomic sample data can be collected according to the sampling method requirements of different environmental samples (such as air, water and soil); or it can be searched in literature databases (such as Web of Science) or public databases (such as NCBI) according to the keywords corresponding to different environmental samples, and samples that do not meet the requirements are manually screened and removed, and basic information such as sampling time, latitude and longitude, and data size of the corresponding samples are compiled.
[0036] 1-2) Use the fastp tool to perform quality control on the second-generation metagenomic samples obtained in step 1-1), remove gene sequences with a length of less than 100 bp and a quality value of less than 20, and obtain the metagenomic dataset.
[0037] 2) Identify high-risk antibiotic resistance genes in environmental samples;
[0038] Step 2), specifically includes:
[0039] 2-1) The metagenomic dataset obtained in step 1) is assembled using the assembly module of the metaWRAP software to obtain the contiguous group sequence of each sample, and the open reading frame is predicted using the prodigal software to obtain the open reading frame sequence.
[0040] 2-2) Align the open reading frame sequences obtained in step 2-1) with the SARG antibiotic resistance gene database to obtain contiguous group sequences carrying antibiotic resistance genes;
[0041] 2-3) The open reading frame sequences obtained in step 2-1) are compared with the mobileOG-db mobile genetic element database to obtain contiguous sequences carrying mobile genetic elements. Antibiotic resistance genes and mobile genetic elements are defined as having mobility when they are in the same contiguous group and the distance between them is less than 5 kb. A list of mobile antibiotic resistance genes is obtained by screening according to the definition.
[0042] 2-4) The contiguous sequences carrying antibiotic resistance genes obtained in step 2-2) are compared with the GTDB database for species annotation. The host of antibiotic resistance genes in the publicly available list of 1,005 clinically relevant species is defined as the pathogen. Based on the definition, contiguous sequences and antibiotic resistance gene lists with the host as the pathogen are obtained.
[0043] 2-5) Define antibiotic resistance genes that are both mobile and host pathogens as high-risk antibiotic resistance genes. Based on the definition, the intersection of the list of mobile antibiotic resistance genes obtained in step 2-3) and the list of antibiotic resistance genes that host pathogens obtained in step 2-4) is recorded to obtain the types of high-risk antibiotic resistance genes in the environmental sample.
[0044] 3) The ARGs-OAP workflow was used to identify and quantify antibiotic resistance genes in the metagenomic dataset obtained in step 1). Parameter settings: e value ≤ 10. -7Similarity ≥ 80%, coverage ≥ 75%, minimum matching length ≥ 25 amino acids, and based on the high-risk antibiotic resistance gene types obtained in step 2), the types and relative abundance of high-risk antibiotic resistance genes in the environmental samples are obtained.
[0045] Step 3), specifically includes:
[0046] 3-1) Use BWA (v0.7.17) and BLASTn to compare the samples with the 16S database "Greengenes" to calculate the number of 16S genes in each sample. Use Diamond (v.2.1.6) to compare the single copy marker gene databases of bacteria and archaea to calculate the number of cells in each sample. The results are recorded in a text file named "metadata.txt".
[0047] 3-2) Using the results of 3-1) as input data, compare the data using BLAST with the SARG database (parameter setting: e value ≤ 10). -7 The criteria for ARGs are: similarity ≥ 80%, coverage ≥ 75%, minimum matching length ≥ 25 amino acids. This allows for the acquisition of annotation information for major ARG classes and subtypes in each sample, and accurate quantification of the abundance of each ARG, expressed in ARG copy number per cell (copy / cell).
[0048] 3-3) Based on the screening of high-risk antibiotic resistance genes obtained in step 2), the types and relative abundance of high-risk antibiotic resistance genes in the environmental samples are obtained.
[0049] 4) Obtain the environmental parameter index data corresponding to the environmental sample described in step 1);
[0050] In step 4), the environmental parameter index data corresponding to the environmental sample specifically includes the sample's physicochemical indicators, the meteorological and hydrological parameters of the sample's location, and the socio-economic and public service indicators of the sample's location. The sample's physicochemical indicators include one or more of temperature, pH value, salinity, dissolved oxygen, total nitrogen, total phosphorus, chemical oxygen demand, and biochemical oxygen demand. The meteorological and hydrological parameters of the sample's location include one or more of rainfall, air pressure, wind speed, particulate matter concentration, runoff, water depth, and hydraulic residence time. The socio-economic and public service indicators of the sample's location include one or more of GDP, population density, labor force, immigration rate, employment-to-population ratio, trade rate, aquaculture production, fertilizer consumption, current per capita health expenditure, number of hospital beds, and antibiotic usage.
[0051] 5) Based on the types and relative abundance of high-risk antibiotic resistance genes obtained in step 3) and the environmental parameter index data obtained in step 4), construct an input dataset for building a machine learning model, and perform missing value processing and normalization on the dataset to obtain the processed dataset.
[0052] 6) Divide the processed dataset obtained in step 5) into a training set and a test set, with the training set comprising 80% and the test set comprising 20%. Train at least two machine learning models based on the training set data, and obtain candidate optimal models through cross-validation, hyperparameter optimization, and feature selection. Analyze the test set data using the candidate optimal models and calculate the coefficient of determination R of the candidate optimal models. 2 Compared with the mean squared error (MSE), select the one with the highest R. 2 The model with the lowest value and the lowest MSE value is taken as the optimal model;
[0053] Step 6), specifically includes:
[0054] 6-1) Use Python to read the dataset processed in step 5) and divide it into a training set (80%) and a test set (20%).
[0055] 6-2) Train at least two machine learning algorithms (random forest model, extreme random tree model, extreme gradient boosting model and lightweight gradient boosting machine model) using the training set, and perform a 10-fold cross-validation step to avoid overfitting;
[0056] 6-3) After hyperparameter optimization and feature selection by the RandomizedSearchCV module, the best model parameters are generated for each algorithm, and the candidate optimal model after training is obtained.
[0057] 6-4) Analyze the test set data using the trained candidate optimal model and calculate the determination coefficient R of the candidate optimal model. 2 Compared with the mean squared error (MSE), the results with the highest R-value are selected. 2 The model with the lowest MSE value is selected as the optimal model.
[0058] 7) Perform SHAP interpretability analysis on the optimal model obtained in step 6), output the ranking results of environmental impact factors of high-risk antibiotic resistance genes, and screen the top three environmental impact factors in the ranking results.
[0059] 8) Using the three factors of antibiotic resistance, mobility, and host pathogenicity, a resistance risk assessment model was established with the following formula. A self-built Python script was used to count the number of contiguous sequences mentioned in step 2), and the antibiotic resistance risk value of bacteria in the environmental samples was calculated according to the formula:
[0060]
[0061] In the formula: Q Struct This indicates the antibiotic resistance risk value of the bacteria; N Contig N represents the total number of contiguous group sequences in the sample; ARG N represents the number of contiguous sequences containing only antibiotic resistance genes; ARG,MGE This indicates the number of contiguous sequences containing both antibiotic resistance genes and mobile genetic elements, specifically the physical co-localization of antibiotic resistance genes and mobile genetic elements within the same contiguous group and less than 5 kb apart; N ARG,MGE,PAT N represents the number of contiguous sequences in a host that is a pathogen and simultaneously carries antibiotic resistance genes and mobile genetic elements; ARG,PAT This indicates the number of contiguous sequences that are pathogens in the host and also carry antibiotic resistance genes.
[0062] 9) Using the antibiotic resistance risk value of bacteria in the environmental sample obtained in step 8) and the top three environmental impact factors obtained in step 7), obtain the comprehensive antibiotic resistance risk assessment result that can be quantified and explained in the environmental sample.
[0063] Example 1: Identification and assessment of high-risk antibiotic resistance genes in ambient air.
[0064] This embodiment uses ambient air as the object to demonstrate the assessment process for high-risk antibiotic resistance factors based on second-generation metagenomics.
[0065] First, samples that did not meet the research objectives were retrieved from the Web of Science and NCBI databases using combined keywords ("air*"), "atmospher*", "PM*", "particulate matter", "total suspended particle", and "aerosol*") and ("metagenom*"), "high-throughput sequencing", and "Illumina"). Then, publicly available second-generation metagenomic data obtained from literature databases and public databases were merged and deredundantd. Finally, ambient air samples from five provincial capitals or municipalities in China were selected, and basic information such as sampling time, latitude and longitude, location, and data volume were compiled.
[0066] The original data underwent quality control, with FPAP used for pruning and quality filtering, and KneadData used to remove eukaryotic influences, resulting in a quality-controlled metagenomic dataset. The quality-controlled metagenomic dataset was then assembled using the assembly module of MetaWRAP software to obtain contiguous sequences for each sample. Open reading frames (OPFs) were predicted using Prodigal software on these contiguous sequences, yielding OPF sequences. These OPF sequences were compared to the SARG antibiotic resistance gene database and the MobileOG-db mobile genetic element database. Based on the co-localization of antibiotic resistance genes and mobile genetic elements within the same contiguous group and a distance of less than 5 kb, the mobility of antibiotic resistance genes was determined. Species annotation was then performed on the OPF sequences, and the annotation results were compared with a publicly available list of pathogens containing 1,005 clinically relevant species to identify whether the host is a pathogen. Finally, antibiotic resistance genes carrying both antibiotic resistance genes and mobile genetic elements (less than 5 kb) in a host pathogen were defined as high-risk antibiotic resistance genes, thus identifying the types of high-risk antibiotic resistance genes in ambient air.
[0067] Table 1 lists the high-risk antibiotic resistance genes in the ambient air of this region. Ten high-risk antibiotic resistance genes were identified in the ambient air of this embodiment, and the types of antibiotics they are resistant to are shown in the table. These antibiotic resistance genes are both mobile and carried by pathogen hosts, posing a high risk of infection to the population.
[0068] Table 1. High-risk drug resistance genes identified in atmospheric environmental samples
[0069] Serial Number High-risk drug resistance genes Classes of resistant antibiotics 1 APH(3'')-Ib Aminoglycosides 2 aadA Aminoglycosides 3 cmlA5 Chloramphenicol 4 fexA Chloramphenicol 5 erm(42) Macrolide-Lincoamide-Streptocin 6 lnuG Macrolide-Lincoamide-Streptocin 7 mphB Macrolide-Lincoamide-Streptocin 8 emrD Multi-drug 9 eptA Polymyxins 10 tetM Tetracyclines
[0070] Example 2: Analysis of Environmental Factors Driving High-Risk Antibiotic Resistance Genes in the Atmosphere
[0071] Based on Example 1, this embodiment collects data on 18 environmental parameters (including temperature, rainfall, fine particulate matter concentration, GDP, labor force, employment-to-population ratio, trade rate, merchandise trade, aquaculture production, fertilizer consumption, population density, number of immigrants, current per capita health expenditure, number of hospital beds, and antibiotic use) through statistical yearbooks (https: / / www.stats.gov.cn / sj / ndsj / ) and World Bank (https: / / data.worldbank.org.cn). Subsequently, the ARGs-OAP workflow was used to identify and quantify airborne ARGs. The specific steps consisted of two parts: 1) Using BWA (v0.7.17) and BLASTn, the samples were compared with the 16S database "Greengenes" to calculate the number of 16S genes in each sample. Diamond (v.2.1.6) was used to compare the single-copy marker gene databases of bacteria and archaea to calculate the number of cells in each sample. The results were recorded in a text file named "metadata.txt"; 2) Using the results from step 1 as input data, BLAST was compared with the SARG database (parameter setting: e value ≤ 10). -7 The criteria for ARGs were: similarity ≥80%, coverage ≥75%, and minimum matching length ≥25 amino acids. Annotation information for the major categories and subtypes of ARGs in each sample was obtained, and the abundance of each ARG was accurately quantified, expressed as ARG copy number per cell (copy / cell). High-risk ARGs were screened based on the high-risk antibiotic resistance gene table obtained in Example 1, and machine learning methods were used to analyze the environmental factors driving high-risk antibiotic resistance genes in the atmosphere.
[0072] The specific steps of machine learning include: using the types and relative abundance of high-risk antibiotic resistance genes and collected environmental parameter data as input data, firstly, after handling missing values and normalizing the data, the dataset is divided into a training set (80%) and a test set (20%), and then a random forest model, an extreme random tree model, and an extreme gradient boosting model are trained respectively. 10-fold cross-validation is used to avoid overfitting. After hyperparameter optimization and feature selection, the optimal model is selected using metrics such as R² and MSE. Finally, SHAP is used for interpretability analysis to output the ranking results of environmental impact factors. Taking the lnuG gene as an example... Figure 2 This is a schematic diagram showing the ranking of environmental factors affecting high-risk antibiotic resistance genes. The lnuG gene is resistant to macrolide-lincosamide-streptomycin antibiotics, and temperature, antibiotic usage, and rainfall are the main environmental factors affecting its existence.
[0073] Example 3: Comparison of antibiotic resistance risk values in typical urban water environments
[0074] To demonstrate the versatility of this invention across various environmental media, this embodiment uses water bodies as the research object, showcasing the comprehensive risk assessment results of antibiotic resistance quantified using three factors: antibiotic resistance, mobility, and host pathogenicity. First, samples were collected from the influent, activated sludge, and effluent of a wastewater treatment plant, as well as samples from 300 meters upstream, near, and 300 meters downstream of the plant's discharge outlet. Using bioinformatics techniques and the resistance risk assessment model established in this invention, the comprehensive risk of antibiotic resistance in typical urban water environments was investigated. Figure 3 This diagram illustrates the risk assessment of bacterial antibiotic resistance in samples from different aquatic environments. The results show that the risk of antibiotic resistance is highest in the influent of wastewater treatment plants. As the wastewater treatment process progresses, the risk gradually decreases in activated sludge and effluent. However, the risk increases sequentially from upstream to downstream of the wastewater treatment plant outlet. Using the antibiotic resistance risk assessment model established in this invention to calculate the bacterial antibiotic resistance risk in aquatic environments allows for a direct assessment of the antibiotic resistance risk in samples from different aquatic environments.
Claims
1. A method for assessing the risk of bacterial antibiotic resistance in environmental samples based on machine learning, characterized in that, Includes the following steps: Step 1: Obtain second-generation metagenomic sample data from environmental samples, and perform quality control to obtain a metagenomic dataset; Step 2: Identify high-risk antibiotic resistance genes in environmental samples, including: 2-1: The metagenomic dataset obtained in step 1 is assembled using the assembly module of the metaWRAP software to obtain the contiguous sequences of each sample, and the open reading frames are predicted using the prodigal software to obtain the open reading frame sequences. 2-2: Align the open reading frame sequences obtained in step 2-1 with the SARG antibiotic resistance gene database to obtain contig sequences carrying antibiotic resistance genes; 2-3: Align the open reading frame sequences obtained in step 2-1 with the mobileOG-db mobile genetic element database to obtain contiguous sequences carrying mobile genetic elements. Define antibiotic resistance genes and mobile genetic elements as having mobility when they are in the same contiguous group and the distance between them is less than 5 kb. Based on the definition, obtain a list of mobile antibiotic resistance genes. 2-4: The contiguous sequences carrying antibiotic resistance genes obtained in step 2-2 are compared with the GTDB database for species annotation. The host of antibiotic resistance genes in the publicly available list of 1,005 clinically relevant species is defined as the pathogen. Based on the definition, contiguous sequences and antibiotic resistance gene lists whose hosts are pathogens are obtained. 2-5: Define antibiotic resistance genes that are both mobile and host pathogens as high-risk antibiotic resistance genes. Based on the definition, the intersection of the list of mobile antibiotic resistance genes obtained in step 2-3 and the list of antibiotic resistance genes that host pathogens obtained in step 2-4 is recorded to obtain the types of high-risk antibiotic resistance genes in the environmental sample. Step 3: Use the ARGs-OAP workflow to identify and quantify antibiotic resistance genes in the metagenomic dataset obtained in Step 1. Parameter setting: e value ≤ 10 -7 Similarity ≥ 80%, coverage ≥ 75%, minimum matching length ≥ 25 amino acids, and based on the types of high-risk antibiotic resistance genes obtained in step 2, the types and relative abundance of high-risk antibiotic resistance genes in the environmental samples are obtained. Step 4: Obtain the environmental parameter index data corresponding to the environmental sample described in Step 1; Step 5: Based on the types and relative abundance of high-risk antibiotic resistance genes obtained in Step 3 and the environmental parameter index data obtained in Step 4, construct the input dataset for building the machine learning model, and perform missing value processing and normalization on the dataset to obtain the processed dataset. Step 6: Divide the processed dataset obtained in Step 5 into a training set and a test set, with the training set comprising 80% and the test set comprising 20%. Train at least two machine learning models based on the training set data, and obtain candidate optimal models through cross-validation, hyperparameter optimization, and feature selection. Analyze the test set data using the candidate optimal models and calculate the coefficient of determination R of the candidate optimal models. 2 Compared with the mean squared error (MSE), select the one with the highest R. 2 The model with the lowest value and the lowest MSE value is taken as the optimal model; Step 7: Perform SHAP interpretability analysis on the optimal model obtained in Step 6, output the ranking results of environmental impact factors of high-risk antibiotic resistance genes, and screen the top three environmental impact factors in the ranking results. Step 8: Using the three factors of antibiotic resistance, mobility, and host pathogenicity, establish a resistance risk assessment model with the following formula. Use a self-built Python script to count the number of contiguous sequences mentioned in Step 2, and calculate the antibiotic resistance risk value of bacteria in the environmental samples according to the formula: ; In the formula: Q Struct This indicates the antibiotic resistance risk value of the bacteria; N Contig N represents the total number of contiguous group sequences in the sample; ARG N represents the number of contiguous sequences containing only antibiotic resistance genes; ARG,MGE This indicates the number of contiguous sequences containing both antibiotic resistance genes and mobile genetic elements, specifically the physical co-localization of antibiotic resistance genes and mobile genetic elements within the same contiguous group and less than 5 kb apart; N ARG,MGE,PAT N represents the number of contiguous sequences in a host that is a pathogen and simultaneously carries antibiotic resistance genes and mobile genetic elements; ARG,PAT This indicates the number of contiguous sequences that are pathogens in the host and also carry antibiotic resistance genes. Step 9: Using the antibiotic resistance risk value of bacteria in the environmental sample obtained in Step 8 and the top three environmental impact factors obtained in Step 7, obtain the comprehensive antibiotic resistance risk assessment result that can be quantified and explained in the environmental sample.
2. The method according to claim 1, characterized in that, The environmental sample mentioned in step 1 includes one of the following media: air, water, soil, and sludge.
3. The method according to claim 1, characterized in that, The environmental parameter index data corresponding to the environmental sample mentioned in step 4 include the sample's physicochemical indicators, the meteorological and hydrological parameters of the sample's location, and the socio-economic and public service indicators of the sample's location. The sample's physicochemical indicators include one or more of the following: temperature, pH value, salinity, dissolved oxygen, total nitrogen, total phosphorus, chemical oxygen demand, and biochemical oxygen demand. The meteorological and hydrological parameters of the sample's location include one or more of the following: rainfall, air pressure, wind speed, particulate matter concentration, runoff, water depth, and hydraulic residence time. The socio-economic and public service indicators of the sample's location include one or more of the following: GDP, population density, labor force, immigration rate, employment-to-population ratio, trade rate, aquaculture production, fertilizer consumption, current per capita health expenditure, number of hospital beds, and antibiotic usage.
4. The method according to claim 1, characterized in that, The machine learning model described in step 6 is selected from at least two of the following: random forest model, extreme random tree model, extreme gradient boosting model, and lightweight gradient boosting machine model.