Antibiotic resistance gene host analysis method based on interpretable machine learning

Through an interpretable machine learning method, the relationship between antibiotic resistance genes and host bacteria is analyzed, and problems such as time-consuming and difficult to quantify contributions in traditional methods are solved, and more accurate and interpretable data analysis is achieved to support the ecological risk management of antibiotic resistance genes.

CN120148651AActive Publication Date: 2025-06-13YANGZHOU UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510284448.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Traditional antibiotic resistance gene detection methods have problems such as long time-consuming, difficulty in culturing all types of bacteria, data analysis is susceptible to human factors, and difficulty in accurately quantifying the contribution of each bacterial genus to antibiotic resistance gene in a multi-bacterial environment.

Method used

The antibiotic resistance gene host analysis method based on interpretability machine learning was used to collect genomic data, remove outliers using Pauta criterion, screen key bacterial genes with random forest algorithms, and CatBoost regression algorithm to build correlation models, and perform interpretability analysis through Shap value to quantify the contribution of different bacterial genes to antibiotic resistance genes.

Benefits of technology

Accurate analysis of the relationship between antibiotic resistance genes and host bacteria in a polybacterial environment provides more accurate and interpretable data analysis results to support the ecological risk assessment and management of antibiotic resistance genes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148651A_ABST
    Figure CN120148651A_ABST
Patent Text Reader

Abstract

The invention discloses an antibiotic resistance gene host analysis method based on interpretable machine learning in the technical field of bioinformatics. Comprising the following steps: 1) genome data collection: collecting a large amount of genome data of a same environmental sample, calculating antibiotic resistance genes, species and genus types and relative abundance in the gene data, and processing the gene data into a matrix, analyzing the relative abundance of the selected antibiotic resistance gene by using a Pauta criterion, removing abnormal values, and dividing into a training set and a test set; 2) constructing an analysis model; and 3) interpretable correlation analysis: quantifying the contribution of different bacteria genus to the antibiotic resistance gene in the model through the Shap value, and judging the influence of the bacteria genus on the antibiotic resistance gene according to the difference of the average Shap value of different bacteria genus. And necessary technical support is provided for ecological risk evaluation and management of antibiotic resistance genes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly to a method for analyzing the host of antibiotic resistance genes. Background Art

[0002] Traditional methods for detecting antibiotic resistance genes mainly rely on culture medium detection, PCR (Polymerase Chain Reaction), and sequencing technologies. These methods can detect and identify antibiotic resistance genes to a certain extent, but there are some significant limitations. The culture medium detection method usually takes a long time and it is difficult to culture all types of bacteria. Although PCR and sequencing technologies can provide relatively accurate gene information, it is difficult to accurately quantify the contribution of each genus of bacteria to antibiotic resistance genes in a multi-genus environment.

[0003] In addition, traditional methods are easily affected by human factors during the data analysis and interpretation process and lack precise quantification criteria. These limitations restrict our in-depth understanding of the transmission pathways and mechanisms of antibiotic resistance genes.

[0004] With the rapid development of high-throughput sequencing technology, researchers can obtain a large amount of genomic data. These data provide a rich information resource for studying the transmission of antibiotic resistance genes. However, how to extract useful information from these massive data and accurately analyze the relationship between antibiotic resistance genes and host bacteria remains a huge challenge. Traditional bioinformatics analysis methods often seem inadequate when dealing with these complex data and are difficult to reveal deep patterns and associations.

[0005] Machine learning, as a powerful data analysis tool, has been widely used in the field of environmental microbiology in recent years. Machine learning algorithms have the ability to process large-scale data and complex relationships and can find potential patterns and associations in massive genomic data. For example, by constructing classification models and regression models, machine learning can help researchers analyze the transmission pathways of antibiotic resistance genes, identify key host bacteria, and evaluate the contribution of different genera of bacteria to antibiotic resistance genes. Compared with traditional methods, machine learning-based methods not only have significant improvements in accuracy and efficiency but also can provide more biological insights. Summary of the Invention

[0006] Aiming at the deficiencies in the prior art, the present invention provides a method for analyzing the host of antibiotic resistance genes based on interpretable machine learning, which can analyze the correlation between multiple genera of bacteria and antibiotic resistance genes, quantify the contribution of different genera of bacteria to antibiotic resistance genes, thereby identifying the host bacteria of antibiotic resistance genes and providing necessary technical support for the ecological risk assessment and management of antibiotic resistance genes.

[0007] The object of the present invention is achieved as follows: A method for analyzing the host of antibiotic resistance genes based on interpretable machine learning, comprising the following steps:

[0008] Step 1) Genomic data collection: Collect a large amount of genomic data of the same environmental samples, calculate the types and relative abundances of antibiotic resistance genes and species genera in the gene data and process them into a matrix, and use the Pauta criterion to analyze the relative abundances of the selected antibiotic resistance genes, remove outliers, and then divide the training set and the test set;

[0009] Step 2) Construction of the analysis model:

[0010] 2-1) First, use the random forest algorithm to establish an initial correlation analysis model for the species genus matrix;

[0011] 2-2) Calculate the percentage increase in mean squared error (%IncMSE) for each genus to measure the importance of the genus. The formula is as follows:

[0012]

[0013] Where: MSE original is the mean squared error when the feature values are not shuffled; MSE random is the mean squared error after shuffling the feature values;

[0014] Perform five 10-fold cross-validations to obtain the relationship between the number of the most important genera for fitting and the average cross-validation error, select the number of genera with the smallest average cross-validation error as the ideal number of variables, and screen out the final genus matrix. The cross-validation error formula is as follows:

[0015]

[0016] Where: k is the number of folds, which is 10; MSEi is the mean squared error of the i-th fold;

[0017] 2-3) Establish an antibiotic resistance gene-bacterial genus correlation model through the CatBoost regression algorithm, use grid search to select the optimal hyperparameters, and construct an antibiotic resistance gene-bacterial genus correlation model based on the RF-CatBoost algorithm;

[0018] The predicted value of the RF-CatBoost model is the weighted sum of the predictions of each decision tree:

[0019]

[0020] Where, is the initial predicted value, usually the mean of the target values; f m (x) is the output of the 𝑚-th tree; is the learning rate; M is the number of trees;

[0021] 2 - 4) Verify the goodness of fit, robustness, and analytical ability of the model;

[0022] Step 3) Explainable correlation analysis: Quantify the contribution of different genera of bacteria to antibiotic resistance genes in the model through Shap values, and judge the influence of genera on antibiotic resistance genes according to the difference in the average Shap values of different genera; The SHAP value is used to decompose the prediction result of the model into the contribution value of each feature, that is:

[0023]

[0024] where, is the predicted value of the model; is the baseline value of the model, usually the average initial value; is the Shap value of the i-th genus of bacteria, indicating the contribution of the genus x i to the model;

[0025] For the genus of bacteria x i , its SHAP value is defined as the weighted average marginal contribution of all possible subsets of genera of bacteria:

[0026]

[0027] where, is a subset of the genus of bacteria, and ; is the predicted value of the model under the subset of the genus of bacteria ; is the predicted value of the model after adding the genus of bacteria x to the subset i ; n is the total number of genera of bacteria, is the size of the subset ; is the coefficient for weighting the contribution of each subset.

[0028] Furthermore, the environmental samples in step 1) include one or more of the following: activated sludge, sewage and wastewater, river water, lake water, air, soil.

[0029] Furthermore, the specific method of dividing the training set and the test set in step 1) is as follows: Arrange the processed data in ascending order of the relative abundance of antibiotic resistance genes, divide 5 data into a group, put the fifth data in each group into the test set, and the remaining data form the training set; The data in the training set are used for model establishment and internal verification, while the data in the test set are used for external verification and performance evaluation of the model.

[0030] Further, the specific method for selecting the optimal hyperparameters using grid search in step 2-3) is as follows: perform grid search on the number of base decision trees, the maximum depth contained in each base decision tree, and the learning rate, and the parameter combination that minimizes the mean absolute error is the best parameter combination.

[0031] Further, the specific method for quantifying the contribution of different genera to antibiotic resistance genes in the model using Shap values in step 3) is as follows: through the output of the CatBoost model, calculate the contribution value of each genus to the analysis result of antibiotic resistance genes, extract the Shap values of each genus in all samples, calculate the average value and distribution of the Shap values of each genus in different samples, and sort the genera according to the absolute average value of the Shap values to identify the genus that contributes the most to antibiotic resistance genes.

[0032] Further, the specific method for judging the influence of genera on antibiotic resistance genes according to the difference in the average Shap values of different genera in step 3) is as follows: by analyzing the changes in the Shap values of each genus, evaluate its positive and negative effects on the model analysis result, and visualize the influence degree of different genera on antibiotic resistance genes. Combine the Shap value analysis result to explain the influence mechanism of the genus, clarify the role of specific genera in the analysis of antibiotic resistance genes, and assist the research and decision-making in related fields.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0034] 1) In the data processing stage, the present invention analyzes the relative abundance of antibiotic resistance genes in genomic data by using the Pauta criterion, thereby removing outliers. This process lays a high-quality data foundation for the subsequent training of the model and directly affects the accuracy and robustness of the model. At the same time, the training set and test set of the data are strictly divided according to the distribution of the relative abundance of antibiotic resistance genes, ensuring the rationality of the data distribution and providing reliable support for the subsequent construction and verification of the model.

[0035] 2) After data preprocessing, in step 2-1), a preliminary correlation model is established by using the random forest algorithm to screen the key genus matrix. The random forest algorithm not only effectively screens out the most influential genera but also significantly reduces the data dimension and the risk of overfitting.

[0036] 3) In step 2-3), a correlation model between antibiotic resistance genes and bacterial genera was constructed through the RF-CatBoost algorithm, and the performance of the model was fully verified through robustness and external validation in step 2-4). Subsequently, in step 3), the Shap value method was used to deeply analyze the model, and the contributions of different genera to antibiotic resistance genes were quantitatively analyzed. The Shap value results not only verified the rationality of the screening and modeling in the previous two steps but also provided a scientific basis for the ecological mechanism of specific genera on antibiotic resistance genes.

[0037] 4) The invention combines the Pauta criterion, random forest, and CatBoost algorithm to solve the problems of unscientific outlier handling, low efficiency of genus screening, and insufficient model interpretability in traditional bioinformatics methods. Especially by organically combining genus screening, model construction, and interpretability analysis, the present invention can more comprehensively reveal the complex relationship between antibiotic resistance genes and bacterial genera. This multi-step linkage optimization strategy provides an innovative technical path for the analysis of antibiotic resistance gene hosts. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0039] Figure 1 It is a flowchart of the dataset in the embodiment of the present invention.

[0040] Figure 2 It is a graph of the fitting degree and verification results of the correlation model in the embodiment of the present invention.

[0041] Figure 3 It is a SHAP value graph of the correlation model in the embodiment of the present invention.

[0042] Figure 4 It is a graph of the importance of the SHAP value of the correlation model in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0044] AsFigure 1 A host analysis method for antibiotic resistance genes based on machine learning is shown as follows, including the following steps.

[0045] Step 1) Genome data collection: Collect 755 metagenomic data of sludge samples in wastewater treatment plants, including different biomass samples such as aerobic sludge, anaerobic sludge, biofilm, and granular sludge; in this embodiment, the annotation of tetracycline antibiotic resistance genes is selected as the analysis object. The antibiotic resistance genes are obtained using the RGI software, and the species distribution characteristics in the genomic data are obtained through the Kaiju software. The antibiotic resistance genes and species annotation results are processed into a matrix using Python scripts; the relative abundances of the selected antibiotic resistance genes are analyzed using the Pauta criterion to remove outliers; the collected data are sorted in ascending order of the relative abundances of the antibiotic resistance genes, with 5 data in a group, and the fifth data in each group is put into the test set, and the remaining data form the training set; the data in the training set are used for model establishment and internal verification, and the data in the test set are used for external verification and performance evaluation of the model.

[0046] Step 2) Construction of the analysis model: Use the species genus matrix in the genome as the independent variable, and establish an initial correlation model based on the random forest algorithm; finally, calculate the percentage increase in mean squared error (%IncMSE) of each descriptor to measure the importance of the descriptor, and perform five 10-fold cross-validations to obtain the relationship between the number of the most important descriptors for fitting and the average cross-validation error. Select the number of descriptors with the smallest average cross-validation error as the ideal number of variables, and retain the top 12 most important genera (Sutterella, Laribacter, Haliscomenobacter, Pusillimonas, Alcaligenes, Entomoplasma, Stenotrophomonas, Saprospira, Alicycliphilus, Arcticibacterium, Ilumatobacter, Thauera) to obtain ideal analysis results;

[0047] Use the final descriptor as the independent variable, and establish a tetracycline resistance gene-species genus correlation model through the classification boosting (CatBoost) regression algorithm; perform grid searches on several hyperparameters such as iterations, depth, and learning_rate to obtain the parameter combination with the smallest mean absolute error as the best parameter combination (iterations = 1000, depth = 6, and learning_rate = 0.05); use the final genus as the independent variable, set the hyperparameter combination as the best parameter combination, and establish a TET resistance gene-species genus correlation model through the classification boosting regression algorithm.

[0048] Use R2 tra to characterize the model fitting degree; use Q2 BOOT to characterize the model robustness; use R2 test to represent the model external analysis ability; judgment basis: R 2 > 0.7, Q 2 > 0.6, R 2 - Q 2 < 0.3, CCC > 0.85; the specific parameters are as follows:

[0049] n tra = 584, R2 tra = 0.976, Q2 BOOT = 0.794; n ext = 147, R2 test = 0.834.

[0050] Among them, n tra and n ext are the number of genomes in the training set and the test set respectively. R2 tra is 0.976, indicating that the model has a high linear fitting ability; the internal validation coefficient Q2 BOOT is 0.794, indicating that the correlation model has good robustness; R2 test of external validation = 0.834, indicating that the model has good external analysis ability. Figure 2 Give the fitting degree and verification results of the model; specifically include:

[0051] 2-1) First, use the random forest algorithm to establish an initial correlation analysis model for the species-genus matrix;

[0052] 2-2) Calculate the percentage increase in mean squared error (%IncMSE) of each genus to measure the importance of the genus. The formula is as follows:

[0053]

[0054] Among them: MSE original is the mean squared error when the feature values are not shuffled; MSE random is the mean squared error after shuffling the feature values;

[0055] Perform five times of 10-fold cross-validation to obtain the relationship between the number of the most important genera for fitting and the average cross-validation error. Select the number of genera with the smallest average cross-validation error as the ideal number of variables, and screen out the final genus matrix. The cross-validation error formula is as follows:

[0056]

[0057] where: k is the number of folds, which is 10; MSEi is the mean squared error of the i-th fold;

[0058] 2-3) Establish an antibiotic resistance gene-bacterial genus correlation model through the CatBoost regression algorithm, select the optimal hyperparameters using grid search, and construct an antibiotic resistance gene-bacterial genus correlation model based on the RF-CatBoost algorithm;

[0059] The predicted value of the RF-CatBoost model is the weighted sum of the predictions of each decision tree:

[0060]

[0061] where, is the initial predicted value, usually the mean of the target values; f m (x) is the output of the m-th tree; is the learning rate; M is the number of trees;

[0062] 2-4) Verify the goodness of fit, robustness, and analytical ability of the model.

[0063] Step 3) Explainable correlation analysis: Use the Shap value to characterize the contributions of different bacterial genera in the tetracycline resistance gene-species genus correlation model ( Figure 3 ). The results show that the contribution rates of Stenotrophomonas, Haliscomenobacter, Alcaligenes, Sutterella, and Entomoplasma to the correlation model all exceed 10%. The above five bacterial genera have a greater impact on the composition of tetracycline resistance genes in sewage treatment plant sludge samples. At the same time, the Shap value results also show that Stenotrophomonas, Alcaligenes, and Sutterella have a positive contribution to the composition of tetracycline resistance genes. Therefore, Stenotrophomonas, Alcaligenes, and Sutterella may be potential hosts of tetracycline resistance genes.

[0064] It should be noted that the purpose of the SHAP value is to decompose the prediction result of the model into the contribution value of each feature (the bacterial genus abundance in this patent), that is:

[0065]

[0066] where, is the predicted value of the model; is the benchmark value of the model, usually the average initial value; is the Shap value of the i-th bacterial genus, indicating the contribution of the x i bacterial genus to the model;

[0067] For genus x i , its SHAP value is defined as the weighted average marginal contribution of all possible subsets of genera:

[0068]

[0069] where is a subset of genera, and ; is the model prediction value under the subset of genera ; is the model prediction value after adding genus x to the subset i ; n is the total number of genera, is the size of the subset ; is the coefficient that weights the contributions of each subset, ensuring balanced consideration of all possible subsets.

[0070] Considering the high dimensionality of the bacterial community, how to select the genera most relevant to the target antibiotic resistance gene from the genus set to identify the hosts of antibiotic resistance genes has become increasingly important for the ecological risk assessment of antibiotic resistance genes; for large and complex genomic datasets, Procrustes analysis and network analysis based on Spearman correlation coefficients have been commonly used to confirm the relationship between antibiotic resistance genes and bacterial communities in the past, but Procrustes analysis cannot confirm the relevant genera of antibiotic resistance genes, and network analysis can only analyze the correlation between a single genus and antibiotic resistance genes, resulting in a relatively high false correlation in the analysis of the correlation between antibiotic resistance genes and genera in complex environmental samples; while the Shap value interpretation method can effectively run on non-linear models of large datasets, and judge the impact of genera on antibiotic resistance genes according to the differences in the average Shap values of different genera. Specifically: by analyzing the changes in the Shap values of each genus, evaluate its positive and negative impacts on the model analysis results, and visualize the impact degree of different genera on antibiotic resistance genes. Combining the Shap value analysis results, explain the impact mechanism of genera, clarify the role of specific genera in the analysis of antibiotic resistance genes, so as to assist the research and decision-making in related fields.

[0071] The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A host analysis method for antibiotic resistance genes based on interpretable machine learning, characterized in that: The following steps are involved: Step 1) Genomic data collection: Collect a large amount of genomic data of the same environmental samples, calculate the types and relative abundances of antibiotic resistance genes, species and genera in the genetic data and process them into a matrix, use the Pauta criterion to analyze the relative abundance of selected antibiotic resistance genes, remove outliers and divide them into training sets and test sets; Step 2) Construction of analytical model: 2-1) First, use the random forest algorithm to establish an initial correlation analysis model for the species genus matrix; 2-2) Calculate the percentage increase in mean square error (%IncMSE) of each genus to measure the importance of the genus. The formula is as follows: ; Where: MSE original is the mean square error when the eigenvalues ​​are not perturbed; MSE random is the mean square error after shuffling the eigenvalues; Five 10-fold cross validations were performed to obtain the relationship between the number of the most important bacterial genera used for fitting and the average cross validation error. The number of bacterial genera with the smallest average cross validation error was selected as the ideal number of variables, and the final bacterial genus matrix was screened out. The cross validation error formula is as follows: ; Where: k is the number of folds, which is 10; MSEi is the mean square error of the i-th fold; 2-3) Establish an antibiotic resistance gene-bacterial genus correlation model through CatBoost regression algorithm, use grid search to select the optimal hyperparameters, and construct an antibiotic resistance gene-bacterial genus correlation model based on RF-CatBoost algorithm; The prediction value of the RF-CatBoost model is the weighted sum of the predictions of each decision tree: ; in, is the initial forecast value, usually the mean of the target value; f m (x) is the output of the 𝑚th tree; is the learning rate; M is the number of trees; 2-4) Verify the model’s fit, robustness, and analytical capabilities; Step 3) Interpretable correlation analysis: The contribution of different bacterial genera to antibiotic resistance genes in the model is quantified by Shap value, and the influence of bacterial genera on antibiotic resistance genes is judged according to the difference in average Shap values ​​of different bacterial genera; SHAP value is used to decompose the prediction results of the model into the contribution value of each feature, that is: ; in, is the predicted value of the model; is the baseline value of the model, usually the average initial value; is the Shap value of the i-th genus, marked by x i Contribution of bacterial genus to the model; For the genus x i , whose SHAP value is defined as the weighted average marginal contribution of all possible genus subsets: ; in, is a subset of the genus Bacteria, and ; In the genus The model prediction value under ; is in the subset Add bacteria genus x i The model prediction value after ; n is the total number of bacterial genera, Is a subset size; is a coefficient that weights the contribution of each subset.

2. The host analysis method of antibiotic resistance genes based on interpretable machine learning according to claim 1, characterized in that: The environmental samples in step 1) include one or more of the following: activated sludge, sewage and wastewater, river water, lake water, air, and soil.

3. The host analysis method of antibiotic resistance genes based on interpretable machine learning according to claim 1, characterized in that: The specific division of the training set and the test set in step 1) is as follows: the processed data are arranged in ascending order according to the relative abundance of antibiotic resistance genes, 5 data are divided into a group, the fifth data in each group is put into the test set, and the remaining data constitute the training set; the data in the training set is used for model establishment and internal verification, and the data in the test set is used for external verification and performance evaluation of the model.

4. The host analysis method of antibiotic resistance genes based on interpretable machine learning according to claim 1, characterized in that: In step 2-3), grid search is used to select the optimal hyperparameters. Specifically, grid search is performed on the number of basic decision trees, the maximum depth of each basic decision tree, and the learning rate, so that the parameter combination with the smallest mean absolute error is the optimal parameter combination.

5. The host analysis method of antibiotic resistance genes based on interpretable machine learning according to claim 1, characterized in that: In step 3), the contribution of different bacterial genera to antibiotic resistance genes in the quantified model using Shap values ​​is specifically as follows: the contribution of each bacterial genus to the antibiotic resistance gene analysis results is calculated through the output of the CatBoost model, the Shap value of each bacterial genus in all samples is extracted, the average Shap value and its distribution of each bacterial genus in different samples are calculated, and the bacterial genera are sorted according to the absolute average Shap value to identify the bacterial genus that contributes the most to antibiotic resistance genes.

6. The host analysis method of antibiotic resistance genes based on interpretable machine learning according to claim 1, characterized in that: In step 3), the influence of bacterial genera on antibiotic resistance genes is judged according to the difference in average Shap values ​​of different bacterial genera. Specifically, the positive and negative effects on the model analysis results are evaluated by analyzing the changes in the Shap values ​​of each bacterial genera, and the influence of different bacterial genera on antibiotic resistance genes is visualized. Combined with the Shap value analysis results, the influence mechanism of the genus is explained, and the role of specific bacterial genera in antibiotic resistance gene analysis is clarified to assist research and decision-making in related fields.

Citation Information

Patent Citations

  • Method for screening important characteristic genes related to drug resistance phenotype of bacteria based on machine learning

    CN114067912A

  • Antibiotic-loaded bacterium outer membrane vesicle as well as preparation method and application thereof

    CN115006367A

  • Graph embedded metagenome binning method and system based on view enhancement

    CN118609662A

  • Method for predicting relative abundance of antibiotic resistance genes in surface water and application

    CN119207586A

  • Methods and systems for determining antibiotic susceptibility

    US20170253917A1