A method for analyzing the drug resistance phenotype of pathogenic microorganisms

By screening drug-resistant gene characteristics using the Mann-Whitney U test and multi-factor analysis of variance method, and building a scoring model with multiple linear analysis and random forest algorithm, the problem of insufficient accuracy of microbial resistance detection in the prior art was solved, and a more comprehensive drug resistance risk assessment and prediction was achieved.

CN118098369BActive Publication Date: 2025-07-25HANGZHOU LUOXI MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410347361.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-07-25
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

The prior art has been present in microbial resistance detection for a long time, is complex in operation, and is unable to effectively consider the combined effects of multiple drug-resistant genes, resulting in insufficient accuracy in predicting drug-resistant phenotypes.

Method used

The Mann-Whitney U test method and multi-factor analysis of variance were used to screen drug resistance gene characteristics, and the drug resistance characteristic score model was constructed by combining multiple linear analysis. The random forest algorithm was used to determine the importance of drug resistance gene characteristics, and the final drug resistance risk score model was constructed, and the drug resistance phenotype of metagenomic sequencing data was comprehensively evaluated through multiple models.

Benefits of technology

It improves the accuracy and comprehensiveness of microbial resistance detection, can better guide clinical drug use, and comprehensively evaluate the relationship between interactions and drug resistance phenotypes between drug-resistant genes through multiple models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118098369B_ABST
    Figure CN118098369B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of pathogenic gene detection, and discloses a method for analyzing the drug resistance phenotype of pathogenic microorganisms, including the following modules: a drug resistance feature screening module, which uses the Mann-Whitney U test method and the multi-factor variance analysis method to screen potential drug resistance gene features affecting the drug resistance phenotype of microorganisms, and uses the multiple linear analysis method to construct a drug resistance risk score model 1; a drug resistance risk score model determination module, which uses the random forest algorithm to determine a drug resistance risk score model 2, and determines the final drug resistance risk score model according to the influence degree; a drug resistance risk assessment module, which detects drug resistance gene features from metagenomic sequencing data and analyzes the drug resistance risk of pathogenic microorganisms. The present invention uses the multiple linear analysis method and the random forest algorithm to jointly construct a drug resistance risk score model by using a variety of drug resistance gene features, and then uses this model to comprehensively evaluate the drug resistance phenotype of metagenomic sequencing data, and can more accurately predict the drug resistance phenotype of microorganisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pathogenic gene detection, and in particular to a method for analyzing pathogenic drug resistance based on metagenomic sequencing data. Background Art

[0002] Pathogenic microorganisms refer to a class of microorganisms that can cause infections in humans or animals, thereby leading to the occurrence of diseases, including bacteria, fungi, viruses, parasites, etc. At present, using antibiotics to treat pathogenic microorganisms is a commonly used clinical method, but the overuse of antibiotics has caused microorganisms to develop drug resistance.

[0003] Currently, microbial drug resistance detection usually mainly relies on drug susceptibility tests, but this method has problems such as long time and complex operation. Metagenomic sequencing technology can quickly and unbiasedly detect the sequence information of all microorganisms in a sample. Through this technology, drug resistance gene information can be detected, thereby assisting in the prediction of drug resistance phenotypes. Patent CN2021114005405 uses the methods of genome-wide association studies and machine learning to screen features related to bacterial drug resistance phenotypes and assigns weight analysis, but does not consider the influence of the combined action of multiple drug resistance genes on microbial drug resistance, and uses the same model for drug resistance prediction, which may reduce the accuracy of prediction.

[0004] Therefore, there is a need for a multi-drug resistance phenotype prediction model that considers the interaction of multiple genes, which can comprehensively analyze the prediction results of multiple models to improve the accuracy of drug resistance phenotype prediction. Summary of the Invention

[0005] Based on the above, the present invention provides a method for analyzing the drug resistance phenotype of pathogenic microorganisms. This method screens out drug resistance gene features related to the drug resistance of microorganisms to specific antibiotics from whole genome data, and constructs a drug resistance risk scoring model using multiple methods (multiple linear analysis method and random forest method) and multiple drug resistance gene features (drug resistance gene families and drug resistance gene subtypes), and then uses multiple models to comprehensively analyze the drug resistance phenotypes of the drug resistance genes detected in metagenomic sequencing data to evaluate the drug resistance phenotypes of microorganisms. This method comprehensively evaluates the drug resistance phenotypes of metagenomic sequencing data through multiple models, can comprehensively predict the drug resistance risk of pathogenic microorganisms, and better guide clinical drug use.

[0006] The specific technical solution of the present invention is as follows:

[0007] A method for analyzing the drug resistance phenotype of pathogenic microorganisms, including a drug resistance feature screening module, a drug resistance risk scoring model determination module, and a drug resistance risk assessment module;

[0008] Drug resistance feature screening module: Detect the drug resistance gene features of the whole genome data of microorganisms, analyze the detected drug resistance gene features using the Mann-Whitney U test method and the multi-factor analysis of variance method, obtain the potential drug resistance gene features that affect the drug resistance of microorganisms to antibiotics, and use the multiple linear analysis method to construct a drug resistance feature scoring model 1 under different drug resistance gene features;

[0009] Drug resistance risk scoring model determination module: Use the random forest algorithm to explore the influence degree of potential drug resistance gene features on drug resistance performance, determine the drug resistance feature scoring model 2 under different drug resistance gene features according to the influence degree, formulate the weight scores of the drug resistance feature scoring model 1 and the drug resistance feature scoring model 2 under the same drug resistance gene features, and construct the final drug resistance risk score model;

[0010] Drug resistance risk assessment module: Detect the drug resistance gene features of the metagenomic sequencing data and analyze the drug resistance risk of pathogenic microorganisms according to the final drug resistance risk score model, formulate a drug resistance risk threshold, and determine the drug resistance phenotype of pathogenic microorganisms.

[0011] Specifically, the drug resistance gene features include drug resistance gene families and drug resistance gene subtypes. Among them, drug resistance gene families such as: KPC, CTX-M, SHV and other gene families, and drug resistance gene subtypes such as: KPC-1, KPC-2, SHV-12, SHV-36, etc.;

[0012] Preferably, the specific steps of the drug resistance feature screening module are as follows:

[0013] Step S11: Collect the whole genome sequences of microorganisms and their corresponding drug resistance phenotypes of antibiotics from the microbial whole genome database;

[0014] Step S12: Use the drug resistance gene database to detect the drug resistance gene feature information from the whole genome data of microorganisms;

[0015] Step S13: Construct a drug resistance gene feature-phenotype relationship model according to the detected drug resistance gene feature information and its corresponding drug resistance phenotype respectively;

[0016] Step S14: Use the Mann-Whitney U test method and the multi-factor analysis of variance method to explore the influence degree of drug resistance gene features and the drug resistance gene features under interaction on drug resistance phenotypes in the drug resistance gene feature-phenotype relationship model respectively;

[0017] Step S15: Comprehensively evaluate the influence degrees obtained by the two methods to obtain a list of potential drug resistance gene features;

[0018] Step S16: Use the multiple linear analysis method to construct a drug resistance feature scoring model 1 under different drug resistance genes;

[0019] Preferably, the drug resistance gene feature-phenotype relationship model in step S13 includes the detection status of drug resistance gene features in each sample and their corresponding drug resistance phenotypes. When a drug resistance gene feature is detected, it is recorded as 1, and when a drug resistance gene feature is not detected, it is recorded as 0;

[0020] Preferably, the calculation formula of the drug resistance feature scoring model 1 described in step 16 is as follows:

[0021]

[0022] where a is the error value of the model, and b i represents the correlation of the i-th drug resistance gene feature in the list of m drug resistance gene features, and X i is the detection status of the i-th drug resistance gene feature;

[0023] Preferably, the specific steps of the drug resistance risk scoring model determination module are as follows:

[0024] Step S21: Encode the list of potential drug resistance gene features screened by the drug resistance feature screening module and their drug resistance phenotypes respectively;

[0025] Step S22: Randomly divide the sample set into five training subsets;

[0026] Step S23: Use the random forest algorithm to establish a drug resistance gene feature-phenotype drug resistance model;

[0027] Step S24: Extract the importance of each drug resistance gene feature for each drug resistance gene feature-phenotype drug resistance model;

[0028] Step S25: According to the drug resistance gene features and their importance in the five training subsets, perform weighted averaging on the importance of the drug resistance gene features to obtain the final drug resistance feature importance of each drug resistance gene feature;

[0029] Step S26: Determine the drug resistance feature scoring model 2 based on different drug resistance gene features constructed by the random forest algorithm;

[0030] Step S27: Determine the model weights under different drug resistance gene features and different model algorithms;

[0031] Step S28: Construct the final drug resistance risk score model.

[0032] Preferably, the coding rules in step S21 include a drug resistance gene characteristic coding rule and a drug resistance phenotype coding rule. The drug resistance gene characteristic coding rule is as follows: If a drug resistance gene characteristic is detected, it is coded as 1; if a drug resistance gene characteristic is not detected, it is coded as 0; for drug resistance gene characteristics under combined action, when the drug resistance gene characteristics under combined action are detected simultaneously, it is coded as 1, and in other cases, it is coded as 0. The drug resistance phenotype coding rule is as follows: If the drug resistance phenotype is drug resistance, it is coded as 1; if the drug resistance phenotype is sensitivity, it is coded as 0.

[0033] Preferably, the calculation formula of the drug resistance characteristic scoring model 2 in step S26 is as follows:

[0034]

[0035] where c i represents the drug resistance characteristic importance of the i-th drug resistance gene characteristic among m drug resistance gene characteristics, and X i is the detection situation of the i-th drug resistance gene characteristic among m drug resistance gene characteristics;

[0036] Preferably, the model weights in step S27 are as follows: The drug resistance characteristic scoring model 1 and the drug resistance characteristic scoring model 2 constructed using drug resistance gene families are respectively given weights of 0.15 - 0.20; the drug resistance characteristic scoring model 1 and the drug resistance characteristic scoring model 2 constructed using drug resistance gene subtypes are respectively given weights of 0.30 - 0.35, and the sum of the weights of the four models is 1;

[0037] Preferably, the calculation formula of the drug resistance risk score model is as follows:

[0038]

[0039] where m 1i and m 2i respectively represent the weights in the drug resistance characteristic scoring model 1 and the drug resistance characteristic scoring model 2 in the i-th drug resistance gene characteristic, and Y 1i and Y 2i respectively represent the scores of the drug resistance characteristic scoring model 1 and the drug resistance scoring model 2 in the i-th drug resistance gene characteristic;

[0040] Preferably, the specific steps of the drug resistance risk assessment module are as follows:

[0041] Step S31: Obtain the metagenomic sequencing data of the sample to be evaluated;

[0042] Step S32: Perform quality control on the metagenomic sequencing data of the sample to be tested, filter out human sequences, and obtain microbial sequencing data;

[0043] Step S33: Detect drug resistance gene characteristics from the microbial sequencing data;

[0044] Step S34: Calculate the comprehensive drug resistance risk score of the sample to be tested according to the drug resistance risk scoring model;

[0045] Step S35: Set the drug resistance risk scoring rules and drug resistance risk thresholds;

[0046] Step S36: Determine the drug resistance risk according to the drug resistance risk threshold, and output the predicted result of the drug resistance phenotype of the sample to be tested.

[0047] Preferably, the drug resistance risk scoring rules and drug resistance risk thresholds described in Step S35 are as follows: when the comprehensive drug resistance risk score < 0.3, it indicates that the sample is sensitive; when 0.3 ≤ comprehensive drug resistance risk score < 0.5, it indicates that the sample may be sensitive; when 0.5 ≤ comprehensive drug resistance risk score < 0.8, it indicates that the sample may be drug-resistant; when the comprehensive drug resistance risk score ≥ 0.8, it indicates that the sample is drug-resistant;

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] A method for analyzing the drug resistance phenotype of pathogenic microorganisms of the present invention first comprehensively evaluates the relationship between each drug resistance gene feature and the drug resistance phenotype using whole genome data, multiple features (drug resistance gene families and drug resistance gene subtypes), and multiple methods (Mann-Whitney U test method and multi-factor variance analysis method). This method can not only effectively evaluate the association between each gene and the drug resistance phenotype, but also explore the relationship between the interaction between drug resistance genes and the drug resistance phenotype; Secondly, the invention constructs multiple models (based on multiple linear analysis method and based on random forest model) using multiple features (drug resistance gene subtypes and drug resistance gene families), which can solve the inaccuracy of the model caused by a single model and analyze the drug resistance of pathogenic microorganisms more comprehensively and accurately. Description of the Drawings

[0050] Figure 1 It is a flowchart of the method for analyzing the drug resistance phenotype of pathogenic microorganisms in the embodiment of the present invention.

[0051] Figure 2 It is a flowchart of the drug resistance feature screening module in the embodiment of the present invention.

[0052] Figure 3 It is a flowchart of the drug resistance risk scoring model determination module in the embodiment of the present invention.

[0053] Figure 4 It is a flowchart of the drug resistance risk assessment module in the embodiment of the present invention. Detailed Embodiments

[0054] The present invention will be further described below in conjunction with embodiments.

[0055] Figure 1 It is the technical roadmap of the whole invention, including a drug resistance feature screening module, a drug resistance risk scoring model determination module, and a drug resistance risk assessment module. Among them, Examples 1 to 3 are specific examples of each module.

[0056] Example 1: Screen the drug resistance gene features related to Klebsiella pneumoniae resistance to meropenem and construct the drug resistance feature scoring model 1, as Figure 2 shown. The specific steps are as follows:

[0057] Step S11: Collect the whole genome information of Klebsiella pneumoniae and its drug resistance phenotype to meropenem from the microbial whole genome database;

[0058] Step S12: According to the whole genome sequence information of Klebsiella pneumoniae, use the drug resistance gene database to detect all the drug resistance gene features it carries, and statistically obtain the KP_MER_AMR_result file;

[0059] Step S13: Construct a drug resistance gene feature-phenotype relationship model based on the detected drug resistance gene feature information of Klebsiella pneumoniae and its drug resistance phenotype to meropenem;

[0060] Step S14: Use the Mann-Whitney U test method and the multivariate analysis of variance method to analyze the drug resistance gene feature-phenotype model respectively, and explore the influence degree of the drug resistance gene features and the drug resistance gene features under interaction on the drug resistance phenotype;

[0061] Step S15: Comprehensively evaluate the influence degrees obtained by the two methods to obtain a list of potential drug resistance gene features;

[0062] Step S16: Use the multiple linear analysis method to comprehensively construct the drug resistance feature scoring model 1 of Klebsiella pneumoniae to meropenem.

[0063] Specifically, the microbial whole genome database described in Step S11 includes the NCBI database and the PATRIC database. A total of 1615 strains of Klebsiella pneumoniae whole genome data are collected, among which 1134 strains are sensitive to meropenem and 481 strains are resistant to meropenem;

[0064] Specifically, the drug resistance gene database described in Step S12 includes the CARD and ResFinder databases;

[0065] Specifically, in step S12, the blast software is used to detect drug-resistant genes in the whole genome sequence. When the parameters Identity and Coverage are greater than 60% and 80% respectively, they are screened as drug-resistant gene features, and the names, Identity, and Coverage information of the drug-resistant genes are counted into KP_MER_AMR_result;

[0066] Specifically, the drug-resistant gene features include drug-resistant gene families and drug-resistant gene subtypes. Among them, drug-resistant gene families such as: KPC, CTX-M, SHV and other gene families, and drug-resistant gene subtypes such as: KPC-1, KPC-2, SHV-12, SHV-36, etc.;

[0067] Specifically, the Mann-Whitney U test method, multi-factor variance analysis method, and multiple linear analysis method used by the drug-resistant feature screening module are analyzed using SPSS software, and the correlation between the drug-resistant gene feature and the drug-resistant phenotype is determined according to the parameter p-value. When p-value ≤ 0.05, it indicates that the drug-resistant gene feature is significantly correlated with the drug-resistant phenotype. When p-value > 0.05, it indicates that the drug-resistant gene feature has no significant correlation with the drug-resistant phenotype:

[0068] Specifically, the potential drug-resistant gene families and drug-resistant gene subtypes that affect the drug-resistant phenotype of Klebsiella pneumoniae against meropenem obtained by the Mann-Whitney U test and multi-factor variance analysis methods in step S14 are shown in Table 1 and Table 2 respectively;

[0069] Table 1: Results of the list of potential gene families affecting the drug-resistant phenotype of Klebsiella pneumoniae against meropenem

[0070] KPC OXA CTX-M dfrA tet qnr fosA KPC, aph(3') dfrA, aph(3') aph(3'), fosA aph(3'), qnr tet, aph(3') KPC, CTX-M dfrA, CTX-M qnr, CTX-M sul, CTX-M KPC, OXA KPC, fosA KPC, tet dfrA, OXA tet, OXA dfrA, fosA dfrA, qnr dfrA, sul sul, fosA sul, qnr tet, qnr

[0071] Table 2: List of potential gene subtypes affecting the drug-resistant phenotype of Klebsiella pneumoniae against meropenem

[0072]

[0073]

[0074] Furthermore, step S16 uses the multiple linear analysis method to analyze the drug-resistant gene features significantly correlated with the drug-resistant phenotype and their influence on the correlation. The results are shown in Table 3;

[0075] Table 3: Influence correlation of genes (families / subtypes) finally affecting the drug-resistant phenotype of Klebsiella pneumoniae against meropenem.

[0076]

[0077]

[0078] Specifically, the drug resistance characteristic scoring model 1 of Klebsiella pneumoniae to meropenem includes the drug resistance characteristic scoring model 1 of drug resistance gene subtypes and the calculation formula of the drug resistance characteristic scoring model 1 of drug resistance gene families as follows:

[0079] Drug resistance characteristic scoring model 1 of drug resistance gene subtypes = 0.038 + 0.829 * x1 + 0.09 * x2 - 0.073 * x3 - 0.283 * x4 + 0.401 * x5 - 0.694 * x6 - 0.499 * x7 - 0.366 * x8 + 0.512 * x9 + 0.476 * x10;

[0080] The drug resistance gene subtypes of x1 to x10 are: KPC-2, OXA-9, TEM-A, KPC-2*qnrB1, OXA-9*dfrA14*TEM1B, TEM-1A*dfrA14*sul1, dfrA14*sul1*TEM-1B, OXA-9*TEM-1B, TEM-1A*qnrB1, dfrA14*sul1;

[0081] Drug resistance characteristic scoring model 1 of drug resistance gene families = 0.039 + 0.975 * x1 + 0.09 * x2 + 0.107 * x3 - 0.175 * x4 + 0.29 * x5 - 0.095 * x6 - 0.065 * x7 + 0.672 * x8 - 0.087 * x9 - 0.168 * x10 - 0.615 * x11 - 0.168 * x12 + 0.079 * x13 - 0.099 * x14 - 0.057 * x15 - 0.531 * x16 + 0.124 * x17 + 0.248 * x18 - 0.135 * x19;

[0082] The drug resistance gene families of x1 to x19 are: KPC, OXA, tet, qnr, aph(3')*qnr, tet*aph(3'), KPC*CTX-M, qnr*CTX-M, sul*CTX-M, KPC*OXA, KPC*fosA, KPC*tet, dfrA*OXA, tet*OXA, dfrA*fosA, dfrA*qnr, sul*qnr, tet*qnr

[0083] Example 2: Determination of the drug resistance characteristic value and the drug resistance characteristic scoring model 2 of Klebsiella pneumoniae to meropenem based on the random forest model, as Figure 3 shown, the specific steps are as follows:

[0084] Step S21: Encode the potential drug resistance gene characteristics and drug resistance phenotypes of Klebsiella pneumoniae;

[0085] Step S22: Randomly divide the samples into five training subsets with equal numbers;

[0086] Step S23: Use the random forest algorithm to establish a drug-resistant gene feature - phenotype model;

[0087] Step S24: Extract the importance values of the drug-resistant gene features in each drug-resistant gene feature - phenotype model;

[0088] Step S25: According to the drug-resistant gene features and their corresponding importance values of the five training subset models, perform weighted averaging on the importance values of the drug-resistant gene features respectively to obtain the final drug-resistant feature importance of each drug-resistant gene.

[0089] Step S26: Determine the drug-resistant feature scoring model 2 based on different drug-resistant gene features constructed by the random forest algorithm;

[0090] Step S27: Determine the drug-resistant weights under different drug-resistant gene features and different model algorithms;

[0091] Step S28: Construct the final drug-resistant risk score model.

[0092] Specifically, the coding rules described in Step S21 include a drug-resistant gene feature coding rule and a drug-resistant phenotype coding rule. The drug-resistant gene feature coding rule is: if the drug-resistant gene feature is detected, it is coded as 1; if the drug-resistant gene feature is not detected, it is coded as 0; for drug-resistant gene features under co-action, when the drug-resistant gene features under co-action are detected simultaneously, it is coded as 1, and in other cases, it is coded as 0; the drug-resistant phenotype coding rule is: if the drug-resistant phenotype is drug-resistant, it is coded as 1; if the drug-resistant phenotype is sensitive, it is coded as 0.

[0093] Specifically, the rules for randomly allocating training subsets described in Step S22 include: (1) the number of sensitive and drug-resistant samples in the drug-resistant phenotypes of the training subsets is the same, (2) the number of samples in each training subset is not less than 120% of the number of samples of the phenotype (drug-resistant or sensitive) with the smaller number. In this embodiment, each training subset has 600 samples, with 300 drug-resistant samples and 300 sensitive samples each.

[0094] Specifically, the feature importance of each drug-resistant gene subtype and drug-resistant gene family and their final drug-resistant feature importance in the established drug-resistant gene feature - phenotype model obtained in Steps S23 - S25 are shown in Tables 4 and 5;

[0095] Table 4: Feature importance values of each drug-resistant gene subtype obtained based on the random forest:

[0096]

[0097]

[0098] Table 5: Feature importance values of drug-resistant gene families obtained based on random forest:

[0099]

[0100] Furthermore, the drug-resistant feature scoring model 2 described in step S26 includes the calculation formulas of the drug-resistant feature scoring model 2 for drug-resistant gene subtypes and the drug-resistant feature scoring model 2 for drug-resistant gene families as follows:

[0101] Drug-resistant feature scoring model 2 for drug-resistant gene subtypes = 0.615*x1 + 0.278*x2 + 0.032*x3 + 0.003*x4 + 0.016*x5 + 0.002*x6 + 0.005*x7 + 0.013*x8 + 0.003*x9 + 0.028*x10;

[0102] The drug-resistant gene subtypes of x1 to x10 are respectively: KPC-2, OXA-9, TEM-A, KPC-2*qnrB1, OXA-9*dfrA14*TEM1B, TEM-1A*dfrA14*sul1, dfrA14*sul1*TEM-1B, OXA-9*TEM-1B, TEM-1A*qnrB1, dfrA14*sul1

[0103] Drug-resistant feature scoring model 2 for drug-resistant gene subtypes = 0.465*x1 + 0.078*x2 + 0.027*x3 + 0.010*x4 + 004*x5 + 0.009*x6 + 0.090*x7 + 0.009*x8 + 0.030*x9 + 0.183*x10 + 0.003*x11 + 0.008*x12 + 0.014*x13 + 0.020*x14 + 0.008*x15 + 0.007*x16 + 0.018*x17 + 0.007*x18 + 0.009*x19

[0104] The drug-resistant gene families of x1 to x19 are respectively: KPC, OXA, tet, qnr, aph(3')*qnr, tet*aph(3'), KPC*CTX-M, qnr*CTX-M, sul*CTX-M, KPC*OXA, KPC*fosA, KPC*tet, dfrA*OXA, tet*OXA, dfrA*fosA, dfrA*qnr, sul*qnr, tet*qnr

[0105] The weights of the drug resistance gene features described in step S27 are as follows: the weights of drug resistance feature scoring model 1 and drug resistance feature scoring model 2 constructed using drug resistance gene families are 0.175 respectively, and the weights of drug resistance feature scoring model 1 and drug resistance feature scoring model 2 constructed using drug resistance gene subtypes are 0.325 respectively;

[0106] Furthermore, the calculation formula of the drug resistance risk score model in step S28 is as follows:

[0107] Drug resistance risk score model = 0.325 * (drug resistance feature scoring model 1 of drug resistance gene subtype + drug resistance feature scoring model 2 of drug resistance gene subtype) + 0.175 * (drug resistance feature scoring model 1 of drug resistance gene family + drug resistance feature scoring model 2 of drug resistance gene family)

[0108] Example 3: Evaluation of the drug resistance of Klebsiella pneumoniae to meropenem in metagenomic sequencing data, as Figure 4 shown, the specific steps are as follows:

[0109] Step S31: Obtain the metagenomic sequencing data of Klebsiella pneumoniae samples S1 - S10;

[0110] Step S32: Perform quality control on the metagenomic sequencing data and filter out human sequences to obtain the microbial sequencing data Mirco_reads of each sample;

[0111] Step S33: Detect drug resistance gene features in Micro_reads to obtain the result AMR_result;

[0112] Step S34: Calculate the comprehensive drug resistance risk scores of the test samples S1 - S10 according to the drug resistance risk scoring model;

[0113] Step S35: Set the drug resistance risk scoring rules and drug resistance risk thresholds;

[0114] Step S36, according to the drug resistance risk threshold, determine the drug resistance risk and output the drug resistance phenotype prediction results of the test samples S1 - S10.

[0115] Specifically, for the quality control in step S32, FastQC, Trimmomatic, and bowtie2 software are used for quality control and human sequence removal respectively;

[0116] Specifically, for step S33, ResFinder and CARD databases are used for drug resistance gene detection to obtain AMR_result;

[0117] Furthermore, evaluating the drug resistance of samples S1 - S10 according to the drug resistance risk score model obtained from the above Example 1 and Example 2, the drug resistance comprehensive scoring results of step S34 are shown in Table 6;

[0118] Table 6: Drug resistance risk assessment and its results of samples S1 - S10

[0119]

[0120]

[0121] Furthermore, the drug resistance risk scoring rule and drug resistance risk threshold described in step S36 are as follows: when the comprehensive drug resistance score is < 0.3, it indicates that the sample is sensitive; when 0.3 ≤ comprehensive drug resistance score < 0.5, it indicates that the sample may be sensitive; when 0.5 ≤ comprehensive drug resistance score < 0.8, it indicates that the sample may be drug - resistant; when the comprehensive drug resistance score ≥ 0.8, it indicates that the sample is drug - resistant;

[0122] Furthermore, according to the drug resistance risk scoring rule and drug resistance risk threshold obtained in step S36, the prediction results are output. Among them, the prediction results and drug susceptibility results of samples S1 - S10 are shown in Table 7.

[0123] Table 7: Prediction results and drug susceptibility results of samples S1 - S10:

[0124] Sample Number Comprehensive Drug Resistance Score Predicted Phenotype Drug Sensitivity Results Conclusion Sample S1 0.03265 Sensitive Sensitive Consistent Sample S2 0.0092 Sensitive Sensitive Consistent Sample S3 0.8755 Resistant Sensitive Inconsistent Sample S4 0.740 Possible Resistance Resistant Consistent Sample S5 0.547 Possible Resistance Resistant Consistent Sample S6 0.878 Resistant Resistant Consistent Sample S7 0.560 Possible Resistance Resistant Consistent Sample S8 0.292 Sensitive Sensitive Consistent Sample S9 0.878 Resistant Resistant Consistent Sample S10 0.740 Possible Resistance Resistant Consistent

[0125] According to the results shown in Table 6 and Table 7, the drug resistance scores of sample S5 in the two detection methods based on drug - resistant gene subtypes are 0.456 and 0.326 respectively, showing that it may be sensitive. However, after incorporating the drug resistance score model based on drug - resistant gene families, the comprehensive drug resistance score is 0.547, and the final prediction result is that it may be drug - resistant, indicating that the present invention can effectively correct the inaccuracy of a single model through phenotypic prediction analysis of multiple models.

Claims

1. A method for analyzing the drug resistance phenotype of pathogenic microorganisms, characterized in that, The specific steps are as follows: S1 Drug Resistance Feature Screening Module: Detect the drug resistance gene features of the whole genome data of microorganisms, analyze the detected drug resistance gene features using the Mann-Whitney U test method and the multi-factor variance analysis method, obtain the potential drug resistance gene features that affect the drug resistance of microorganisms to antibiotics, and construct a drug resistance feature scoring model under different drug resistance gene features using the multiple linear analysis method a is the error value of the model, bi represents the correlation of the i-th drug resistance gene feature in the list of m drug resistance gene features, and Xi is the detection situation of the i-th drug resistance gene feature; S2 Drug resistance risk scoring model determination module: Encode the candidate drug resistance gene characteristics and phenotypes. The encoding rules for drug resistance gene characteristics are as follows: If a drug resistance gene characteristic is detected, it is encoded as 1; if a drug resistance gene characteristic is not detected, it is encoded as 0; for drug resistance gene characteristics under combined action, when all drug resistance gene characteristics under combined action are detected simultaneously, it is encoded as 1, and in other cases, it is encoded as 0. The encoding rules for drug resistance phenotypes are as follows: If the drug resistance phenotype is drug resistance, it is encoded as 1; if the drug resistance phenotype is sensitivity, it is encoded as 0; Randomly divide the sample set into 5 training subsets, use the random forest algorithm to explore the influence degree of potential drug resistance gene characteristics on drug resistance phenotypes, and determine the drug resistance characteristic scoring model under different drug resistance gene characteristics according to the influence degree c i represents the importance of the drug resistance characteristics of the i-th drug resistance gene characteristic among m drug resistance gene characteristics, X i is the detection situation of the i-th drug resistance gene characteristic among m drug resistance gene characteristics. Set the weight scores of the drug resistance characteristic scoring model 1 and the drug resistance characteristic scoring model 2 under the same drug resistance gene characteristics, and construct the final drug resistance risk score model m 1i and m 2i respectively represent the weights in the drug resistance characteristic scoring model 1 and the drug resistance characteristic scoring model 2 in the i-th drug resistance gene characteristic, Y 1i and Y 2i respectively represent the scores of the drug resistance characteristic scoring model 1 and the drug resistance scoring model 2 in the i-th drug resistance gene characteristic; The assignment conditions for the weight scores are as follows: The drug resistance characteristic scoring model 1 and drug resistance characteristic scoring model 2 constructed using drug resistance gene families are each assigned a weight of 0.15 - 0.20; the drug resistance characteristic scoring model 1 and drug resistance characteristic scoring model 2 constructed using drug resistance gene subtypes are each assigned a weight of 0.30 - 0.35, and the sum of the weights of the four models is 1; S3 Drug resistance risk assessment module: Detect drug resistance gene characteristics from metagenomic sequencing data and analyze the drug resistance risk of pathogenic microorganisms according to the final drug resistance risk score model, formulate drug resistance risk scoring rules and drug resistance risk thresholds, and determine the drug resistance phenotype of pathogenic microorganisms. The drug resistance risk scoring rules and drug resistance risk thresholds are as follows: When the comprehensive drug resistance risk score < 0.3, it indicates that the sample is sensitive; when 0.3 ≤ comprehensive drug resistance risk score < 0.5, it indicates that the sample may be sensitive; when 0.5 ≤ comprehensive drug resistance risk score < 0.8, it indicates that the sample may be drug resistant; when the comprehensive drug resistance risk score ≥ 0.8, it indicates that the sample is drug resistant; The drug resistance gene characteristics include drug resistance gene families and drug resistance gene subtypes.

2. The method for analyzing the drug-resistant phenotype of pathogenic microorganisms according to claim 1, wherein The specific steps of the drug resistance characteristic screening module are as follows: Step S11: Collect the whole genome sequences of microorganisms and their corresponding drug resistance phenotypes of antibiotics from the microbial whole genome database; Step S12: Detect drug resistance gene characteristic information from the whole genome data of microorganisms using the drug resistance gene database; Step S13: Construct a drug resistance gene characteristic - phenotype relationship model according to the detected drug resistance gene characteristic information and its corresponding drug resistance phenotype respectively; Step S14: Use the Mann - Whitney U - test method and multi - factor analysis of variance method respectively on the drug resistance gene characteristic - phenotype relationship model to explore the influence degree of drug resistance gene characteristics and drug resistance gene characteristics under interaction on the drug resistance phenotype; Step S15: Comprehensively evaluate the influence degrees obtained by the two methods to obtain a list of potential drug resistance gene characteristics; Step S16: Use multiple linear analysis method to construct drug resistance characteristic scoring model 1 under different drug resistance gene characteristics.

3. The method for analyzing the drug resistance phenotype of pathogenic microorganisms according to claim 1, characterized in that The specific steps of the drug resistance risk scoring model determination module are as follows: Step S21: Encode the list of potential drug resistance gene characteristics screened by the drug resistance characteristic screening module and their drug resistance phenotypes respectively; Step S22: Randomly divide the sample set into five training subsets; Step S23: Use the random forest algorithm to establish a drug resistance gene characteristic - phenotype drug resistance model; Step S24: Extract the importance of each drug resistance gene characteristic for each drug resistance gene characteristic - phenotype drug resistance model; Step S25: According to the drug resistance gene features and their importance of the five training subsets, perform a weighted average on the importance of the drug resistance gene features to obtain the final drug resistance feature importance of each drug resistance gene feature; Step S26: Determine the drug resistance feature scoring model 2 under different drug resistance gene features constructed based on the random forest algorithm; Step S27: Determine the model weights under different drug resistance gene features and different model algorithms; Step S28: Construct the final drug resistance risk score model.

4. The method for analyzing the drug resistance phenotype of pathogenic microorganisms according to claim 1, characterized in that, The specific steps of the drug resistance risk assessment module are as follows: Step S31: Obtain the metagenomic sequencing data of the sample to be tested; Step S32: Perform quality control on the metagenomic sequencing data of the sample to be tested and filter out human-derived sequences to obtain microbial sequencing data; Step S33: Detect drug resistance gene features from the microbial sequencing data; Step S34: Calculate the comprehensive drug resistance risk score of the sample to be tested according to the drug resistance risk scoring model; Step S35: Set the drug resistance risk scoring rules and drug resistance risk thresholds; Step S36: Determine the drug resistance risk according to the drug resistance risk threshold and output the drug resistance phenotype prediction result of the sample to be tested.

Citation Information

Patent Citations

  • Method and application of pan-tumor targeted drug sensitivity state evaluation model constructed based on high-throughput sequencing data and clinical phenotypes

    CN111640508A

  • Method for screening important characteristic genes related to drug resistance phenotype of bacteria based on machine learning

    CN114067912A