Detection system for rapidly identifying drug resistance phenotypes of pathogenic bacteria based on machine learning in combination with MALDI-TOF MS (matrix-assisted laser desorption ionization-time of flight mass spectrometry)
By selecting optimal detection conditions and analyzing data from multiple time points, combined with machine learning models, the technical problems of rapid drug susceptibility testing technology in the prior art have been solved. It has been achieved that bacterial drug resistance can be determined within 2 hours, improving the accuracy and adaptability of the test, supporting rapid and reliable drug susceptibility results in clinical practice, and supporting early and accurate anti-infection treatment decisions.
Patent Information
- Application Number
- CN202511344652.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-30
AI Technical Summary
Existing rapid antimicrobial susceptibility testing technologies have shortcomings in terms of speed, accuracy (especially for intermediate resistance), ease of operation, result stability, information dimensions provided, and standardization. They fail to effectively utilize the dynamic growth information of bacteria under antibiotic treatment provided by MALDI-TOF MS, particularly for intermediate resistance detection technologies. Furthermore, they fail to effectively utilize the dynamic growth information of bacteria under antibiotic treatment provided by MALDI-TOF MS, combined with powerful machine learning models, for automated, high-precision, and interpretable resistance prediction.
By screening for optimal detection conditions, pathogenic bacteria are grouped for culture and sampled at multiple time points. After mass spectrometry data acquisition and preprocessing, an enhanced dynamic relative growth characteristic vector matrix is calculated. Data analysis is performed using a machine learning model, and the optimal detection model is screened through cross-validation to predict the drug resistance phenotype of bacteria.
It enables the determination of bacterial resistance within 2 hours, significantly improving the accuracy of resistance detection, especially in distinguishing intermediate resistance phenotypes. It has good scalability and adaptability, and can be applied to other types of bacteria and different combinations of antibiotics, supporting rapid and reliable drug sensitivity results in clinical practice, and greatly supporting early and precise anti-infective treatment decisions.
Smart Images

Figure CN121237227A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of clinical microbiological detection, and particularly relates to a system for rapid identification of pathogenic bacteria drug resistance phenotype based on machine learning combined with MALDI-TOF MS. BACKGROUND
[0002] The problem of antibiotic resistance is increasingly serious worldwide, and it is urgent to develop a rapid and reliable drug resistance detection method. Traditional drug sensitivity test methods such as paper disc diffusion method and micro-broth dilution method, although are the current gold standard detection method, have inherent defects such as long detection period (usually 12-24 hours) and complicated operation steps, which seriously limit their application in scenarios requiring rapid clinical decision-making. As an alternative, molecular detection technologies such as reverse transcription polymerase chain reaction (RT-PCR) can accelerate the discovery of drug resistance, but they have higher operating costs and rely on pre-set known drug resistance gene targets, making it difficult to comprehensively evaluate the complex phenotype of bacterial drug resistance spectrum.
[0003] Matrix-Assisted Laser Desorption Ionization Time of Flight Mass Spectrometry (MALDI-TOF MS) has been widely used in microbial identification due to its rapid, economical and high-throughput characteristics. In recent years, this technology has also been tried for drug sensitivity detection. For example, Lang et al. proposed a MALDI-TOF MS antibiotic sensitivity test rapid determination method (MBT-ASTRA), which compares the mass spectrum differences of bacteria after short-term incubation in drug-containing and drug-free media to evaluate their drug resistance. Although this method significantly shortens the detection time to a few hours, it still has problems such as complex operation, insufficient result stability, inability to provide accurate MIC values, and limited detection capability for complex drug resistance mechanisms. Wilhelm et al. improved MBT-ASTRA by manually selecting 6 characteristic peaks for semi-quantitative analysis, which simplified the operation to some extent, but the manual peak selection was highly subjective and had poor universality, making it difficult to achieve standardization and limiting its clinical application.
[0004] The development of artificial intelligence technology provides a new way for drug resistance phenotype detection. Machine learning can automatically extract features and mine potential patterns from high-dimensional mass spectrum data, which is expected to achieve more accurate drug resistance prediction. However, existing researches are mostly based on static mass spectrum data, which fails to systematically capture the dynamic response process of bacteria under antibiotic pressure; mass spectrum analysis is mostly limited to overall spectrum, lacking in-depth analysis of the correlation between specific m / z intervals and drug resistance mechanisms; the model has poor interpretability, making it difficult to trace key decision features; in addition, existing methods generally ignore the discrimination ability of "intermediate" drug resistance phenotype, and the application scope is mostly limited to one or a few pathogen-antibiotic combinations, lacking in clinical universality.
[0005] For example, the paper CN119438596A, entitled "A rapid detection method for bacterial drug resistance combining bacterial metabolic fingerprinting after short-term antibiotic stimulation and machine learning," has many shortcomings:
[0006] (1) Lack of refined analysis in dynamic time dimension: Only a single metabolic fingerprint profile is obtained through short incubation (10 minutes to 2 hours), without dynamic sampling at different time points of antibiotic action. This single-point detection cannot capture the changes in bacterial metabolic pathways over time, making it difficult to reflect the differences in growth kinetics between drug-resistant and sensitive bacteria, and is very likely to miss key drug resistance phenotypic characteristics.
[0007] (2) Mass-to-charge ratio (m / z) interval analysis is too broad: only the overall metabolic fingerprint is used as the analysis object. However, since metabolites of different molecular weights (such as small molecule metabolites, peptides, and proteins) may have different associations with drug resistance mechanisms, this overall analysis method will lead to insufficient feature extraction and thus cannot focus on the biomarker region directly related to drug resistance.
[0008] (3) The ability to detect intermediate drug resistance phenotypes has not been validated: the method only classifies "drug-resistant" and "susceptible" strains and does not cover the "intermediate drug resistance" phenotypes commonly seen in clinical practice. Since the metabolic characteristics of intermediate strains may overlap with those of susceptible / drug-resistant strains, this method may face the risk of missed detection or misjudgment in practical applications, and lacks comprehensive coverage of complex drug resistance states in clinical practice. In addition, only achieving qualitative classification of "drug resistance / susceptibility" cannot provide quantitative data on the degree of drug resistance like the minimum inhibitory concentration (MIC) test. The lack of MIC values will limit clinicians' precise adjustment of drug dosage, and has low guiding value when dealing with severe infections or borderline drug-resistant strains.
[0009] (4) Deficiencies in data processing and model building: Data preprocessing only includes basic steps such as background subtraction and baseline smoothing, without establishing systematic quality control standards. This easily leads to the inclusion of low-quality mass spectrometry data, thereby introducing noise and affecting the reliability of subsequent classification models. In addition, the trained model was not subjected to interpretability analysis (such as evaluation of the contribution of key features), making it impossible to clarify which metabolite features are the basis for drug resistance prediction. This makes it difficult to verify the scientific validity of the model and is not conducive to the subsequent biological explanation of drug resistance mechanisms.
[0010] (5) Narrow scope of application and lack of universality verification: The only implementation is only for the Escherichia coli-ceftriaxone sodium system and has not been extended to other common clinical bacteria (such as Staphylococcus aureus and Klebsiella pneumoniae) or antibiotic types (such as quinolones and aminoglycosides). The adaptability of its technical solution to different bacterial metabolic characteristics and drug resistance mechanisms has not been verified, making it difficult to promote to complex clinical scenarios.
[0011] In summary, existing rapid drug susceptibility testing technologies generally have shortcomings in terms of speed, accuracy (especially for intermediate resistance), ease of operation, result stability, information dimensionality, and standardization. In particular, they fail to effectively utilize the dynamic bacterial growth information provided by MALDI-TOF MS under antibiotic treatment, combined with powerful machine learning models, for automated, high-precision, and interpretable drug resistance prediction. Therefore, developing novel rapid drug resistance detection methods and systems that integrate dynamic mass spectrometry data analysis and machine learning has significant clinical implications and application value. Summary of the Invention
[0012] The purpose of this invention is to address the shortcomings of existing technologies by providing a method and system for rapid identification of drug-resistant phenotypes of pathogenic bacteria based on machine learning. This method involves screening for optimal detection conditions, grouping and culturing clinical isolates of pathogenic bacteria under these conditions, and sampling at multiple time points. After mass spectrometry data acquisition and preprocessing, the enhanced dynamic relative growth (RBD-RG) feature vector matrix of the clinical isolates of pathogenic bacteria is calculated. Multiple machine learning models are trained using this RBD-RG feature vector matrix, and the optimal model is selected through cross-validation for drug resistance phenotype prediction. This enables rapid and accurate prediction of bacterial drug resistance phenotypes (such as sensitive, intermediate, and resistant).
[0013] The objective of this invention is achieved through the following approach:
[0014] A method for rapidly identifying drug-resistant phenotypes of pathogenic bacteria based on machine learning combined with MALDI-TOF MS includes the following steps:
[0015] 1) First, using standard strains of pathogenic bacteria, combined with the MALDI rapid assay for biotype antibiotic susceptibility, the optimal antibiotic concentration and optimal culture time for distinguishing the drug resistance phenotypes of pathogenic bacteria were determined.
[0016] 2) Collect and screen clinical isolates of pathogenic bacteria, and determine the drug susceptibility labels of the clinical isolates of pathogenic bacteria using standard drug susceptibility testing methods;
[0017] 3) During the process of culturing pathogenic bacterial clinical isolates at the optimal antibiotic concentration and optimal culture time, the dynamic relative growth values of the mass spectrometry peaks of the pathogenic bacterial clinical isolates at different time points are dynamically acquired to form multiple enhanced dynamic relative growth feature vector matrices corresponding to different time points. These matrix matrices are then combined with the drug sensitivity labels of the pathogenic bacterial clinical isolates to form a dataset of pathogenic bacterial clinical isolates.
[0018] 4) Train multiple machine learning models using clinical isolates of pathogenic bacteria, and use performance evaluation metrics to select the optimal machine learning model as a drug resistance phenotype classification model.
[0019] 5) When identifying the drug resistance phenotype of clinical pathogenic bacteria samples, the enhanced dynamic relative growth feature vector matrix of the clinical pathogenic bacteria samples is extracted, and the enhanced dynamic relative growth feature vector matrix is input into the drug resistance phenotype classification model to obtain the drug resistance phenotype of the clinical pathogenic bacteria samples.
[0020] Preferably, in step 1), the determination of the optimal antibiotic concentration and optimal culture time specifically includes:
[0021] 1-1) Set several target antibiotic concentrations C (1,2,...,N) and several target culture times T (1,2,...,M) in a gradient.
[0022] 1-2) The target antibiotic concentration C1 is correlated with the target culture times T1, T2, ..., T... M Combine them separately, and pair the target antibiotic concentration C2 with T1, T2, ..., T... M Combine them separately, ..., to achieve the target antibiotic concentration C N With T1, T2, ..., T M Several target culture conditions were obtained by combining them separately, and each target culture condition is defined by (target antibiotic concentration C). N Target training time T M Record in the form of ), where N and M are both natural numbers;
[0023] 1-3) Using several target culture conditions, the pathogenic bacterial standard strains were cultured for a short time, and after reaching the target culture time, the cultures were processed to obtain mass spectrometry samples that could be analyzed by MALDI-TOF MS.
[0024] 1-4) Use MALDI-TOF MS to detect the mass spectra of the standard strains of pathogenic bacteria under each target culture condition, and select the optimal antibiotic concentration and optimal culture time based on the mass spectra.
[0025] This invention screens and optimizes antibiotic concentrations and culture times based on the differentiation of different phenotypic states of standard pathogenic bacterial strains, in order to obtain the dynamic changes in the growth of pathogenic bacteria in the shortest time. Combined with machine learning methods, it avoids the subjectivity of traditional manual selection of characteristic peaks or setting conditions based on experience, making the experimental results comparable when different laboratories use the same parameters, and solving the problem of result differences caused by inconsistent operating standards in traditional methods.
[0026] Preferably, in step 2), the collection and screening of clinical isolates of pathogenic bacteria, and the determination of the drug susceptibility label of the clinical isolates of pathogenic bacteria using standard drug susceptibility testing methods, specifically includes:
[0027] 2-1) Several strains were isolated from clinical specimens and purified and cultured at 37°C and 5% CO2.
[0028] 2-2) The confidence scores of matching items for several strains were calculated multiple times using MALDI-TOF MS. The average confidence scores of matching items for several strains were then calculated. Strains with an average confidence score of matching items greater than or equal to 2.0 were taken as clinical isolates of pathogenic bacteria.
[0029] 2-3) Use standard drug susceptibility testing methods to distinguish the drug resistance phenotypes of clinical isolates of pathogenic bacteria, and use them as drug susceptibility labels for clinical isolates of pathogenic bacteria.
[0030] Preferably, in step 3), the dynamic acquisition of the dynamic relative growth values of mass spectrometry peaks of clinical isolates of pathogenic bacteria at different time points constitutes an enhanced dynamic relative growth feature vector matrix, which is then combined with the drug sensitivity labels of the clinical isolates of pathogenic bacteria to form a dataset of clinical isolates of pathogenic bacteria, specifically including:
[0031] 3-1) Prepare standard bacterial suspensions of clinical isolates of pathogenic bacteria, and divide them into experimental group (containing the optimal concentration of antibiotics) and control group (without antibiotics);
[0032] 3-2) Samples were taken from the experimental group and the control group at multiple sampling time points within the optimal culture time, and MALDI-TOF mass spectrometry samples of the experimental group and the control group at each sampling time point were prepared.
[0033] 3-3) Based on the MALDI-TOF mass spectrometry samples of the experimental group and the control group, the mass spectrometry peak diagrams of the experimental group and the control group at each sampling time point were obtained;
[0034] 3-4) Data cleaning was used to preprocess the mass spectrometry peak images of the experimental group and the control group at each sampling time point to improve the quality of the mass spectrometry peak images;
[0035] 3-5) Divide the bacterial proteome regions in the mass spectrometry peaks of the experimental group and the control group at each sampling time point into continuous and fixed-width mass-to-charge ratio sub-intervals, and calculate the dynamic relative growth values of the experimental group and the control group at each sampling time point according to the mass-to-charge ratio sub-intervals, thus forming several enhanced dynamic relative growth feature vector matrices.
[0036] 3-6) Combine several enhanced dynamic relative growth feature vector matrices with the drug sensitivity labels of pathogenic bacteria clinical isolates to form a dataset of pathogenic bacteria clinical isolates.
[0037] Preferably, in steps 3-4), the preprocessing is data cleaning, specifically including:
[0038] 3-4-1) Global intensity correction and peak alignment were performed on each mass spectrum peak of the experimental group and the control group using internal standards.
[0039] 3-4-2) Detect and extract signals of bacterial characteristic protein peaks from each mass spectrum peak in the experimental group and each mass spectrum peak in the control group, respectively;
[0040] 3-4-3) In accordance with the quality control procedure, low-quality data of each mass spectrum peak in the experimental group and each mass spectrum peak in the control group were screened / removed.
[0041] Preferably, in steps 3-5), the enhanced dynamic relative growth feature vector is constructed in the following manner:
[0042] 3-5-1) Divide the bacterial proteome region in each mass spectrum peak of the experimental group and the control group into several continuous, fixed-width mass-to-charge ratio sub-intervals, and number each mass-to-charge ratio sub-interval.
[0043] 3-5-2) Exclude all low-information-content and mass-to-charge ratio sub-intervals containing internal standard signals from each mass spectrum peak in the experimental and control groups;
[0044] 3-5-3) Calculate the relative growth values of each mass-to-charge ratio sub-interval in the mass spectrum peak diagram of the experimental group and the corresponding mass-to-charge ratio sub-interval in the mass spectrum peak diagram of the control group at each sampling time point;
[0045] 3-5-4) Using the relative growth values calculated in step 3-5-3) and the mass-to-charge ratio sub-interval numbers calculated in step 3-5-1), a matrix is formed, and this matrix is used as the enhanced dynamic relative growth feature vector.
[0046] Preferably, in step 4), the classification model is obtained in the following manner:
[0047] 4-1) Select several machine learning models, train each model using a dataset of clinical isolates of pathogenic bacteria, and optimize hyperparameters by combining five-fold cross-validation and random grid search.
[0048] 4-2) Use performance evaluation metrics to evaluate the performance of different machine learning models, and use the best-performing machine learning model as the final classification model;
[0049] 4-3) Apply model interpretability reasoning techniques to the classification model to analyze the contribution or correlation score of each enhanced dynamic relative growth value to the final classification result, in order to enhance the interpretability of the model and discover drug resistance-related biomarkers.
[0050] Preferably, in step 4-2), the various machine learning models include logistic regression, decision tree, XGBoost, and multilayer perceptron, covering different algorithm characteristics and adapting to the needs of detecting drug resistance phenotypes of pathogenic bacteria.
[0051] Preferably, in step 4-3), the performance evaluation metrics include area under the curve, accuracy, precision, recall, and F1 score, which constructs a multi-dimensional, standardized, and clinically oriented evaluation system that can avoid the limitations of a single metric and ensure the scientific and reliable nature of machine learning model performance evaluation.
[0052] The advantages of this invention are as follows:
[0053] ① This invention significantly improves the speed of drug resistance detection. By screening for the optimal culture time (the optimized total culture time window is only about 2 hours) and combining it with rapid MALDI-TOF mass spectrometry detection technology, the determination of bacterial drug resistance can be completed within 2 hours. Compared with traditional phenotypic drug susceptibility testing methods that require 24 to 48 hours or even longer, the speed of drug resistance detection in this invention represents a qualitative leap. It can provide rapid and reliable drug susceptibility results for clinical use, greatly supporting early and accurate anti-infection treatment decisions, and helping to improve patient prognosis and control the spread of infection.
[0054] ② This invention significantly improves prediction accuracy, particularly in distinguishing intermediate resistance phenotypes. By introducing an innovative enhanced dynamic relative growth (RBD-RG) feature extraction strategy, this invention no longer relies solely on overall growth differences or changes in a few manually selected feature peaks. Instead, it precisely captures the complex dynamic relative changes in the bacterial proteome across multiple mass-to-charge ratio (m / z) sub-intervals and multiple time points under antibiotic treatment. This multi-dimensional, dynamic feature information contains richer biological connotations. By combining it with machine learning models, it can more effectively learn and identify subtle but crucial patterns associated with different resistance phenotypes (sensitive, intermediate, and resistant). Compared to the traditional threshold-based MBT-ASTRA method, this invention significantly improves overall accuracy (from 87.6% to over 94.4%), especially in accurately identifying difficult-to-distinguish "intermediate" resistant strains.
[0055] ③ This invention has good scalability and adaptability. By simply optimizing key experimental parameters (such as antibiotic concentration, culture time, and sampling points) and training the corresponding machine learning model, it can be applied to other types of bacteria and combinations of different antibiotics, showing good application potential.
[0056] Glossary
[0057] Robust dynamic relative growth (RBD-RG) eigenvector matrix: refers to the matrix composed of RG values dynamically collected at different time points in different mass-to-charge ratio intervals, used to measure the dynamic growth trend of microorganisms over a period of time.
[0058] MALDI-TOF mass spectrometry sample: refers to a co-crystallized thin layer formed by mixing the analyte microorganism (such as pathogenic bacteria cells or their lysates) with a matrix solution (such as α-cyano-4-hydroxycinnamic acid, HCCA), spotting the mixture onto a dedicated metal target plate, and drying it at room temperature or under specific conditions. This sample is ready for direct loading into a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer for analysis. In this invention, it specifically refers to a mixed crystal containing microbial proteins, internal standards (such as RNase B), and matrix molecules.
[0059] RG value (Relative Growth Value) refers to the ratio of the growth rate of microorganisms under specific culture conditions (such as in a culture medium containing antibiotics) to that of a control group without antibiotics. It is used to quantify the inhibitory effect of antibiotics on microbial growth.
[0060] AiMedLab MS: An automated system built for this invention, designed to integrate and automate the key data processing and analysis steps in this invention (such as enhanced dynamic relative growth (RBD-RG) feature calculation and extraction, machine learning model construction and prediction, etc.), improve efficiency and standardization, and includes communication interfaces with laboratory information systems (LIS) and hospital information systems (HIS) to facilitate clinical data transmission.
[0061] Prognosis refers to the prediction of disease progression and possible outcomes, including whether the disease can be cured, the speed of recovery, the risk of complications, the length of survival, and the quality of life. The "improved prognosis" in this invention refers to shortening infection control time, reducing the spread of drug-resistant bacteria, lowering the probability of severe illness, and ultimately enabling patients to recover faster and reduce sequelae.
[0062] The MALDI Biotype Rantibiotic Susceptibility Test Rapid Assay (MBT-ASTRA) is a semi-quantitative analytical method based on matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) technology, designed to rapidly and accurately determine the susceptibility of bacteria to antibiotics. Attached Figure Description
[0063] Figure 1 This describes the overall workflow of the present invention;
[0064] Figure 2 This is a schematic diagram showing the screening results of the optimal detection conditions for the MBT-ASTRA method based on RG values. Figure 2 A shows the screening results of the MBT-ASTRA method at 2h, 4h, and 6h. Figure 2 B is a schematic diagram showing the dynamic changes of RG values of pathogenic bacteria with different drug-resistant phenotypes at different time points within the optimal reaction system;
[0065] Figure 3 This is a diagram illustrating the predictive performance of a machine learning model. Figure 3 A is the sample distribution graph. Figure 3 B represents the classification accuracy of the four machine learning models in 10 repeated experiments. Figure 3 C is a schematic diagram illustrating the internal test results of the four machine learning models. Figure 3 D is a schematic diagram of the external test results of the four machine learning models;
[0066] Figure 4 This diagram illustrates the interpretability analysis of MALDI-TOF mass spectrometry quality control, characteristic peak distribution, RBD-RG value calculation, and AMR prediction results. Figure 4 A is a schematic diagram of the MALDI-TOF MS quality control automated assessment process. Figure 4 B is a schematic diagram showing the characteristic peak distribution of Escherichia coli and RNase B, and the calculation of RBD-RG values. Figure 4 C is a schematic diagram illustrating the impact of the average |correlation score| of the seven m / z sub-intervals at different time points on the decision outcome. Figure 4 D is a heatmap visualization of the average |correlation score| across four different m / z sub-intervals at four time points;
[0067] Figure 5 This is a schematic diagram of the login page of Ai MedLab MS in this embodiment;
[0068] Figure 6 This is a schematic diagram of the homepage of Ai MedLab MS in this embodiment;
[0069] Figure 7 This is a schematic diagram illustrating the screening of the patient list based on registration criteria in this embodiment;
[0070] Figure 8 This is a schematic diagram illustrating the selection of the corresponding patient to enter the analysis interface in this embodiment;
[0071] Figure 9 This is a schematic diagram illustrating the report review and submission process in the data analysis workstation in this embodiment;
[0072] Figure 10 This is a flowchart of the method of the present invention; Detailed Implementation
[0073] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0074] Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, and Staphylococcus aureus are all important clinical pathogens that frequently cause various nosocomial and community infections. Clinically, appropriate antibiotics are commonly used to treat these pathogens. For example, levofloxacin (LEV) is used to treat Escherichia coli infections, meropenem (MEM) is used to treat infections caused by carbapenem-sensitive strains of Klebsiella pneumoniae, Pseudomonas aeruginosa, and Acinetobacter baumannii, and oxacillin (OXA) is used to treat methicillin-sensitive Staphylococcus aureus infections.
[0075] Example 1: Optimal detection conditions for screening strains
[0076] 1) Within the range of 1–128 mg / L, set the target concentrations of antibiotics (levofloxacin, meropenem, oxacillin) as 1 mg / L, 4 mg / L, 16 mg / L, 64 mg / L, and 128 mg / L, and set the target incubation times as 2 hours, 4 hours, and 6 hours.
[0077] 2) By combining the target concentrations of each antibiotic with different target culture times, 15 distinct culture conditions were created:
[0078] (1 mg / L, 2 hours), (1 mg / L, 4 hours), (1 mg / L, 6 hours);
[0079] (4 mg / L, 2 hours), (4 mg / L, 4 hours), (4 mg / L, 6 hours);
[0080] (16 mg / L, 2 hours), (16 mg / L, 4 hours), (16 mg / L, 6 hours);
[0081] (64 mg / L, 2 hours), (64 mg / L, 4 hours), (64 mg / L, 6 hours);
[0082] (128 mg / L, 2 hours), (128 mg / L, 4 hours), (128 mg / L, 6 hours);
[0083] 3) Standard strains were cultured for a short time under different culture conditions, and the standard strains covered different susceptibility types of the corresponding antibiotics: sensitive, intermediate, resistant or sensitive and resistant;
[0084] 4) The MBT-ASTRA method based on RG value calculation was used to distinguish the phenotypes (sensitive, intermediate, and resistant) of standard strains under different culture conditions. Under the premise of ensuring that different resistant phenotype strains can be distinguished with good accuracy, the culture conditions with the shortest time and lowest concentration were taken as the optimal culture conditions, the antibiotic concentration was taken as the optimal antibiotic concentration, and the culture time was taken as the optimal culture time.
[0085] like Figure 2 As shown in Figure A, taking the standard strain of Escherichia coli as an example, when the concentration of levofloxacin is 64 mg / L and the culture time is 2 h, the three strain types of sensitive, intermediate and resistant can be distinguished with good accuracy. Therefore, the optimal detection conditions are selected as a culture time of 120 min (optimal culture time) and a concentration of levofloxacin of 64 mg / L (optimal antibiotic concentration).
[0086] Standard strains were rapidly cultured under optimal detection conditions, and samples were collected every 30 minutes. The phenotypes of the Escherichia coli standard strains were distinguished using the MBT-ASTRA method based on RG values, such as... Figure 2 As shown in Figure B, it can be observed that with increasing culture time, the growth trends of susceptible and resistant Escherichia coli strains show significant differences, gradually transitioning from being indistinguishable to having clear boundaries. At 120 min, the boundaries between the three strains are clearly discernible, and the critical values are calculated and visualized in the figure. RG < 0.101 indicates susceptible strains, RG > 0.396 indicates resistant strains, and those in between are intermediate strains.
[0087] For each target pathogen-antibiotic combination, prior to large-scale clinical isolate testing, pre-experiments are necessary to systematically optimize and determine its specific key experimental parameters. This embodiment aims to screen for the optimal antibiotic working concentration and optimal dynamic monitoring time window (including total culture time and multiple sampling time points within the window) that maximize the differentiation of different drug resistance phenotypes for a specific pathogen-antibiotic combination. These optimized parameters will serve as standard conditions for subsequent processing of clinical isolates of this pathogen species.
[0088] This embodiment uses the typical combination of "Escherichia coli-levofloxacin" as an example, and conducts preliminary experiments using standard strains of Escherichia coli (covering different levofloxacin susceptibility types: sensitive, intermediate, and resistant) to systematically optimize and explore key experimental parameters in subsequent detection procedures. This step aims to screen out the working concentration of antibiotics that can maximize the differentiation of different resistance phenotypes and the time window for dynamic monitoring (including total culture time and multiple sampling time points within it) for processing subsequent clinical isolates.
[0089] Example 2: Collection and Screening of Strains
[0090] 1) Escherichia coli clinical isolates were purified and cultured on Columbia blood agar medium for 16-18 hours at 37℃ and 5% CO2.
[0091] The Escherichia coli clinical isolates used in this embodiment were collected from four medical institutions in Chongqing. They were obtained by collecting clinical Escherichia coli samples (including urine, sputum, blood, pus and secretions) from patients infected with the pathogen. The isolates were purified and cultured in the Clinical Microbiology Laboratory of Dazu Hospital Affiliated to Chongqing Medical University using Columbia Blood Agar medium.
[0092] 2) The clinical isolates of Escherichia coli were identified using MALDI-TOF mass spectrometry (Zhongyuan Huiji EX-Accuspec V2 system), and strains with an average score of ≥2.0 from three repeated identifications were screened.
[0093] 3) The selected strains were subjected to levofloxacin susceptibility testing using the VITEK 2Compact system according to the 2024 CLSI standards. The strains were classified as sensitive, intermediate, and resistant, and 20 clinical isolates were selected for each category to construct a standardized database.
[0094] In this embodiment, the specific process of the drug sensitivity test is as follows:
[0095] ① Preparation of bacterial suspension: Select 2-3 colonies of the target bacteria, approximately 3 mm in size, after 18-24 hours of pure culture, and dilute them in a test tube containing 3 mL of 0.45% physiological saline. Measure the bacterial suspension concentration using a standard turbidimeter. Adjust the turbidity by adding physiological saline or bacterial colonies until the bacterial suspension concentration reaches the standard required by the test card, generally 0.5 McFarland units (approximately 1.5 × 10⁻⁶). 8 (CFU / mL), place the test tube on the sample rack.
[0096] ② Prepare the drug susceptibility test card: Take the drug susceptibility test card suitable for levofloxacin from the refrigerator and place it at room temperature for 2-3 minutes to allow it to reach the same temperature as room temperature, avoiding condensation from affecting the test.
[0097] ③ Card Filling: Place the test card on the card holder and immerse the sample delivery tube in the standard tube containing the bacterial solution to be tested. Place the card holder into the instrument's filling chamber and press the "START FILL" contact switch. Filling will be completed in approximately 70 seconds.
[0098] ④ On-machine testing: After filling is complete, wait for the blue indicator light to flash, remove the card holder, and place it into the VITEK 2Compact system within 10 minutes.
[0099] ⑤ Incubation and Detection: The instrument automatically incubates the bacterial solution while simultaneously monitoring bacterial growth in real time using optical sensors. Based on the bacteria's tolerance to levofloxacin, the instrument automatically generates corresponding growth curves.
[0100] ⑥ Result Interpretation and Reporting: The system automatically analyzes bacterial growth data based on the built-in 2024 CLSI standard, obtains the minimum inhibitory concentration (MIC) value of levofloxacin, and determines the sensitivity of bacteria to levofloxacin. The results are divided into sensitive (S), intermediate (I) and resistant (R).
[0101] Example 3 Sample Processing and Mass Spectrometry Data Acquisition
[0102] 1) Fresh, pure colonies were selected from the screened clinical isolates of Escherichia coli, and two bacterial suspensions were prepared. The turbidity was adjusted to 0.5 McFarland units as the standard. One tube was the control group (without antibiotics), and the other tube was the experimental group (antibiotic-treated group). Levofloxacin was added to the experimental group to achieve a final concentration of 64 mg / L.
[0103] 2) In order to capture the dynamic process of bacterial growth, multiple sampling time points were set within the optimal culture time, specifically 30 minutes, 60 minutes, 90 minutes and 120 minutes after the start of incubation. Samples were taken from the control group and the experimental group at these four time points.
[0104] 3) Samples taken from the control group and experimental group were prepared as MALDI-TOF mass spectrometry samples for the control group and MALDI-TOF mass spectrometry samples for the experimental group, respectively. The specific process for preparing MALDI-TOF mass spectrometry samples is as follows:
[0105] 3-1) First, centrifuge the sample at high speed (13,000 rpm, 3 minutes) to collect the bacterial cells in the sample, wash with sterile water, centrifuge again and discard the supernatant.
[0106] 3-2) Next, treat the bacterial cells with an appropriate amount of 75% ethanol and then air dry.
[0107] 3-3) In order to effectively lyse the bacterial cell wall and release proteins, add a small amount of 70% formic acid, followed by 100% acetonitrile solution containing internal standard protein (RNase B with a final concentration of 20 g / L), and mix thoroughly.
[0108] 3-4) Centrifuge again, and take the supernatant (usually 1 μL) and spot it onto the designated well in the metal target plate.
[0109] 3-5) After the supernatant is removed from each sample spot, add an equal volume of matrix solution (α-cyano-4-hydroxycinnamic acid, HCCA) to the covered spot and allow it to air dry at room temperature.
[0110] 4) Set the mass-to-charge ratio detection range of the MALDI-TOF mass spectrometer (Zhongyuan Huiji EXS3000) to 2000-20000 Da, the characteristic peak search range to 3400-20000 Da, and use a 60 Hz solid-state laser with the laser energy set to 10%.
[0111] 5) Load the MALDI-TOF mass spectrometry samples (i.e., metal target plates) of the control group and the experimental group into the MALDI-TOF mass spectrometer (Zhongyuan Huiji EXS3000) for mass spectrometry data acquisition, and obtain 4 mass spectrometry peak diagrams of the control group and 4 mass spectrometry peak diagrams of the experimental group.
[0112] To ensure data reliability and reproducibility, instrument calibration and quality control were performed using a standard strain (Escherichia coli ATCC25922) before each batch of experiments. Multiple biological replicates (6 lysis buffers) and technical replicates (6 target sites per lysis buffer, 2 measurements per target plate) were prepared for each strain, and sampling paths were randomized. The acquired raw mass spectrometry data (mass spectrum peaks) were typically pre-processed by the instrument software (e.g., baseline correction, smoothing) before being exported.
[0113] Example 4: Data Preprocessing and Quality Control
[0114] The obtained raw mass spectrum peaks require further preprocessing and quality control, primarily using a programming language (Python 3.11) and a related scientific computing library (SciPy 1.13.1) for data processing. Key steps include:
[0115] 1) Based on the signal intensity and area of the internal standard protein (RNase B, whose characteristic peaks are located at 7400-7800 m / z and 14800-15600 m / z), global intensity correction and alignment were performed on all mass spectrometry peaks of the control group and the experimental group to eliminate systematic bias.
[0116] 2) Identify the main protein peaks of Escherichia coli (concentrated in the range of 3400-14800 m / z) in each mass spectrometry peak diagram of the control group and the experimental group;
[0117] 3) Establish an automated quality control process, and score each mass spectrometry peak of the control group and the experimental group according to the number and quality of internal standard peaks, and eliminate low-quality data (using the Levey-Jennings 3σ principle).
[0118] Example 5: Calculation of Enhanced Dynamic Relative Growth (RBD-RG) Characteristics
[0119] 1) Divide the main distribution area (3400-14800 m / z) of Escherichia coli characteristic protein peaks in each mass spectrometry peak image of the control group and the experimental group into multiple continuous sub-intervals with fixed widths, and set the width of each sub-interval to 1000 m / z.
[0120] 2) Ranges containing characteristic peaks of the internal standard protein (RNase B) in the mass spectra of both the control and experimental groups (e.g., 7400-7800 m / z) were excluded because the signals in these ranges do not directly reflect the growth status of the bacteria. Simultaneously, ranges with extremely weak or absent signals in the vast majority of samples (e.g., in this study, over 90% of the data in the 10800-14800 m / z range were 0 or infinite, indicating low information content) were also excluded to improve computational efficiency and reduce noise interference. After screening, a series of effective m / z sub-ranges for analysis were finally determined. For example, in this embodiment, there are seven effective m / z sub-ranges: 3400–4400 m / z, 4400–5400 m / z, 5400–6400 m / z, 6400–7400 m / z, 7800–8800 m / z, 8800–9800 m / z, and 9800–10800 m / z.
[0121] 3) Calculate the RBD-RG value of each selected m / z fraction in the mass spectrometry peak diagram corresponding to each sampling time point (30 minutes, 60 minutes, 90 minutes, and 120 minutes) during the incubation process, specifically including:
[0122] The RBD-RG value intuitively reflects the relative impact of antibiotic treatment on bacterial protein synthesis or stability at a specific time point and within a specific molecular weight range. This value is calculated by comparing the characteristic peak areas of the LEV-treated sample (treatment group) and the untreated sample (control group) within the same time point and the same m / z sub-interval. To eliminate systematic biases caused by experimental factors such as sample size, extraction efficiency, and ionization efficiency, internal standard proteins are used for normalization during the calculation.
[0123] 3-1) First, calculate the total peak area of the treatment group in this sub-interval and divide it by the total area of the peaks within the treatment group;
[0124] 3-2) Then, calculate the total peak area of the control group in this sub-interval and divide it by the total peak area of the control group;
[0125] 3-3) Finally, divide the normalized signal value of the treatment group by the normalized signal value of the control group to obtain the RBD-RG value of the m / z sub-interval at that time point.
[0126] 4) The calculated m / z sub-interval RBD-RG values are integrated into a feature representation that can represent the complete dynamic response pattern of the sample. Specifically, this is integrated into a matrix, where rows represent different m / z sub-intervals and columns represent different sampling time points (e.g., a 7×4 matrix). This matrix is usually flattened into a single long vector, forming a 1-dimensional feature vector containing 28 elements.
[0127] This 28-dimensional RBD-RG feature vector, incorporating temporal and spatial (m / z) dimensions, constitutes the final input feature for training and predicting the levofloxacin resistance phenotype of Escherichia coli. Compared to traditional single RG values, this multi-dimensional dynamic feature representation can more comprehensively and meticulously characterize the differences in bacterial responses to antibiotics, thus potentially improving the accuracy and robustness of resistance prediction, especially for difficult-to-distinguish intermediate-resistant strains.
[0128] Example 6: Machine Learning Model Construction, Training, and Evaluation
[0129] 1) Mass spectrometry data of some Escherichia coli clinical isolates were divided into training and testing sets in an 8:2 ratio. Another part of the mass spectrometry data of Escherichia coli clinical isolates was used as an external validation set. The RBD-RG feature vector calculated in the data preprocessing stage was used as the input for model training. The classification target was the drug resistance spectrum (sensitive / intermediate / resistant) of the original isolates.
[0130] 2) Select multiple machine learning models, including logistic regression (LR), decision tree (DT), XGBoost (XGB), and multilayer perceptron (MLP), train all models using the training set, and combine five-fold cross-validation and random grid search to optimize hyperparameters in order to explore the best model.
[0131] like Figure 3 As shown in C and 3D, the four models were first tested using a test set, and then validated using data from hospitals other than the one used in the training process. This was to verify the cross-hospital performance and clinical practice results of the four models. The performance of the four models is as follows:
[0132] Internal testing: LR, XGB, and MLP all performed excellently in all three categories. XGB showed overall stability and strong performance in the "LEV-S" and "LEV-R" categories, but had some shortcomings in identifying "LEV-I" samples. In contrast, MLP performed the best. DT performed the worst, struggling to distinguish between "LEV-S" and "LEV-I" strains.
[0133] External Testing: Compared to internal testing, LR showed a significant gap in external validation, exhibiting poor classification performance for "LEV-I" samples. DT also performed poorly, frequently misclassifying "LEV-S" samples. Compared to other models, MLP demonstrated higher robustness to external data.
[0134] The classification accuracy (mean ± std) of the four machine learning models in 10 repeated experiments is as follows: Figure 3 As shown in B, where mean represents the average and std represents the standard deviation. Combining the analysis of mean and variance, the MLP model exhibits the strongest generalization ability, demonstrating excellent performance in both internal testing and external validation.
[0135] 3) After using the test set to input the model and generate prediction results, calculate the area under the curve (AUROC), accuracy, precision, recall, and F1 score of each machine learning model to evaluate the performance of different machine learning models, and select the machine learning model with the best performance as the final classification model.
[0136] MLP achieved the highest classification accuracy (0.944±0.040) and performed well in recalling intermediate strains (0.953±0.056). XGBoost demonstrated overall performance with low recall rates for sensitive strains (0.893±0.073) and
[0137] The AUROC score was relatively high (0.985±0.035). LR performed well, but its generalization ability was weak, resulting in poor performance in external validation. DT performed poorly, especially on sensitive and resistant strains, with recall and precision both below 0.8, and overall precision <0.8. In external validation, MLP continued to outperform other models, achieving an AUROC score of 0.984±0.014 for intermediate strains. The F1 score reached 0.982±0.016. After comprehensive evaluation, MLP was selected as the final classification model.
[0138] 5) Perform external validation on the finalized classification model. Specifically, input external validation set data from independent data sources into the model and comprehensively evaluate its predictive performance, including calculating multiple key performance indicators such as the area under the curve (AUROC). Through this external validation results, the actual predictive efficacy and generalization level of the model on cross-center and cross-batch data are finally confirmed.
[0139] 6) Use LRP correlation scores to analyze the contribution of each RBD-RG interval as input feature to AMR prediction and identify potential biomarkers within specific m / z sub-intervals.
[0140] RBD-RG features contain dynamic information about bacterial responses to specific antibiotics in different molecular weight regions. To gain a deeper understanding of the model's decision-making mechanism or to uncover potential biomarker clues, model interpretive inference techniques (hierarchical correlation propagation (LRP) or other similar methods) can be applied to the trained optimal machine learning model. Figure 4 As shown in C and 4D. By analyzing the contribution or correlation score of each input feature (i.e., the RBD-RG value of a specific m / z sub-interval at a specific time point) to the final prediction result, key features that play a decisive role in distinguishing different drug resistance phenotypes can be identified (e.g., determining the most important early dynamic changes in certain low m / z intervals). This not only enhances the credibility of the model's prediction results but may also reveal potential biomarker regions related to drug resistance mechanisms, providing valuable clues for subsequent research on drug resistance mechanisms.
[0141] It is worth noting that after the classification model is trained, for pathogenic bacteria samples with unknown drug susceptibility phenotypes, the enhanced dynamic relative growth (RBD-RG) feature vector is obtained through the same steps described above. This vector is then input into the trained optimal model to obtain the drug resistance prediction result for the sample. Figure 1 As shown.
[0142] Example 7: Construction of Ai MedLab MS
[0143] Traditional MALDI-TOF MS-based drug resistance detection methods (such as MBT-ASTRA and its improved methods) suffer from cumbersome operation and reliance on manual selection of characteristic peaks. These issues introduce subjectivity and human error, leading to insufficient stability and universality of the results. To overcome these problems, this embodiment develops the intelligent mass spectrometry analysis system Ai MedLab MS, whose construction logic can be broken down into a four-layer architecture of "hardware-data-model-application," specifically including:
[0144] ① Hardware layer: This layer focuses on "equipment customization + process automation" and is responsible for integrating standardized mass spectrometry equipment with automated culture.
[0145] a. Mass spectrometry equipment configuration: Zhongyuan Huiji EX-Accuspec V2 equipment is used for pathogenic bacteria identification, and EXS3000 equipment is used for mass spectrometry data acquisition; through the internal standard (RNase B) calibration module, the mass spectrometry peak intensity of different time points and different samples is normalized to eliminate systematic bias.
[0146] b. Automated Culture and Sampling: An integrated gradient antibiotic concentration culture module (such as 64 mg / L levofloxacin for Escherichia coli in Example 1) is used, with four dynamic sampling time points set at 30 / 60 / 90 / 120 minutes. Sampling of the experimental group (containing antibiotics) and the control group (without antibiotics) is completed simultaneously via a robotic arm or peristaltic pump, reducing human error. The sampling progress can be visualized through the system's "CurrentProgress" timeline (corresponding to...). Figure 8 a).
[0147] c. Standardized sample preparation: Built-in automated sample processing workflow, including 75% ethanol lysis of bacterial cells and formic acid-acetonitrile extraction of bacterial proteins, ensuring consistency of bacterial protein extraction, and finally outputting standardized metal target plate mass spectrometry samples that can be directly used for MALDI-TOF MS analysis.
[0148] ② Data Layer: This layer focuses on "dynamic feature extraction + data quality control" and is responsible for constructing the core analysis data RBD-RG feature matrix.
[0149] a. Construction of RBD-RG feature matrix:
[0150] Time dimension: Differences in bacterial growth kinetics were captured through four sampling points at 30 / 60 / 90 / 120 minutes (e.g., ...). Figure 2 As shown in Figure B, the RG value of sensitive strains remained below 0.1, while the RG value of resistant strains was above 0.4.
[0151] Spatial dimension: The bacterial proteome mass spectrum peaks were divided into 7 continuous mass-to-charge ratio sub-intervals from 3400 to 10800 m / z (excluding the 7400 to 7800 m / z interval where the internal standard RNase B is located). The peak area ratio of the experimental group to the control group in each interval was calculated to form a 7×4 feature matrix.
[0152] b. Data quality control: The Levey-Jennings 3σ principle is used to remove low-quality mass spectra; the system monitors data quality in real time through the "QC" tag, and the quality control results can be visualized.
[0153] ③ Model layer: This layer focuses on "multi-model competition + interpretability + iterative optimization" to achieve accurate prediction of drug resistance phenotypes.
[0154] a. Multi-model training and optimization: Integrating four machine learning models: Logistic Regression (LR), Decision Tree (DT), XGBoost, and Multilayer Perceptron (MLP); the optimal model is selected through five-fold cross-validation and hyperparameter grid search (such as optimizing the number of hidden layer nodes and learning rate of MLP).
[0155] b. Model interpretability: Based on hierarchical correlation propagation (LRP) correlation analysis, key features for predicting drug resistance phenotypes are identified (e.g., the RG value in the 3400-4400 m / z range at 60 minutes contributes the most to the intermediate drug resistance phenotype), and a visual heatmap is generated to assist in clinical interpretation.
[0156] c. Incremental learning mechanism: After each test, the system automatically stores the new data (including test sample information, RBD-RG feature matrix, and drug sensitivity label) into the knowledge base, and iteratively optimizes the model through reinforcement learning to improve the adaptability to different pathogenic bacteria-antibiotic combinations.
[0157] ④ Application layer: This layer focuses on "clinical usability," achieving deep integration of the system with clinical scenarios.
[0158] a. Human-computer interaction interface: The system dashboard realizes three major functions: sample screening (such as filtering by department and sampling time), data visualization (such as mass spectrometry peak comparison between experimental group and control group), and automatic report generation (such as labeling drug resistance phenotype results such as "Predicted as LEV-S").
[0159] b. LIS / HIS system integration: Data is transmitted via the HL7 interface, including basic patient information and drug resistance test results, and test reports are automatically archived to the electronic medical record system.
[0160] c. Hierarchical permission management: Differentiate between "audit" (can modify data and audit reports) and "view" (can only monitor the detection progress, no right to modify), which complies with HIPAA privacy protection requirements.
[0161] When using Ai MedLab MS, users need to first access the dashboard on the login screen, such as... Figure 5 As shown, identity verification is then performed. Different user accounts have different levels of permissions, which determine the scope of operations. User permissions are divided into two categories: "Audit" and "View". Users with "Audit" permissions can access the data analysis workstation within the dashboard to analyze and audit samples that have completed data collection, and send the final report to the LIS system. Users with only "View" permissions can only monitor the current progress of the workflow and have no right to modify patient data. These accounts are typically logged into the computer connected to the data display terminal to show the current workflow status to laboratory technicians, patients, etc.
[0162] After successfully logging in, users can access the system homepage interface, such as... Figure 6 As shown, the core interactive area of the system's homepage is divided into three parts: the information panel, the operation panel, and the data panel. The information panel displays the system's operating status parameters in real time; the operation panel integrates control functions such as data filtering and analysis initiation; and the data panel presents the workflow node status in a visual format. These three elements work together to achieve human-computer interaction.
[0163] In the patient data processing stage, users can use a filtering interface, such as... Figure 7 As shown, click the "Filter" button to enter the filtering interface, and filter the patient list from four dimensions based on preset registration criteria. After selecting the target patient, click the "Data Analysis" button (e.g., ...). Figure 8 (As shown in a) Upon starting the data analysis workstation, the system automatically transmits patient identification information and mass spectrometry data from four time points to this module. After receiving the data, the data analysis workstation sequentially executes an automated processing flow including characteristic peak detection, mass spectrometry quality control, and dynamic relative growth measurement. The processing results are visualized using different colors to distinguish mass spectrometry peaks before and after antibiotic intervention (e.g., ...). Figure 8 (as shown in b).
[0164] like Figure 9 As shown, during the report generation phase, users trigger data inference calculations by clicking the "Calculate" button. After the calculation is completed, the status indicator light in the top bar turns green. After manual verification, the structured report can be transferred to the system dashboard, ultimately completing the data integration with the LIS system.
[0165] The Ai MedLab MS system, independently developed in this embodiment, significantly improves the automation, standardization, and ease of operation of the testing process. This system integrates the entire analysis workflow, from data preprocessing and quality control to RBD-RG feature calculation and machine learning model prediction, minimizing manual steps and potential subjective errors, and improving the stability and repeatability of results. A rigorous automated quality control module further ensures the reliability of input data. Simultaneously, it achieves data interoperability with Laboratory Information Systems (LIS) and Hospital Information Systems (HIS) through ports, making result reporting and data management more efficient and allowing for rapid integration into existing clinical laboratory workflows.
[0166] In summary, this embodiment, by innovatively combining dynamic multi-interval mass spectrometry feature analysis with machine learning technology and supplemented by automated system design, successfully overcomes the limitations of existing bacterial resistance detection technologies in terms of speed, accuracy (especially for intermediate resistance), automation, information depth, and standardization. It provides a rapid, accurate, reliable, automated new approach for resistance detection with potential biomarker discovery, possessing significant clinical application value and broad development prospects, and contributing to the promotion of precision medicine and the rational use of antimicrobial drugs. Although this invention has been described in detail using *Escherichia coli* as an example, the core idea of this invention—namely, predicting resistance phenotypes through dynamic multi-interval mass spectrometry feature peaks combined with machine learning—has broad applicability. Those skilled in the art can refer to the steps disclosed in this invention (such as...) Figure 10 As shown in the figure, by performing routine experiments to screen the optimal antibiotic concentration and culture time in step 1), this method can be extended to the detection of drug resistance in other pathogens (such as Enterococcus spp. and Streptococcus spp.) and other classes of antibiotics (such as aminoglycosides and macrolides) without requiring creative work.
[0167] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications made to the present invention by those skilled in the art without departing from the spirit of the present invention shall fall within the protection scope of the present invention.
Claims
1. A method for rapid identification of pathogenic bacteria drug resistance phenotype based on machine learning combined with MALDI-TOF MS, characterized in that, The method comprises the following steps: 1) Firstly, the optimal antibiotic concentration and the optimal culture time for distinguishing the drug resistance phenotype of pathogenic bacteria are determined by using standard strains of pathogenic bacteria and combining with the MALDI biotype antibiotic sensitivity test rapid determination method; 2) Collect and screen clinical isolates of pathogenic bacteria, and determine the drug sensitivity label of the clinical isolates of pathogenic bacteria by a standard drug sensitivity test method; 3) Culture the clinical isolates of pathogenic bacteria under the optimal antibiotic concentration and the optimal culture time, dynamically obtain the mass spectrum peak graph dynamic relative growth value of the clinical isolates of pathogenic bacteria at different time points, constitute a plurality of enhanced dynamic relative growth characteristic vector matrices corresponding to different time points, and combine with the drug sensitivity label of the clinical isolates of pathogenic bacteria to construct a clinical isolate data set of pathogenic bacteria; 4) Train a plurality of machine learning models by using the clinical isolate data set of pathogenic bacteria, and select the optimal machine learning model as a drug resistance phenotype classification model by using a performance evaluation index; 5) When identifying the drug resistance phenotype of a clinical test sample of pathogenic bacteria, extract the enhanced dynamic relative growth characteristic vector matrix of the clinical test sample of pathogenic bacteria, and input the enhanced dynamic relative growth characteristic vector matrix into the drug resistance phenotype classification model to output the drug resistance phenotype of the clinical test sample of pathogenic bacteria.
2. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 1, characterized in that, In step 1), the determination of the optimal antibiotic concentration and the optimal culture time specifically comprises: 1-1) Gradiently set a plurality of target antibiotic concentrations C(1, 2,..., N) and a plurality of target culture times T(1, 2,..., M); 1-2) combine the target antibiotic concentration C1 with T1, T2,..., T M respectively, the target antibiotic concentration C2 with T1, T2,..., T M respectively,..., the target antibiotic concentration C N with T1, T2,..., T M respectively, to obtain several target culture conditions; 1-3) Culture the standard strains of pathogenic bacteria by using a plurality of target culture conditions to obtain mass spectrum samples available for MALDI-TOF MS analysis; 1-4) Detect the mass spectrum peak graphs of the standard strains of pathogenic bacteria under each target culture condition by using MALDI-TOF MS, and select the optimal antibiotic concentration and the optimal culture time.
3. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 1, characterized in that, In step 2), the collection and screening of clinical isolates of pathogenic bacteria and the determination of the drug sensitivity label of the clinical isolates of pathogenic bacteria by a standard drug sensitivity test method specifically comprise: 2-1) Isolate a plurality of strains from clinical specimens, and purify and culture the plurality of strains at 37°C in a 5% CO2 environment; 2-2) Calculate the matching item confidence score of the plurality of strains by using MALDI-TOF MS multiple times, and then calculate the average value of the matching item confidence scores of the plurality of strains, wherein the strains with an average value of the matching item confidence scores greater than or equal to 2.0 are used as the clinical isolates of pathogenic bacteria; 2-3) Distinguish the drug resistance phenotype of the clinical isolates of pathogenic bacteria by using a standard drug sensitivity test method, and use the drug resistance phenotype as the drug sensitivity label of the clinical isolates of pathogenic bacteria.
4. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 1, characterized in that, In step 3), the dynamic acquisition of the mass spectrum peak graph dynamic relative growth value of the clinical isolates of pathogenic bacteria at different time points, the constitution of the enhanced dynamic relative growth characteristic vector matrix, and the combination with the drug sensitivity label of the clinical isolates of pathogenic bacteria as the clinical isolate data set of pathogenic bacteria specifically comprise: 3-1) Prepare a standard bacterial suspension of the clinical isolates of pathogenic bacteria, and divide the standard bacterial suspension into an experimental group and a control group; 3-2) At multiple sampling time points within the optimal culture time, samples are taken from the experimental group and the control group respectively, and MALDI-TOF mass spectrum samples of the experimental group and the control group at each sampling time point are prepared; 3-3) According to the MALDI-TOF mass spectrum samples of the experimental group and the control group, mass spectrum peak maps of the experimental group and the control group at each sampling time point are obtained; 3-4) The mass spectrum peak maps of the experimental group and the control group at each sampling time point are preprocessed in a data cleaning manner to improve the quality of the mass spectrum peak maps; 3-5) The bacterial proteome region in the mass spectrum peak maps of the experimental group and the control group at each sampling time point is divided into continuous and fixed-width m / z subintervals, and the dynamic relative growth values of the experimental group and the control group at each sampling time point are calculated according to the m / z subintervals, to form several enhanced dynamic relative growth feature vector matrices; 3-6) The several enhanced dynamic relative growth feature vector matrices are combined with the drug sensitivity labels of the pathogenic bacteria clinical isolates to serve as a pathogenic bacteria clinical isolate dataset.
5. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 4, characterized in that, In step 3-4), the preprocessing is data cleaning, specifically including: 3-4-1) The internal standard is used to perform global intensity correction and peak alignment on each mass spectrum peak map of the experimental group and each mass spectrum peak map of the control group respectively; 3-4-2) The signals of the bacterial characteristic protein peaks in each mass spectrum peak map of the experimental group and each mass spectrum peak map of the control group are detected and extracted respectively; 3-4-3) According to the quality control program, low-quality data in each mass spectrum peak map of the experimental group and each mass spectrum peak map of the control group are screened out or removed.
6. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 4, characterized in that, In step 3-5), the enhanced dynamic relative growth feature vector is formed in the following manner: 3-5-1) The bacterial proteome region in each mass spectrum peak map of the experimental group and the control group is divided into several continuous and fixed-width m / z subintervals, and each m / z subinterval is numbered; 3-5-2) All m / z subintervals with low information content and containing internal standard signals in each mass spectrum peak map of the experimental group and the control group are excluded; 3-5-3) The relative growth values of each m / z subinterval in the mass spectrum peak map of the experimental group and the corresponding m / z subinterval in the mass spectrum peak map of the control group at each sampling time point are calculated respectively; 3-5-4) A matrix is formed using the relative growth values calculated in step 3-5-3) and the m / z subinterval numbers calculated in step 3-5-1), and the matrix is taken as an enhanced dynamic relative growth feature vector.
7. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 1, characterized in that, In step 4), the classification model is obtained in the following manner: 4-1) Several machine learning models are selected, each model is trained using the pathogenic bacteria clinical isolate dataset, and hyperparameter optimization is performed combined with five-fold cross-validation and random grid search; 4-2) The performance of different machine learning models is evaluated using performance evaluation indicators, and the machine learning model with the best performance is taken as the final classification model; 4-3) Model interpretability reasoning techniques are applied to the classification model to analyze the contribution or correlation score of each enhanced dynamic relative growth value to the final classification result, to enhance the model interpretability and discover drug resistance-related biomarkers.
8. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 7, characterized in that, In step 4-2), the several machine learning models include logistic regression, decision tree, XGBoost, and multilayer perceptron.
9. The method for rapid identification of pathogenic bacteria drug resistance phenotype according to claim 7, characterized in that, In step 4-3), the performance evaluation indicators include area under the curve, accuracy, precision, recall, and F1 score.
Citation Information
Patent Citations
Rapid detection method for bacterial drug resistance by combining bacterial metabolism fingerprint spectrum after short-time antibiotic stimulation and machine learning
CN119438596A