Machine Learning-Based Risk Prediction Methods and Equipment for Clinical Mass Spectrometry

By constructing a 25-dimensional feature space based on machine learning and a hybrid ensemble model, the lag and rigidity problems of the LC-MS system were solved, enabling real-time risk assessment and early warning of the mass spectrometry system, and improving the accuracy and efficiency of detection.

CN122084815AActive Publication Date: 2026-05-26SHANGHAI CLINICAL LAB CENT
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI CLINICAL LAB CENT
Filing Date
2026-04-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing LC-MS mass spectrometry systems suffer from lag, bias, and rigidity during detection, leading to frequent false alarms and an inability to effectively handle complex nonlinear coupling relationships in multidimensional feature dimensions, thus affecting the accuracy and efficiency of detection results.

Method used

A machine learning-based approach is used to construct a 25-dimensional feature space. Random forest and XGBoost models are used for risk prediction. Virtual training data is generated by combining random sampling techniques with multiple probability distributions. The risk discrimination logic is dynamically adjusted to achieve deep coupling analysis of multi-dimensional features and real-time risk assessment.

Benefits of technology

It significantly improves the accuracy and foresight of the early warning system of mass spectrometry, reduces the false alarm rate, realizes the transformation from post-event error correction to pre-event prevention, enhances the adaptability and accuracy of detection, and reduces sample waste and detection delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122084815A_ABST
    Figure CN122084815A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of clinical laboratory medicine and artificial intelligence, and provides a risk prediction method and device based on machine learning for clinical mass spectrometry. The risk prediction method includes: defining N+M dimensions of features based on a liquid chromatography-tandem mass spectrometry system to obtain an N+M dimension feature vector structure; obtaining a virtual training dataset based on the N+M dimension feature vector structure, acquiring parameter values ​​of each dimension of the current batch through a data acquisition interface, and assembling them into N+M dimension feature vector data; performing format verification and invalid value filtering on the N+M dimension feature vector data using the N+M dimension feature vector structure to obtain a feature matrix; and obtaining a standardized real-time risk score based on the feature matrix and a trained fusion-integrated risk prediction model. This invention utilizes a machine learning model to uncover the complex nonlinear relationship between configuration parameters and dynamic parameters, thereby achieving more accurate and forward-looking risk warnings than traditional single-threshold methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of clinical laboratory medicine and artificial intelligence, and in particular to a risk prediction method and device for clinical mass spectrometry based on machine learning. Background Technology

[0002] Liquid chromatography-mass spectrometry (LC-MS) is widely recognized as the "gold standard" for small molecule compound detection due to its extremely high sensitivity, specificity, and ability to simultaneously detect multiple components. It has been widely used in clinical endocrine hormone detection, therapeutic drug monitoring, newborn screening for inherited metabolic diseases, and vitamin detection. However, the LC-MS system is a precise and complex system, and the accuracy of its final detection results is affected by the nonlinear coupling of multiple factors, including the mass spectrometer hardware (such as vacuum level and ion source), chromatographic conditions (column life and mobile phase), and sample matrix (phospholipids, protein precipitates).

[0003] Existing Westgard rules or SST (Self-Testing and Testing) are essentially "post-mortem" quality control. The system only alarms when instrument performance has significantly deteriorated (e.g., column collapse or severe ion source contamination), leading to substantial deviations from quality control results or SST failure. This often means that the entire batch of valuable clinical samples needs to be retested, resulting in significant waste of reagents and delays in reporting. Therefore, the lag in post-mortem alarms in existing technologies is a problem that urgently needs to be addressed.

[0004] Current technologies treat each parameter in isolation, and existing linear models cannot handle complex nonlinear coupling relationships involving multiple dimensions of features. Therefore, the problem of how to handle complex nonlinear coupling relationships involving multiple dimensions of features urgently needs to be solved.

[0005] Current technologies rely on static, fixed thresholds. However, LC-MS is a dynamically evolving system, and using rigid, fixed thresholds can easily generate a large number of false positive alarms during normal system fluctuations, leading to alarm fatigue among researchers and causing them to overlook real risks. Therefore, the rigidity of fixed thresholds is a problem that urgently needs to be addressed.

[0006] Existing technologies still suffer from three major flaws: lag, one-sidedness, and rigidity. Summary of the Invention

[0007] The purpose of this invention is to provide a risk prediction method and device for clinical mass spectrometry based on machine learning.

[0008] To address the above problems, this invention provides a risk prediction method for clinical mass spectrometry based on machine learning, comprising: Based on the built-in N+M dimensional feature definition of the LC-MS system (liquid chromatography-tandem mass spectrometry system), the N+M dimensional feature vector structure is obtained; A virtual training dataset is obtained based on an N+M dimensional feature vector structure; Based on the N+M dimensional feature vector structure, the parameter values ​​of each dimension of the current batch are obtained through the data acquisition interface and assembled into N+M dimensional feature vector data; The N+M dimensional feature vector data is subjected to format validation and invalid value filtering to obtain the feature matrix; Based on the virtual training dataset, a fully trained fusion and integration risk prediction model is obtained; Based on the pure feature matrix and the trained fusion-integrated risk prediction model, a standardized real-time risk score is obtained.

[0009] Furthermore, in the above method, based on the N+M dimensional feature definition built into the LC-MS system, an N+M dimensional feature vector structure, such as 25 dimensions, is obtained, including: Based on the N+M dimensional feature definition built into the LC-MS system, N static configuration parameters are defined, each of which corresponds to a fixed system health baseline, resulting in N health baselines. Based on the N+M dimensional feature definition built into the LC-MS system, M dynamic operating parameters are defined. Each dynamic operating parameter is used to capture the changes in the operating state of the LC-MS system in real time, and is used as M real-time state variables respectively. Based on N health baselines and M real-time state variables, an N+M dimensional feature vector structure is initialized.

[0010] Furthermore, in the above method, the N static configuration parameters include: Overall precision, recovery range, matrix effect coefficient, retention time baseline, retention time allowable window, column design life, tailing factor limit, initial pump pressure, signal-to-noise ratio threshold, upper limit of linear range, correlation coefficient threshold, internal standard stability coefficient, pretreatment complexity (quantification level), isomer interference, internal standard type, and clinical risk level. M real-time state variables, including: number of days from source cleaning, cumulative number of injections, current pump pressure, retention time drift, actual tailing factor, current signal-to-noise ratio, intra-batch precision, system bias, and current standard curve correlation coefficient.

[0011] Furthermore, in the above method, based on the N+M dimensional feature vector structure, a virtual training dataset is obtained, including: For the physical characteristics of different parameters in the N+M dimensional feature vector structure, random sampling is performed using normal distribution, uniform distribution, or exponential distribution respectively. At the same time, preset LC-MS system fault modes are embedded in the sampling process, and finally a high-fidelity virtual dataset containing normal operation state and typical fault modes is generated, which serves as a virtual training dataset that conforms to the physical laws of LC-MS for model training.

[0012] Furthermore, in the above method, based on the detection data exported from the mass spectrometry workstation and the N+M dimensional feature vector structure, the parameter values ​​of each dimension of the current batch are obtained through the data acquisition interface and assembled into N+M dimensional feature vector data, including: Use the detection data file of the current batch exported from the mass spectrometry workstation as the import file; Call the file parsing interface to read the parameter data in the imported file, i.e., the values ​​corresponding to each parameter field. The parameter data includes: the measured values ​​of static configuration parameters and the real-time values ​​of dynamic running parameters. Using a preset field mapping algorithm, the parameter fields in the imported file are matched one by one with the N+M feature dimensions of the feature vector structure. After a successful match, the parameter data of the corresponding parameter field is filled into the corresponding dimension of the N+M feature vector structure to obtain the mapped and filled N+M feature vector data.

[0013] Furthermore, in the above method, the N+M dimensional feature vector data undergoes format validation and invalid value filtering to obtain a feature matrix, including: The system verifies row by row whether the number of dimensions of each feature vector record in the N+M dimensional feature vector data is consistent with the preset N+M dimensions, and selects the feature vector records with matching number of dimensions as the feature vector records that pass the verification. From the verified feature vector records, delete records containing null values, missing values, or invalid outliers to ensure the purity of the input data and obtain valid feature vector records. The effective feature vectors are recorded and arranged in rows to construct a standard feature matrix. The original physical dimensions of each dimension parameter are preserved, resulting in a feature matrix after format verification and invalid sample removal.

[0014] Furthermore, in the above method, based on the virtual training dataset, the trained fusion and integration risk prediction model is obtained, including: Using the virtual training dataset as input, a random forest sub-model containing 25 decision trees is trained. A bagging parallel ensemble strategy is adopted to reduce the model variance through double random sampling of samples and features. Using the virtual training dataset as input, an XGBoost sub-model containing 25 decision trees is trained with a maximum depth of 3 and a learning rate of 0.1. A sequential iteration strategy is adopted to reduce model bias by gradually correcting the residuals. The prediction weights of the random forest sub-model and the XGBoost sub-model are set; based on the prediction weights of the random forest sub-model and the XGBoost sub-model, a linear weighted fusion rule is obtained; based on the prediction results of the random forest sub-model and the XGBoost sub-model and the linear weighted fusion rule, the prediction results output by the random forest sub-model and the XGBoost sub-model are weighted and integrated to obtain the final risk score, thus obtaining the trained fusion risk prediction model, which includes: a random forest sub-model, an XGBoost sub-model, and a linear weighted fusion rule.

[0015] Furthermore, in the above method, based on the feature matrix and the trained fusion-integrated risk prediction model, a standardized real-time risk score is obtained, including: The pure feature matrix is ​​input into the trained fusion risk prediction model. The random forest sub-model and the XGBoost sub-model inside the fusion risk prediction model perform independent inference on the input data and output their respective initial risk prediction values. The linear weighted fusion rule built into the fusion and integration risk prediction model is invoked to perform weighted calculation on the initial predicted values ​​output by the random forest sub-model and the XGBoost sub-model, and the weighted calculation result is obtained. The weighted calculation results are normalized and mapped to a continuous interval of 0–100 to obtain a standardized real-time risk score.

[0016] Furthermore, in the above method, after obtaining the standardized real-time risk score based on the feature matrix, it also includes: Based on preset thresholds, standardized real-time risk scores are mapped to risk levels; risk levels are then matched with corresponding color indicators and operational suggestions.

[0017] According to another aspect of the present invention, a computer-readable storage medium is also provided, having stored thereon computer-executable instructions, wherein when executed by a processor, the computer-executable instructions cause the processor to perform the method described in any of the preceding claims.

[0018] Compared with existing technologies, this invention proposes a method and system for early warning of operational risks in mass spectrometry analysis systems based on multi-dimensional feature fusion. It constructs a 25-dimensional feature space including configuration parameters and dynamic parameters, and utilizes a hybrid ensemble learning model for risk prediction. Compared with existing technologies, this invention has the following significant advantages: First, this application breaks through the limitations of a single dimension, significantly improving the accuracy and foresight of early warnings. The system pre-defines a 25-dimensional feature space, integrating "configuration parameters" describing the baseline state (such as column design life and recovery range) and "dynamic operating parameters" describing the real-time state (such as current pump pressure and retention time drift). This application changes the lag inherent in traditional mass spectrometry quality control, which relies solely on a single indicator (such as checking whether the quality control sample is within the 2SD range). The model can capture the nonlinear coupling relationship between parameters (e.g., identifying the combined risk of the column life approaching its limit and a slight increase in initial pump pressure), thereby issuing early warnings before a failure occurs (i.e., during the "latency period"), achieving a shift from "post-event correction" to "pre-event prevention."

[0019] Secondly, this application addresses the cold start challenge by enabling model initialization with zero samples. It employs random sampling simulation technology based on multiple probability distributions (normal, uniform, and exponential distributions) to simulate the physical laws of LC-MS and generate training data. This effectively overcomes the industry pain point of lacking real fault samples (labeled data) in the early stages of machine learning model application. By using high-fidelity simulations of different physical parameter characteristics (e.g., using a normal distribution for the mechanical pulsation of pump pressure and an exponential distribution for the asymmetric decay of R²), the system can possess a warning model with good generalization ability even before accumulating a large amount of real clinical data, thus lowering the deployment threshold.

[0020] Furthermore, this application balances model stability and sensitivity, enhancing its robustness to noise. It constructs a hybrid ensemble model integrating Random Forest (Bagging strategy) and XGBoost (Boosting strategy), employing linear weighted fusion. This application combines the low variance and strong noise resistance of Random Forest (suitable for handling high-dimensional redundant features) with the low bias and strong fitting ability of XGBoost (suitable for capturing subtle residuals). This hybrid architecture prevents overfitting caused by individual outliers (suppressed by Random Forest) while also sensitively detecting minor performance drifts (captured by XGBoost), thus outputting a more robust risk score than a single model.

[0021] Finally, this application presents a clear and intuitive risk grading system with strong guidance. Specifically, it maps a continuous risk score (0-100) to four warning levels (low, medium, high, and critical), and provides specific operational suggestions for each level (such as monitoring drift and pausing sample injection). This application reduces the user's workload by transforming complex algorithm outputs into decision-making instructions that clinical operators can directly understand and execute, effectively guiding laboratory personnel to rationally plan maintenance schedules and avoid sample waste or testing delays caused by sudden instrument malfunctions. Attached Figure Description

[0022] Figure 1 This is a flowchart of a machine learning-based clinical mass spectrometry risk prediction method according to an embodiment of the present invention. Detailed Implementation

[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] To ensure the reliability of clinical test results, the industry currently primarily follows the CLSI (Clinical Laboratory Standards Institute) C62-A guideline for quality management and monitoring. The current mainstream technical situation is as follows: In the first scenario, quality management and monitoring are based on the System Suitability Test (SST) according to the CLSI C62 guideline. According to the CLSIC 62-A guideline, a System Suitability Test must be performed before clinical sample analysis. In practice, laboratories typically use one or more injections of a standard at a specific concentration to monitor its retention time, peak area, peak tailing factor, and signal-to-noise ratio. Then, a "pass / fail" binary judgment mode is used for evaluation. For example, the retention time drift may not exceed 2.5% of the baseline value, or the tailing factor may be <2.0. Only after the SST is passed is subsequent clinical sample cohort analysis permitted.

[0025] In the second scenario, quality management and monitoring are conducted based on the "bracketing" model of quality control materials. This is currently the most stringent process control method implemented in clinical laboratories.

[0026] In practice, according to the C62-A guidelines, laboratories insert quality control samples of low, medium, and high concentrations at the beginning and end of the analysis batch, as well as at intervals of a certain number (e.g., 20-50) of patient samples in between.

[0027] This method employs the classic Westgard multirule control theory as its algorithmic principle. The laboratory constructs Levey-Jennings control charts and calculates the bias and coefficient of variation of the quality control samples. When the quality control results violate rules such as 13s (single point exceeding 3SD) or 22s (two consecutive points exceeding 2SD), the batch is deemed "out of control," and all clinical reports for that batch are blocked.

[0028] In the third scenario, quality management and monitoring are conducted based on matrix effects and recovery monitoring using isotope internal standards. To address complex matrix interferences in biological samples (such as ion inhibition or enhancement caused by phospholipids), LC-MS methods commonly employ stable isotope-labeled internal standards for correction.

[0029] In practice, the absolute response value (peak area) of the internal standard in each clinical sample is monitored.

[0030] Then, a broad, fixed threshold range is set as the judgment criterion. For example, CLSI and FDA guidelines typically recommend that the recovery rate of the internal standard should be within 50% to 150% of the mean of the previous batch or method validation. If it exceeds this range, it may indicate a serious matrix effect or injection failure.

[0031] In the fourth scenario, quality management and monitoring are carried out through preventative maintenance records of the instrument hardware.

[0032] Currently, hardware maintenance relies primarily on manual record-keeping. In practice, this includes recording the cumulative number of injections into the chromatographic column, the cleaning date of the ion source, and the replacement cycle of the vacuum pump oil. This type of data is usually static, serving only as a reference for periodically replacing consumables, and rarely directly participates in the real-time quality algorithms for daily batches.

[0033] The existing technologies based on the CLSI C62 guidelines and statistical process control constitute the cornerstone of current clinical mass spectrometry quality management. However, in the face of the high complexity and dynamic changes of the LC-MS system, the existing technologies still have three major defects: lag, one-sidedness, and rigidity.

[0034] The following shortcomings exist in existing technologies for quality control and monitoring of clinical test results.

[0035] First, existing Westgard rules or SST (Self-Testing and Testing) are essentially "post-mortem" quality control. The system only alarms when instrument performance has significantly deteriorated (e.g., column collapse or severe ion source contamination), leading to substantial deviations from quality control results or SST failure. At this point, it often means that the entire batch of valuable clinical samples needs to be retested, resulting in a huge waste of reagents and consumables and delays in reporting cycles.

[0036] This invention aims to capture subtle, cumulative trends in data using machine learning algorithms, issuing early warnings of risks before the system experiences substantial "loss of control," thus shifting from corrective to predictive maintenance. Therefore, this application solves the problem of delayed alerts in existing technologies, achieving proactive risk prediction. It also enables deep coupling analysis of multi-dimensional features.

[0037] Secondly, existing technologies view each parameter in isolation. For example, the C62 guideline only considers whether the retention time is acceptable or whether the internal standard recovery rate meets the standard. However, in actual operation, risks often arise from the combined effect of multiple seemingly acceptable parameters. For instance, a "column design life" (configuration parameter) that is still within its validity period, combined with a slightly elevated "matrix effect coefficient" (dynamic parameter) and a "tailing factor" that is on the verge of expiration, may collectively indicate an impending sensitivity avalanche. Existing linear models cannot handle this complex nonlinear coupling relationship involving 25 characteristic dimensions (16 configuration parameters + 9 dynamic parameters).

[0038] Finally, existing technologies rely on static, fixed thresholds (e.g., internal standard range 50%-150%, retention time ±0.1 min). However, LC-MS is a dynamically evolving system. For example, as the column is used more frequently, the retention time will naturally shift forward, and the background pressure will gradually increase; this is normal physical aging, not a malfunction. Using rigid, fixed thresholds can easily generate a large number of false positive alarms during normal system fluctuations, leading to alarm fatigue among researchers and causing them to overlook real risks.

[0039] This invention utilizes the adaptive capabilities of machine learning to dynamically adjust the risk discrimination logic based on the instrument's current configuration status (such as the cumulative lifespan of the chromatographic column and maintenance history), accurately distinguishing between normal system drift and abnormal fault signals, significantly reducing false alarm rates and improving monitoring accuracy. Therefore, this application also addresses the rigidity of fixed thresholds in existing technologies, enhancing the adaptability of detection.

[0040] Example 1 like Figure 1 As shown, this embodiment provides a risk prediction method for clinical mass spectrometry based on machine learning, including: Step 1: Construct and load the system's preset multidimensional feature space: Based on the N+M dimensional feature definition built into the LC-MS system (liquid chromatography-tandem mass spectrometry system), obtain the N+M dimensional feature vector structure; Step 11: Obtain the N+M dimensional (e.g., 25 dimensional) feature definitions built into the LC-MS system, such as: N configuration parameters (e.g., 16) + M dynamic parameters (e.g., 9). Step 12: Based on the N+M dimensional feature definition built into the LC-MS system, define N static configuration parameters, such as column design life, recovery range, matrix effect coefficient, retention time baseline value, etc. Each static configuration parameter corresponds to a fixed system health baseline, and N health baselines are obtained respectively. Step 13: Based on the N+M dimensional feature definition built into the LC-MS system, define M dynamic operating parameters, such as the current pump pressure, retention time drift, cumulative number of injections, and number of days from ion source cleaning. Each dynamic operating parameter is used to capture the changes in the operating status of the LC-MS system in real time, and serves as M real-time state variables. Step 14: Based on N health baselines and M real-time state variables, initialize an N+M dimensional feature vector structure (V), such as a 25-dimensional feature vector structure, and embed it in the software logic. This 25-dimensional feature vector structure (V) will subsequently serve as a unified data structure template for data filling and model input in the following step 2.

[0041] In step 1, to address the issues of single feature dimensions and lack of correlation analysis in existing technologies, this invention pre-defines a standardized 25-dimensional feature space within the system. This feature space is defined by the developers based on the methodological principles of clinical mass spectrometry and is embedded in the software logic, eliminating the need for manual definition by the user. The system initializes a 25-dimensional feature vector, which specifically includes configuration parameters and dynamic operating parameters, the roles of which in the model are described below.

[0042] There are 16 configuration parameter dimensions. These parameters describe the baseline state and hardware design limits of the assay, serving as a static background for the model to assess risk. Specifically, these include: overall precision, recovery range, matrix effect coefficient, baseline retention time, allowable retention time window, column design life, tailing factor limit, initial pump pressure, signal-to-noise ratio threshold, upper limit of linear range, correlation coefficient threshold, internal standard stability coefficient, pretreatment complexity (quantification level), isomer interference, internal standard type, and clinical risk level (1-4). These parameters provide the model with a healthy baseline for each assay. For example, column design life combined with initial pump pressure helps the model understand whether the current pressure increase is due to normal aging or abnormal blockage.

[0043] There are nine dynamic operating parameters, which describe the real-time status of the current batch and serve as dynamic variables for the model to assess risk. These include: distance from source cleaning in days, cumulative injection count, current pump pressure, retention time drift (the absolute value of the difference between the measured and baseline values), actual tailing factor, current signal-to-noise ratio, intra-batch precision, system bias, and the correlation coefficient of the current standard curve. These parameters reflect the real-time trends of the system. The model predicts potential failure risks by learning the coupling relationship between these dynamic parameters and configuration parameters (e.g., the rate of change of system bias as distance from source cleaning increases).

[0044] Step 2, Simulation data acquisition based on multi-probability distribution random sampling: In view of the scarcity of real sample data in the early stage of model training and the fact that the system is not yet connected to the physical interface, Step 2 adopts a data acquisition technique that combines multi-probability distribution random sampling simulation with offline standardized import.

[0045] Step 21, Training Branch: Based on the N+M dimensional feature vector structure, obtain the virtual training dataset; Based on the N+M dimensional (preferably 25 dimensional) feature vector structure (V) output from step 1, random sampling is performed using normal, uniform, or exponential distributions for different parameters in the feature vector structure (V) based on their physical characteristics (such as parameters with continuous fluctuations, parameters with uncertain ranges, and parameters with asymmetric decay). Simultaneously, preset LC-MS system fault modes (such as column blockage, ion source contamination, and abnormal matrix effects) are embedded during the sampling process. Finally, a high-fidelity virtual dataset containing normal operating conditions and typical fault modes is generated. This dataset can accurately simulate the real operating fluctuations and fault evolution of the LC-MS system and serves as a virtual training dataset that conforms to the physical laws of LC-MS for model training.

[0046] In fields such as mass spectrometry, chromatography, instrument condition monitoring, fault diagnosis, and risk prediction, using simulation and modeling data is standard practice. For precision instruments like liquid chromatography-tandem mass spectrometry, which operate normally most of the time, data on actual malfunctions, high risks, and anomalies is extremely rare, difficult to collect, or even nonexistent. Therefore, without real negative samples, models can only be trained and validated using simulated abnormal data. Since real high-risk, fault samples are difficult to obtain and extremely costly to acquire, this invention uses virtual simulation data to complete model training and preliminary validation. The model structure and fusion strategy are universal and can be directly transferred to real-world data scenarios.

[0047] Specifically, for example, parameters with continuous fluctuations can be sampled using a normal distribution; parameters with uncertain ranges can be sampled using a uniform distribution; and parameters with asymmetric decay can be sampled using an exponential distribution. Here, the simulation generation of training data (a technical means in the model training phase) is described: To complete the initial training of the AI ​​model, the system has a built-in simulation data generator based on stochastic processes. This simulation data generator no longer relies on a single distribution assumption, but instead constructs a reproducible random number generation engine by initializing a determined random number seed, and performs multi-dimensional mixed sampling simulation using normal, uniform, and exponential distributions according to the characteristics of different physical parameters of LC-MS.

[0048] The specific implementation logic includes: 1) Simulation of natural fluctuations and system noise based on normal distribution: For continuous variables affected by environmental thermodynamic fluctuations or mechanical precision, the system uses a Gaussian distribution model for simulation.

[0049] Pump pressure simulation: Simulate the mechanical pulsation of a liquid phase pump, generate random noise that follows a normal distribution with a mean of 0 and a specific standard deviation (e.g., 5 Bar), and superimpose it onto the reference pressure; Retention time simulation: Simulates retention time drift caused by small disturbances in column temperature or flow rate, generating normally distributed data with a mean of 0 and a small standard deviation (e.g., 0.005 min); Matrix effect simulation: Simulates background interference from different patient samples to generate normally distributed data that fluctuates around a baseline of 1.0 and whose standard deviation is controlled by configuration parameters.

[0050] 2) Uncertainty in the simulation range based on uniform distribution: For parameters where only the upper and lower thresholds are known but the internal probability density is unknown, the system uses a uniform distribution model for simulation.

[0051] Recovery rate simulation: Based on the preset methodology validation range, uniform random sampling is performed within a specific closed interval to simulate the unbiased fluctuation of the sample recovery rate.

[0052] 3) Simulating asymmetric performance degradation based on exponential distribution: For indicators that exhibit long-tail effects or asymmetric bias characteristics, the system employs an exponential distribution model for simulation.

[0053] Linear correlation coefficient (R2) simulation: Considering that the standard curve R2 value is usually close to 1.0 and the closer it is to 1, the more difficult it is to achieve, the system uses an exponential distribution to generate small deviation values, thereby simulating the real performance degradation characteristics that are good in most cases but occasionally have large deviations.

[0054] By employing the aforementioned multi-probability distribution hybrid sampling technique, this invention can construct a high-fidelity virtual dataset that conforms to the physical laws and statistical characteristics of LC-MS without the need for real physical sampling, effectively solving the technical problem of lacking labeled fault samples in the cold start phase of machine learning models.

[0055] Step 22, Application Branch: Based on the N+M dimensional feature vector structure, obtain the parameter values ​​of each dimension of the current batch through the data acquisition interface, and assemble them into N+M dimensional feature vector data; The current batch of detection data files in .csv / .xlsx format exported from the mass spectrometry workstation is used as the import file; the file parsing interface is called to read the parameter data in the import file, i.e., the values ​​corresponding to each parameter field. The parameter data includes: the measured values ​​of static configuration parameters and the real-time values ​​of dynamic operating parameters; a preset field mapping algorithm is used to match the parameter fields (i.e., column names) in the import file with the N+M feature dimensions (such as current pump pressure, cumulative number of injections) of the feature vector structure (V) in step 1 one by one. After a successful match, the parameter data of the corresponding parameter field is filled into the corresponding dimension of the N+M feature vector structure (V) to obtain the mapped and filled N+M dimensional (25-dimensional) feature vector data; Here, the offline import of the test data is performed: the system is equipped with a standardized file parsing interface. Users can import the result file exported from the mass spectrometry workstation or the table file recording the running parameters (supporting .csv and .xlsx formats) after organizing it into the system's preset template format. Specifically, the system reads the imported file, uses a field mapping algorithm to map the column data in the file and fill it into the 25-dimensional feature vector described in step 1.

[0056] Step 3, Data Cleaning and Feature Input: Perform format verification and invalid value filtering on the N+M dimensional feature vector data to obtain the feature matrix; Step 31, perform format parsing and integrity verification: verify line by line whether the number of dimensions of each feature vector record in the N+M dimensional feature vector data after the mapping and filling of the output of the application branch in Step 2 is completely consistent with the preset N+M dimensions, and filter out the feature vector records with matching number of dimensions as the feature vector records that pass the verification. Here, the system parses the input data line by line, verifies whether the feature dimensions of each record are consistent, and checks whether all feature values ​​are valid values.

[0057] Step 32, perform invalid sample removal: delete records containing null values ​​(NaN), missing values ​​or invalid outliers from the verified feature vector records to ensure the purity of the input data and obtain valid feature vector records; Here, for sample records containing non-numeric (NaN) or missing data, the system adopts a row-wide elimination strategy, retaining only complete and valid samples to construct the training set, ensuring the purity of the model input.

[0058] Step 33, perform feature matrix reconstruction: record the effective feature vectors, arrange them in rows to construct a standard feature matrix, retain the original physical dimensions of each dimension parameter, do not perform normalization or standardization processing, and obtain the feature matrix after format verification and invalid sample removal.

[0059] Here, the validated data is not normalized or standardized, but its original physical dimensions (such as pressure value, retention time, etc.) are directly retained and used as the feature matrix to be input into the fusion and integration risk prediction model for training.

[0060] Step 4, construct an integrated risk prediction model based on bagging and lifting: based on the virtual training dataset and the trained integrated risk prediction model, obtain the trained integrated risk prediction model. Step 41, Train the Random Forest (RF) sub-model: Using the virtual training dataset output from the training branch in Step 2 as input, train a random forest sub-model containing 25 decision trees. The bagging parallel ensemble strategy is adopted to reduce the model variance and improve stability through double random sampling of samples and features. Step 42, Train the Gradient Boosting (XGBoost) sub-model: Using the same virtual training dataset as input, train an XGBoost sub-model containing 25 decision trees, set the maximum depth to 3 and the learning rate to 0.1, and adopt the boosting serial iteration strategy to reduce model bias and improve fitting accuracy by gradually correcting the residuals; Step 43, Sub-model fusion: Preset linear weighted fusion rules, set the prediction weight of the random forest sub-model to 0.6 and the prediction weight of the XGBoost sub-model to 0.4 to define the fused linear weighted fusion rules; based on the prediction results of the two sub-models and the linear weighted fusion rules, the prediction results output by the random forest sub-model and the XGBoost sub-model are weighted and integrated to obtain the final risk score, thus obtaining the trained fused risk prediction model. The trained fused risk prediction model includes the random forest sub-model, the XGBoost sub-model, and the linear weighted fusion rules.

[0061] Here, the optimal weight coefficients can be determined based on the training performance of the two sub-models (such as validation set accuracy, F1 score, or mean squared error), namely the prediction weights of the random forest sub-model and the prediction weights of the XGBoost sub-model.

[0062] To balance model stability (low variance) and accuracy (low bias), this invention constructs a hybrid ensemble model that combines Random Forest and XGBoost. By employing dual random sampling of samples and features, variance is reduced, robustness to 25-dimensional high-dimensional features is enhanced, and overfitting is suppressed. The model's structure and parameters are shown in Table 1 below.

[0063] Table 1

[0064] Model training method: The simulated dataset generated in step S2 is used as the training set. The classification task is treated as a regression task, and the loss function is defined as mean squared error (MSE). ; Where y_true represents the true risk label value of each sample (in the virtual training dataset, it is a continuous value of 0~100 based on the preset fault mode label, for example, normal samples are labeled 0 and critical samples are labeled 100). y_pred: Represents the model's risk prediction for this sample (output of the sub-model or fusion model); (y_true−y_pred): Represents the prediction residual for a single sample; ∑: represents the summation of the squared residuals over all samples.

[0065] The XGBoost model progressively optimizes the residuals using gradient descent, while the random forest model splits and grows by minimizing the mean squared error of nodes. Finally, the ensemble model outputs a comprehensive risk score. The model fusion scheme obtains the final risk prediction value through a linear weighted fusion rule. It is obtained by linearly weighted fusion of the outputs of the two sub-models: , in, This represents the output of the random forest submodel; This represents the output of the XGBoost sub-model; = 0.6, = 0.4 is the corresponding fusion weight (the default value in this embodiment).

[0066] Specifically, 1) During the sub-model training phase: When training the Random Forest sub-model and the XGBoost sub-model separately, the risk prediction values ​​(intermediate outputs) of these two sub-models are used as the loss function formula. y_pred in the text.

[0067] The loss function at this time It is used to optimize the two sub-models separately: to make the risk prediction values ​​of Random Forest and XGBoost as close as possible to the true label y_true (the continuous risk prediction value from 0 to 100).

[0068] 2) After the two sub-models are trained, the final risk prediction value output by the fusion model is obtained through a linear weighted fusion rule (0.6 × Random Forest risk prediction value + 0.4 × XGBoost risk prediction value). This final score is then used as the loss function formula. The final risk prediction value y_pred can be used to calculate the risk using the formula... The overall loss is calculated again to evaluate the final performance of the entire hybrid ensemble model and verify whether the fusion effect is better than that of a single sub-model.

[0069] Step 5, Real-time risk score calculation: Based on the pure feature matrix, a standardized real-time risk score is obtained; Step 51: Input the clean feature matrix output in Step 3 into the fusion and integration risk prediction model obtained in Step 4. The random forest sub-model and the XGBoost sub-model inside the fusion and integration risk prediction model perform independent reasoning on the input data and output their respective initial risk prediction values. Step 52: Call the built-in linear weighted fusion rule (random forest weight 0.6, XGBoost weight 0.4) of the fusion and integration risk prediction model to perform weighted calculation on the initial predicted values ​​output by the random forest sub-model and the XGBoost sub-model to obtain the weighted calculation result; Step 53: Normalize the weighted calculation results and map them to a continuous interval of 0–100 to obtain a standardized real-time risk score.

[0070] Here, in the application phase, the 25-dimensional feature vectors of the current batch imported in step 2 are input into the hybrid ensemble model trained in step 4. After internal tree group calculation and weighted fusion, the model outputs a continuous scalar value between 0 and 100, which is the risk score.

[0071] Step 6: Generate tiered early warning information: Based on the preset threshold, the standardized real-time risk score output in step 5 is mapped to four risk levels, such as low, medium, high, and critical. The four risk levels are matched with corresponding color labels (green, yellow, orange, and red) and operational suggestions to obtain four-level early warning results, corresponding labels, and operational suggestions.

[0072] Based on the calculated risk score, as shown in Table 2 below, the system determines the risk level through its built-in decision-making logic.

[0073] Table 2

[0074] In the following example, a glycated hemoglobin (HbA1c) test conducted in a reference measurement laboratory of a clinical testing center was selected as the application object. The detection platform was an AB Sciex 5500 QTRAP triple quadrupole mass spectrometer system, and the detection date was May 27, 2025. This batch contained 25 samples, including 17 samples, 6 calibrators, and 2 quality control samples. The equipment status parameters, chromatographic behavior parameters, and some characteristic parameters generated during the batch's operation were used to perform risk prediction analysis using the system of this invention.

[0075] 1. Feature Data Acquisition and Assembly Based on the 25-dimensional feature vector structure defined in step 1 of this invention, the static configuration parameters and dynamic operating parameters corresponding to this batch are extracted and assembled. The static configuration parameters can be derived from methodology validation documents, instrument maintenance records, and project basic configuration files; the dynamic operating parameters are derived from the instrument workstation exported files, quality control records, and standard curve results for this batch.

[0076] In this embodiment, some of the extracted parameters are shown in Table 3.

[0077] Table 3. Some characteristic parameters of this actual batch

[0078] The above parameters are filled into a 25-dimensional feature vector structure to form the feature vector data corresponding to the current batch.

[0079] 2. Data Validation and Feature Matrix Construction In this case, the original number of records was 100. After format validation, 100 records were retained. After filtering for invalid values, 96 valid records were finally retained, and the corresponding feature matrix was constructed.

[0080] 3. Risk Prediction Model Reasoning The aforementioned feature matrix is ​​input into the trained fusion risk prediction model. This model includes a random forest sub-model and an XGBoost sub-model, where the random forest sub-model outputs a risk prediction value of RF = 52, and the XGBoost sub-model outputs a risk prediction value of XGB = 25.

[0081] According to the preset linear weighted fusion rule of the present invention: Final risk prediction = [0.6] × Random Forest sub-model output + [0.4] × XGBoost sub-model output; Therefore, the weighted calculation result in this embodiment is: Final risk forecast = 41; Further mapping the result to a standardized range of 0 to 100 yielded a real-time risk score of 41 for this batch.

[0082] 4. Risk Level Assessment and Early Warning Output According to the risk classification rules preset in this invention: When the risk score is less than 30, it is considered low risk; When the risk score is 30 ≤ risk score < 60, it is considered a medium risk. When the risk score is 60 ≤ risk score < 80, it is considered high risk; When the risk score is ≥ 80, it is judged as a critical risk.

[0083] In this embodiment, the real-time risk score for this batch is 41 points, therefore the system determines that this batch belongs to medium risk. Simultaneously, a corresponding yellow indicator is output, along with the following operational suggestion: Given that the column has accumulated 500 injections and it has been 60 days since the last ion source cleaning, it is recommended to perform preventative maintenance as planned.

[0084] 5. Explanation of the correspondence between the results and actual batch operation results During subsequent actual operation of this batch, researchers observed that the retention time drift continued to increase, the column required more frequent cleaning, and the system needed to be paused for maintenance. After rinsing the column and cleaning the ion source, the correlation coefficient of the standard curve increased. This actual operational performance is consistent with the medium-risk warning result output by this invention, indicating that this invention can effectively predict the potential risks of clinical mass spectrometry systems under real batch data conditions and provide actionable warning information for laboratory personnel.

[0085] 6. Implementation Results Description As can be seen from this embodiment, the present invention can complete steps such as multi-dimensional feature extraction, format verification, feature matrix construction, model reasoning and risk classification output based on real batch detection data. It can transform the complex clinical mass spectrometry system status into a standardized risk score and express it intuitively in the form of risk level and operation suggestions, thereby providing technical support for predictive maintenance and quality risk prevention and control in clinical laboratories.

[0086] This invention provides a method and device for early warning of operational risks in clinical mass spectrometry systems based on multi-dimensional feature fusion. It can be implemented using computer software technology and machine learning algorithms and runs on a computer with computing capabilities. Through the above steps, this invention utilizes a machine learning model to uncover the complex nonlinear relationship between configuration parameters and dynamic parameters, thereby achieving more accurate and forward-looking risk warnings than traditional single-threshold methods.

[0087] According to another aspect of the present invention, a computer-readable storage medium is also provided, having stored thereon computer-executable instructions, wherein when executed by a processor, the computer-executable instructions cause the processor to perform the method described in any of the preceding claims.

[0088] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0089] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0090] Obviously, those skilled in the art can make various modifications and variations to the invention without departing from the spirit and scope of the invention. Therefore, if these modifications and variations fall within the scope of the claims of the invention and their equivalents, the invention is also intended to include these modifications and variations.

Claims

1. A risk prediction method for clinical mass spectrometry based on machine learning, characterized in that, include: Based on the unique N+M dimensional feature definition of the liquid chromatography-tandem mass spectrometry system, an N+M dimensional feature vector structure is obtained; A virtual training dataset is obtained based on an N+M dimensional feature vector structure; Based on the N+M dimensional feature vector structure, the parameter values ​​of each dimension of the current batch are obtained through the data acquisition interface and assembled into N+M dimensional feature vector data; the N+M dimensional feature vector data is then subjected to format verification and invalid value filtering to obtain the feature matrix. Based on the virtual training dataset, a fully trained fusion and integration risk prediction model is obtained; Based on the feature matrix and the trained fusion-integrated risk prediction model, a standardized real-time risk score is obtained.

2. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 1, characterized in that, Based on the unique N+M dimensional feature definition of liquid chromatography-tandem mass spectrometry systems, an N+M dimensional feature vector structure is obtained, including: Based on the unique N+M dimensional feature definition of the liquid chromatography-tandem mass spectrometry system, N static configuration parameters are defined, each static configuration parameter corresponds to a fixed system health baseline, and N health baselines are obtained respectively. Based on the unique N+M dimensional feature definition of the liquid chromatography-tandem mass spectrometry system, M dynamic operating parameters are defined. Each dynamic operating parameter is used to capture the changes in the operating state of the liquid chromatography-tandem mass spectrometry system in real time, and is respectively used as M real-time state variables. Based on N health baselines and M real-time state variables, an N+M dimensional feature vector structure is initialized.

3. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 2, characterized in that, N static configuration parameters, including: Overall precision, recovery range, matrix effect coefficient, retention time baseline, retention time allowable window, column design life, tailing factor limit, initial pump pressure, signal-to-noise ratio threshold, upper limit of linear range, correlation coefficient threshold, internal standard stability coefficient, pretreatment complexity, isomer interference, internal standard type, and clinical risk level. M real-time state variables, including: number of days from source cleaning, cumulative number of injections, current pump pressure, retention time drift, actual tailing factor, current signal-to-noise ratio, intra-batch precision, system bias, and current standard curve correlation coefficient.

4. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 1, characterized in that, Based on the N+M dimensional feature vector structure, a virtual training dataset is obtained, including: For the physical characteristics of different parameters in the N+M dimensional feature vector structure, random sampling is performed using normal distribution, uniform distribution, or exponential distribution, respectively. At the same time, preset liquid chromatography-tandem mass spectrometry system fault modes are embedded in the sampling process, and finally a high-fidelity virtual dataset containing normal operation status and typical fault modes is generated, which serves as a virtual training dataset that conforms to the physical laws of liquid chromatography-tandem mass spectrometry.

5. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 2, characterized in that, Based on an N+M dimensional feature vector structure, the parameter values ​​of each dimension of the current batch are obtained through a data acquisition interface and assembled into N+M dimensional feature vector data, including: Use the detection data file of the current batch exported from the mass spectrometry workstation as the import file; Call the file parsing interface to read the parameter data in the imported file, i.e., the values ​​corresponding to each parameter field. The parameter data includes: the measured values ​​of static configuration parameters and the real-time values ​​of dynamic running parameters. Using a preset field mapping algorithm, the parameter fields in the imported file are matched one by one with the N+M feature dimensions of the feature vector structure. After a successful match, the parameter data of the corresponding parameter field is filled into the corresponding dimension of the N+M feature vector structure to obtain the mapped and filled N+M feature vector data.

6. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 1, characterized in that, The N+M dimensional feature vector data is subjected to format validation and invalid value filtering to obtain a feature matrix, including: The system verifies row by row whether the number of dimensions of each feature vector record in the N+M dimensional feature vector data is consistent with the preset N+M dimensions, and selects the feature vector records with matching number of dimensions as the feature vector records that pass the verification. From the verified feature vector records, delete records containing null values, missing values, or invalid outliers to ensure the purity of the input data and obtain valid feature vector records. The effective feature vectors are recorded and arranged in rows to construct a standard feature matrix. The original physical dimensions of each dimension parameter are preserved, resulting in a feature matrix after format verification and invalid sample removal.

7. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 1, characterized in that, Based on the virtual training dataset, a trained fusion-integrated risk prediction model is obtained, including: Using the virtual training dataset as input, a random forest sub-model containing 25 decision trees is trained. A bagging parallel ensemble strategy is adopted to reduce the model variance through double random sampling of samples and features. Using the virtual training dataset as input, an XGBoost sub-model containing 25 decision trees is trained with a maximum depth of 3 and a learning rate of 0.

1. A sequential iteration strategy is adopted to reduce model bias by gradually correcting the residuals. The prediction weights of the random forest sub-model and the XGBoost sub-model are set. Based on the prediction weights of the random forest sub-model and the XGBoost sub-model, a linear weighted fusion rule is obtained. Based on the prediction results of the random forest sub-model and the XGBoost sub-model and the linear weighted fusion rule, the prediction results output by the random forest sub-model and the XGBoost sub-model are weighted and integrated to obtain the final risk score, thus obtaining the trained fusion integrated risk prediction model. The trained fusion integrated risk prediction model includes: a random forest sub-model, an XGBoost sub-model, and a linear weighted fusion rule.

8. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 7, characterized in that, Based on the aforementioned feature matrix and the trained fusion-integrated risk prediction model, a standardized real-time risk score is obtained, including: The feature matrix is ​​input into the trained fusion risk prediction model. The random forest sub-model and the XGBoost sub-model within the fusion risk prediction model perform independent reasoning on the input data and output their respective initial risk prediction values. The linear weighted fusion rule built into the fusion and integration risk prediction model is invoked to perform weighted calculation on the initial predicted values ​​output by the random forest sub-model and the XGBoost sub-model, and the weighted calculation result is obtained. The weighted calculation results are normalized and mapped to a continuous interval of 0–100 to obtain a standardized real-time risk score.

9. The risk prediction method for clinical mass spectrometry based on machine learning as described in claim 1, characterized in that, After obtaining the standardized real-time risk score based on the feature matrix, the following steps are also included: Based on preset thresholds, standardized real-time risk scores are mapped to risk levels; risk levels are then matched with corresponding color indicators and operational suggestions.

10. A computer-readable storage medium having stored thereon computer-executable instructions, wherein, When the computer-executable instructions are executed by the processor, the processor causes the processor to perform the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Deep learning-based medicine compliance risk early warning and decision suggestion method and system

    CN119940900A

  • Metabolite target interaction prediction method and system for myocardial injury

    CN121075411A

  • Depression risk assessment method and system based on intestinal flora characteristics

    CN121460180A

  • Clinical condition deterioration risk prediction and early warning system and method based on machine learning

    CN121705706A

  • Method for Predicting Benchmark Value of Unit Equipment Based on XGBoost Algorithm and System thereof

    US20230213895A1