Respirator machine-on time prediction method and system based on machine learning, electronic equipment and storage medium
The MIMIC-III data is preprocessed and model training based on machine learning, which solves the problems of low accuracy and poor generalization of time prediction of ventilator in the prior art, and accurately predicts the use time of ventilator in patients, providing reliable support for clinical practice.
Patent Information
- Application Number
- CN202510195051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing ventilator time prediction technology has problems with low accuracy and poor generalization, which is difficult to meet the needs of clinical precision.
Using a machine learning-based approach, we use preprocessing of MIMIC-III data, build a dedicated data set, train and optimize multiple machine learning regression models, and select the best-performing model for ventilator time prediction.
Accurate prediction of the patient's ventilator time is achieved, which not only predicts whether the ventilator needs to be used, but also accurately predicts the usage time, providing reliable decision-making support for clinicians.
Smart Images

Figure CN120126804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ventilator intubation time prediction, and particularly relates to a method, system, electronic device and storage medium for ventilator intubation time prediction based on machine learning. Background Art
[0002] Respiratory failure is a serious clinical syndrome caused by various factors that lead to disorders in pulmonary ventilation and gas exchange functions, resulting in the body's inability to effectively perform gas exchange. If this condition is not treated promptly, it will rapidly progress to cause multiple organ failure throughout the body and even endanger life. As a key life support device in the treatment of respiratory failure, the ventilator plays an indispensable role. Through mechanical ventilation, it can provide necessary oxygen supply for patients, expel carbon dioxide, maintain gas exchange and acid-base balance in the body, and gain time for treating the primary disease. However, different patients have different demands and responses to the ventilator, and the ventilator intubation time varies according to individual conditions. Therefore, how to scientifically and reasonably predict the ventilator intubation time of different patients is of great significance.
[0003] Currently, the methods for predicting the ventilator intubation time mainly include two categories: traditional statistical analysis and basic machine learning models. Among them, the traditional statistical analysis methods include: 1) Based on patients' physiological indicators and disease scores, doctors make judgments and estimations according to their own experience; 2) Based on the linear regression model, a linear relationship is established using a small number of features such as patients' age and basic vital signs for ventilator intubation time prediction. The basic machine learning models include: using early basic machine learning models such as decision trees or shallow neural networks for ventilator intubation time prediction.
[0004] However, the method of judging and estimating the ventilator intubation time based on doctors' own experience is greatly affected by subjective factors. The linear regression model-based methods can only establish a linear relationship between a small number of features of patients and the ventilator intubation time, and it is difficult to capture complex non-linear associations. For the early basic machine learning models, their feature selection and model capacity are limited, and they are prone to ignoring the deep interactions of multi-modal data. It can be seen that due to problems such as limited data utilization, insufficient model complexity, and lack of time series analysis ability in the existing technologies, the prediction accuracy of the ventilator intubation time is low and the generalization ability is poor, making it difficult to meet the clinical precision requirements. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the existing technologies, the present invention provides a method, system, electronic device and storage medium for ventilator intubation time prediction based on machine learning, which solves at least the problems of low accuracy and poor generalization ability existing in the existing ventilator intubation time prediction technologies.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the present invention is implemented through the following technical solutions:
[0009] In a first aspect, the present application first proposes a method for predicting the ventilator intubation time based on machine learning, and the method includes:
[0010] Preprocessing the pre-acquired MIMIC-III data;
[0011] Constructing a dedicated data set based on the preprocessed MIMIC-III data, where the dedicated data set includes patient vital signs and respiratory function status indicators, and the ventilator intubation time corresponding one-to-one to the patient vital signs and respiratory function status indicators;
[0012] Training and optimizing a variety of preselected machine learning regression models based on the dedicated data set;
[0013] Predicting the ventilator intubation time based on the machine learning regression model with the optimal performance after optimization.
[0014] In one embodiment, the preprocessing includes data cleaning and data standardization, and the data cleaning includes processing of data missing values and outliers.
[0015] Preferably, the data standardization includes Z-score standardization; the processing of data missing values includes k-nearest neighbor imputation and mean imputation.
[0016] In one embodiment, the patient vital signs and respiratory function status indicators include: basic physiological indicators of the patient, severity of illness score, ventilator parameters, and blood gas analysis results.
[0017] In one embodiment, the method further includes: when training and optimizing the machine learning regression model, judging the training and optimization effects by monitoring the loss function of the model and the evaluation indicators on the validation set, and evaluating the generalization ability of the machine learning regression model by the method of cross-validation.
[0018] Preferably, the machine learning regression model includes a linear regression model, a support vector regression model, and a random forest regression model.
[0019] More preferably, the evaluation indicators include the mean square error between the predicted value and the true value and the goodness of fit of the model to the data;
[0020] The calculation formula of the mean square error MSE is:
[0021]
[0022] Among them, n is the number of samples; y i is the true value, is the predicted value;
[0023] The goodness of fit R 2 The calculation formula of is as follows:
[0024]
[0025] Wherein, is the mean value of the true value.
[0026] In a second aspect, the present application further provides a ventilator intubation time prediction system based on machine learning, and the system includes:
[0027] A data preprocessing module, configured to preprocess the pre-acquired MIMIC-III data;
[0028] A dedicated dataset acquisition module, configured to construct a dedicated dataset based on the preprocessed MIMIC-III data, where the dedicated dataset includes patient vital signs and respiratory function status indicators, and the ventilator intubation time corresponding one-to-one to the patient vital signs and respiratory function status indicators;
[0029] A model training module, configured to train and optimize a variety of preselected machine learning regression models based on the dedicated dataset;
[0030] An intubation time prediction module, configured to predict the ventilator intubation time based on the machine learning regression model with the optimal performance after optimization.
[0031] In a third aspect, the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for predicting the ventilator intubation time based on machine learning as described in any one of the above are implemented.
[0032] In a fourth aspect, the present application finally provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the method for predicting the ventilator intubation time based on machine learning as described in any one of the above are implemented.
[0033] (III) Beneficial effects
[0034] The present invention provides a method, a system, an electronic device, and a storage medium for predicting the ventilator intubation time based on machine learning. Compared with the prior art, the following beneficial effects are achieved:
[0035] A technology for predicting the ventilator usage time based on machine learning proposed in this application uses advanced machine learning and deep learning algorithms, combines a large amount of clinical case data, trains a prediction model that can accurately capture the relationship between the patient's condition characteristics and the ventilator usage time, and then uses this prediction model to deeply analyze and mine multi-dimensional information such as the patient's clinical data and physiological parameters to achieve accurate prediction of the patient's ventilator usage time. This technology can not only predict whether a patient needs to use a ventilator, but also accurately predict its usage duration, providing reliable decision-making support for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of a method for predicting the ventilator usage time based on machine learning according to the present invention;
[0038] Figure 2 It is a flowchart of a method for predicting the ventilator usage time based on machine learning in an embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of data extraction in an embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of predicting the ventilator usage time based on a data set and machine learning in an embodiment of the present invention;
[0041] Figure 5 It is a fitting result diagram of twelve models for data in an embodiment of the present invention;
[0042] Figure 6 It is a cross-validation result diagram of three relatively optimal models in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0044] According to the statistical data of the World Health Organization, the number of people dying from respiratory failure globally each year is as high as millions. Moreover, due to factors such as population aging, the increasing incidence of chronic diseases, and environmental pollution, the prevalence of respiratory failure shows an increasing trend year by year. In some regions of Asia, respiratory failure is also one of the important diseases threatening people's health. Especially in the Intensive Care Unit (ICU), patients with respiratory failure account for a relatively large proportion.
[0045] As a key life support device in the treatment of respiratory failure, the ventilator plays an indispensable role. Through mechanical ventilation, it can provide necessary oxygen supply for patients, discharge carbon dioxide, maintain gas exchange and acid-base balance in the body, and gain time for treating the primary disease. However, there are significant differences in the needs and responses of different patients to the ventilator, and the intubation time of the ventilator varies according to individual conditions. Therefore, there is an urgent need to propose a technology for scientifically and reasonably predicting the intubation time of the ventilator for different patients.
[0046] Currently, the prediction methods for the intubation time of the ventilator mainly rely on traditional statistical analysis and basic machine learning models, specifically including:
[0047] 1) Clinical experience judgment: Doctors make empirical estimates based on patients' physiological indicators (such as blood oxygen saturation, heart rate) and disease scores (such as APACHE II, SOFA), but this type of technology is greatly affected by subjective factors; 2) Linear regression model: Establish a linear relationship using a small number of features such as patients' age and basic vital signs, but this type of technology is difficult to capture complex non-linear associations; 3) Early machine learning methods: Such as decision trees or shallow neural networks, with limited feature selection and model capacity, but this type of technology is prone to ignoring the deep interactions of multi-modal data (such as laboratory indicators, ventilator parameters).
[0048] In summary, existing technologies often have low prediction accuracy and poor generalization due to limited data utilization, insufficient model complexity, and lack of time series analysis ability, making it difficult to meet the clinical precision requirements.
[0049] By providing a method, system, electronic device, and storage medium for predicting the intubation time of the ventilator based on machine learning in the embodiments of the present application, at least the problems of low accuracy and poor generalization existing in the existing ventilator intubation time prediction technology are solved, and the problem of scientifically and reasonably predicting the intubation time of the ventilator for different patients is realized.
[0050] To better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the accompanying drawings of the specification and specific implementation manners.
[0051] Example 1:
[0052] In the first aspect, see Figure 1-2, the present invention first proposes a method for predicting the ventilator usage time based on machine learning, and the method includes:
[0053] S1. Preprocess the pre-acquired MIMIC-III data;
[0054] S2. Construct a dedicated data set based on the preprocessed MIMIC-III data, where the dedicated data set includes patient vital signs and respiratory function status indicators, and the ventilator usage time corresponding one-to-one to the patient vital signs and respiratory function status indicators;
[0055] S3. Train and optimize a variety of preselected machine learning regression models based on the dedicated data set;
[0056] S4. Predict the ventilator usage time based on the machine learning regression model with the optimal performance after optimization.
[0057] A method for predicting the ventilator usage time based on machine learning proposed in this embodiment first preprocesses the pre-acquired MIMIC-III data, then constructs a dedicated data set based on the preprocessed MIMIC-III data, then trains and optimizes a variety of preselected machine learning regression models based on the dedicated data set, and finally predicts the ventilator usage time based on the machine learning regression model with the optimal performance after optimization. This technical method can not only predict whether a patient needs to use a ventilator, but also accurately predict its usage duration, providing reliable decision support for clinicians.
[0058] The following combines the attached Figure 1-6 , and the explanations of the specific steps of S1-S4, to detail the implementation process of an embodiment of the present invention.
[0059] S1. Preprocess the pre-acquired MIMIC-III data.
[0060] Obtain the MIMIC-III (Medical Information Mart for Intensive Care, MIMIC-III, Intensive Care Medical Information Market) data set, and clean and preprocess the data in the MIMIC-III data to ensure the quality and usability of the data. Specifically:
[0061] Data cleaning mainly includes handling missing values and outliers. Standard Query Language (SQL) is used to extract patient data into a table with a four-hour time window.
[0062] Data standardization is an important step in data preprocessing. Through standardization, the impact caused by inconsistent data dimensions can be eliminated, making data with different features comparable and improving the training effect and convergence speed of the model. In this example, the standardization method used is Z-score standardization, which can be expressed by the formula:
[0063]
[0064] where \(x\) is the original data, \(u\) is the mean of the data, and \(\sigma\) is the standard deviation of the data.
[0065] For missing values, different processing methods are adopted according to the characteristics and distribution of the data. If less than 30% of the data is missing, k-nearest neighbor (KNN) interpolation with \(k = 3\) is used. For data with 30% to 95% missing, the sample and hold method of the time window is used. The initial value is adopted and used to replace the following values until a new value is reached or the limit is reached. When the initial value is missing, mean imputation is performed. Finally, for data with more than 95% missing, the variable is deleted from the state space.
[0066] This embodiment utilizes advanced data preprocessing and missing value processing strategies, which can ensure the high quality and usability of the data, and thus improve the accuracy of predicting the ventilator use time.
[0067] S2. Construct a dedicated data set based on the preprocessed MIMIC-III data. The dedicated data set includes vital signs and respiratory function status indicators, as well as the corresponding ventilator use time.
[0068] Existing linear regression models only establish linear relationships using a small number of features such as patient age and basic vital signs, and it is difficult to capture complex non-linear associations. To solve such problems, the prediction model proposed in this embodiment has the ability to integrate multi-dimensional and high-granularity data. Based on this, a dedicated data set with multi-dimensional and high-granularity data is constructed.
[0069] Select indicators that can reflect the patient's vital signs and respiratory function status, including the patient's physiological indicators such as blood oxygen saturation, blood pressure, heart rate, respiratory rate and other indicators, as well as the corresponding patient ventilator use time data. Specifically:
[0070] First, combining clinical knowledge and experience, exclude some features that are relevant but difficult to obtain in actual applications or have little impact on the prediction results of the ventilator use time, and ensure that the selected features have practical clinical significance and application value. Such as Figure 3As shown, it is a schematic diagram of data extraction. These feature data include the patient's basic physiological indicators (such as blood oxygen, blood pressure, heart rate, etc.), severity of illness scores (such as APACHE II, SOFA scores), ventilator parameters (such as tidal volume, inspiratory time, respiratory rate, etc.), and blood gas analysis results (such as pH value, PaO2 / FiO2 ratio, etc.). Specifically as follows:
[0071] Demographics: age, gender, weight, readmission, Elixhauser score;
[0072] Vital signs: SOFA, SIRS, GCS, heart rate, systemic blood pressure, ambulatory blood pressure, mean blood pressure, shock index, body temperature spo2;
[0073] Laboratory values: potassium, sodium, chloride, glucose, blood sugar, creatinine, magnesium, carbon dioxide, hemoglobin, white blood cell count, platelet count, prothrombin time, clotting time, international normalized ratio, pH, partial pressure of carbon dioxide, base excess, bicarbonate;
[0074] Body fluids: urine output, pressure device, intravenous infusion, cumulative fluid balance.
[0075] Then, calculate the time from each patient's state to the end of ventilation, and extract the intubation time corresponding to each patient's state.
[0076] S3. Train and optimize multiple preselected machine learning regression models based on the dedicated dataset.
[0077] Build a suitable machine learning regression model for the task of predicting the intubation time of patients. As Figure 4 shown, it is a schematic diagram of predicting the intubation time of a ventilator based on a dataset and machine learning. During the model training process, monitor the loss function of the model and evaluation metrics on the validation set, such as mean squared error (MSE), mean absolute error (MAE), etc. When the loss function converges or the evaluation metrics on the validation set no longer improve, stop training. Specifically:
[0078] Considering that the relationship between the patient's state and the duration of ventilator use may be relatively complex, select multiple machine learning regression models (such as linear regression, support vector regression, random forest regression, etc.) at the same time, and use the patient's data for training to optimize the model parameters. In addition, considering the temporal characteristics and non-linear relationships of the data, deep learning methods can be introduced for comparison. For example, long short-term memory networks (LSTM) or gated recurrent units (GRU) can process time series data and capture the dynamic changes of the patient's state.
[0079] During the model training process, the cross-validation method is used to evaluate the generalization ability of the model. The above-mentioned dedicated dataset is divided into a training set, a validation set, and a test set, usually in the ratio of 70%, 15%, 15%. The model is trained on the training set, and the hyperparameters of different models are adjusted on the validation set. Through multiple cross-validations, the model parameters that perform best on the validation set are selected to avoid overfitting or underfitting of the model. The gradient descent optimization algorithm is used to update the parameters of the model so that the model can better fit the training data. The specific steps are as follows:
[0080] Initialize parameters: First, the parameters of the model (such as weights) need to be initialized. Usually, the parameters are initialized to zero or small random values.
[0081] Set the learning rate and the maximum number of iterations: The learning rate is a hyperparameter that controls the step size and determines the magnitude of the parameter change in each update. Generally speaking, the learning rate should not be too large, otherwise it may skip the optimal solution; nor should it be too small, otherwise it will lead to a too slow convergence rate. At the same time, a maximum number of iterations also needs to be set to prevent the algorithm from falling into an infinite loop.
[0082] Define the loss function: The loss function (such as the mean squared error) is used to measure the gap between the model prediction value and the true value. The goal of gradient descent is to continuously adjust the model parameters so that the value of the loss function is as small as possible.
[0083] Calculate the gradient: In each iteration, calculate the derivative (gradient) of the loss function with respect to the model parameters. The gradient represents the change trend of the loss function at the current parameter values. By calculating the gradient, it can be known how to adjust the parameters to make the value of the loss function decrease.
[0084] Update the parameters: According to the direction and magnitude of the gradient, update the model parameters. The update method is: the parameter minus the learning rate multiplied by the gradient. In this way, the model parameters are adjusted in the direction that can reduce the loss function.
[0085] Check the convergence situation: After each parameter update, check whether the loss function has converged (that is, the change becomes very small) or the gradient has become close to zero. If a certain stopping condition is met (such as the change in the loss function is less than a certain threshold), stop the algorithm.
[0086] Output the optimized parameters: After several iterations, the parameters of the model will be gradually optimized, and finally a set of parameters that can minimize the loss function will be obtained.
[0087] During the training process, monitor the loss function of the model and the evaluation metrics on the validation set, such as the mean squared error (MSE), mean absolute error (MAE), etc. When the loss function converges or the evaluation metrics on the validation set no longer improve, stop the training.
[0088] In this embodiment, to address the complex relationship between the patient's condition and the duration of ventilator use, models such as linear regression, support vector regression (SVR), and random forest are trained simultaneously. The optimal model is selected through cross-validation, which can avoid the limitations of a single algorithm.
[0089] S4. Predict the ventilator usage time based on the machine learning regression model with the optimal performance after optimization.
[0090] Indicators such as mean squared error (MSE) and R2 are used to evaluate the performance of the model, compare the prediction effects of different algorithms, and select the model with the optimal evaluation index as the final prediction model. Specifically:
[0091] The mean squared error MSE measures the average squared error between the predicted value and the true value. The formula is:
[0092]
[0093] where n is the number of samples; y i is the true value, is the predicted value. The smaller the value of MSE, the closer the prediction result of the model is to the true value, and the better the performance of the model. The R 2 index is used to measure the goodness of fit of the model to the data. The formula is:
[0094]
[0095] where, is the mean of the true values. The closer the value of R 2 is to 1, the better the fitting effect of the model to the data, and the model can explain most of the variations in the data.
[0096] Based on indicators such as mean squared error (MSE) and R2, compare the prediction effects of different model algorithms after the above training and optimization. By running different models on the test set, calculate their evaluation indicators, and select the model with the optimal evaluation index as the final prediction model. Conduct a visual analysis of the prediction results of the model, draw scatter plots, residual plots, etc. of the predicted values and the true values, and intuitively observe the prediction effect and error distribution of the model. For example, in this embodiment, twelve widely used machine learning regression models are selected for training. The results are as Figure 5 shown. It can be seen that the three models with the highest goodness of fit scores are Decision Tree, RFR, and Bagging, indicating that their fitting effects are among the top three. Select these three relatively good models for cross-validation. The results are as Figure 6 shown. It can be seen that the RFR model has the lowest RMSE value, which can verify that it has the best fitting effect on this data set.
[0097] In this embodiment, the distribution of predicted values and true values is shown through charts, making the evaluation model more intuitive and reliable. Through these analyses, the performance of the model can be further understood, problems existing in the model can be discovered, and a basis for improving the model can be provided.
[0098] Finally, from all the optimized models, the machine learning regression model with the best performance is selected to predict the ventilator use time of patients.
[0099] Thus, all the processes of a method for predicting ventilator use time based on machine learning in this embodiment are completed.
[0100] Embodiment 2:
[0101] Secondly, the present invention also provides a system for predicting ventilator use time based on machine learning, and the system includes:
[0102] A data preprocessing module, which is used to preprocess the pre-acquired MIMIC-III data;
[0103] A dedicated dataset acquisition module, which is used to construct a dedicated dataset based on the preprocessed MIMIC-III data, and the dedicated dataset includes patient vital signs and respiratory function status indicators, as well as the ventilator use time corresponding one-to-one to the patient vital signs and respiratory function status indicators;
[0104] A model training module, which is used to train and optimize a variety of preselected machine learning regression models based on the dedicated dataset;
[0105] A ventilator use time prediction module, which is used to predict the ventilator use time based on the machine learning regression model with the best performance after optimization.
[0106] Among them, when the data preprocessing module, the dedicated dataset acquisition module, the model training module, and the ventilator use time prediction module process data, they execute the steps of the method for predicting ventilator use time based on machine learning proposed in the first aspect above, and the steps mainly include:
[0107] S1. Preprocess the pre-acquired MIMIC-III data;
[0108] S2. Construct a dedicated dataset based on the preprocessed MIMIC-III data, and the dedicated dataset includes patient vital signs and respiratory function status indicators, as well as the ventilator use time corresponding one-to-one to the patient vital signs and respiratory function status indicators;
[0109] S3. Train and optimize a variety of preselected machine learning regression models based on the dedicated dataset;
[0110] S4. Predict the ventilator intubation time based on the machine learning regression model with the optimal performance after optimization.
[0111] It can be understood that the ventilator intubation time prediction system based on machine learning provided by the embodiments of the present invention corresponds to the above-mentioned ventilator intubation time prediction method based on machine learning. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the ventilator intubation time prediction method based on machine learning, which will not be elaborated here.
[0112] Embodiment 3:
[0113] Thirdly, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the steps of the ventilator intubation time prediction method based on machine learning as described in any of the above embodiments. The method includes:
[0114] S1. Preprocess the pre-acquired MIMIC-III data;
[0115] S2. Construct a dedicated data set based on the preprocessed MIMIC-III data. The dedicated data set includes the patient's vital signs and respiratory function status indicators, as well as the ventilator intubation time corresponding to the patient's vital signs and respiratory function status indicators one by one;
[0116] S3. Train and optimize a variety of preselected machine learning regression models based on the dedicated data set;
[0117] S4. Predict the ventilator intubation time based on the machine learning regression model with the optimal performance after optimization.
[0118] It can be understood that the storage medium for ventilator intubation time prediction based on machine learning provided by the embodiments of the present invention corresponds to the above-mentioned ventilator intubation time prediction method based on machine learning. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the ventilator intubation time prediction method based on machine learning, which will not be elaborated here.
[0119] Embodiment 4:
[0120] Fourthly, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it realizes the steps of the ventilator intubation time prediction method based on machine learning as described in any of the above embodiments. The method includes:
[0121] S1. Preprocess the pre-acquired MIMIC-III data;
[0122] S2. Construct a dedicated dataset based on the preprocessed MIMIC-III data. The dedicated dataset includes the patient's vital signs and respiratory function status indicators, as well as the ventilator intubation time corresponding one-to-one to the patient's vital signs and respiratory function status indicators.
[0123] S3. Train and optimize a variety of preselected machine learning regression models based on the dedicated dataset.
[0124] S4. Predict the ventilator intubation time based on the machine learning regression model with the optimal performance after optimization.
[0125] It can be understood that the electronic device for predicting the ventilator intubation time based on machine learning provided by the embodiments of the present invention corresponds to the above-mentioned method for predicting the ventilator intubation time based on machine learning. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the method for predicting the ventilator intubation time based on machine learning, which will not be elaborated here.
[0126] In summary, compared with the prior art, the following beneficial effects are achieved:
[0127] 1. A technology for predicting the ventilator intubation time based on machine learning proposed in this application uses advanced machine learning and deep learning algorithms, combines a large amount of clinical case data, trains a prediction model that can accurately capture the relationship between the patient's condition characteristics and the ventilator intubation time, and then uses this prediction model to deeply analyze and mine multi-dimensional information such as the patient's clinical data and physiological parameters to achieve accurate prediction of the patient's ventilator intubation time. This technology can not only predict whether a patient needs to use a ventilator, but also accurately predict its usage duration, providing reliable decision-making support for clinicians.
[0128] 2. A technology for predicting the ventilator intubation time based on machine learning proposed in this application preprocesses the original data using advanced data preprocessing and missing value processing strategies, which can improve the quality and usability of the data, and further improve the accuracy of predicting the ventilator intubation time.
[0129] 3. A technology for predicting the ventilator intubation time based on machine learning proposed in this application preselects a variety of machine learning regression models including linear regression, support vector regression (SVR), random forest, etc. for training and optimization at the same time, and selects the optimal model through cross-validation for intubation time prediction, avoiding the limitations of a single algorithm.
[0130] 4. A technology for predicting the ventilator intubation time based on machine learning proposed in this application constructs a dedicated dataset as a multi-feature dataset, and at the same time its prediction model has the ability to integrate multi-dimensional and high-granularity data, significantly improving the accuracy and generalization of predicting the ventilator intubation time.
[0131] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0132] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting ventilator time based on machine learning, characterized in that: The method comprises: Preprocess the pre-acquired MIMIC-III data; Constructing a dedicated data set based on the preprocessed MIMIC-III data, wherein the dedicated data set includes patient vital signs and respiratory function status indicators, and ventilator on-machine time corresponding to the patient vital signs and respiratory function status indicators; Training and optimizing a plurality of pre-selected machine learning regression models based on the dedicated data set; The ventilator time was predicted based on the optimized machine learning regression model with the best performance.
2. The method according to claim 1, characterized in that The preprocessing includes data cleaning and data standardization, and the data cleaning includes data missing value processing and outlier processing.
3. The method according to claim 2, characterized in that The data standardization includes Z-score standardization; the data missing value processing includes k-nearest neighbor interpolation and mean interpolation.
4. The method according to claim 1, characterized in that The patient's vital signs and respiratory function status indicators include: the patient's basic physiological indicators, disease severity score, ventilator parameters, and blood gas analysis results.
5. The method according to claim 1, characterized in that The method also includes: when training and optimizing the machine learning regression model, judging the training and optimization effects by monitoring the loss function of the model and the evaluation indicators on the validation set, and evaluating the generalization ability of the machine learning regression model by a cross-validation method.
6. The method according to claim 5, characterized in that The machine learning regression model includes a linear regression model, a support vector regression model, and a random forest regression model.
7. The method according to claim 5, characterized in that The evaluation indicators include the mean square error between the predicted value and the true value and the goodness of fit of the model to the data; The calculation formula for mean square error MSE is: Where n is the number of samples; y i is the true value, is the predicted value; Goodness of fit R 2 The calculation formula is: in, is the mean of the true values.
8. A ventilator time prediction system based on machine learning, characterized in that: The system comprises: A data preprocessing module, used to preprocess the pre-acquired MIMIC-III data; A dedicated data set acquisition module, used to construct a dedicated data set based on the preprocessed MIMIC-III data, wherein the dedicated data set includes the patient's vital signs and respiratory function status indicators, and the ventilator on-machine time corresponding to the patient's vital signs and respiratory function status indicators; A model training module, used for training and optimizing a plurality of pre-selected machine learning regression models based on the dedicated data set; The on-machine time prediction module is used to predict the ventilator on-machine time based on the optimized machine learning regression model with the best performance.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for predicting ventilator on-time based on machine learning as described in any one of claims 1 to 7 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for predicting ventilator on-time based on machine learning as described in any one of claims 1 to 7 are implemented.