ARDS prediction method and system based on multi-modal data and deep learning
By combining multimodal analysis of CT imaging and clinical data, and using 2.5D U-Net and XGBoost models, an ARDS prediction model was constructed. This model addresses the shortcomings of existing models in early warning and dynamic progression warning, enabling precise treatment of ARDS patients and reducing the risk of death.
Patent Information
- Application Number
- CN202610136275.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing ARDS prediction models fail to fully integrate clinical indicators and biological markers, lacking multi-dimensional comprehensive prediction, resulting in insufficient early warning and dynamic progression warning, making it difficult to achieve precision treatment and reduce the risk of death.
By fusing CT imaging features with clinical data, a multimodal prediction model for the severity of ARDS patients was constructed using a 2.5D U-Net segmentation model and the LE-Model algorithm for image data processing, combined with the XGBoost model and the K-means clustering algorithm.
It enables accurate prediction of the severity of ARDS patients' conditions, improves the accuracy and stability of early diagnosis, reduces the risk of death, and optimizes the allocation of medical resources and public health emergency response capabilities.
Smart Images

Figure CN122000068A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer-aided diagnostic technology, and more specifically, to an ARDS prediction method and system based on multimodal data and deep learning. Background Technology
[0002] ARDS, a severe clinical manifestation of acute hypoxic respiratory failure, is a critical respiratory disease induced by a combination of intrapulmonary and / or extrapulmonary factors. Its core pathophysiological feature is refractory hypoxemia, often secondary to severe infections, trauma, burns, shock, and other critical illnesses. This syndrome is characterized by insidious onset, rapid progression, and high clinical mortality. The 2012 Berlin definition clearly defines its diagnostic criteria as: progressively worsening dyspnea within one week of the precipitating factor, diffuse infiltrates in both lungs on chest X-ray (CXR) or CT, and dyspnea that cannot be explained by other causes such as atelectasis, lung tumors, or pulmonary nodules. In recent years, with the global pandemic of COVID-19, the total number of ARDS patients caused by COVID-19 has increased significantly, and its high mortality rate has attracted widespread attention in the global critical care medicine field. Related research data shows that the mortality rate of COVID-19-related ARDS is approximately 26%-50%, with about 25% being mild cases and about 75% progressing to moderate to severe ARDS. Meanwhile, ARDS patients account for about 5% of all ICU patients receiving mechanical ventilation, indicating a very high disease burden within the ICU population. Given the difficulty in establishing diagnostic criteria for ARDS (such as increased pulmonary vascular permeability and diffuse alveolar damage) in clinical practice, we believe that precise CT imaging segmentation can significantly advance the standardization of intrapulmonary assessment of ARDS. Combined with clinical parameters, this can enable early and accurate warning of ARDS, becoming a crucial element in optimizing clinical treatment strategies and reducing mortality.
[0003] Currently, the focus of ARDS research has gradually shifted from simple supportive treatment to early identification and precise intervention in high-risk patients. Early and accurate identification of ARDS and its tendency to worsen in ICU inpatients can provide a basis for decision-making regarding the early implementation of lung-protective ventilation strategies and conservative fluid management protocols, which has significant clinical implications. Existing research has confirmed that various laboratory biomarkers and quantitative imaging features can be used to construct ARDS prediction models. For example, the feasibility of monitoring ARDS from CT images based on radiomics or traditional quantitative analysis methods has been validated; studies have indicated that the area under the curve (AUC) of the lung-affected region on CT at admission is statistically significant in predicting whether ARDS patients require extracorporeal membrane oxygenation (ECMO). Furthermore, calculation methods based on CT image data and gas exchange parameters have been used to deduce the mechanical ventilation requirements during the lung function recovery process of ARDS patients. A nationwide study constructed a deep learning (DL) model using quantitative CT features and initial clinical indicators to predict the severity of COVID-19, demonstrating higher accuracy and stability in disease grading. Another study proposed a rapid and accurate method for pneumonia inflammation segmentation, which significantly reduces the computational complexity of the model and greatly improves segmentation accuracy by decomposing the 3D segmentation problem into three 2D problems. Research on ARDS prediction based on imaging features and DL has yielded initial success. Experimental results show that the prediction model built based on image features achieves an accuracy of 82.4%, effectively assessing the severity of ARDS and providing reliable intelligent support for clinical diagnosis and treatment decisions.
[0004] However, current research on CT imaging-based artificial intelligence (AI) both domestically and internationally focuses primarily on lesion segmentation and diagnosis, and has not yet fully integrated clinical indicators and biological markers to construct multi-dimensional comprehensive predictive models. This, to some extent, limits the application value of AI technology in clinical early warning and prospective intervention. Most existing early warning models are "static diagnosis" or "severity grading" models, failing to effectively predict the dynamic progression of diseases. Furthermore, most AI models remain in the algorithm validation stage and have not yet been effectively packaged into early warning platforms that can be directly used in clinical practice, lacking a practical validation phase combined with large-scale databases from clinical treatment centers.
[0005] This invention aims to quantitatively obtain key parameters such as lung volume and inflammatory volume from imaging indicators such as CT imaging, and combine them with clinical diagnosis and treatment data and laboratory indicators for multimodal information joint analysis to construct an accurate predictive model of the severity of ARDS patients' conditions. This will assist clinicians in making more reasonable treatment decisions for ARDS patients, such as the indications for the use of high-flow oxygen therapy or mechanical ventilation, the timing of respiratory support implementation, and the optimization of ventilator parameters. Ultimately, this will achieve early intervention and precise treatment of the disease, minimizing the risk of death for patients. Summary of the Invention
[0006] The technical problem to be solved by this invention is:
[0007] The core objective of this invention is to provide a multimodal information analysis method based on the fusion of CT imaging features and clinical data. This method uses computer algorithms to intelligently predict the severity of ARDS patients' conditions, thereby assisting clinicians in implementing early intervention and individualized treatment to minimize the risk of death for patients.
[0008] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0009] This invention provides an ARDS prediction method based on multimodal data and deep learning, comprising the following steps:
[0010] S100: Collect CT imaging data and clinical diagnosis and treatment data of ARDS patients;
[0011] S200. The CT image data obtained in step S100 is normalized, and then a 2.5D U-Net segmentation model is used for data segmentation, that is, the 3D segmentation task is simplified into 2D segmentation tasks from different perspectives, and then the 2D segmentation tasks are fused to obtain the final segmentation result.
[0012] S300. After completing the segmentation of CT image data, the sub-visual recognition algorithm LE-Model is applied to remove blood vessels and trachea from the CT images for subsequent feature enhancement and feature extraction.
[0013] S400: Based on the XGBoost model, the ARDS prediction model is constructed and predictive analysis is performed by fusing the feature-enhanced CT image data from step S300 with the clinical data collected in step S100.
[0014] Furthermore, in step S100, the data collection targets are patients diagnosed with ARDS according to the Berlin diagnostic criteria and who have been admitted to the ICU for more than 24 hours.
[0015] Furthermore, in step S100, the clinical data involved includes patient demographic indicators and ICU stay duration, wherein the patient demographic indicators include gender, age, height, and weight; laboratory test indicators, including D-dimer count, white blood cell count, absolute neutrophil count, absolute lymphocyte count, C-reactive protein, IL-6, and procalcitonin levels; respiratory support-related parameters, including respiratory support mode on the day of treatment, duration of mechanical ventilation, ventilator settings, concentration and flow parameters of high-flow oxygen therapy, arterial partial pressure of oxygen, and arterial oxygenation index; and patient prognostic information, including whether the patient was successfully weaned off the ventilator and clinical outcome.
[0016] Further, in step S200, data normalization preprocessing is performed, including two dimensions: spatial normalization and signal normalization. Spatial normalization is used to unify the resolution and pixel size of CT image data and eliminate spatial dimension deviations caused by differences in scanning equipment. Signal normalization is based on the lung window settings of the CT scanner and standardizes the signal intensity of each voxel to ensure the consistency of signal amplitude.
[0017] Furthermore, in step S200, the 2.5D U-Net segmentation model employs a three-view collaborative segmentation strategy to perform 2D segmentation operations, specifically the cross-section, coronal plane, and sagittal plane.
[0018] Further, in step S300, 20,000 voxels are randomly extracted from the remaining lung parenchyma region obtained from each CT scan. Using the baseline CT image as a reference, the median CT signal of the sampled voxels is extracted, and the standard deviation of the CT signal in the healthy lung parenchyma region is calculated. In the process of calculating the standard deviation, in order to reduce noise interference and remove outliers, the top 20% of the largest and smallest CT values are discarded. Based on the median CT signal and the standard deviation σ, the optimal window width and window level parameters for observing sub-visual lesions are determined.
[0019] For each CT image, by truncating the CT signal, the enhanced CT involves linear scaling of the original CT signal, effectively eliminating scan-level bias;
[0020]
[0021] Among them, enhanced CT refers to enhanced CT; original CT refers to original CT; and baseline refers to the baseline.
[0022] Furthermore, in step S400, the K-means clustering algorithm is introduced for collaborative prediction. The prediction is completed based on the label information of the training samples that are closest to the new data point by locating a preset number of training samples. If it is a regression prediction scenario, the average of the predicted values of the K nearest neighbor samples is used as the final prediction result to improve the prediction stability and accuracy.
[0023] An ARDS prediction system based on multimodal data and deep learning has program modules corresponding to the steps described above. When running, the system executes the steps in the ARDS prediction method based on multimodal data and deep learning.
[0024] A computer-readable storage medium storing a computer program configured to, when invoked by a processor, implement steps of an ARDS prediction method based on multimodal data and deep learning.
[0025] Compared with the prior art, the beneficial effects of the present invention are:
[0026] This invention proposes a method to predict the severity of ARDS patients by integrating multimodal data that combines clinical and imaging features with machine learning models, and constructs an early warning platform for ARDS. The experimental methods include a segmentation algorithm for lung inflammatory lesions, an enhancement method to improve segmentation results, and a prediction method based on clinical and imaging features. The 2.5D U-Net segmentation model decomposes the 3D segmentation task into 2D segmentation tasks for different views, achieving a better balance between performance and model complexity. The LE-Model algorithm enhances the segmentation results by removing blood vessels and trachea from lung images.
[0027] This invention employs the XGBoost model, an efficient and scalable machine learning algorithm based on gradient boosting, to predict regression values. Combined with distance discriminant analysis, it uses the K-Means clustering method to jointly predict the severity of ARDS.
[0028] Experimental results show that the method proposed in this invention achieves a high accuracy in predicting the severity of ARDS, with a Dice coefficient of 90.2% for lung inflammation segmentation and an accuracy of 86.5% for predicting the severity of ARDS.
[0029] Addressing key issues in ARDS diagnosis and treatment, such as delayed early identification and limited assessment dimensions, this project systematically integrates resources including imaging data, clinical physiological indicators, and biomarkers to construct an intelligent early warning and analysis system. Through innovative multi-task, multi-modal feature deep fusion modeling methods, it overcomes the limitations of traditional single-dimensional analysis, achieving accurate identification of severe ARDS risk, which is of great significance for the early diagnosis and treatment of ARDS. Clinically validated, this project can improve the accuracy and stability of ARDS diagnosis, providing reliable technical support for clinical decision-making, and enabling early and accurate disease control and preventative treatment. The method proposed in this invention can help clinicians assess the condition of ARDS patients earlier and more accurately, taking more effective treatment measures, thereby reducing ARDS mortality. It can not only promote the optimal allocation of medical resources but also enhance public health emergency response capabilities, while solving the problem of inconsistent treatment of severe cases caused by uneven distribution of medical resources and lack of standardized treatment protocols among different regions and levels of medical institutions, thus improving the success rate of ARDS treatment. Attached Figure Description
[0030] Figure 1 This is a flowchart of an ARDS prediction method based on multimodal data and deep learning, as described in an embodiment of the present invention.
[0031] Figure 2 This is a data normalization graph in an embodiment of the present invention;
[0032] Figure 3 This is a diagram showing the result of 2D segmentation from three angles in an embodiment of the present invention;
[0033] Figure 4 This is a visualization comparison chart of the predicted pneumonia area using the LE-Model in an embodiment of the present invention;
[0034] Figure 5 This is a comparison chart of the mean and standard deviation of different models under normal conditions in the embodiments of the present invention;
[0035] Figure 6 This is a comparison chart of classification metrics for different models in embodiments of the present invention;
[0036] Figure 7 This is a predicted distribution diagram of all samples after 5x cross-validation in an embodiment of the present invention. Detailed Implementation
[0037] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0038] Specific Implementation Plan 1: Combining Figures 1 to 7As shown, this invention provides an ARDS prediction method based on multimodal data and deep learning, comprising the following steps:
[0039] S100: Collect CT imaging data and clinical data of ARDS patients.
[0040] Clinical and imaging data were collected from 250 ARDS patients in the Intensive Care Unit of the First Affiliated Hospital of Harbin Medical University. Multiple sets of CT imaging data and clinical data were included for each patient. All subjects were diagnosed with ARDS strictly according to the Berlin diagnostic criteria and required at least 24 hours of ICU admission. This project has been approved by the Ethics Committee of the First Affiliated Hospital of Harbin Medical University (Ethics Review No.: 2023XS22-02), and informed consent has been obtained from all patients.
[0041] This study system requires the collection of multi-dimensional data, covering the following core categories: Baseline demographic and clinical characteristics: including basic information such as gender, age, height, and weight; underlying diseases (comorbidities); and initial severity grading of COVID-19; Laboratory indicators: covering blood routine related indicators: white blood cell count (WBC), neutrophil percentage (NEUT%), lymphocyte count (LYMPH), lymphocyte percentage (LYM%), red blood cell distribution width (RDW), and platelet count (PLT); Inflammatory marker C-reactive protein (CRP); Cytokine profile: IL-6, TNF-α; Liver and kidney function indicators: alanine aminotransferase (ALT), aspartate aminotransferase (AST), serum creatinine (SCr), albumin (ALB), total bilirubin (TB), direct bilirubin (DBIL), indirect bilirubin (IBIL), and coagulation function related indicators (D-dimer); Disease severity and prognostic indicators: Sequential organ failure assessment calculated based on dynamic disease progression data (…). SOFA score, Acute Physiology and Chronic Health Evaluation II (APACHE II) score; Imaging data: Chest CT and chest X-ray images throughout the treatment; Respiratory support parameters: Respiratory support mode, duration of mechanical ventilation, ventilator settings, concentration and flow rate of high-flow oxygen therapy on the day of treatment, arterial partial pressure of oxygen, arterial oxygenation index; and prognostic information for the patient: whether the patient was successfully weaned off the ventilator and clinical outcome.
[0042] To ensure the integrity, standardization, and normalization of the data, Python data processing tools and SPSS statistical software were used to systematically preprocess the collected data. The specific steps included: data cleaning (removing logical contradictions, outliers, and invalid records), duplicate data removal (removing duplicates based on the patient's unique identifier to ensure the uniqueness of each record), and missing value handling (adopting appropriate filling strategies or marking missing values according to data type to avoid the impact of data bias on subsequent analysis).
[0043] S200: The CT image data acquired in step S100 is segmented using a 2.5D U-Net segmentation model.
[0044] Since acquired CT image data typically comes from different CT scanners with varying parameters, their format standards may differ, making them unsuitable for direct segmentation. Furthermore, one of the main bottlenecks of deep learning-based computer-aided diagnostic methods is the need for training on specific datasets, limiting their applicability to data with specific parameters and hindering generalization to other datasets. To address these two issues, we normalize the images, converting them into a single, standardized form. A normalization diagram is shown below. Figure 2 As shown, the normalization method comprises two parts: spatial normalization and signal normalization. Spatial normalization unifies the resolution and size of the CT scan data; signal normalization, based on the lung window settings of the CT scanner, standardizes the signal intensity of each voxel to ensure the consistency of signal amplitude. This method embeds lung CT scan data into a standard space, enabling the model to use heterogeneous datasets as input and enhancing the model's generalization ability.
[0045] CT scan data is represented in 3D tensor form. The most intuitive method for human organ segmentation is to apply 3D deep learning models, such as 3D Convolutional Neural Networks (CNNs) and 3D U-Nets. However, while these 3D models achieve organ segmentation, they still have the following problems: training 3D models involves a large number of parameters, resulting in high memory requirements and slow convergence speed. Most human tissues can be identified using 2D images. For example, when segmenting pneumonia from CT slices, experienced radiologists only need to obtain information from the previous, current, and next slices to identify pneumonia. However, when segmenting on CT slices, it is not necessary to input the entire CT scan data, but a single CT slice lacks spatial information between slices, which leads to lower segmentation accuracy. Therefore, this invention adopts a 2.5D U-Net segmentation model, which simplifies the 3D segmentation task into a 2D segmentation task from different perspectives, and then fuses these 2D segments to obtain the final segmentation result, achieving a better balance between performance and model complexity.
[0046] When radiologists manually annotate pneumonia, they typically do so on the xy-plane; when unsure whether a particular voxel is part of a target object, they refer to images in the yz and xz-planes to make a final determination; therefore, 2D images from these three planes contain basic information about whether voxels are inflamed, decomposing the 3D segmentation problem into three 2D segmentation problems, such as... Figure 3 As shown; the 2D U-Net segmentation model performs 2D segmentation from three angles: xy (cross-section), yz (coronal plane), and xz (sagittal plane);
[0047] The STL file format is a storage format used to describe the geometric information of 3D objects, using triangular meshes to represent the surface structure of the objects. This invention utilizes the Visualization Toolkit (VTK) framework to convert STL format 3D models into MHA format to extract richer data in the field of medical imaging. In addition, this invention combines the functions of PyRadiomics with an image loading tool for MHA format to extract image features from MHA files, including morphological features, density features, and texture features.
[0048] This invention successfully extracted 62 imaging features and 9 clinical features from 250 collected data, totaling 71 features. The clinical features include age, D-dimer, white blood cell count, absolute neutrophil count, absolute lymphocyte count, C-reactive protein, IL-6 and procalcitonin, inhaled oxygen concentration, oxygen partial pressure, oxygenation index, etc.
[0049] Using Dice coefficient and HD95 as evaluation metrics for segmentation results, the segmentation performance of 2D U-Net, 2.5D U-Net and 3DU-Net segmentation models for lung inflammation was compared, and their parameter computation was also compared. As shown in Table 1, the 2.5DU-Net segmentation model had the best segmentation performance and fewer parameters.
[0050] Table 1 Comparison of Pneumonia Segmentation Accuracy
[0051]
[0052] S300: After completing the segmentation of the CT scan data, based on the sub-visual recognition algorithm LE-Model (LaplacianEigenmap Model, LE-Model), the blood vessels and trachea in the CT image are segmented and removed from the CT image for feature enhancement.
[0053] The CT values of the trachea, blood vessels, and visible lesions are all higher than those of normal lung parenchyma. Therefore, after obtaining the segmentation of pulmonary blood vessels, trachea, and visible pneumonia lesions based on the 2.5D U-Net segmentation model, pulmonary blood vessels, trachea, and visible pneumonia lesions can be removed while preserving normal lung parenchyma. The specific process is as follows: 20,000 voxels are randomly extracted from the remaining parenchyma of each CT scan (20,000 voxels can cover most of the lung parenchyma). The baseline CT is the median CT signal of the sampled voxels, i.e., the median CT signal. When removing blood vessels, trachea, and visible pneumonia lesions from the lung parenchyma, there are still a small amount of other tissues, such as bronchioles and lymph nodes; however, compared with the entire lung parenchyma, the volume of these tissues is very small and can be basically ignored. Using the median CT signal can effectively remove the noise generated by these tissues and provide a baseline CT value for healthy lung parenchyma. At the same time, the standard deviation of healthy lung parenchyma is calculated. When calculating the standard deviation, in order to reduce the influence of noise, the top 20% of the largest and smallest CT values are discarded to remove outliers. By using the baseline CT and the standard deviation σ, the optimal window width and window level for observing sub-visual lesions are determined.
[0054] For each CT image, the CT signal is truncated, and in this invention, two enhanced versions are provided to radiologists: one using [baseline, baseline + 3σ] and the other using [baseline - 3σ, baseline]. On the other hand, enhanced CT involves linear scaling of the original CT signal, effectively eliminating scan-level bias.
[0055]
[0056] Among them, enhanced CT refers to enhanced CT; original CT refers to original CT; and baseline refers to the baseline.
[0057] Combination Figure 4 As shown, the left image represents the original lung CT scan, with the red and blue areas representing segmented pulmonary vessels and trachea, respectively; the right image shows only the lung parenchyma after the vessels and trachea have been removed, and the signal values have been adjusted so that the healthy lung parenchyma has a signal of 0.
[0058] The feature enhancement method of this invention was combined with 2D U-Net, 2.5D U-Net and 3D U-Net segmentation models respectively, and pneumonia segmentation experiments were conducted. The experiments showed that the sub-visual recognition algorithm LE-Model method can significantly improve the segmentation effect after removing blood vessels and trachea from lung images.
[0059] The accuracy of pneumonia segmentation after applying the sub-visual recognition algorithm LE-Model is shown in Table 2 below:
[0060] Table 2 Comparison of Pneumonia Segmentation Accuracy
[0061]
[0062] Combination Figure 4 As shown, the mean (μ) and standard deviation (σ) of different methods under normal conditions were compared; from Figure 5 It can be observed that the mean and standard deviation of the four methods—2.5D U-Net, 2.5D U-Net+LE-Model, Manual, and Label—are different. Specifically, the mean of 2.5D U-Net+LE-Model is the highest at 1.34, while the mean of 2.5D U-Net is the lowest at 1.088. The standard deviation of 2.5D U-Net+LE-Model is the highest at 0.72283, while the standard deviation of 2.5D U-Net is the lowest at 0.77113. In the cross-validation classification data, the classification fit of the 2.5D U-Net model lags behind that of 2.5D U-Net+LE-Model, while the latter is closer to Label, demonstrating superior performance.
[0063] S400: The enhanced CT scan data obtained in step S300 and the clinical data collected in step S100 are used to make predictions based on the XGBoost model.
[0064] The XGBoost model is an efficient and scalable machine learning algorithm based on gradient boosting. During training, each weak classifier is iteratively trained on the data, continuously correcting the errors of the previous weak classifiers and gradually improving the overall model performance. To further improve the prediction results of the XGBoost model, this invention adopts the K-means algorithm. The purpose of using K-means is to perform cluster analysis on these prediction results, so as to aggregate similar samples and exclude as many outlier samples as possible, thereby reducing the number of false positives (FP) and improving accuracy.
[0065] When making predictions, the input lung CT images are first preprocessed, including standardization, feature extraction and structured encoding, to obtain image feature vectors for regression analysis; then the image feature vectors are input into the machine learning regression module to predict the severity of ARDS.
[0066] Among them, the K-Nearest Neighbors (KNN) algorithm is a simple yet stable non-parametric method. Its basic idea is that for a new sample to be predicted, KNN will find the K nearest labeled training samples in the feature space and infer the final predicted value of the new sample based on the label values of these K neighbors. In regression tasks, prediction is usually accomplished by averaging the labels of these K neighbors, thus obtaining a stable and robust estimate.
[0067] In this invention, the distance metric is Euclidean distance, and the number of neighbors is set to K=5. That is, for each test sample, the system will automatically retrieve the 5 training samples that are closest to it, and use the average of the labels of these 5 samples as the final regression prediction result. This method has the advantages of simple implementation, stable prediction, and low requirements for feature distribution assumptions, and is suitable for medical scenarios where the number of image features is limited but the correlation is strong.
[0068] Specific implementation scheme two: The present invention provides an ARDS prediction system based on multimodal data and deep learning. The system has program modules corresponding to the above steps, and executes the steps in the above ARDS prediction method based on multimodal data and deep learning when running.
[0069] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0070] Specific Implementation Scheme 3: The present invention is a computer-readable storage medium storing a computer program configured to implement, when called by a processor, the steps of an ARDS prediction method based on multimodal data and deep learning.
[0071] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0072] experiment
[0073] Experimental environment
[0074] This experiment was implemented in Python 3.8, ran on Ubuntu 22.04.3, and trained on a single NVIDIA Tesla A40 GPU with 48GB of memory. Image reading was performed using SimpleITK and VTK, image feature extraction was performed using Radiomics, and data processing was performed using the NumPy and Pandas frameworks. A five-fold cross-validation method was used to evaluate the model's performance.
[0075] In the implementation of this invention, multi-source clinical and imaging data were collected from target patients to establish a multimodal ARDS database. This project collected data from 250 ARDS patients at the First Affiliated Hospital of Harbin Medical University, including clinical data and lung imaging data. Clinical data included baseline information such as age, gender, height, and weight to comprehensively reflect individual patient characteristics; it also included clinical diagnosis and treatment data and prognostic indicators related to severity to reflect the severity of the patient's disease. The collected imaging data consisted of lung CT images, with CT scan slices 1 mm thick, used for automatic segmentation and quantitative analysis of lung lesion areas, and further extraction of radiomics features such as density, texture, and morphology. In addition, chest X-ray images were acquired to obtain imaging scores of pulmonary edema, thereby helping to reflect the extent and severity of lung involvement.
[0076] In addition to imaging data, multiple clinical physiological indicators and disease severity scores, including APACHE II and SOFA scores, were recorded. These scores encompass various physiological data such as body temperature, hemodynamic parameters, respiratory and blood gas parameters, electrolyte concentrations, renal function indicators, hematological parameters, and level of consciousness, comprehensively reflecting the patient's overall health status and the severity of the disease. Furthermore, laboratory test results were also obtained, including white blood cell count, C-reactive protein, creatinine, and electrolyte levels such as sodium and potassium, to reflect the degree of inflammation, infection status, and vital organ function. The patient's past chronic disease status was also included in the data system of this invention to assess underlying risk factors that may influence disease development.
[0077] Regarding disease progression-related information, this invention records whether the patient developed ARDS during hospitalization, whether respiratory support treatment was required, and key clinical time points, which serve as the basis for model prediction. Simultaneously, high-dimensional features automatically extracted based on imaging data and deep learning algorithms are also incorporated into the analysis to further improve the model's representational capabilities and predictive performance.
[0078] To ensure the reliability of model evaluation, this invention divides the dataset into a training set containing 200 samples and a test set consisting of 50 samples, and uses a five-fold cross-validation method for comparative analysis. During model performance evaluation, a ten-fold cross-validation method is further employed to ensure the stability and generalizability of the results across different subsets of the data.
[0079] Through the systematic integration of the above multi-dimensional data, this invention constructs a comprehensive data system covering imaging, physiological indicators, laboratory indicators, basic characteristics, disease course and outcome, as well as features automatically extracted by the model, providing complete and reliable data support for subsequent prediction model construction and performance verification.
[0080] Evaluation indicators
[0081] Accuracy, precision, and recall were used as evaluation metrics for the experiment.
[0082]
[0083]
[0084]
[0085] TP represents a true positive; FP represents a false positive; and FN represents a false negative.
[0086] result
[0087] Medical image segmentation was performed using the 2.5D U-Net segmentation model, followed by XGBoost model for prediction and classification, achieving an accuracy of 78.4%. This finding demonstrates that the 2.5D U-Net segmentation model can effectively extract image features, which is beneficial for subsequent classification tasks. The integration of the LE-Model with the 2.5D U-Net segmentation model not only improved segmentation accuracy but also increased the classification accuracy of the XGBoost model to 83.1%, indicating the positive impact of local enhancement models on classification performance.
[0088] Further refinement was achieved by integrating clinical data into the 2.5D U-Net+LE-Model. With the addition of clinical information, the segmentation results improved slightly, and the XGBoost model achieved a classification accuracy of 83.2%. This indicates that fusing clinical data helps improve the performance of the classification model. The accuracy of different features is shown in Table 3.
[0089] Using XGBoost prediction models
[0090] An experiment was conducted using the XGBoost model to classify pneumonia, and the precision, recall, and accuracy metrics for predicting severe pneumonia were calculated. The results are shown in Table 5 below. The experiment demonstrates that the 2.5D U-Net + LE-Model algorithm achieved the highest accuracy. Figure 6 As shown, ROC curves for three categories are plotted: 0, 1, and 2, as shown in (a), (b), and (c) respectively.
[0091] Table 3 Accuracy of different features
[0092]
[0093] Table 4 Evaluation Indicators for Different Methods
[0094]
[0095] * This indicates a subsequent classification method based on inflammation region segmentation.
[0096] Significance test
[0097] Wilcoxon signed-rank tests were performed on 2.5D U-Net and 2.5D U-Net+LE-Model to assess whether there are significant differences between the two algorithms. Prior to the Wilcoxon signed-rank test, the Bowman-Shenton test was used to verify the symmetry of the data, as the Wilcoxon test requires symmetrically distributed data. R was used to perform the Shapiro-Wilk normality test and the Bowman-Shenton symmetry test. The results are shown in Table 5 below, indicating whether the distribution of the differences is normal and symmetric, respectively.
[0098] Table 5. Comparison of significance tests between the 2.5D U-Net and 2.5D U-Net+LE-Model methods.
[0099]
[0100] The test results show that both sets of data meet the conditions for performing the Wilcoxon signed-rank test. We performed the Wilcoxon signed-rank test on the results of the 2.5D U-Net algorithm and the 2.5D U-Net+LE-Model algorithm to compare the performance of the two algorithms. The test results show that the 2.5D U-Net+LE-Model algorithm is statistically significantly better than the 2.5D U-Net algorithm, indicating that the 2.5D U-Net+LE-Model algorithm is more accurate or performs better than the 2.5D U-Net algorithm. Specifically, the Wilcoxon signed-rank test produced a very low p-value (p=7.974*10). -8 The results provided sufficient statistical evidence to reject the null hypothesis, therefore it is concluded that the results of the 2.5D U-Net+LE-Model algorithm are significantly better than those of the 2.5D U-Net algorithm.
[0101] Using KNN (k=5), the results are shown in Table 6.
[0102] Table 6. Accuracy of KNN with Different Features
[0103]
[0104] In the experiments of this invention, a five-fold cross-validation method was used for model training and evaluation on all 250 patient samples. Specifically, the dataset was randomly divided into five subsets, with each subset serving as the validation set in turn, and the remaining four subsets serving as the training set for model training. For each subset's validation set, the model outputs a predicted value for each sample, reflecting the predicted probability of its disease risk or pathological characteristics.
[0105] After the five-fold cross-validation is completed, the predicted values of all samples are summarized to construct a prediction distribution map, which is then combined with... Figure 7 As shown, this distribution plot illustrates the predicted probability distribution of the model across different samples, providing a clear picture of the model's predictive performance and consistency across all samples. Analysis of the prediction distribution plot reveals: the model's ability to distinguish between high-risk and low-risk samples; the consistency and stability of validation results across different folds; and the distribution characteristics of potential prediction biases or outlier samples.
[0106] This method enables a systematic evaluation of the model's generalization performance under five-fold cross-validation, ensuring that the model's prediction results are stable and reliable across different data subsets, and providing reference probability distribution information for subsequent clinical applications.
[0107] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. An ARDS prediction method based on multimodal data and deep learning, characterized in that, Includes the following steps: S100: Collect CT imaging data and clinical diagnosis and treatment data of ARDS patients; S200. The acquired CT image data is normalized, and then a 2.5D U-Net segmentation model is used for data segmentation, which simplifies the 3D segmentation task into 2D segmentation tasks from different perspectives. Then, the 2D segmentation tasks are fused to obtain the final segmentation result. S300. After completing the segmentation of CT image data, the sub-visual recognition algorithm Laplacian feature mapping model is applied to delete blood vessels and trachea in the CT image for subsequent feature enhancement and feature extraction. S400: Based on the XGBoost model, the ARDS prediction model is constructed and predictive analysis is performed by fusing the feature-enhanced CT image data from step S300 with the clinical data collected in step S100.
2. The ARDS prediction method based on multimodal data and deep learning according to claim 1, characterized in that: In step S100, the data collection subjects are limited to patients diagnosed with ARDS according to the Berlin diagnostic criteria and who have been admitted to the intensive care unit for more than 24 hours.
3. The ARDS prediction method based on multimodal data and deep learning according to claim 2, characterized in that: The clinical data in step S100 includes patient demographic indicators and ICU stay duration, including gender, age, height, and weight; laboratory test indicators, including D-dimer count, white blood cell count, absolute neutrophil count, absolute lymphocyte count, C-reactive protein, IL-6, and procalcitonin levels; respiratory support-related parameters, including respiratory support mode on the day of treatment, duration of mechanical ventilation, ventilator settings, concentration and flow rate of high-flow oxygen therapy, arterial partial pressure of oxygen, and arterial oxygenation index; and patient prognostic information, including successful weaning and clinical outcome.
4. The ARDS prediction method based on multimodal data and deep learning according to claim 3, characterized in that: The data normalization preprocessing in step S200 includes two dimensions: spatial normalization and signal normalization. Spatial normalization is used to unify the resolution and pixel size of CT image data and eliminate spatial dimension deviations caused by differences in scanning equipment. Signal normalization is based on the lung window settings of the CT scanner and standardizes the signal intensity of each voxel to ensure the consistency of signal amplitude.
5. The ARDS prediction method based on multimodal data and deep learning according to claim 4, characterized in that: In step S200, the 2.5D U-Net segmentation model uses a three-view collaborative segmentation strategy to perform 2D segmentation operations. The three views are specifically the cross-section, coronal plane, and sagittal plane.
6. The ARDS prediction method based on multimodal data and deep learning according to claim 5, characterized in that: In step S300, 20,000 voxels are randomly extracted from the remaining lung parenchyma region obtained from each CT scan. Using the baseline CT image as a reference, the median CT signal of the sampled voxels is extracted, and the standard deviation of the CT signal in the healthy lung parenchyma region is calculated. In the process of calculating the standard deviation, in order to reduce noise interference and remove outliers, the top 20% of the largest and smallest CT values are discarded. Based on the median CT signal and the standard deviation σ, the optimal window width and window horizontal parameters for observing sub-visual lesions are determined. For each CT image, by truncating the CT signal, the enhanced CT involves linear scaling of the original CT signal, effectively eliminating scan-level bias; Among them, enhanced CT refers to enhanced CT; original CT refers to original CT; and baseline refers to the baseline.
7. The ARDS prediction method based on multimodal data and deep learning according to claim 6, characterized in that: In step S400, the K-means clustering algorithm is further introduced for collaborative prediction. The prediction is completed based on the label information of the training samples that are closest to the new data point by locating a preset number of training samples. In regression prediction scenarios, the mean of the predicted values of the K nearest neighbor samples is used as the final prediction result to improve prediction stability and accuracy.
8. An ARDS prediction system based on multimodal data and deep learning, characterized in that: The system comprises a series of functional program modules, which are adapted to the steps of the prediction method described in any one of claims 1-7. When the system is running, all the steps of the ARDS prediction method based on multimodal data and deep learning can be fully executed through the coordinated scheduling of the functional modules.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to implement, when invoked by a processor, the steps of the ARDS prediction method based on multimodal data and deep learning as described in any one of claims 1-7.