Hypoxemia risk prediction system based on multi-modal data

By constructing a hypoxemia risk prediction model with multimodal data, combining facial images and structured data, the problem of poor hypoxemia prediction effect in painless gastroenteroscopy is solved, and higher prediction accuracy is achieved, providing new technical means for clinical practice.

CN120148844APending Publication Date: 2025-06-13THE FIRST AFFILIATED HOSPITAL OF ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204098.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has poor prediction effect of hypoxemia in painless gastroenteroscopy. The traditional method relies on factors such as age and BMI, and the prediction effect based on deep learning combined with facial images has not yet been clarified.

Method used

A hypoxemia risk prediction system based on multimodal data is adopted. By obtaining a sample set including facial images and structured data, a hypoxemia risk prediction model is constructed. The model includes an image processing module, a structured data processing module, a feature fusion module, a fully connected layer and an output layer. The convolutional neural network and residual module are used for feature extraction and fusion, and finally the hypoxemia prediction results are generated through the fully connected layer and the output layer.

Benefits of technology

The prediction accuracy of hypoxemia risk in painless gastroenteroscopy has been significantly improved, and the AUC value of the ROC curve reaches more than 0.8, which is better than the existing diagnostic methods and provides a more accurate prediction method for clinical practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148844A_ABST
    Figure CN120148844A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biomedicine, and relates to a hypoxemia risk prediction system based on multi-modal data, the hypoxemia risk prediction system comprises an input module and a hypoxemia prediction module; the input module is used for inputting data information of a patient subjected to painless gastrointestinal endoscopy, and the data information comprises a facial image and structured data of the patient subjected to painless gastrointestinal endoscopy; the hypoxemia prediction module is used for processing and analyzing the input data information to obtain a hypoxemia prediction result of a patient subjected to painless gastrointestinal endoscopy and outputting the hypoxemia prediction result; wherein a hypoxemia risk prediction model is carried in the hypoxemia prediction module. According to the hypoxemia risk prediction system, the facial image features of the patient subjected to painless gastrointestinal endoscopy and the structured data information can be combined to carry out hypoxemia risk prediction, and the prediction accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biomedicine, and particularly relates to a hypoxemia risk prediction system based on multimodal data. Background Art

[0002] Digestive endoscopy has currently become one of the gold standards for diagnosing gastrointestinal diseases and an important means of treatment. However, while it provides convenience for patients, it also brings pain and discomfort. The emergence of painless gastroscopy and colonoscopy has greatly reduced the pain and discomfort of patients during the operation, and greatly improved the compliance of patients and the quality of endoscopic operations by doctors. However, painless gastroscopy and colonoscopy need to be performed under general anesthesia, so many complications may occur during the examination, including hypoxemia. Untreated hypoxemia may cause serious consequences such as myocardial infarction and cerebral infarction. Therefore, it is particularly important to predict the occurrence of hypoxemia in patients during painless gastroscopy and colonoscopy. Traditional prediction methods include predicting based on factors such as age, body mass index (BMI), history of habitual snoring, and neck circumference. More importantly, these factors can often be reflected on the face. However, the effect of predicting hypoxemia in patients during painless gastroscopy and colonoscopy based on deep learning combined with facial images is still unclear at present. Summary of the Invention

[0003] Aiming at the problems and deficiencies existing in the prior art, the purpose of the present invention is to provide a hypoxemia risk prediction system based on multimodal data.

[0004] To achieve the purpose of the invention, the technical solution adopted by the present invention is as follows:

[0005] In the first aspect of the present invention, a training method for a hypoxemia risk prediction model based on multimodal data is provided. The training method includes the following steps:

[0006] S1: Obtain a sample set, the sample set includes a plurality of samples, each sample includes the data information of a patient undergoing painless gastroscopy and colonoscopy and the corresponding true classification label of hypoxemia. The data information includes the facial image and structured data of the patient undergoing painless gastroscopy and colonoscopy; divide the sample set into a training set and a validation set according to a ratio;

[0007] S2: Input the data information of each sample in the training set into a pre-constructed hypoxemia risk prediction model in turn for training, update the network parameters of the hypoxemia risk prediction model, and obtain a trained hypoxemia risk prediction model;

[0008] S3: Use the test set to test the trained hypoxemia risk prediction model, evaluate the performance of the trained hypoxemia risk prediction model in predicting the hypoxemia risk, and continue the test until the accuracy of the trained hypoxemia risk prediction model in predicting the hypoxemia risk is greater than or equal to the preset threshold, so as to obtain the trained hypoxemia risk prediction model.

[0009] According to the above training method, preferably, the hypoxemia risk prediction model includes an image processing module, a structured data processing module, a feature fusion module, a fully connected layer, and an output layer; the image processing module is used to extract features from the facial image of the input sample to obtain the vectorized features of the facial image and output them. The input of the image processing module is the facial image of the sample, and the output of the image processing module is the vectorized features of the facial image; the structured data processing module is used to process and extract features from the structured data of the input sample to obtain the feature vector of the structured data and output it; the feature fusion module is used to fuse the vectorized features of the facial image output by the image processing module with the feature vector of the structured data output by the structured data processing module to obtain a fused feature vector and output it; the fully connected layer is used to perform dimensionality reduction and feature extraction on the fused feature vector output by the fusion feature module to obtain a one-dimensional feature vector and output it; the output layer is used to convert the one-dimensional feature vector output by the fully connected layer into a probability distribution to obtain the hypoxemia prediction classification result corresponding to the input sample.

[0010] According to the above training method, preferably, the image processing module consists of an input module, a convolution module, a max pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, a global average pooling layer, and a custom fully connected layer connected in sequence; the convolution module is used to perform convolution processing on the facial image of the input sample to obtain the convolution feature map of the facial image and output it, and the max pooling layer is used to perform a pooling operation on the convolution feature map output by the convolution module to obtain a pooled feature map and output it; the first residual module is used to extract features from the pooled feature map output by the max pooling layer to obtain a first residual feature map and output it; the second residual module is used to extract features from the first residual feature map to obtain a second residual feature map and output it; the third residual module is used to extract features from the second residual feature map to obtain a third residual feature map and output it; the fourth residual module is used to extract features from the third residual feature map to obtain a fourth residual feature map and output it; the global average pooling layer is used to perform global average pooling on the fourth residual feature map to obtain a one-dimensional feature vector; the custom fully connected layer is used to process the one-dimensional feature vector output by the global average pooling layer to obtain the vectorized features of the facial image of the input sample.

[0011] According to the above training method, preferably, the output layer uses a Sigmoid activation function to output classification probabilities; a ReLU activation function is provided in the custom fully connected layer to introduce non-linearity, thereby helping the network learn more complex features.

[0012] According to the above training method, preferably, the convolution module includes a convolution layer with a convolution kernel of 7×7 and a stride of 2.

[0013] According to the above training method, preferably, the first residual module includes three residual units, each residual unit is composed of three convolution layers, the convolution kernels of the three convolution layers are 1×1, 3×3, and 1×1 respectively, and the number of convolution kernels of the three convolution layers are 64, 64, and 256 respectively.

[0014] According to the above training method, preferably, the second residual module includes four residual units, each residual unit is composed of three convolution layers, the convolution kernels of the three convolution layers are 1×1, 3×3, and 1×1 respectively, and the number of convolution kernels of the three convolution layers are 28, 28, and 512 respectively.

[0015] According to the above training method, preferably, the third residual module includes six residual units, each residual unit is composed of three convolution layers, the convolution kernels of the three convolution layers are 1×1, 3×3, and 1×1 respectively, and the number of convolution kernels of the three convolution layers are 256, 256, and 1024 respectively.

[0016] According to the above training method, preferably, the fourth residual module includes three residual units, each residual unit is composed of three convolution layers, the convolution kernels of the three convolution layers are 1×1, 3×3, and 1×1 respectively, and the number of convolution kernels of the three convolution layers are 512, 512, and 2048 respectively.

[0017] According to the above training method, preferably, the specific operation of step S2 is as follows: sequentially input the data information of each sample in the training set into the pre-constructed hypoxemia risk prediction model for training to obtain the hypoxemia prediction result of each sample; compare the hypoxemia prediction result of the sample with the true classification label of hypoxemia corresponding to the sample to construct a loss function; then use the backpropagation method to update the network parameters of the hypoxemia risk prediction model to reduce the value of the loss function until the loss function converges, and obtain the trained hypoxemia risk prediction model.

[0018] According to the above training method, preferably, the structured data includes the age (actual age, unit: years) of patients undergoing painless gastroscopy and colonoscopy, BMI (body mass index, calculated by dividing the patient's weight (kg) by the square of the height (m)), gender, ASA classification (physical status classification of the American Society of Anesthesiologists, with the classification range from I to V), whether smoking, whether drinking alcohol, whether having habitual snoring, whether having a history of hypertension, and whether having a history of diabetes.

[0019] According to the above training method, preferably, the gender is represented by binary coding, with female represented as 0 and male represented as 1; whether smoking is represented by binary coding, with non-smoking represented as 0 and smoking represented as 1; whether drinking alcohol is represented by binary coding, with non-drinking represented as 0 and drinking represented as 1; whether having habitual snoring is represented by binary coding, with no habitual snoring represented as 0 and having habitual snoring represented as 1; whether having a history of hypertension is represented by binary coding, with no history of hypertension represented as 0 and having a history of hypertension represented as 1; whether having a history of diabetes is represented by binary coding, with no history of diabetes represented as 0 and having a history of diabetes represented as 1.

[0020] According to the above training method, preferably, the acquisition method of each sample facial image in the sample set is: using a high-definition camera device to acquire the frontal facial image of a patient undergoing painless gastroscopy and colonoscopy, ensuring uniform light during shooting, no occlusion of the face, and the patient being in a static state. After acquiring the facial image of the sample, preprocessing of the facial image is required. The preprocessing method includes image scaling and data augmentation processing. First, adjust the facial image of the hypoxemia patient to a unified size (preferably 224×224), and then use data augmentation techniques (such as rotation, translation, scaling, flipping, etc.) to increase the diversity of image data and reduce overfitting.

[0021] According to the above training method, preferably, the collection method of each sample structured data in the sample set is: obtaining the structured data of the patient through the hospital's information management system (HIS). After acquiring the structured data of the sample, preprocessing of the structured data is required. The preprocessing methods include: data cleaning, normalization or standardization processing, etc.

[0022] According to the above training method, preferably, in step S2, before training the pre-constructed hypoxemia risk prediction model using the training set, the pre-constructed hypoxemia risk prediction model is pre-trained using the ImageNet dataset, and then the pre-trained hypoxemia risk prediction model is trained using the training set. By pre-training the pre-constructed hypoxemia risk prediction model using the ImageNet dataset, when the pre-trained hypoxemia risk prediction model is trained using the training set constructed by the present invention, the pre-trained hypoxemia risk prediction model can retain the pre-trained weights to capture general features in the images, and can avoid overfitting problems caused by training the model from scratch and small samples.

[0023] The second aspect of the present invention provides a hypoxemia risk prediction system based on multimodal data. The hypoxemia risk prediction system includes an input module and a hypoxemia prediction module. The input module is used to input data information of a patient undergoing painless gastroscopy and colonoscopy, and the data information includes the facial image and structured data of the patient undergoing painless gastroscopy and colonoscopy. The hypoxemia prediction module is used to process and analyze the input data information, obtain the hypoxemia prediction result of the patient undergoing painless gastroscopy and colonoscopy and output it. Among them, a hypoxemia risk prediction model is loaded in the hypoxemia prediction module. The hypoxemia risk prediction model is a hypoxemia risk prediction model trained using the training method described in the first aspect above.

[0024] According to the above hypoxemia risk prediction system, preferably, the hypoxemia risk prediction system further includes an alarm and feedback module. The alarm and feedback module is used to analyze the prediction result output by the hypoxemia prediction module to determine whether it is necessary to trigger a hypoxemia risk alarm. If the prediction result output by the hypoxemia prediction module analyzed belongs to the presence of hypoxemia risk, the alarm is triggered.

[0025] According to the above hypoxemia risk prediction system, preferably, the structured data includes the age (actual age, unit: years) of the patient undergoing painless gastroscopy and colonoscopy, BMI (body mass index, calculated by dividing the patient's weight (kg) by the square of the height (m)), gender, ASA classification (physical status classification of the American Society of Anesthesiologists, classification range from I to V), whether smoking, whether drinking alcohol, whether having habitual snoring, whether having a history of hypertension, whether having a history of diabetes.

[0026] According to the above hypoxemia risk prediction system, preferably, the gender is represented by binary coding, with female represented as 0 and male represented as 1; whether smoking is represented by binary coding, non-smoking is represented as 0, and smoking is represented as 1; whether drinking alcohol is represented by binary coding, non-drinking is represented as 0, and drinking alcohol is represented as 1; whether having habitual snoring is represented by binary coding, no habitual snoring is represented as 0, and having habitual snoring is represented as 1; whether having a history of hypertension is represented by binary coding, no history of hypertension is represented as 0, and having a history of hypertension is represented as 1; whether having a history of diabetes is represented by binary coding, no history of diabetes is represented as 0, and having a history of diabetes is represented as 1.

[0027] The third aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the training method described in the first aspect above is implemented.

[0028] The fourth aspect of the present invention provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and it is characterized in that when the computer program is executed by a processor, the training method described in the first aspect above is implemented.

[0029] Compared with the prior art, the positive and beneficial effects achieved by the present invention are as follows:

[0030] (1) The image processing module in the hypoxemia risk prediction model of the present invention is composed of an input module, a convolution module, a max pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, a global average pooling layer, and a custom fully connected layer connected in sequence. This image processing module can capture low-level features such as edges and textures through the first residual module and the second residual module, and capture high-level semantic features through the third residual module and the fourth residual module. Therefore, it can achieve multi-level feature extraction of facial images and obtain richer facial vectorized features.

[0031] (2) The hypoxemia risk prediction model in the present invention consists of an image processing module, a structured data processing module, a feature fusion module, a fully connected layer, and an output layer. The image processing module is used to extract features from the facial images of the input samples to obtain the vectorized features of the facial images. The structured data processing module is used to process the structured data of the input samples to obtain structured data features. Then, the feature fusion module fuses the vectorized features of the facial images with the feature vectors of the structured data. After being processed by the fully connected layer and the output layer, the hypoxemia prediction result of the input sample is obtained. Therefore, the hypoxemia risk prediction model of the present invention can combine the facial image features of patients undergoing painless gastroscopy and colonoscopy with structured data (such as the age, BMI, gender, ASA classification, smoking status, drinking habits, presence of habitual snoring, hypertension, and diabetes of the subject to be tested) information for hypoxemia risk prediction, which can improve the accuracy of hypoxemia risk prediction. The existing prediction method using the NOSOS questionnaire (including five variables: neck circumference, obesity, snoring, age, and gender) to predict the hypoxemia risk of patients undergoing painless gastroscopy and colonoscopy has an AUC of 0.7340 for the ROC curve of diagnosing and differentiating hypoxemia. The AUC of the ROC curve of using the hypoxemia risk prediction model of the present invention to predict the hypoxemia risk of patients undergoing painless gastroscopy and colonoscopy reaches above 0.8, which is significantly higher than the existing diagnostic method and has high accuracy, providing a new technical means for the prediction of clinical hypoxemia. Description of the Drawings

[0032] Figure 1 It is a schematic diagram of the network architecture of the hypoxemia risk prediction model based on multi-modal data in the present invention;

[0033] Figure 2 It is the ROC curve of the hypoxemia risk prediction model under K-fold cross-validation. Detailed Embodiment

[0034] The following further elaborates on the present invention through specific embodiments, but does not limit the scope of the present invention.

[0035] Embodiment 1:

[0036] A training method for a hypoxemia risk prediction model based on multi-modal data, the specific steps are as follows:

[0037] S1: Collect the facial images and data information of multiple patients undergoing painless gastroscopy and colonoscopy, construct a sample set according to the collected data information. The sample set includes multiple samples, and each sample includes the data information of the patient undergoing painless gastroscopy and colonoscopy and the corresponding true classification label of hypoxemia. Randomly divide the sample set into a training set and a validation set according to a ratio of 7:3.

[0038] Among them, the acquisition method of each sample facial image in the sample set is as follows: Use a high-definition camera device to acquire the frontal facial image of the patient undergoing painless gastroscopy and colonoscopy, ensure uniform light during shooting, no occlusion of the face, and the patient is in a static state. After acquiring the facial image of the sample, it is necessary to preprocess the facial image. The preprocessing method includes image scaling and data augmentation. First, adjust the facial image of the hypoxemia patient to a unified size (preferably 224×224), and then use data augmentation techniques (such as rotation, translation, scaling, flipping, etc.) to increase the diversity of image data and reduce overfitting.

[0039] The method for collecting the structured data of each sample in the sample set is as follows: Obtain the structured data of the patient through the hospital's information management system (HIS). After collecting the structured data of the sample, it is necessary to preprocess the structured data. The preprocessing methods include: data cleaning, normalization or standardization processing, etc.

[0040] The structured data includes the age (actual age, unit: years) of the patient undergoing painless gastroscopy and colonoscopy, BMI (body mass index, calculated by dividing the patient's weight (kg) by the square of the height (m)), gender, ASA classification (physical status classification of the American Society of Anesthesiologists, the classification range is from I to V), whether smoking, whether drinking alcohol, whether having habitual snoring, whether having a history of hypertension, whether having a history of diabetes. The gender is represented by binary coding, 0 represents female, and 1 represents male; whether smoking is represented by binary coding, 0 represents non-smoking, and 1 represents smoking; whether drinking alcohol is represented by binary coding, 0 represents non-drinking, and 1 represents drinking; whether having habitual snoring is represented by binary coding, 0 represents no habitual snoring, and 1 represents having habitual snoring; whether having a history of hypertension is represented by binary coding, 0 represents no history of hypertension, and 1 represents having a history of hypertension; whether having a history of diabetes is represented by binary coding, 0 represents no history of diabetes, and 1 represents having a history of diabetes.

[0041] S2: Input the data information of the samples in the training set into the pre-constructed hypoxemia risk prediction model (determine the initial network parameters of the model, such as learning rate, batch size, number of layers, number of neurons, etc.) for training to obtain the hypoxemia prediction result of each sample; compare the hypoxemia prediction result of the sample with the true classification label of hypoxemia corresponding to the sample to construct a loss function; then use the backpropagation method to update the network parameters of the hypoxemia risk prediction model to reduce the value of the loss function until the loss function converges to obtain the trained hypoxemia risk prediction model.

[0042] Among them, the hypoxemia risk prediction model (such as Figure 1As shown, it includes an image processing module, a structured data processing module, a feature fusion module, a fully connected layer, and an output layer.

[0043] The image processing module is used to extract features from the facial image of the input sample, obtain the vectorized features of the facial image and output them. The input of the image processing module is the facial image of the sample, and the output of the image processing module is the vectorized features of the facial image. The image processing module consists of an input module, a convolution module, a max pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, a global average pooling layer, and a custom fully connected layer connected in sequence. The convolution module is used to perform convolution processing on the facial image of the input sample (preferably 224×224×3), obtain the convolution feature map of the facial image (preferably 112×112×64) and output it; preferably, the convolution module contains 1 convolution layer, and the convolution layer contains 64 7×7 convolution kernels with a stride of 2. The max pooling layer is used to perform a pooling operation on the convolution feature map output by the convolution module (preferably 112×112×64), obtain the pooled feature map (preferably 56×56×64) and output it; preferably, the pooling window of the max pooling layer is 3×3 and the stride is 2. The first residual module is used to extract features from the pooled feature map output by the max pooling feature layer (preferably 56×56×64), obtain the first residual feature map (preferably 56×56×256) and output it; preferably, the first residual module contains 3 residual units, and each residual unit consists of 3 convolution layers. The convolution kernel sizes of the 3 convolution layers are 1×1, 3×3, and 1×1 respectively, and the convolution kernel numbers of the 3 convolution layers are 64, 64, and 256 respectively. The second residual module is used to extract features from the first residual feature map (preferably 56×56×256), obtain the second residual feature map (preferably 28×28×512) and output it; preferably, the second residual module contains 4 residual units, and each residual unit consists of 3 convolution layers. The convolution kernel sizes of the 3 convolution layers are 1×1, 3×3, and 1×1 respectively, and the convolution kernel numbers of the 3 convolution layers are 128, 128, and 512 respectively. The third residual module is used to extract features from the second residual feature map (preferably 28×28×512), obtain the third residual feature map (preferably 14×14×1024) and output it; preferably, the third residual module contains 6 residual units, and each residual unit consists of 3 convolution layers. The convolution kernel sizes of the 3 convolution layers are 1×1, 3×3, and 1×1 respectively, and the convolution kernel numbers of the 3 convolution layers are 256, 256, and 1024 respectively.The fourth residual module is used to extract features from the third residual feature map (preferably 14×14×1024), obtain a fourth residual feature map (preferably 7×7×2048) and output it; preferably, the fourth residual module contains 3 residual units, and each residual unit consists of 3 convolutional layers. The kernel sizes of the 3 convolutional layers are 1×1, 3×3, and 1×1 respectively, and the numbers of kernels of the 3 convolutional layers are 512, 512, and 2048 respectively. The global average pooling layer is used to perform global average pooling on the fourth residual feature map (preferably 7×7×2048) to obtain a one-dimensional feature vector (the length of the one-dimensional feature vector is 2048). The neurons in the custom fully connected layer are connected to the one-dimensional feature vector output by the global average pooling layer. By learning the weights and biases, the one-dimensional feature vector input to the custom fully connected layer is mapped to the output category space to achieve predictions of different categories, and the vectorized features of the facial image of the input sample are obtained (the number of channels of the vectorized features is preferably 1024). Preferably, the custom fully connected layer is provided with a ReLU activation function to introduce non-linearity, thereby helping the network learn more complex features.

[0044] The structured data processing module is used to process the feature vector of the structured data of the input sample, obtain structured data features and output them; preferably, the structured data processing module contains Dense Layer 1 and Dense Layer 2. The number of neurons in Dense Layer 1 is 64, and Dense Layer 1 is provided with an activation function ReLU. The number of neurons in Dense Layer 2 is 32, and Dense Layer 2 is provided with an activation function ReLU.

[0045] The feature fusion module is used to fuse the vectorized features of the facial image output by the image processing module and the structured data features output by the structured data processing module to obtain a fused feature vector and output it.

[0046] The fully connected layer is used to perform dimensionality reduction and feature extraction on the fused feature vector output by the fusion feature module to obtain a one-dimensional feature vector and output it. Preferably, the fully connected layer contains Dense Layer 1 and Dense Layer 2. The number of neurons in Dense Layer 1 is 128, and Dense Layer 1 is provided with an activation function ReLU. The number of neurons in Dense Layer 2 is 64, and Dense Layer 2 is provided with an activation function ReLU.

[0047] The output layer is used to convert the one-dimensional feature vector output by the fully connected layer into a probability distribution, and obtain the hypoxemia prediction classification result corresponding to the input sample. Preferably, the output layer uses the Sigmoid activation function to output the classification probability.

[0048] Furthermore, when using the training set to train the pre-constructed hypoxemia risk prediction model, the network parameters of the hypoxemia risk prediction model adjusted and updated include the learning rate, batch size, optimizer, and regularization parameters, etc. Among them, in the process of adjusting the learning rate, different learning rates (such as 0.1, 0.01, 0.001, etc.) can be tried to find the most suitable learning rate for the model; in addition, a learning rate scheduler (such as ReduceLROnPlateau, Cosine Annealing, etc.) can be used to dynamically adjust the learning rate during training. In the process of adjusting the batch size, different batch sizes (such as 32, 64, 128, etc.) can be tried to balance the training speed and the model performance. In the process of adjusting the optimizer, different optimizers (such as SGD, Adam, RMSprop, etc.) can be selected and their parameters (such as momentum, decay rate, etc.) can be adjusted. The regularization parameter adjustment includes the Dropout rate and the L2 regularization coefficient. The Dropout rate adjusts the dropout rate of the Dropout layer (such as 0.2, 0.5, etc.) to prevent overfitting, and the L2 regularization coefficient adjusts the weight coefficient of the L2 regularization to limit the complexity of the model.

[0049] Furthermore, before using the training set to train the pre-constructed hypoxemia risk prediction model, the pre-constructed hypoxemia risk prediction model can be pre-trained using the ImageNet dataset first, and then the training set can be used to train the pre-trained hypoxemia risk prediction model. By pre-training the pre-constructed hypoxemia risk prediction model using the ImageNet dataset, when using the training set constructed by the present invention to train the pre-trained hypoxemia risk prediction model, the pre-trained hypoxemia risk prediction model can retain the pre-trained weights to capture the general features in the image, and can avoid the overfitting problem caused by training the model from scratch and small samples.

[0050] S3: Use the test set to test the trained hypoxemia risk prediction model, evaluate the performance of the trained hypoxemia risk prediction model in predicting the hypoxemia risk, until the accuracy of the trained hypoxemia risk prediction model in predicting the hypoxemia risk is greater than or equal to the preset threshold, and obtain the trained hypoxemia risk prediction model.

[0051] Example 2:

[0052] A hypoxemia risk prediction system based on multimodal data, the hypoxemia risk prediction system includes an input module and a hypoxemia prediction module; the input module is used to input the data information of the patients undergoing painless gastroscopy, and the data information includes the facial images and structured data of the patients undergoing painless gastroscopy; the hypoxemia prediction module is used to process and analyze the input data information, obtain the hypoxemia prediction result of the patients undergoing painless gastroscopy and output it; wherein, a hypoxemia risk prediction model is carried in the hypoxemia prediction module; the hypoxemia risk prediction model is a trained hypoxemia risk prediction model obtained by using the training method described in Embodiment 1 above.

[0053] Further, the hypoxemia risk prediction system further includes an alarm and feedback module, and the alarm and feedback module is used to analyze the prediction result output by the hypoxemia prediction module to judge whether it is necessary to trigger a hypoxemia risk alarm. If the prediction result output by the hypoxemia prediction module belongs to the existence of hypoxemia risk after analysis, an alarm is triggered.

[0054] Embodiment 3:

[0055] An electronic device includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the training method of the hypoxemia risk prediction model described in Embodiment 1 is implemented.

[0056] Embodiment 4:

[0057] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the training method of the hypoxemia risk prediction model described in Embodiment 1 is implemented.

[0058] Performance test of the hypoxemia risk prediction model based on multimodal data of the present invention for hypoxemia risk prediction:

[0059] Collect clinical samples to perform a performance test on the trained hypoxemia risk prediction model obtained by using the training method described in Embodiment 1 of the present invention.

[0060] 1. Clinical sample source: A total of 1003 patients who received painless gastroscopy and colonoscopy at the First Affiliated Hospital of Zhengzhou University from October 2022 to October 2023 were collected. After screening and exclusion (excluding patients with facial makeup and plastic surgery history and patients who could not remove ornaments, excluding patients receiving oxygen supply through laryngeal mask and tracheal catheter, excluding patients who could not understand and cooperate with the researchers), 361 samples were finally included. Facial images and structured data of the 361 included samples were collected. Among the 361 samples, the number of patients with hypoxemia after general anesthesia for painless gastroscopy and colonoscopy was 45 (recorded as the hypoxemia group, hypoxemia), and the number of patients with normal blood oxygen was 316 (recorded as the normal blood oxygen group, Normal). The statistical results of the baseline data of the structured data of the samples in the hypoxemia group and the normal blood oxygen group are shown in Table 1.

[0061] Table 1 Statistical results of baseline data of structured data of samples in the hypoxemia group and the normal blood oxygen group

[0062]

[0063]

[0064] 2. K-fold cross-validation:

[0065] The trained hypoxemia risk prediction model obtained by using the training method described in Example 1 of the present invention was used to perform K-fold cross-validation (K = 5) on the 361 collected samples. Then, an ROC curve was made according to the results of the K-fold cross-validation, and the accuracy (Accuracy, that is, the proportion of the number of correctly classified samples in the total number of samples), precision (Precision, that is, the proportion of the number of samples predicted as positive for hypoxemia by the hypoxemia risk prediction model in the number of samples in the hypoxemia group), recall (Recall, that is, the proportion of the number of samples in the hypoxemia group correctly predicted as positive for hypoxemia by the hypoxemia risk prediction model), specificity (Specificity, that is, the proportion of the number of samples in the normal blood oxygen group correctly predicted as negative for hypoxemia by the hypoxemia risk prediction model), and F1 score (F1 Score, that is, the harmonic mean of precision and recall) of the trained hypoxemia risk prediction model were calculated.

[0066] Among them, the ROC curve of the hypoxemia risk prediction model under K-fold cross-validation is as Figure 2 shown.

[0067] By Figure 2It can be seen that the AUC values of the ROC curves of the five validation results of the hypoxemia risk prediction model after training of the present invention are all above 0.8. Thus, it shows that the hypoxemia risk prediction model obtained by training with the training method of the present invention has good stability and accuracy, and can be used to predict the risk of hypoxemia after general anesthesia for painless gastroenteroscopy examination.

[0068] The accuracy, precision, recall, specificity, and F1 score of the hypoxemia risk prediction model under K-fold cross-validation were statistically analyzed, and the statistical results are shown in Table 2.

[0069] Table 2 shows the statistical results of the accuracy, precision, recall, specificity, and F1 score of the hypoxemia risk prediction model under K-fold cross-validation.

[0070] AUC Accuracy Precision Recall Specificity F1 Score The 1st Cross-Validation 0.8112 0.6953 0.9085 0.4529 0.9640 0.6045 The 2nd Cross-Validation 0.8464 0.7688 0.8171 0.6752 0.8571 0.7394 The 3rd Cross-Validation 0.8631 0.7719 0.7198 0.8271 0.7246 0.7697 The 4th Cross-Validation 0.8757 0.8000 0.8318 0.7884 0.8136 0.8095 The 5th Cross-Validation 0.8642 0.7734 0.8959 0.6188 0.9281 0.7320 Average Evaluation Index 0.8521 0.7618 0.8346 0.6724 0.8574 0.7310

[0071] As can be seen from Table 2, the accuracy of the five validation results of the hypoxemia risk prediction model after training of the present invention is greater than 0.69, the precision is greater than 0.71, the specificity is greater than 0.72, and the F1 score is greater than 0.6. Thus, it shows that the hypoxemia risk prediction model after training of the present invention has good prediction accuracy and can be used for the prediction of hypoxemia after general anesthesia.

[0072] Finally, it should be noted that the above embodiments are only preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may make changes or modifications by using the above technical content as inspiration. These equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical concept of the present invention still fall within the protection scope of the claims of the present invention.

Claims

1. A method for training a hypoxemia risk prediction model based on multimodal data, characterized in that: The following steps are involved: S1: Obtain a sample set, the sample set includes multiple samples, each sample includes data information of patients undergoing painless gastrointestinal endoscopy and corresponding hypoxemia true classification labels, the data information includes facial images and structured data of patients undergoing painless gastrointestinal endoscopy; divide the sample set into a training set and a validation set in proportion; S2: inputting the data information of each sample in the training set into a pre-built hypoxemia risk prediction model in sequence for training, updating the network parameters of the hypoxemia risk prediction model, and obtaining a trained hypoxemia risk prediction model; S3: Use the test set to test the trained hypoxemia risk prediction model to evaluate the performance of the trained hypoxemia risk prediction model in predicting hypoxemia risk, until the accuracy of the trained hypoxemia risk prediction model in predicting hypoxemia risk is greater than or equal to a preset threshold, thereby obtaining a trained hypoxemia risk prediction model.

2. The training method according to claim 1, characterized in that: The hypoxemia risk prediction model includes an image processing module, a structured data processing module, a feature fusion module, a fully connected layer and an output layer; the image processing module is used to extract features from the facial image of the input sample, obtain the vectorized features of the facial image and output them; The structured data processing module is used to perform feature extraction and processing on the structured data of the input sample, obtain the feature vector of the structured data and output it; the feature fusion module is used to perform feature fusion on the facial image vectorization features output by the image processing module and the structured data feature vector output by the structured data processing module, obtain the fused feature vector and output it; the fully connected layer is used to perform dimensionality reduction and feature extraction processing on the fused feature vector output by the fusion feature module, obtain the one-dimensional feature vector and output it; the output layer is used to convert the one-dimensional feature vector output by the fully connected layer into a probability distribution, and obtain the hypoxemia prediction classification result corresponding to the input sample.

3. The training method according to claim 2, characterized in that: The image processing module consists of an input module, a convolution module, a maximum pooling layer, a first residual module, a second residual module, a third residual module, a fourth residual module, a global average pooling layer, and a custom fully connected layer, which are connected in sequence.

4. The training method according to claim 3, characterized in that: The output layer uses a Sigmoid activation function to output classification probability, and a ReLU activation function is provided in the custom fully connected layer.

5. The training method according to claim 4, characterized in that: The convolution module includes a convolution layer with a convolution kernel of 7×7 and a step size of 2; the first residual module shown includes three residual units, the second residual module shown includes four residual units, the third residual module shown includes six residual units, and the fourth residual module shown includes three residual units. Each residual unit is composed of three convolution layers, and the convolution kernels of the three convolution layers are 1×1, 3×3, and 1×1, respectively.

6. The training method according to claim 5, characterized in that: The structured data includes the age, BMI (body mass index), gender, ASA grade, smoking status, drinking status, habitual snoring, history of hypertension, and history of diabetes of the patient undergoing painless gastroenteroscopy; the gender is represented by binary coding, with female being represented by 0 and male being represented by 1; the smoking status is represented by binary coding, with non-smoking being represented by 0 and smoking being represented by 1; Whether or not the patient drinks alcohol is represented by binary coding, with no drinking represented by 0 and drinking represented by 1; whether or not the patient has habitual snoring is represented by binary coding, with no habitual snoring represented by 0 and habitual snoring represented by 1; whether or not the patient has a history of hypertension is represented by binary coding, with no history of hypertension represented by 0 and a history of hypertension represented by 1; whether or not the patient has a history of diabetes is represented by binary coding, with no history of diabetes represented by 0 and a history of diabetes represented by 1.

7. A hypoxemia risk prediction system based on multimodal data, characterized in that: The hypoxemia risk prediction system includes an input module and a hypoxemia prediction module; the input module is used to input data information of patients undergoing painless gastrointestinal endoscopy, and the data information includes facial images and structured data of patients undergoing painless gastrointestinal endoscopy; the hypoxemia prediction module is used to process and analyze the input data information, obtain and output hypoxemia prediction results for patients undergoing painless gastrointestinal endoscopy; wherein the hypoxemia prediction module is equipped with a hypoxemia risk prediction model; the hypoxemia risk prediction model is a trained hypoxemia risk prediction model obtained by training using any training method described in claims 1-5.

8. The hypoxemia risk prediction system according to claim 7, characterized in that: The hypoxemia risk prediction system also includes an alarm and feedback module, which is used to analyze the prediction results output by the hypoxemia prediction module to determine whether a hypoxemia risk alarm needs to be triggered. If the prediction results output by the hypoxemia prediction module are determined to indicate a hypoxemia risk, an alarm is triggered.

9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor implements the training method according to any one of claims 1 to 6 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method according to any one of claims 1 to 6 is implemented.