Newborn pain classification method and system based on multi-modal evidence

Through the neonatal pain classification method based on multimodal evidence, combined with Dirichlet distribution and evidence theory, dynamically evaluate the credibility of multimodal data, the problems of low pain classification accuracy and unbalanced data quality in neonatal babies are solved, and higher pain assessment accuracy and classification accuracy are achieved.

CN120036730APending Publication Date: 2025-05-27NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510185422.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the pain classification accuracy of neonatal children is low, and the multimodal data quality is unbalanced, resulting in low reliability.

Method used

A newborn pain classification method based on multimodal evidence was adopted to construct a multimodal data set by collecting video clips and ECG signals, and data processing and decision-making fusion were used using Dirichlet distribution and evidence theory to dynamically evaluate the credibility of each modal data.

Benefits of technology

Improves the accuracy and classification accuracy of pain assessment, enhances the reliability and interpretability of the model, and provides more reliable pain classification results in complex medical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036730A_ABST
    Figure CN120036730A_ABST
Patent Text Reader

Abstract

The invention provides a neonatal pain classification method and system based on multi-modal evidence, and the method comprises the steps: collecting a neonatal video clip and an electrocardiosignal synchronized with the video, and constructing a neonatal pain multi-modal data set containing a pain label; constructing a neonatal pain classification model based on multi-modal evidence, wherein the model comprises a data processing module, an evidence extraction module, a category confidence and classification uncertainty estimation module and a decision fusion module; using a sample training model in the newborn pain multi-modal data set; and carrying out pain classification on newly input newborn videos and electrocardiosignals by utilizing the trained model. According to the method, by combining Dirichlet distribution and an evidence theory, pain category probability distribution and classification uncertainty of expressions, postures and electrocardio modes are predicted respectively, decision fusion is carried out on prediction results, classification uncertainty evaluation is provided while pain classification is realized, and the reliability and interpretability of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for classifying neonatal pain based on multi-modal evidence, and belongs to the technical field of pain classification. Background Art

[0002] Research shows that fetuses already have the ability to perceive pain in the second trimester of pregnancy, and the pain tolerance of newborns is much lower than that of children of other ages, and their pain perception is more intense and persistent. The number of pain stimuli experienced by each newborn ranges from dozens to hundreds. Repeated pain stimuli may cause permanent negative effects such as developmental retardation, emotional disorders, central nervous system damage, cognitive impairment, inattention, and learning difficulties in newborns. However, newborns cannot express pain through language or use gestures for non-verbal communication. Therefore, accurately assessing the pain level of newborns is crucial, which not only helps medical staff understand the specific situation of patients, but also provides a basis for taking appropriate analgesic measures and formulating reasonable treatment plans.

[0003] Existing intelligent pain assessment technologies are mostly based on single modality. The most common one is to identify pain through facial expressions. However, newborns usually show a series of complex behavioral and physiological reactions when in pain, such as distorted facial expressions, crying, large swings of limbs. Physiological reactions include, but are not limited to: abnormal heart rate, changes in brain waves and cerebral oxygen saturation, increased skin conductivity, etc. Since newborns cannot express pain through language or gestures, and due to the complexity and individual differences of pain reactions, relying solely on single-modal data for assessment has limitations. At the same time, various interference factors in the clinical environment may cause large data noise or loss in a certain modality, thus affecting the accuracy of the assessment results. Therefore, the fusion of multi-modal data is particularly important in neonatal pain assessment. This method can not only integrate multiple response modes, but also improve the clinical applicability and accuracy of pain assessment.

[0004] Although multi-modal fusion can effectively improve the accuracy of pain classification, in the medical scenario, it is crucial to ensure the reliability of multi-modal integration and the final decision, especially when the data is noisy, damaged, or there are large differences between modalities. However, the data quality in the medical environment is often uneven. Therefore, it is necessary to dynamically evaluate the credibility of each modality data to achieve more reliable multi-modal fusion. In view of this, it is indeed necessary to propose a method and system for classifying neonatal pain based on multi-modal evidence. Summary of the Invention

[0005] The purpose of the present invention is to provide a method and system for classifying neonatal pain based on multi-modal evidence to solve the problems of low accuracy of single-modal neonatal pain classification and uneven data quality and low credibility in multi-modal neonatal pain classification existing in the prior art.

[0006] The technical solution of the present invention is as follows:

[0007] A neonatal pain classification method based on multi-modal evidence, comprising the following steps:

[0008] S1. Collect neonatal video clips and electrocardiogram signals synchronized with the video, and construct a neonatal pain multi-modal data set containing pain labels;

[0009] S2. Construct a neonatal pain classification model based on multi-modal evidence. The neonatal pain classification model based on multi-modal evidence includes a data processing module, an evidence extraction module, a class confidence and classification uncertainty estimation module, and a decision fusion module; wherein:

[0010] Data processing module: Preprocess the input neonatal video and electrocardiogram signals to obtain an expression image sequence, a posture image sequence, and an electrocardiogram signal spectrogram;

[0011] Evidence extraction module: Extract evidence e of each modality m supporting pain classification from the expression image sequence, the posture image sequence, and the electrocardiogram signal spectrogram m , including expression modality evidence e 1 , posture modality evidence e 2 , and electrocardiogram modality evidence e 3 ;

[0012] Class confidence and classification uncertainty estimation module: Construct a Dirichlet distribution, i.e., a Dirichlet distribution, for the extracted expression modality evidence e 1 , posture modality evidence e 2 , and electrocardiogram modality evidence e 3 respectively to generate belief functions of each modality where b m is the class confidence of modality m, and u m is the classification uncertainty of modality m;

[0013] Decision fusion module: Fuse the belief functions of each modality to generate a joint belief function where b is the class confidence of the overall model, and u is the classification uncertainty of the overall model; finally, calculate the joint evidence e of each class and the corresponding joint Dirichlet distribution parameters according to the joint belief function and output the final pain class prediction probability;

[0014] S3. Use the neonatal pain multi-modal data set to train the neonatal pain classification model based on multi-modal evidence, and adjust the parameters of the neonatal pain classification model based on multi-modal evidence to the optimal through the error backpropagation algorithm to obtain the trained model;

[0015] S4. Use the trained model to classify the pain of newly input neonatal videos and electrocardiogram signals.

[0016] Furthermore, in step S2, the data processing module includes an expression data processing module, a posture data processing module, and an electrocardiogram data processing module.

[0017] Expression data processing module: Extract a video frame sequence with a length of q frames from the neonatal video within t seconds at the same time interval, where t ranges from 20 to 40. Perform face detection, cropping, alignment, and normalization on each image in the video frame sequence to obtain an expression image sequence with a length of q frames. where w 1 and h 1 are the width and height of the expression image.

[0018] Posture data processing module: Extract a video frame sequence with a length of q frames from the neonatal video at the same time interval. Perform limb detection, cropping, alignment, and normalization on each image in the video frame sequence to obtain a posture image sequence with a length of q frames. where w 2 and h 2 are the width and height of the posture image.

[0019] Electrocardiogram data processing module: Denoise, correct baseline drift, and normalize the input n-lead electrocardiogram signal to obtain a one-dimensional electrocardiogram signal with a length of l. Perform short-time Fourier transform to obtain the electrocardiogram signal spectrogram. where w 3 and h 3 are the width and height of the electrocardiogram spectrogram.

[0020] Furthermore, the evidence extraction module includes an expression evidence extraction module, a posture evidence extraction module, and an electrocardiogram evidence extraction module.

[0021] Expression evidence extraction module: The input expression image sequence sequentially passes through p groups of convolutional modules and a hybrid attention module, and then sequentially passes through a global average pooling layer, a fully connected layer, and a Softplus activation layer to extract the expression modality evidence e 1 , where p is an integer between 4 and 10.

[0022] Posture evidence extraction module: The input posture image sequence sequentially passes through p groups of convolutional modules and a hybrid attention module, and then sequentially passes through a global average pooling layer, a fully connected layer, and a Softplus activation layer to extract the posture modality evidence e 2 , where p is an integer between 4 and 10.

[0023] ECG evidence extraction module: Extract the ECG modality evidence e from the input ECG signal spectrogram 3 。

[0024] Furthermore, the facial expression evidence extraction module and the posture evidence extraction module adopt 3D convolutional neural networks with the same structure but different parameters.

[0025] Furthermore, the ECG evidence extraction module includes a 2D convolutional layer, a normalization layer, an activation layer, a pooling layer, a hybrid attention module, a global pooling layer, a fully connected layer, and a Softplus activation layer connected in sequence. The input ECG signal spectrogram I 3 Extracts ECG features through the 2D convolutional layer. The ECG features are sequentially passed through the normalization layer, the activation layer, and the pooling layer, and then the feature map F is output. The feature map F is enhanced by the hybrid attention module in the channel domain and the spatial domain in sequence, and then passed through the global pooling layer and the fully connected layer in sequence. The Softplus activation layer uses the Softplus activation function to map the features output by the previous layer into the ECG modality evidence e 3 。

[0026] Furthermore, the hybrid attention module includes a channel domain attention module and a spatial domain attention module connected in series

[0027] Channel domain attention module: First, the input feature map Where w and h are the width and height of the feature map respectively, and c is the number of channels of the feature map. Perform max pooling and average pooling respectively to aggregate spatial information and generate two intermediate feature maps Then add the two intermediate feature maps and input them into the fully connected layer. After Sigmoid activation, generate the channel attention weight map Finally, multiply the channel attention weight map F CA With the input feature map F channel by channel to obtain the enhanced feature map

[0028] Spatial domain attention module: First, perform max pooling and average pooling on the input feature map Along the channel dimension to generate two intermediate feature maps respectively Then stack the two intermediate feature maps and perform a convolution operation. After Sigmoid activation, generate the spatial attention weight map Finally, multiply the spatial attention weight map F SA With the input feature map Perform a dot product operation to obtain the enhanced feature map Where w and h are the width and height of the feature map respectively, and c is the number of channels of the feature map.

[0029] Furthermore, in the category confidence and classification uncertainty estimation module, specifically,

[0030] (1) Each modality m uses the extracted evidence e m to construct a Dirichlet distribution respectively as the conjugate prior of the pain category distribution and generate the predicted distribution of the category. The Dirichlet distribution parameter α of modality m m :

[0031]

[0032] Among them, represents the Dirichlet parameter of the 1-Kth pain category in modality m, that is, the Dirichlet parameter, and the Dirichlet parameter of the kth pain category in modality m Among them, represents the evidence of the kth pain category in modality m;

[0033] (2) Based on the subjective logic theory, for the K-classification problem, the belief functions of each modality m are modeled respectively Among them, the single-modal category confidence and the single-modal classification uncertainty u m , for modality m, the K category confidences and the classification uncertainty u m are all non-negative and their sum is 1, as shown in equation (2):

[0034]

[0035] According to the evidence e m calculate the single-modal classification uncertainty u m and the single-modal category confidence b m , as shown in equations (3) and (4):

[0036]

[0037] Among them, the Dirichlet strength

[0038] Among them, represents the Dirichlet parameter of the kth pain category in modality m; represents the evidence of the kth pain category in modality m;

[0039] (3) Calculate the category prediction probability of modality m through the mean value of the Dirichlet distribution of modality m Among them, represents the prediction probability of modality m for the 1-Kth pain category, as shown in equation (5):

[0040]

[0041] Among them, represents the predicted probability of modality m for the k-th category.

[0042] Furthermore, in the decision fusion module, specifically,

[0043] (1) Use Dempster's fusion rule to fuse the belief functions of each modality obtained to generate a joint belief function The joint belief function includes the class confidence b = [b 1 , b 2 , …, b K and classification uncertainty u, where b 1 , b 2 , …, b K represents the credibility that the model input sample belongs to the 1-K pain categories. The fusion rules are shown in equations (6), (7), and (8):

[0044]

[0045] Among them, represents the belief function when modality m is 1, 2, 3 representing the facial expression modality, posture modality, and electrocardiogram modality respectively; ⊕ represents Dempster's fusion, i.e., Dempster fusion; respectively represent the credibility that the input sample belongs to pain category k when modality m is 1, 2, 3; u 1 , u 2 , u 3 respectively represent the classification uncertainties of modalities 1, 2, 3; b k represents the credibility that the model as a whole has for the input sample belonging to pain category k, represents the confidence conflict between different modalities;

[0046] (2) Calculate the joint evidence e = [e , e 1 , …, e 2 , …, e K for each category according to the joint belief function 1 , e 2 , …, e K represents the joint evidence of the 1-K pain categories, and construct the corresponding joint Dirichlet distribution, and calculate the parameters α = [α 1 , α 2 , …, α K and the intensity S of the joint Dirichlet distribution, where α1 , α 2 , …, α K represent the Dirichlet parameters of the 1-Kth pain categories, as shown in equations (9) and (10):

[0047]

[0048] where, e k represents the joint evidence of the kth pain category, and α k represents the Dirichlet parameter of the kth pain category;

[0049] (3) Calculate the final pain category prediction probability p = [p 1 , p 2 , …, p K according to the joint Dirichlet distribution parameter α and the joint Dirichlet distribution intensity S, where p 1 , p 2 , …, p K represent the prediction probabilities of the model for the 1-Kth pain categories respectively, as shown in equation (11):

[0050]

[0051] where, p k is the prediction probability that the input sample belongs to the kth pain category.

[0052] Furthermore, in step S3, the following loss function is used to adjust the parameters of the pain classification model to the optimal by the error backpropagation algorithm:

[0053]

[0054] where, y k represents the true probability that the training sample including the neonatal video segment and the electrocardiogram signal synchronized with the video belongs to the kth category, p k is the prediction probability that the training input sample belongs to the kth pain category, represents the prediction probability of modality m for the kth category, S is the joint Dirichlet distribution intensity, and S m is the Dirichlet intensity. The loss function consists of two terms. The first term is the joint classification loss term, which is used to constrain the overall decision-making accuracy of the model. The second term is the single-modal classification loss term, which is used to constrain the classification accuracy of a single modality during independent learning. λ is a hyperparameter that adjusts the weight of the single-modal loss and takes values between 0 and 1.

[0055] A neonatal pain classification system based on multimodal evidence for implementing the method described in any of the above, comprising a data acquisition module, a model construction module, a model training module, and a pain classification module.

[0056] Data acquisition module: Collect neonatal video clips and electrocardiogram signals synchronized with the video, and construct a neonatal pain multimodal dataset containing pain labels.

[0057] Model construction module: Construct a neonatal pain classification model based on multimodal evidence. The neonatal pain classification model based on multimodal evidence includes a data processing module, an evidence extraction module, a class confidence and classification uncertainty estimation module, and a decision fusion module. Among them:

[0058] Data processing module: Preprocess the input neonatal video and electrocardiogram signals to obtain an expression image sequence, a posture image sequence, and an electrocardiogram signal spectrogram.

[0059] Evidence extraction module: Extract evidence e for pain classification from the expression image sequence, the posture image sequence, and the electrocardiogram signal spectrogram for each modality m m , including expression modality evidence e 1 , posture modality evidence e 2 , and electrocardiogram modality evidence e 3 ;

[0060] Class confidence and classification uncertainty estimation module: Construct Dirichlet distributions, i.e., Dirichlet distributions, for the extracted expression modality evidence e 1 , posture modality evidence e 2 , and electrocardiogram modality evidence e 3 respectively to generate belief functions for each modality where b m is the class confidence of modality m, and u m is the classification uncertainty of modality m;

[0061] Decision fusion module: Fuse the belief functions of each modality to generate a joint belief function where b is the class confidence of the overall model, and u is the classification uncertainty of the overall model; Finally, calculate the joint evidence e for each class and the corresponding joint Dirichlet distribution parameters according to the joint belief function and output the final pain class prediction probability;

[0062] Model training module: Use the neonatal pain multimodal dataset to train the neonatal pain classification model based on multimodal evidence, and adjust the parameters of the neonatal pain classification model based on multimodal evidence to the optimal through the error backpropagation algorithm to obtain the trained model.

[0063] Pain classification module: Use the trained model to classify the pain of newly input neonatal videos and electrocardiogram signals.

[0064] The beneficial effects of the present invention are as follows:

[0065] First, for the neonatal pain classification method and system based on multi-modal evidence, by combining the Dirichlet distribution and the evidence theory, the pain category probability distribution and classification uncertainty of the expression, posture, and electrocardiogram modalities are predicted respectively, and the prediction results are fused for decision-making. While achieving pain classification, it provides an assessment of classification uncertainty, enhancing the reliability and interpretability of the model. It can overcome the limitations of single-modal pain classification methods. Different from the identification methods based on a single data source, the present invention comprehensively considers multiple clinical manifestation forms of neonatal pain by integrating multiple modal data, which can improve the accuracy of pain assessment and the precision of pain classification.

[0066] Second, the present invention adopts a multi-modal fusion method based on the Dempster-Shafer evidence theory, which can dynamically evaluate the sample quality and uncertainty of different modalities, avoiding the influence brought by uneven multi-modal data quality in a complex medical environment. This method not only improves the robustness of data fusion but also gives an assessment of the uncertainty of the current decision during the final classification, providing better model interpretability and decision credibility. By combining the evidence theory and the subjective logic theory, this method can dynamically evaluate the quality of each modal sample, improving the accuracy and robustness of neonatal pain classification. Description of the Drawings

[0067] Figure 1 is a schematic flow chart of the neonatal pain classification method based on multi-modal evidence in an embodiment of the present invention;

[0068] Figure 2 is a schematic illustration of the neonatal pain classification model based on multi-modal evidence in the embodiment;

[0069] Figure 3 is a schematic illustration of the expression evidence extraction module in the embodiment;

[0070] Figure 4 is a schematic illustration of the electrocardiogram evidence extraction module in the embodiment;

[0071] Figure 5 is a schematic illustration of the neonatal pain classification system based on multi-modal evidence in the embodiment. Detailed Embodiment

[0072] The preferred embodiments of the present invention will be described in detail below with reference to the drawings.

[0073] An embodiment provides a neonatal pain classification method based on multimodal evidence, such as Figure 1 , including the following steps,

[0074] S1. Collect neonatal video clips and electrocardiogram signals synchronized with the video, and construct a neonatal pain multimodal dataset containing pain labels.

[0075] In step S1, a camera and a multi-parameter monitor are used to collect neonatal video clips and electrocardiogram signals respectively, and a neonatal pain multimodal dataset containing pain labels is constructed. In this embodiment, the pain is divided into 4 categories according to the pain level, that is, K = 4. The four pain labels are calm, crying, mild pain, and severe pain. In practice, other pain classifications and category labels can also be used.

[0076] S2. Construct a neonatal pain classification model based on multimodal evidence. The neonatal pain classification model based on multimodal evidence includes a data processing module, an evidence extraction module, a category confidence and classification uncertainty estimation module, and a decision fusion module; such as Figure 2 , where:

[0077] Data processing module: Preprocess the input neonatal video and electrocardiogram signals to obtain an expression image sequence, a pose image sequence, and an electrocardiogram signal spectrogram. The data processing module includes an expression data processing module, a pose data processing module, and an electrocardiogram data processing module:

[0078] Expression data processing module: Used to preprocess the neonatal video and images within t seconds of input, where t takes values between 20 and 40; in this embodiment, the model inputs 30s of video and electrocardiogram signals each time, the video frame rate is 30FPS, and the electrocardiogram signal sampling rate is 500Hz; the expression data processing module extracts 1 frame from every 10 frames of the input video using the FFmpeg software to obtain an image sequence with a length of 90 frames, and uses the libfacedetection method to perform face detection, cropping, alignment, and normalization on each image in the image sequence to obtain an expression image sequence with a length of 90 frames and an image size of 224*224

[0079] Pose data processing module: Extract 1 frame from every 10 frames of the input video using the FFmpeg software to obtain an image sequence with a length of 90 frames, and use the YOLOv5 human detection algorithm to perform limb detection, cropping, alignment, and normalization on each image in the image sequence to obtain a pose image sequence with a length of 90 frames and an image size of 224*224

[0080] ECG data processing module: Preprocess the input 3-lead ECG signals. In this embodiment, the Daubechies wavelet 6 filters method is used to denoise the ECG signals and correct the baseline drift, obtaining a 3-lead one-dimensional ECG signal with a length of 15,000. Perform short-time Fourier transform on it to obtain the ECG signal spectrogram. where w 3 , h 3 are the width and height of the ECG spectrogram.

[0081] Evidence extraction module: Extract the evidence e of each modality m supporting pain classification from the expression image sequence, pose image sequence, and ECG signal spectrogram respectively. m , including expression modality evidence e 1 , pose modality evidence e 2 , and ECG modality evidence e 3 . The evidence of each modality m k ∈ [1, 4], where m takes 1, 2, and 3 respectively representing the expression modality, pose modality, and ECG modality. The evidence is a quantitative index used to characterize different pain categories. The larger the value, the stronger the correlation between modality m and pain category k.

[0082] The evidence extraction module includes an expression evidence extraction module, a pose evidence extraction module, and an ECG evidence extraction module. The expression evidence extraction module and the pose evidence extraction module adopt 3D convolutional neural networks with the same structure but different parameters, including the number of convolutional kernels, the size of convolutional kernels, or the size of pooling kernels in the convolutional module.

[0083] Expression evidence extraction module: After the input expression image sequence sequentially passes through p groups of convolutional modules and a hybrid attention module, it then passes through a global average pooling layer, a fully connected layer, and a Softplus activation layer in sequence to extract the expression modality evidence e 1 , where p takes an integer between 4 and 10. For example Figure 3 , each convolutional module contains a 3D convolutional layer, a normalization layer, an activation layer, and a pooling layer connected in sequence. The 3D convolution selects m 1 convolutional kernels of k 1 *k 1 *k 1 to perform convolution operations on the output of the previous layer. Among them, m 1 is selected from the values 4, 8, 16, 32, 64, 128, and k 1 is selected from the values 3, 5, 7. The pooling layer selects a pooling kernel of k 2 *k 2 *k 2 to perform downsampling operations on the output of the previous convolutional layer. Among them, k2 Select from the values 1, 2, and 3. In this embodiment, p is taken as 5, and the convolutional layers of the 5 convolutional modules respectively select 8 3×3×3 convolutional kernels, 16 3×3×3 convolutional kernels, 32 3×3×3 convolutional kernels, 64 3×3×3 convolutional kernels, and 128 3×3×3 convolutional kernels. The pooling layer all selects a 2×2×2 max pooling kernel to perform downsampling on the output of the previous convolutional layer; the hybrid attention module includes a channel-domain attention module and a spatial-domain attention module connected in series, which is used to enhance the important features of the features output by each convolutional module in the channel domain and the spatial domain; the fully connected layer contains 4 neurons, and the features output by the previous layer are mapped into evidence through the Softplus activation layer.

[0084] Pose evidence extraction module: After the input pose image sequence sequentially passes through p groups of convolutional modules and the hybrid attention module, it then sequentially passes through the global average pooling layer, the fully connected layer, and the Softplus activation layer to extract the pose modality evidence e 2 , where p is an integer between 4 and 10;

[0085] ECG evidence extraction module: Extract the ECG modality evidence e from the input ECG signal spectrogram. 3 . The ECG evidence extraction module includes a 2D convolutional layer, a normalization layer, an activation layer, a pooling layer, a hybrid attention module, a global pooling layer, a fully connected layer, and a Softplus activation layer connected in sequence. The input ECG signal spectrogram I 3 extracts ECG features through the 2D convolutional layer. The ECG features sequentially pass through the normalization layer, the activation layer, and the pooling layer and then output the feature map F. The feature map F is enhanced by the hybrid attention module in the channel domain and the spatial domain in sequence and then sequentially passes through the global pooling layer and the fully connected layer containing K = 4 neurons. The Softplus activation layer uses the Softplus activation function to map the features output by the previous layer into the ECG modality evidence e 3 . Among them, the 2D convolutional layer: uses m 2 convolutional kernels of size k 3 ×k 3 , where m 2 is selected from the values 16, 32, 64, and 128, and k 3 is selected from the values 1, 2, and 3; the pooling layer uses a pooling kernel of size k 4 ×k 4 to perform downsampling on the output of the previous layer, where k 4 is selected from the values 1, 2, and 3.

[0086] The hybrid attention module includes a channel-domain attention module and a spatial-domain attention module connected in series:

[0087] Channel-domain attention module: First, the input feature map where w and h are the width and height of the feature map respectively, and c is the number of channels of the feature map, are respectively subjected to max pooling and average pooling to aggregate spatial information and generate two intermediate feature maps Then, the two intermediate feature maps are added and input into a fully connected layer, and after Sigmoid activation, a channel attention weight map is generated Finally, the channel attention weight map F CA is multiplied with the input feature map F channel by channel to obtain an enhanced feature map

[0088] Spatial-domain attention module: First, the input feature map is subjected to max pooling and average pooling along the channel dimension to generate two intermediate feature maps respectively Then, the two intermediate feature maps are stacked and passed through a convolutional operation, and after Sigmoid activation, a spatial attention weight map is generated Finally, the spatial attention weight map F SA is multiplied with the input feature map to perform a dot product operation to obtain an enhanced feature map where w and h are the width and height of the feature map respectively, and c is the number of channels of the feature map.

[0089] Category confidence and classification uncertainty estimation module: For the expression modality evidence e 1 , pose modality evidence e 2 , and electrocardiogram modality evidence e 3 extracted from each modality, Dirichlet distributions, i.e., Dirichlet distributions, are respectively constructed to generate belief functions for each modality where b m is the category confidence of modality m, u m is the classification uncertainty of modality m, and the larger the value of u m , the more uncertain the classification of modality m. In the category confidence and classification uncertainty estimation module, specifically,

[0090] (1) Each modality m uses the extracted evidence e m to respectively construct a Dirichlet distribution as the conjugate prior of the pain category distribution and generate a prediction distribution of the category. The Dirichlet distribution parameter α m of modality m is:

[0091]

[0092] Among them Denote the Dirichlet parameters of the 1st - 4th pain categories in modality m, i.e., the Dirichlet parameter of the kth pain category in modality m Among them Denote the evidence of the kth pain category in modality m;

[0093] (2) Based on the subjective logic theory, for a 4 - classification problem, model the belief functions of each modality m Among them, the single - modality class confidence And the single - modality classification uncertainty u m , for modality m, the 4 class confidences And the classification uncertainty u m Are all non - negative and sum to 1, as shown in equation (2):

[0094]

[0095] According to the evidence e m Calculate the single - modality classification uncertainty u m And the single - modality class confidence b m , as shown in equations (3) and (4):

[0096]

[0097]

[0098] Among them, the Dirichlet strength

[0099] Among them Denote the Dirichlet parameter of the kth pain category in modality m; Denote the evidence of the kth pain category in modality m;

[0100] (3) Calculate the class prediction probability of modality m through the mean of the Dirichlet distribution of modality m Among them Denote the prediction probabilities of modality m for the 1st - 4th pain categories, as shown in equation (5):

[0101]

[0102] Among them Denote the prediction probability of modality m for the kth category.

[0103] Decision fusion module: fuse the belief functions of each modality To generate a joint belief function Among them b = [b 1,b 2 ,b 3 ,b 4 is the class confidence of the overall model, u is the classification uncertainty of the overall model, and b is the credibility that the overall model has for the input sample belonging to pain class k k >0, u>0; Finally, according to the joint belief function Calculate the joint evidence e = [e 1 ,e 2 ,e 3 ,e 4 and the corresponding joint Dirichlet distribution parameters, and output the final pain class prediction probability; In the decision fusion module, specifically,

[0104] (1) Use Dempster's fusion rule, that is, Dempster's fusion rule, to fuse the belief functions of each modality obtained to generate a joint belief function The joint belief function includes the class confidence b = [b 1 ,b 2 ,b 3 ,b 4 and the classification uncertainty u, where, b 1 ,b 2 ,b 3 ,b 4 represents the credibility that the model input sample belongs to the 1st - 4th pain classes. The fusion rules are shown in equations (6), (7), and (8):

[0105]

[0106] Among them, represents the belief functions when modality m is 1, 2, and 3, representing the facial expression modality, the posture modality, and the electrocardiogram modality respectively; ⊕ represents Dempster's fusion, that is, Dempster's fusion; respectively represent the credibility that the input sample belongs to pain class k when modality m is 1, 2, and 3; u 1 ,u 2 ,u 3 respectively represent the classification uncertainties of modalities 1, 2, and 3; b k represents the credibility that the overall model has for the input sample belonging to pain class k, represents the confidence conflict between different modalities;

[0107] (2) According to the joint belief function calculate the joint evidence e = [e 1 ,e 2 ,e 3 ,e4 , where e 1 , e 2 , e 3 , e 4 represents the combined evidence of the 1st - 4th pain categories, constructs the corresponding combined Dirichlet distribution, and calculates the parameters α = [α 1 , α 2 , α 3 , α 4 and the combined Dirichlet distribution intensity S, where α 1 , α 2 , α 3 , α 4 represents the Dirichlet parameters of the 1st - 4th pain categories, as shown in equations (9) and (10):

[0108]

[0109] where e k represents the combined evidence of the kth pain category, and α k represents the Dirichlet parameter of the kth pain category;

[0110] (3) Calculate the final predicted probability p = [p 1 , p 2 , p 3 , p 4 of the pain category based on the combined Dirichlet distribution parameters α and the combined Dirichlet distribution intensity S, where p 1 , p 2 , p 3 , p 4 respectively represent the predicted probabilities of the model for the 1st - 4th pain categories, as shown in equation (11):

[0111]

[0112] where p k is the predicted probability that the input sample belongs to the kth pain category.

[0113] S3. Train the neonatal pain classification model based on multimodal evidence using the neonatal pain multimodal dataset, and adjust the parameters of the neonatal pain classification model based on multimodal evidence to the optimal through the error backpropagation algorithm to obtain the trained model.

[0114] In step S3, the following loss function is used to adjust the parameters of the pain classification model to the optimal through the error backpropagation algorithm:

[0115]

[0116] Among them, y k represents the true probability that the training samples include neonatal video clips and the electrocardiogram signals synchronized with the video belong to the k-th category, and p k is the predicted probability that the training input samples belong to the k-th pain category. represents the predicted probability of modality m for the k-th category, S is the intensity of the joint Dirichlet distribution, and S m is the Dirichlet intensity, and the loss function consists of two terms. The first term is the joint classification loss term, which is used to constrain the overall decision-making accuracy of the model. The second term is the single-modal classification loss term, which is used to constrain the classification accuracy of a single modality during independent learning. λ is a hyperparameter that adjusts the weight of the single-modal loss and takes values between 0 and 1.

[0117] S4. Use the trained model to perform pain classification on newly input neonatal videos and electrocardiogram signals.

[0118] The embodiment also provides a neonatal pain classification system based on multi-modal evidence for implementing the method described in any one of the above, including a data acquisition module, a model construction module, a model training module, and a pain classification module, as Figure 5 follows:

[0119] Data acquisition module: Collect neonatal video clips and electrocardiogram signals synchronized with the video, and construct a neonatal pain multi-modal data set containing pain labels;

[0120] Model construction module: Construct a neonatal pain classification model based on multi-modal evidence. The neonatal pain classification model based on multi-modal evidence includes a data processing module, an evidence extraction module, a class confidence and classification uncertainty estimation module, and a decision fusion module; among them:

[0121] Data processing module: Preprocess the input neonatal videos and electrocardiogram signals to obtain an expression image sequence, a pose image sequence, and an electrocardiogram signal spectrogram;

[0122] Evidence extraction module: Extract the evidence e m supported by each modality m for pain classification from the expression image sequence, the pose image sequence, and the electrocardiogram signal spectrogram, including expression modality evidence e 1 , pose modality evidence e 2 , and electrocardiogram modality evidence e 3 ;

[0123] Class confidence and classification uncertainty estimation module: For the expression modality evidence e 1 , pose modality evidence e 2 , and electrocardiogram modality evidence e3 Construct the Dirichlet distribution, i.e., the Dirchlet distribution, respectively, to generate the belief functions of each modality. where b m is the class confidence of modality m, and u m is the classification uncertainty of modality m;

[0124] Decision fusion module: fuse the belief functions of each modality to generate a joint belief function where b is the class confidence of the overall model, and u is the classification uncertainty of the overall model; finally, calculate the joint evidence e of each class and the corresponding joint Dirichlet distribution parameters according to the joint belief function and output the final pain class prediction probability;

[0125] Model training module: use the neonatal pain multi-modal dataset to train the neonatal pain classification model based on multi-modal evidence, and adjust the parameters of the neonatal pain classification model based on multi-modal evidence to the optimal through the error backpropagation algorithm to obtain the trained model;

[0126] Pain classification module: use the trained model to classify the pain of newly input neonatal videos and electrocardiogram signals.

[0127] This neonatal pain classification method and system based on multi-modal evidence collects neonatal video clips and electrocardiogram signals, and establishes a neonatal pain multi-modal dataset containing pain class labels; constructs a neonatal pain classification model based on multi-modal evidence, including a data processing module, an evidence extraction module, a class confidence and classification uncertainty estimation module, and a decision fusion module; uses the samples in the neonatal pain multi-modal dataset to train the model; uses the trained model to classify the pain of newly input neonatal videos and electrocardiogram signals. By combining the Dirichlet distribution with the evidence theory, this method predicts the pain class probability distribution and classification uncertainty of multi-modal data, including facial expression images, posture images, and electrocardiogram signals, respectively, and fuses the prediction results, providing a classification uncertainty assessment while realizing pain classification, and enhancing the reliability and interpretability of the model.

[0128] This neonatal pain classification method and system based on multi-modal evidence can provide a classification uncertainty assessment while realizing pain classification, enhancing the reliability and interpretability of the model. It can overcome the limitations of single-modal pain classification methods. Different from the identification methods with a single data source, the present invention comprehensively considers multiple modalities of data, fully takes into account various clinical manifestations of neonatal pain, can improve the accuracy of pain assessment, and improve the pain classification accuracy.

[0129] The neonatal pain classification method and system based on multi-modal evidence adopts a multi-modal fusion method based on Dempster-Shafer evidence theory, which can dynamically evaluate the sample quality and uncertainty of different modalities, avoiding the influence brought by uneven multi-modal data quality in complex medical environments. This method not only improves the robustness of data fusion, but also can give an uncertainty evaluation of the current decision during final classification, providing better model interpretability and decision credibility. By combining evidence theory with subjective logic theory, this method can dynamically evaluate the quality of each modal sample, improving the accuracy and robustness of neonatal pain classification.

[0130] The above are only specific implementation manners in the present invention, but the protection scope of the present invention is not limited thereto. Any transformation or replacement that can be understood and conceived by those familiar with the technology within the technical scope disclosed by the present invention should be covered within the scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for classifying neonatal pain based on multimodal evidence, characterized in that: The following steps are included: S1, collect neonatal video clips and ECG signals synchronized with the videos, and construct a neonatal pain multimodal dataset containing pain labels; S2. Construct a neonatal pain classification model based on multimodal evidence. The neonatal pain classification model based on multimodal evidence includes a data processing module, an evidence extraction module, a category confidence and classification uncertainty estimation module, and a decision fusion module; wherein: Data processing module: pre-process the input newborn video and ECG signal to obtain expression image sequence, posture image sequence and ECG signal spectrum diagram; Evidence extraction module: extracts evidence e from each modality m that supports pain classification from expression image sequences, posture image sequences, and ECG signal spectrograms. m , including expression modality evidence 1 , Posture modal evidence 2 and ECG modality evidence 3 ; Category confidence and classification uncertainty estimation module: extract the expression modality evidence 1 , Posture modal evidence 2 , ECG modality evidence 3 Construct Dirichlet distribution and generate belief functions for each mode where b m is the category confidence of mode m, u m is the classification uncertainty of mode m; Decision fusion module: The belief function of each modality Fusion generates joint belief function Where b is the category confidence of the model as a whole, and u is the classification uncertainty of the model as a whole; finally, according to the joint belief function Calculate the joint evidence e of each category and the corresponding joint Dirichlet distribution parameters, and output the final prediction probability of the pain category; S3. Using the neonatal pain multimodal dataset to train a neonatal pain classification model based on multimodal evidence, adjusting the parameters of the neonatal pain classification model based on multimodal evidence to the optimal value through an error back propagation algorithm to obtain a trained model; S4. Use the trained model to classify pain in newly input neonatal videos and ECG signals.

2. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 1, characterized in that: In step S2, the data processing module includes an expression data processing module, a posture data processing module and an electrocardiogram data processing module. Expression data processing module: extract a video frame sequence of length q frames at the same time interval from the input newborn video within t seconds, where t is between 20 and 40, perform face detection, cropping, alignment and normalization on each image in the video frame sequence, and obtain an expression image sequence of length q frames. Where w1 and h1 are the width and height of the expression image; Posture data processing module: extract a video frame sequence of length q frames from the newborn video at the same time interval, perform limb detection, cropping, alignment and normalization on each image in the video frame sequence, and obtain a posture image sequence of length q frames Where w2 and h2 are the width and height of the posture image; ECG data processing module: denoises, corrects baseline drift and normalizes the input n-lead ECG signal to obtain a one-dimensional ECG signal with a length of l Perform short-time Fourier transform to obtain the ECG signal spectrum Where w3 and h3 are the width and height of the ECG spectrum.

3. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 1, characterized in that: The evidence extraction module includes expression evidence extraction module, posture evidence extraction module and electrocardiogram evidence extraction module. Expression evidence extraction module: The input expression image sequence passes through p groups of convolution modules and hybrid attention modules, and then passes through the global average pooling layer, the fully connected layer and the Softplus activation layer to extract the expression modality evidence e 1 , where p is an integer between 4 and 10; Posture evidence extraction module: The input posture image sequence passes through p groups of convolution modules and hybrid attention modules, and then passes through the global average pooling layer, the fully connected layer and the Softplus activation layer to extract the posture modal evidence e 2 , where p is an integer between 4 and 10; ECG evidence extraction module: extract ECG modal evidence from the input ECG signal spectrum 3 .

4. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 3, characterized in that: The expression evidence extraction module and the posture evidence extraction module use 3D convolutional neural networks with the same structure but different parameters.

5. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 3, characterized in that: The ECG evidence extraction module includes a sequentially connected 2D convolutional layer, a normalization layer, an activation layer, a pooling layer, a hybrid attention module, a global pooling layer, a fully connected layer and a Softplus activation layer. The input ECG signal spectrum I3 is extracted by the 2D convolutional layer. The ECG features are sequentially passed through the normalization layer, the activation layer and the pooling layer to output the feature map F. The feature map F is enhanced in the channel domain and the spatial domain by the hybrid attention module, and then passes through the global pooling layer and the fully connected layer in turn. The Softplus activation layer uses the Softplus activation function to map the features output by the previous layer into ECG modal evidence e 3 .

6. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 5, characterized in that: The hybrid attention module consists of a serially connected channel domain attention module and a spatial domain attention module. Channel domain attention module: First, the input feature map Where w and h are the width and height of the feature map, respectively, and c is the number of channels of the feature map. Maximum pooling and average pooling are performed to aggregate spatial information and generate two intermediate feature maps respectively. Then the two intermediate feature maps are added and input into the fully connected layer, and the channel attention weight map is generated after Sigmoid activation. Finally, the channel attention weight map F CA Perform channel-by-channel product with the input feature map F to obtain the enhanced feature map Spatial domain attention module: First, the input feature map Perform maximum pooling and average pooling along the channel dimension to generate two intermediate feature maps respectively The two intermediate feature maps are then stacked and convolved to generate a spatial attention weight map after Sigmoid activation. Finally, the spatial attention weight map F SA With the input feature map Perform point multiplication to obtain the enhanced feature map Where w and h are the width and height of the feature map, respectively, and c is the number of channels of the feature map.

7. A method for classifying neonatal pain based on multimodal evidence according to any one of claims 1 to 6, characterized in that: In the category confidence and classification uncertainty estimation module, specifically, (1) Each modality m uses the extracted evidence e m Construct Dirichlet distributions as conjugate priors for pain category distributions and generate predictive distributions for categories. The Dirichlet distribution parameter α of mode m is m : in, The Dirichlet parameters of the 1st to Kth pain categories in mode m are the Dirichlet parameters, and the Dirichlet parameters of the kth pain category in mode m are in, represents the evidence for the kth pain category in modality m; (2) Based on subjective logic theory, for the K classification problem, the belief function of each mode m is modeled separately Among them, the unimodal category confidence and the unimodal classification uncertainty u m , for mode m, K categories of confidence and classification uncertainty u m are all non-negative numbers and their sum is 1, as shown in formula (2): According to the evidence m Calculate the unimodal classification uncertainty u m and unimodal category confidence b m , as shown in equations (3) and (4): Among them, Dirichlet intensity is Dirichlet intensity in, represents the Dirichlet parameter of the kth pain category in mode m; represents the evidence for the kth pain category in modality m; (3) The category prediction probability of mode m is obtained by calculating the mean of the Dirichlet distribution of mode m in, represents the predicted probability of mode m for the 1st-Kth pain category, as shown in formula (5): in, Represents the predicted probability of mode m for the kth category.

8. A method for classifying neonatal pain based on multimodal evidence according to any one of claims 1 to 6, characterized in that: In the decision fusion module, specifically, (1) Using the Dempster fusion rule, the belief functions of each mode are obtained Fusion generates joint belief function Joint Belief Function Including the category confidence of the model as a whole b = [b1, b2, ..., b K ] and classification uncertainty u, where b1,b2,…,b K It indicates the credibility of the model input sample belonging to the 1-Kth pain category. The fusion rules are shown in (6), (7), and (8): in, Indicates the belief function when the modality m is 1, 2, and 3, respectively, representing the expression modality, posture modality, and electrocardiogram modality; It means Dempster fusion, that is, Dempster fusion; They represent the credibility of the input sample belonging to pain category k when the mode m is 1, 2, and 3 respectively; u 1 、u 2 、u 3 Represents the classification uncertainty of modes 1, 2, and 3 respectively; b k Indicates the overall model's confidence that the input sample belongs to pain category k, Indicates confidence conflicts between different modalities; (2) Based on the joint belief function Calculate the joint evidence for each category e = [e1, e2, …, e K ], where e1, e2, …, e K Represent the joint evidence of the 1-K pain categories, construct the corresponding joint Dirichlet distribution, and calculate the parameters of the joint Dirichlet distribution α = [α1, α2, …, α K ] and the joint Dirichlet distribution strength S, where α1, α2, …, α K represents the Dirichlet parameters of the 1-Kth pain categories, as shown in (9) and (10): Among them, e k represents the joint evidence of the k-th pain category, α k represents the Dirichlet parameter of the kth pain category; (3) The final pain category prediction probability p = [p1, p2, …, p K ], where p1, p2, …, p K They represent the model's predicted probabilities for the 1st to Kth pain categories, as shown in formula (11): Among them, p k is the predicted probability that the input sample belongs to the kth pain category.

9. The method for classifying neonatal pain based on multimodal evidence as claimed in claim 8, characterized in that: In step S3, the following loss function is used The parameters of the pain classification model are adjusted to the optimal value through the error back propagation algorithm: Among them, y k represents the true probability that the training sample includes the video clip of the newborn and the ECG signal synchronized with the video and belongs to the kth category, p k is the predicted probability that the training input sample belongs to the kth pain category, represents the predicted probability of mode m for the kth category, S is the strength of the joint Dirichlet distribution, S m is the Dirichlet intensity, the loss function It consists of two items. The first item is the joint classification loss item, which is used to constrain the overall decision accuracy of the model. The second item is the single-modal classification loss item, which is used to constrain the classification accuracy of a single modality during independent learning. λ is a hyperparameter for adjusting the single-modal loss weight, which takes values ​​between 0 and 1.

10. A neonatal pain classification system based on multimodal evidence implementing the method of any one of claims 1 to 9, characterized in that: It includes data acquisition module, model building module, model training module and pain classification module. Data acquisition module: collects neonatal video clips and ECG signals synchronized with the videos, and constructs a neonatal pain multimodal dataset containing pain labels; Model building module: construct a neonatal pain classification model based on multimodal evidence. The neonatal pain classification model based on multimodal evidence includes a data processing module, an evidence extraction module, a category confidence and classification uncertainty estimation module, and a decision fusion module; among which: Data processing module: pre-process the input newborn video and ECG signal to obtain expression image sequence, posture image sequence and ECG signal spectrum diagram; Evidence extraction module: extracts evidence e from each modality m that supports pain classification from expression image sequences, posture image sequences, and ECG signal spectrograms. m , including expression modality evidence 1 , Posture modal evidence 2 and ECG modality evidence 3 ; Category confidence and classification uncertainty estimation module: extract the expression modal evidence from each modality 1 , Posture modal evidence 2 , ECG modality evidence 3 Construct Dirichlet distribution and generate belief functions for each mode where b m is the category confidence of mode m, u m is the classification uncertainty of mode m; Decision fusion module: The belief function of each modality Fusion generates joint belief function Where b is the category confidence of the model as a whole, and u is the classification uncertainty of the model as a whole; finally, according to the joint belief function Calculate the joint evidence e of each category and the corresponding joint Dirichlet distribution parameters, and output the final prediction probability of the pain category; Model training module: Use the neonatal pain multimodal dataset to train the neonatal pain classification model based on multimodal evidence, and adjust the parameters of the neonatal pain classification model based on multimodal evidence to the optimal value through the error back propagation algorithm to obtain the trained model; Pain classification module: Use the trained model to classify pain in newly input neonatal videos and ECG signals.

Citation Information

Patent Citations

  • A method for assessing the degree of neonatal pain

    CN109124604A

  • Systems and methods of pain treatment

    CN113196410A

  • Intelligent ICU nursing system based on patient video recognition

    CN113257440A

  • Target detection method based on robust sampling and mixed attention pyramid

    CN114841244A

  • Image classification model training method and device and storage medium

    CN117809126A

Cited By

  • Multi-modal feature dynamic fusion method and system based on uncertainty estimation

    CN120257217A