Newborn pain classification and identification method and system based on expression and posture feature fusion
By integrating the expression and posture characteristics of the newborn, and using the pain classification method of self-attention mechanism and cross-attention mechanism, the problem of low accuracy of pain classification in newborns is solved, achieving higher classification accuracy and robustness.
Patent Information
- Application Number
- CN202510098748.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the accuracy of pain classification in neonatal babies is low, the evaluation of single-modal data is limited, and it is difficult to consider modal features and complementarity in multimodal fusion.
A pain classification method for neonatal children based on the fusion of expression and posture features is adopted. Through data processing, multimodal feature extraction, self-attention mechanism and feature fusion module, a pain classification model is constructed to improve the weight and complementarity of modal features.
It significantly improves the accuracy and robustness of pain classification in neonatal children, overcomes the limitations of single-modal information, and provides richer clinical information.
Smart Images

Figure CN120014349A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method and system for classifying and identifying pain in newborns based on the fusion of facial expression and posture features, belonging to the technical field of pain identification. Background Art
[0002] Pain is an uncomfortable feeling, often described as a warning sign from the body that there may be injury or potential damage. The International Society for the Study of Pain defines pain as "a personal experience of an unpleasant, sensory and emotional experience associated with actual or potential tissue damage, or described as being associated with such damage." Among them, newborns have a much lower tolerance for pain than people of other age groups, and newborns cannot express their pain through words. Medical staff usually need to rely on more complex scales for assessment, which involves cumbersome and time-consuming operations and may be subjective. At the same time, clinical compliance is low, resulting in the identification of pain may not be timely enough. Therefore, accurate classification of pain levels is a guarantee for efficient and safe pain management and treatment of newborns during hospitalization.
[0003] Most of the existing intelligent pain assessment technologies are based on a single modality, such as recognizing facial expressions or electrocardiogram signals when in pain. Relying solely on single-modality data for assessment has certain limitations. Compared with single-modality, the fusion of multimodal data (such as expressions, postures, etc.) can provide richer and more extensive information, and has less impact on data missing. However, direct multimodal fusion cannot take into account the important characteristics of the modality itself and improve the complementarity between different modalities.
[0004] How to give full play to the advantages of multimodal fusion in the task of neonatal pain classification while improving the weight of important features of each modality and the complementarity between different modalities to improve the accuracy of the task of neonatal pain classification is still an open topic facing challenges. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for classifying and identifying pain in newborns based on the fusion of facial expression and posture features to solve the problem of low accuracy in classifying pain in newborns in the prior art.
[0006] The technical solution of the present invention is:
[0007] A method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features, comprising the following steps:
[0008] S1. Obtain a neonatal pain dataset, which includes neonatal videos and corresponding pain labels;
[0009] S2. Construct a neonatal pain classification model based on the fusion of expression and posture features. The neonatal pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and the posture data of the input neonatal video to obtain a neonatal expression image sequence and a neonatal posture image sequence respectively; the multimodal feature extraction module extracts expression features F from the input neonatal expression image sequence and the neonatal posture image sequence respectively. a and posture feature F b Output to the self-attention mechanism module; the self-attention mechanism module first inputs the expression feature F a and posture feature F b Perform dimension transformation to obtain the expression feature F after dimension transformation p and posture feature F q , and then introduce the self-attention mechanism to the features after the two-dimensional transformation respectively, and obtain the expression feature F1 and posture feature F2 that are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2, and obtains the first pain feature and the second pain feature with complementary information of expression mode and posture mode, and obtains the spliced feature after splicing, and the spliced feature is flattened to obtain the pain feature vector v; the classification layer classifies and identifies the pain feature vector v to obtain the pain category;
[0010] S3. Using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of expression and posture features are optimized and trained by the error back propagation algorithm to obtain the trained neonatal pain classification model;
[0011] S4. Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
[0012] Furthermore, in step S1, the pain labels include four category labels: calm, mild pain, moderate pain and severe pain.
[0013] Furthermore, in step S2, the data processing module includes an expression data processing unit and a posture data processing unit.
[0014] Expression data processing unit: extract a video frame sequence with a frame length of L from the newborn video, perform face detection and cropping on each image in the video frame sequence to obtain a newborn expression image with a height of H and a width of W, and then align and normalize the newborn expression image to obtain a newborn expression image sequence with a frame length of L;
[0015] Posture data processing unit: extract a video frame sequence with a frame length of L from the newborn video, perform limb detection and cropping on each image in the video frame sequence to obtain a newborn posture image with a height of H and a width of W, and then align and normalize the newborn posture image to obtain a newborn posture image sequence with a frame length of L.
[0016] Further, in step S2, the multimodal feature extraction module includes an expression feature extraction unit and a posture feature extraction unit, the expression feature extraction unit and the posture feature extraction unit respectively include n convolution modules connected in sequence, each convolution module includes one or more 3D convolution layers, a nonlinear activation function layer ReLU and a pooling layer, each 3D convolution layer is connected to a nonlinear activation function layer ReLU, the 3D convolution layer selects m1 k1*k1*k1 convolution kernels to perform convolution operations on the output of the previous layer, and the pooling layer selects k2*k2*k2 pooling kernels to perform downsampling operations on the output of the previous nonlinear activation function layer ReLU, wherein n is selected from 3, 4, 5, 6, 7, and 8, m1 is selected from 64, 128, 256, and 512, k1 is selected from 3, 5, and 7, and k2 is selected from 1 and 2.
[0017] Further, in step S2, the self-attention mechanism module includes an expression self-attention mechanism unit and a posture self-attention mechanism unit;
[0018] The expression of the expression self-attention mechanism unit is:
[0019]
[0020] Among them, Q p , K p 、V p denote the query, key, and value in the expression self-attention mechanism unit, respectively, and F p represents the facial expression features after dimension transformation, They represent three trainable weight matrices, F1 represents expression features, Softmax represents the normalized exponential function, T represents the transposed sign, and C represents the number of channels;
[0021] The expression of the posture self-attention mechanism unit is:
[0022]
[0023] Among them, Q p , K p 、V p denote the query, key, and value in the expression self-attention mechanism unit, respectively, and F q is the posture feature after dimension transformation, They represent three trainable weight matrices respectively, and F2 represents the posture feature.
[0024] Furthermore, in step S2, the feature fusion module introduces two cross attention mechanisms to the input expression feature F1 and posture feature F2 to obtain a first pain feature and a second pain feature having complementary information of expression modality and posture modality, specifically,
[0025]
[0026] Among them, F 1_2 The first pain feature, F 2_1 is the second pain feature, Q2 represents the query in the first cross-attention mechanism, K1 and V1 represent the key and value in the first cross-attention mechanism, respectively. is a trainable weight matrix; Q1 represents the query in the second cross attention mechanism, K2 and V2 represent the key and value in the second cross attention mechanism. is a trainable weight matrix.
[0027] Furthermore, in step S3, when optimizing the parameters of the neonatal pain classification model based on the fusion of expression and posture features through the error back propagation algorithm, the following loss function is used:
[0028]
[0029] in, represents the mean square loss function, K is the number of pain categories, y k Indicates the true value of the newborn video belonging to the kth category of pain label, Indicates the model prediction value that the newborn video belongs to the kth pain label.
[0030] A neonatal pain classification system based on the fusion of facial expression and posture features for implementing any of the above methods comprises a data acquisition module, a model building module, a model training module and a classification and recognition module.
[0031] Data collection module: obtain the neonatal pain dataset, which includes neonatal videos and corresponding pain labels;
[0032] Model construction module: construct a neonatal pain classification model based on the fusion of expression and posture features. The neonatal pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and posture data of the input neonatal video to obtain a neonatal expression image sequence and a neonatal posture image sequence respectively; the multimodal feature extraction module extracts expression features F from the input neonatal expression image sequence and neonatal posture image sequence respectively. a and posture feature F b Output to the self-attention mechanism module; the self-attention mechanism module first inputs the expression feature F a and posture feature F b Perform dimension transformation to obtain the expression feature F after dimension transformation p and posture feature F q , and then introduce the self-attention mechanism to the features after the two-dimensional transformation respectively, and obtain the expression feature F1 and posture feature F2 that are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2, and obtains the first pain feature and the second pain feature with complementary information of expression mode and posture mode, and obtains the spliced feature after splicing, and the spliced feature is flattened to obtain the pain feature vector v; the classification layer classifies and identifies the pain feature vector v to obtain the pain category;
[0033] Model training module: using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of expression and posture features are optimized and trained through the error back propagation algorithm to obtain the trained neonatal pain classification model;
[0034] Classification and recognition module: Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
[0035] The beneficial effects of the present invention are:
[0036] 1. Compared with the existing technology, this method and system for classifying and identifying pain in newborns based on the fusion of expression and posture features can take into account the various manifestations of newborns at the clinical level in many aspects by extracting features from multimodal data including expression and posture, and can provide the model with richer information. Compared with a single data source, it can significantly improve the accuracy of newborn pain classification and effectively overcome the limitation of single information in unimodal pain recognition.
[0037] 2. The present invention adopts the self-attention mechanism to selectively focus on the features in each modality that are more conducive to pain identification, and further enhances the correlation and complementarity between the features of different modalities through the cross-attention mechanism in the feature fusion module, effectively improving the accuracy and robustness of neonatal pain classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flow chart of a method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features according to an embodiment of the present invention;
[0039] Figure 2 is a schematic diagram illustrating a neonatal pain classification and recognition model based on the fusion of facial expression and posture features in an embodiment;
[0040] Figure 3 is a schematic diagram illustrating a neonatal pain classification and recognition system based on the fusion of facial expression and posture features in an embodiment; DETAILED DESCRIPTION
[0041] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0042] The embodiment provides a method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features, comprising the following steps: Figure 1 :
[0043] S1. Obtain a neonatal pain dataset, which includes neonatal videos and corresponding pain labels.
[0044] In step S1, pain labels are divided into four categories, namely calm, mild pain, moderate pain, and severe pain. In practical applications, other pain classifications and category labels may also be used.
[0045] S2. Construct a newborn pain classification model based on the fusion of expression and posture features. The newborn pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and posture data of the input newborn video to obtain a newborn expression image sequence and a newborn posture image sequence respectively; the multimodal feature extraction module extracts expression features from the input newborn expression image sequence and the newborn posture image sequence respectively. and posture features Among them, H1 and W1 represent the height and width of the two feature tensors, C represents the number of channels, and is output to the self-attention mechanism module; the self-attention mechanism module performs dimensional transformation to obtain the expression features after dimensional transformation And the posture features after dimension transformation Among them, N=H1W1, and then the self-attention mechanism is introduced into the features after the two-dimensional transformation respectively, so as to obtain the expression feature F1 and the posture feature F2 which are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms into the input expression feature F1 and the posture feature F2, so as to obtain the first pain feature and the second pain feature with complementary information of the expression modality and the posture modality, and obtain the spliced feature after splicing, and obtain the pain feature vector v by flattening the spliced feature; the classification layer classifies and identifies the pain feature vector v to obtain the pain category.
[0046] In step S2, the data processing module is used to pre-process each input newborn video within t seconds, where t is between 10 and 15. The data processing module includes an expression data processing unit and a posture data processing unit.
[0047] Expression data processing unit: Use FFmpeg software to perform frame processing on the input video, extract a video frame sequence with a frame length of L, such as 60 frames, from the newborn video, use the YOLO model to perform face detection and cropping on each image in the video frame sequence, and obtain a newborn expression image with a height of H and a width of W, such as an image size of 224*224, and then align and normalize the newborn expression image to obtain a newborn expression image sequence with a frame length of L;
[0048] Posture data processing unit: Use FFmpeg software to perform frame processing on the input video, extract a video frame sequence with a frame length of L, such as 60 frames, from the newborn video, use the YOLO model to perform limb detection and cropping on each image in the video frame sequence, and obtain a newborn posture image with a height of H and a width of W, for example, an image size of 224*224, and then align and normalize the newborn posture image to obtain a newborn posture image sequence with a frame length of L.
[0049] In step S2, the multimodal feature extraction module includes an expression feature extraction unit and a posture feature extraction unit. The expression feature extraction unit and the posture feature extraction unit use a 3D convolutional neural network with the same structure but different parameters. The 3D convolutional neural network includes n convolution modules connected sequentially. Taking n as 4 as an example, the following description is made: the first convolution module includes 1 3D convolution layer, a nonlinear activation function layer ReLU and 1 pooling layer. The last three convolution modules include 2 3D convolution layers, 2 nonlinear activation function layers ReLU and 1 pooling layer. The 4 3D convolution layers use 64 3×3×3 convolution kernels, 128 3×3×3 convolution kernels, 256 3×3×3, and 512 3*3*3 convolution kernels to perform convolution operations on the input feature tensor, respectively. The convolution step size is 1, the zero-padding edge length is 1, and the convolution is subjected to ReLU nonlinear mapping. The pooling layer selects a 2×2×2 maximum pooling kernel and performs a downsampling operation on the feature tensor with a step size of 2.
[0050] In step S2, the self-attention mechanism module includes an expression self-attention mechanism unit and a posture self-attention mechanism unit. The self-attention mechanism is introduced to selectively focus on the expression feature F after the dimension transformation. p And the posture feature F after dimension transformation q The important parts of the two modal features are used to extract facial expression features that are more conducive to pain recognition. and posture characteristics
[0051] Expression self-attention mechanism unit, the input is the expression feature F output by the expression feature extraction unit a , the facial expression feature F a Perform dimension transformation to obtain the expression feature F after dimension transformation p , F p Respectively with three trainable weight matrices Multiply them together to get three matrices for calculating expression self-attention Then Q p and the transposed K p Multiply and pass through the Softmax function to get the self-attention matrix Finally, the self-attention matrix A p With V p Multiply to get more pain-recognizing facial features
[0052] The expression of the expression self-attention mechanism unit is:
[0053]
[0054] Among them, Q p , K p 、V p They represent the query, key, and value in the expression self-attention mechanism unit, respectively. p represents the facial expression features after dimension transformation, They represent three trainable weight matrices, F1 represents expression features, Softmax represents the normalized exponential function, T represents the transposed sign, and C represents the number of channels;
[0055] Posture self-attention mechanism unit, the input is the expression feature F output by the posture feature extraction unit b , the posture feature F b Perform dimension transformation to obtain the posture feature F after dimension transformation q , F q Respectively with three trainable weight matrices Multiply them together to get three matrices for calculating posture self-attention Then Q q and the transposed K q Multiply and pass through the Softmax function to get the self-attention matrix Finally, the self-attention matrix A q With V q Multiplying them together yields a more pain-recognizing posture feature F2.
[0056] The expression of the posture self-attention mechanism unit is:
[0057]
[0058] Among them, Q p , K p 、V p denote the query, key, and value in the expression self-attention mechanism unit, respectively, and F q is the posture feature after dimension transformation, They represent three trainable weight matrices, F2 represents posture features, T represents the transpose sign, and C represents the number of channels.
[0059] In step S2, the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2 to obtain the first pain feature with complementary information of expression modality and posture modality and secondary pain characteristics The specific formula is:
[0060]
[0061] in, K1 and V1 represent the key and value in the first criss-cross attention mechanism, Q2 represents the query in the first criss-cross attention mechanism, is a trainable weight matrix; K2 and V2 represent the key and value in the second cross-attention mechanism, Q1 represents the query in the second cross-attention mechanism, is a trainable weight matrix;
[0062] The first pain feature F extracted with complementary information of two modalities 1_2 and the second pain feature F 2_1 Perform splicing to obtain the splicing feature F c , splicing feature F c After flattening, the pain feature vector v is obtained, and the specific expression is:
[0063]
[0064] v=FL(F c )
[0065] in, represents a concatenation operation, and FL represents a flattening operation.
[0066] In step S2, the self-attention mechanism and the cross-attention mechanism are introduced. The self-attention mechanism emphasizes the important features of the modality itself by assigning weights and suppresses the secondary features; the cross-attention mechanism combines different modalities including expression modality and posture modality to help the model understand the relationship between them and improve the model's ability to obtain complementary information between different features. Through the self-attention mechanism and the cross-attention mechanism, the modality's ability to select its own key features is enhanced, and the complementarity between different modalities is improved, thereby improving the accuracy and robustness of neonatal pain classification.
[0067] S3. Using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of facial expression and posture features are optimized and trained through the error back propagation algorithm to obtain the trained neonatal pain classification model.
[0068] In step S3, when optimizing the parameters of the neonatal pain classification model based on the fusion of expression and posture features through the error back propagation algorithm, the following loss function is used:
[0069]
[0070] in, represents the mean square loss function, K is the number of pain categories, y k Indicates the true value of the newborn video belonging to the kth category of pain label, Indicates the model prediction value that the newborn video belongs to the kth pain label.
[0071] S4. Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
[0072] Compared with the existing technology, this method for neonatal pain classification and recognition based on the fusion of expression and posture features can take into account the various manifestations of neonates at the clinical level by extracting features from multimodal data including expression and posture, and can provide the model with richer information. Compared with a single data source, it can significantly improve the accuracy of neonatal pain classification and effectively overcome the limitation of single information in unimodal pain recognition.
[0073] The present invention adopts the self-attention mechanism to selectively focus on the features in each modality that are more conducive to pain identification, and further enhances the correlation and complementarity between the features of different modalities through the cross-attention mechanism in the feature fusion module, effectively improving the accuracy and robustness of neonatal pain classification.
[0074] like Figure 3 The embodiment also provides a neonatal pain classification system based on the fusion of facial expression and posture features to implement any of the above methods, including a data acquisition module, a model building module, a model training module and a classification and recognition module.
[0075] Data collection module: obtain the neonatal pain dataset, which includes neonatal videos and corresponding pain labels;
[0076] Model construction module: construct a neonatal pain classification model based on the fusion of expression and posture features. The neonatal pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and posture data of the input neonatal video to obtain a neonatal expression image sequence and a neonatal posture image sequence respectively; the multimodal feature extraction module extracts expression features F from the input neonatal expression image sequence and neonatal posture image sequence respectively. a and posture feature F b Output to the self-attention mechanism module; the self-attention mechanism module first inputs the expression feature F a and posture feature F b Perform dimension transformation to obtain the expression feature F after dimension transformation p and posture feature F q , and then introduce the self-attention mechanism to the features after the two-dimensional transformation respectively, and obtain the expression feature F1 and posture feature F2 that are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2, and obtains the first pain feature and the second pain feature with complementary information of expression mode and posture mode, and obtains the spliced feature after splicing, and the spliced feature is flattened to obtain the pain feature vector v; the classification layer classifies and identifies the pain feature vector v to obtain the pain category;
[0077] Model training module: using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of expression and posture features are optimized and trained through the error back propagation algorithm to obtain the trained neonatal pain classification model;
[0078] Classification and recognition module: Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
[0079] This method and system for neonatal pain classification and recognition based on the fusion of expression and posture features uses a multimodal fusion method in the neonatal pain classification task, while introducing an attention mechanism. It selectively focuses on the key features of the two modalities through the self-attention mechanism. At the same time, it uses the cross-attention mechanism to enhance the model's ability to extract complementary information from the expression modality and posture modality, effectively captures the correlation information between the two different modalities, and improves the accuracy of pain classification.
[0080] The above are only specific implementations of the present invention, but the protection scope of the present invention is not limited to this. Any person familiar with the technology can understand and think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features, characterized in that: The following steps are included: S1. Obtain a neonatal pain dataset, which includes neonatal videos and corresponding pain labels; S2. Construct a neonatal pain classification model based on the fusion of expression and posture features. The neonatal pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and the posture data of the input neonatal video to obtain a neonatal expression image sequence and a neonatal posture image sequence respectively; the multimodal feature extraction module extracts expression features F from the input neonatal expression image sequence and the neonatal posture image sequence respectively. a and posture feature F b Output to the self-attention mechanism module; the self-attention mechanism module first inputs the expression feature F a and posture feature F b Perform dimension transformation to obtain the expression feature F after dimension transformation p and posture feature F q , and then introduce the self-attention mechanism to the features after the two-dimensional transformation, and obtain the expression feature F1 and posture feature F2 that are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2, and obtains the first pain feature and the second pain feature with complementary information of expression mode and posture mode, and obtains the spliced feature after splicing, and the spliced feature is flattened to obtain the pain feature vector v; The classification layer classifies and identifies the pain feature vector v to obtain the pain category; S3. Using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of expression and posture features are optimized and trained by the error back propagation algorithm to obtain the trained neonatal pain classification model; S4. Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
2. The method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features as claimed in claim 1, characterized in that: In step S1, the pain labels include four category labels: calm, mild pain, moderate pain and severe pain.
3. The method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features as claimed in claim 1, characterized in that: In step S2, the data processing module includes an expression data processing unit and a posture data processing unit. Expression data processing unit: extract a video frame sequence with a frame length of L from the newborn video, perform face detection and cropping on each image in the video frame sequence to obtain a newborn expression image with a height of H and a width of W, and then align and normalize the newborn expression image to obtain a newborn expression image sequence with a frame length of L; Posture data processing unit: extract a video frame sequence with a frame length of L from the newborn video, perform limb detection and cropping on each image in the video frame sequence to obtain a newborn posture image with a height of H and a width of W, and then align and normalize the newborn posture image to obtain a newborn posture image sequence with a frame length of L.
4. A method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features as claimed in any one of claims 1 to 3, characterized in that: In step S2, the multimodal feature extraction module includes an expression feature extraction unit and a posture feature extraction unit, the expression feature extraction unit and the posture feature extraction unit respectively include n convolution modules connected in sequence, each convolution module includes one or more 3D convolution layers, a nonlinear activation function layer ReLU and a pooling layer, each 3D convolution layer is connected to a nonlinear activation function layer ReLU, the 3D convolution layer selects m1 k1*k1*k1 convolution kernels to perform convolution operation on the output of the previous layer, and the pooling layer selects k2*k2*k2 pooling kernels to perform downsampling operation on the output of the previous nonlinear activation function layer ReLU, wherein n is selected from 3, 4, 5, 6, 7, and 8 values, m1 is selected from 64, 128, 256, and 512 values, k1 is selected from 3, 5, and 7 values, and k2 is selected from 1 and 2 values.
5. The method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features according to any one of claims 1 to 3, characterized in that: In step S2, the self-attention mechanism module includes an expression self-attention mechanism unit and a posture self-attention mechanism unit; The expression of the expression self-attention mechanism unit is: Among them, Q p , K p 、V p denote the query, key, and value in the expression self-attention mechanism unit, respectively, and F p represents the facial expression features after dimension transformation, They represent three trainable weight matrices, F1 represents expression features, Softmax represents the normalized exponential function, T represents the transposed sign, and C represents the number of channels; The expression of the posture self-attention mechanism unit is: Among them, Q p , K p 、V p denote the query, key, and value in the expression self-attention mechanism unit, respectively, and F q is the posture feature after dimension transformation, They represent three trainable weight matrices respectively, and F2 represents the posture feature.
6. The method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features according to any one of claims 1 to 3, characterized in that: In step S2, the feature fusion module introduces two cross attention mechanisms to the input expression feature F1 and posture feature F2 to obtain the first pain feature and the second pain feature with complementary information of expression modality and posture modality, specifically, Among them, F 1_2 The first pain feature, F 2_1 is the second pain feature, Q2 represents the query in the first cross-attention mechanism, K1 and V1 represent the key and value in the first cross-attention mechanism, respectively. is a trainable weight matrix; Q1 represents the query in the second cross attention mechanism, K2 and V2 represent the key and value in the second cross attention mechanism. is a trainable weight matrix.
7. The method for classifying and identifying pain in newborns based on the fusion of facial expression and posture features according to any one of claims 1 to 3, characterized in that: In step S3, when optimizing the parameters of the neonatal pain classification model based on the fusion of expression and posture features through the error back propagation algorithm, the following loss function is used: in, represents the mean square loss function, K is the number of pain categories, y k Indicates the true value of the newborn video belonging to the kth category of pain label, Indicates the model prediction value that the newborn video belongs to the kth pain label.
8. A neonatal pain classification system based on the fusion of facial expression and posture features that implements the method described in any one of claims 1 to 7, characterized in that: It includes data acquisition module, model building module, model training module and classification and recognition module. Data collection module: obtain the neonatal pain dataset, which includes neonatal videos and corresponding pain labels; Model construction module: construct a neonatal pain classification model based on the fusion of expression and posture features. The neonatal pain classification model based on the fusion of expression and posture features includes a data processing module, a multimodal feature extraction module, a self-attention mechanism module, a feature fusion module and a classification layer. The data processing module processes the expression data and posture data of the input neonatal video to obtain a neonatal expression image sequence and a neonatal posture image sequence respectively; the multimodal feature extraction module extracts expression features F from the input neonatal expression image sequence and neonatal posture image sequence respectively. a and posture feature F b Output to the self-attention mechanism module; the self-attention mechanism module first inputs the expression feature F a and posture feature F b Perform dimension transformation to obtain the expression feature F after dimension transformation p and posture feature F q , and then introduce the self-attention mechanism to the features after the two-dimensional transformation, and obtain the expression feature F1 and posture feature F2 that are more conducive to pain recognition; the feature fusion module introduces two cross-attention mechanisms to the input expression feature F1 and posture feature F2, and obtains the first pain feature and the second pain feature with complementary information of expression mode and posture mode, and obtains the spliced feature after splicing, and the spliced feature is flattened to obtain the pain feature vector v; The classification layer classifies and identifies the pain feature vector v to obtain the pain category; Model training module: using the neonatal pain dataset, the parameters of the neonatal pain classification model based on the fusion of expression and posture features are optimized and trained through the error back propagation algorithm to obtain the trained neonatal pain classification model; Classification and recognition module: Use the trained neonatal pain classification model to perform pain classification and recognition on the newly input test video.
Citation Information
Patent Citations
Neonatal pain identification method based on facial expression analysis
CN107491740A
Child pain multi-modal data fusion evaluation method based on deep learning
CN118452821A
Multi-modal dialogue emotion recognition method
CN119293740A