Pain expression assessment method based on Batchformer
By introducing the Batchformer module into the neural network for facial pain expression evaluation, the local features of different samples are weighted and combined to generate sample relationship characteristics, which solves the problem of data imbalance and insufficient attention to the boundary ambiguity samples, and significantly improves the accuracy of pain level evaluation.
Patent Information
- Application Number
- CN202311130360.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-09-04
AI Technical Summary
The prior art has data imbalance in the evaluation of facial pain expressions, especially the small amount of high-grade pain data, which leads to poor evaluation performance and insufficient attention to boundary ambiguity samples.
Using the Batchformer-based pain expression evaluation method, the Batchformer module is introduced into the neural network, and the local features of different samples in the same batch are weighted and combined to generate sample relationship characteristics, supporting the propagation of feature information of different samples in the batch, thereby overcoming the problem of insufficient attention to data imbalance and boundary ambiguity of samples.
This improves the accuracy of facial pain assessment, overcomes the problems of data imbalance and insufficient learning of boundary ambiguity samples, and significantly improves the performance of pain level assessment.
Smart Images

Figure CN117197868B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and further relates to a pain expression assessment method based on Batchformer in the field of pattern recognition technology. The present invention can be used to extract facial expression image features from facial expression images of human faces for pain assessment. Background Art
[0002] Facial pain expression assessment technology uses computers to extract facial expression image features, adaptively extracts discriminative features based on the mapping relationship between data and labels, and mines the mapping relationship between different facial representations and pain levels, thereby achieving the purpose of assessing the pain of facial expression image features. In the facial pain assessment task, since high-level pain is unbearable for patients or volunteers providing pain data, high-level pain data is far less than other pain level data, making the facial pain dataset have the characteristics of data imbalance. However, using unbalanced data for training and learning is very challenging, which makes the performance of facial pain assessment tasks not very ideal. This requires the network to do some processing on the unbalanced data, and existing technical solutions do not take into account the factor of data imbalance.
[0003] Pau Rodriguez et al. proposed a method for evaluating the pain of facial expression image features using a deep learning model trained with raw frames in their paper "Deep pain: Exploiting long short-term memory networks for facial expression classification," (IEEE, 2022, pp. 3314–3324). The implementation steps of this method are: ① Use a convolutional neural network to learn facial features from VGG_Faces; ② Link the features to long short-term memory to utilize the temporal relationship between video frames; ③ Perform facial pain assessment. The shortcoming of this method is that it only utilizes the temporal relationship between video frames, ignoring the fact that in practice, pain frames account for a small proportion of the entire video frames, resulting in the problem that the number of pain-free images is far greater than the number of images of other pain levels.
[0004] Dong Huang et al. proposed a method for evaluating the pain of facial expression image features using a pain-aware multistream convolutional neural network for pain estimation in their paper "Pain-awareness multistream convolutional neural network for pain estimation," (Journal of Electronic Imaging, 2019, pp. 1, 7). The implementation steps of this method are: ① Separate the area most relevant to pain expression; ② Use multi-stream CNN to learn the corresponding pain perception features; ③ The features are combined into adaptive weights. The shortcoming of this method is that the corresponding pain perception features are learned only based on the multi-stream convolutional neural network, resulting in a data imbalance problem in the pain dataset for evaluating the pain of facial expression image features.
[0005] Shaoxing Cui et al. proposed a method for evaluating the pain of facial expression image features based on a pain assessment model based on a multi-scale attention network in their paper "Multi-Scale Regional Attention Networks for Pain Estimation" (New York, NY, USA: Association for Computing Machinery, 2021, p. 1–8.). The implementation steps of this method are: ① Crop the input image to obtain k images of different scales; ② Input the k images into k convolutional neural networks to extract features; ③ Establish a self-attention module, input the extracted k features into k fully connected layers and sigmoid activation functions to calculate the attention coefficient of each feature, so as to weight the features by the attention coefficient; ④ Establish a mutual attention module, cascade the k features together, input them into k fully connected layers and sigmoid activation functions to calculate the attention coefficient of each feature, so as to explore the relationship between features of different scales and weight the features by the mutual attention coefficient. The shortcomings of this method are that although it uses the attention mechanism to weight features of different scales, it lacks the correlation analysis between pain-related local facial motor units and their intensity. At the same time, it does not consider the ambiguity in facial expressions. It is difficult to classify boundary samples using a classification model, which easily leads to low prediction accuracy. Summary of the invention
[0006] In view of the deficiencies of the above-mentioned prior art, the present invention proposes a pain expression assessment method based on Batchformer, so as to solve the data imbalance that the number of pain-free images is far greater than the number of images of other pain levels, and the problem of insufficient attention to samples with fuzzy boundaries.
[0007] The technical idea for achieving the purpose of the present invention is: when the present invention performs pain feature learning of facial expression images on a neural network, the batchformer module in the constructed pain expression evaluation neural network is used to perform weighted combination of local features of different samples in the same batch to obtain sample relationship features of the batch, support the propagation of feature information of different sample images in the batch, and solve the problem that features of different sample images are independent of each other; the sample relationship features of a batch are used to train the neural network, so that the majority class samples in the batch can also contribute to the learning of the minority class samples, overcome the problem of insufficient number of minority class samples, and solve the problem of data imbalance in the prior art.
[0008] The specific steps for achieving the purpose of the present invention are as follows:
[0009] Step 1: Generate training set:
[0010] A sample set is formed by facial images of at least 25 subjects, which include four types of pain expressions: no pain, mild pain, moderate pain, and severe pain. Each image in the sample set is normalized, and all normalized samples and their corresponding pain expression labels form a training set.
[0011] Step 2, build a neural network consisting of a feature extractor, a batchformer module, and a pain level classifier connected in series and set the parameters of the network;
[0012] Step 3: Learn and train the neural network on the pain features of facial expression images:
[0013] The sample images of the training set are input into the neural network in batches, and the local features of each image block after segmentation are output by the feature extractor; the weighted combined features of different samples in each batch are calculated by the batchformer module, and the sample relationship features of the batch are output; the pain level classifier outputs the pre-classification results and predicted labels of each sample image in each batch, calculates the cross entropy loss function and performs back propagation, uses the Adam optimizer to optimize the training network, and iteratively updates the network parameters until the network loss function converges to obtain a trained neural network;
[0014] Step 4: Evaluate the pain expression of the facial image:
[0015] The same method as step 1 is used to normalize the face images to be evaluated, and the normalized images are input into the trained neural network to output the pain level corresponding to each face image.
[0016] Compared with the prior art, the present invention has the following advantages:
[0017] First, the present invention adds a Batchformer module between the feature extractor and the classifier of the traditional pain assessment neural network to solve the problem of unbalanced data processing in the prior art. There is no need to change the original feature extractor and classifier network structure, which overcomes the defect of the prior art that a complex neural network needs to be designed. The present invention has the advantage of a simple network structure and improves the efficiency of pain level assessment corresponding to facial images.
[0018] Second, the present invention uses a batchformer module to calculate the weighted combination features of different samples in each batch, and outputs the sample relationship features of the batch, so that the majority class samples in the same batch can also contribute to the learning of the minority class samples, overcoming the problem of poor evaluation performance caused by insufficient learning of minority class samples in the prior art, so that the present invention improves the accuracy of facial pain assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of the present invention;
[0020] Figure 2 Schematic diagram of the Batchformer module constructed for the present invention. DETAILED DESCRIPTION
[0021] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0022] Reference Figure 1 , further describing the implementation steps of the embodiment of the present invention.
[0023] Step 1: Generate training set and test set.
[0024] Step 1.1, an embodiment of the invention is to collect and annotate facial images of 25 subjects from the UNBC-McMaster Shoulder Pain Expression Archive Database, which are four types of facial expression images, namely painless, mild pain, moderate pain, and severe pain, including 40,029 painless images, 5,260 mild pain images, 2,456 moderate pain images, and 653 severe pain images. A total of 48,398 facial expression images were collected to form a sample set.
[0025] Step 1.2, use a sampling resolution of 224×224 to perform bilinear sampling on each image in the sample set, and normalize the sampled images to obtain a normalized sample set.
[0026] In step 1.3, based on the patient ID, the facial image samples of the 25 subjects in the normalized sample set are divided into 25 parts, and a 25-fold cross-validation is performed, in which each validation uses the face of one subject as the test set and the facial images of the remaining subjects as the training set.
[0027] Step 2: Establish a pain assessment neural network based on Batchformer.
[0028] Build a pain assessment neural network consisting of a feature extractor, a Batchformer module, and a pain level classifier connected in series.
[0029] The structure of the feature extractor is composed of an input layer, a convolution layer, a maximum pooling layer, a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block, a ninth convolution block, a tenth convolution block, an eleventh convolution block, a twelfth convolution block, a thirteenth convolution block, a fourteenth convolution block, a fifteenth convolution block, and a sixteenth convolution block in series. The dimension size of the input layer is set to b×3×224×224, and the value of b is equal to the number of samples selected by the convolutional neural network input once. In the embodiment of the present invention, b=32. The number of convolution kernels of the convolution layer is set to 64, the convolution kernel size is set to 7×7, and the step size is set to 2; the pooling window of the maximum pooling layer is set to 3×3, and the step size is set to 2.
[0030] The structures of the first to sixteenth convolution blocks in the feature extractor are the same, and are composed of a 1×1 convolution layer, a 3×3 convolution layer, a 1×1 convolution layer, a normalization layer, and an activation layer connected in series in sequence. The number of convolution kernels of the convolution layers in the first to third convolution blocks is set to 64, 64, and 256, respectively, the number of convolution kernels of the convolution layers in the fourth to seventh convolution blocks is set to 128, 128, and 512, respectively, the number of convolution kernels of the convolution layers in the eighth to thirteenth convolution blocks is set to 256, 256, and 1024, respectively, and the number of convolution kernels of the convolution layers in the fourteenth to sixteenth convolution blocks is set to 512, 512, and 2048, respectively. The step size of the convolution layers in the first to sixteenth convolution blocks is set to 2. The eps of the normalization layer is set to 1e-5, the momentums is set to 0.1, and the activation layer is implemented using the nonlinear activation function ReLu.
[0031] The Batchformer module consists of a transformer unit and a feed-forward fully connected layer in series.
[0032] The Transformer unit is composed of a self-attention layer, a first random dropout layer, a first fully connected layer, a second random dropout layer, a second fully connected layer, a third random dropout layer, a first normalization layer, and a second normalization layer in series. The self-attention layer contains three fully connected layers. The number of attention heads in its multi-head attention mechanism is the number of segmented images, denoted by nhead. In the embodiment of the present invention, nhead=32. The first to third random dropout layers all use the dropout function to set each neuron to 0 with a probability of ρ. In the embodiment of the present invention, ρ=0.1. The inputs of the first and second fully connected layers are set to 2048 dimensions, and the outputs are set to 2048 dimensions. The first and second normalization layers use Formula, x is the normalization layer input, y is the normalization layer output, E[x] is the expectation, Var[x] is the variance, eps is the stability coefficient, and in the embodiment of the present invention, eps=1e-5.
[0033] The feed-forward fully connected layer is composed of the first fully connected layer and the second fully connected layer in series. The input and output dimensions of the first and second fully connected layers are set to 2048.
[0034] The pain level classifier is composed of a fully connected layer and an output layer connected in series. The input of the fully connected layer is set to 1024 dimensions and the output is set to 4 dimensions. The output layer is implemented by a softmax function, which maps the output of the fully connected layer of the classifier to the probability of belonging to a category, and the sum of the probabilities of all categories is 1.
[0035] The probability calculation formula for the output category of the fully connected layer is as follows:
[0036]
[0037] Among them, Softmax(z a ) represents the probability that the input image belongs to the ath category, ∑ is the summation operation, e (·) represents the exponential operation with the natural constant e as the base, z a represents the ath output of the fully connected layer, i∈[1,4], N represents the total number of dimensions of the fully connected layer output, in the embodiment of the present invention, N=4, the first category is labeled 0, i.e. no pain, the second category is labeled 1, i.e. mild pain, the third category is labeled 2, i.e. moderate pain, the fourth category is labeled 3, i.e. severe pain, z c represents the output of the cth dimension of the fully connected layer, c∈[1,N], and in the embodiment of the present invention, c∈[1,4].
[0038] Step 3: Learning and training the pain features of facial expression images on the neural network.
[0039] Step 3.1: Input the sample images of the training set into the neural network in batches, and the feature extractor outputs the local features of each image block after segmentation:
[0040] f ij =E(x ij )
[0041] Among them, x ij represents the jth image block of the i-th sample in a batch, f ij represents the local features of the jth image block of the i-th sample in a batch, and E(·) represents the feature extraction operation. In the embodiment of the present invention, i∈[1,32], j∈[1,32*32].
[0042] Step 3.2, the batchformer module calculates the weighted combined features of different samples in each batch and outputs the sample relationship features of the batch:
[0043]
[0044] Among them, f represents the sample relationship feature of a batch, n represents the total number of image blocks segmented by a batch of samples, softmax(·) represents the normalization operation, and q ij , k ij 、v ij represents the query feature, key feature, and value feature of the jth image block of the ith sample extracted by the three fully connected layers of the self-attention layer in the Transformer unit, d k Represents the normalization coefficient of the stable gradient.
[0045] Reference Figure 2 , further describe step 3.2, f ij represents the jth image block of the i-th sample, i∈[1,m], j∈[1,n*n], f ij The arrows indicate the feature extraction operations of the three fully connected layers of the self-attention layer, q ij , k ij 、v ij f ij The query features, key features, and value features extracted by the three fully connected layers of the self-attention layer are q ij With k ij The arrows indicate the operation of dot multiplication and then division by the normalization coefficient to obtain the attention weight With value feature v ij After weighted summation, we get the sample relationship feature f of the batch.
[0046] Step 3.3, the pain level classifier outputs the pre-classification result and predicted label of each sample image in each batch, calculates the cross entropy loss function and performs back propagation, uses the Adam optimizer to optimize the training network, and iteratively updates the network parameters until the network loss function converges to obtain a trained neural network;
[0047] The cross entropy loss function is as follows:
[0048] L = -log(ρ(y i |x i ))
[0049] Where L represents the loss value between the predicted label output by the pain level classifier and the real pain expression label, log represents the logarithmic operation with base 10, and y i represents the i-th sample x i The corresponding real pain expression label, ρ(y i |x i ) represents the classifier for sample x i The predicted label of .
[0050] Step 4: Evaluate the pain expression of the facial image:
[0051] The facial images to be evaluated are normalized, and the normalized images are input into the trained neural network to output the pain level corresponding to each facial image.
[0052] The effect of the present invention can be further demonstrated through the following simulation.
[0053] 1. Simulation experimental conditions.
[0054] The hardware platform of the simulation experiment of the present invention is as follows: the graphics processor is GeForce GTX 2080Ti GPU, and the video memory is 11G.
[0055] The software platform for the simulation experiment of the present invention is: Windows 10 operating system, Python 3.6, and TensorFlow deep learning development framework.
[0056] The data of the simulation experiment of the present invention are collected from the UNBC-McMaster shoulder pain dataset of the facial pain dataset. The subjects of the UNBC-McMaster shoulder pain expression archive database are 129 adults with shoulder pain. They are tested in their shoulder range of motion and 200 image sequences are collected, with a total of 48,398 frames, including 8,369 pain frames. All images in the dataset are identified by the facial coding system, and the facial AU unit intensity score and PSPI score are provided for each image. Based on the PSPI score, the patient's pain is graded and divided into 17 levels (0-16) according to the intensity of the pain, among which level 0 pain is the weakest and level 16 pain is the strongest. According to the intensity of the pain, it is divided into 4 categories, among which level 0 is painless, level 1 to level 2 is mild pain, level 3 to level 5 is moderate pain, and level 6 to level 16 is severe pain. After classification, 40,029 painless images, 5,260 mild pain images, 2,456 moderate pain images, and 653 severe pain images are obtained. Using a sampling resolution of 224×224, bilinear sampling is performed on each image in the sample set, and the sampled images are normalized. All normalized facial expression images and their corresponding labels constitute the data set in the simulation experiment of the present invention.
[0057] 2. Simulation content and results analysis:
[0058] The simulation experiment of the present invention adopts the present invention and three existing technologies (Deep pain, Multistream CNN, MSRAN) to respectively perform pain assessment on face images in the UNBC-McMaster shoulder pain expression archive database under simulation conditions to obtain the assessment results of each method.
[0059] In the simulation experiment, the three existing technologies used are:
[0060] The Deep pain evaluation method refers to the pain expression evaluation method proposed by Pau Rodriguez et al. in "Deep pain: Exploiting longshort-term memory networks for facial expression classification, IEEE, 2022, pp. 3314–3324", referred to as Deep pain.
[0061] The Multistream CNN evaluation method refers to the pain expression evaluation method proposed by Dong Huang et al. in their paper "Pain-awareness multistream convolutional neural network for pain estimation, Journal of Electronic Imaging, 2019, pp. 1, 7", referred to as Multistream CNN.
[0062] The MSRAN evaluation method refers to the pain expression evaluation method proposed in the paper "Multi-Scale Regional Attention Networks for Pain Estimation, New York, NY, USA: Association for Computing Machinery, 2021, p. 1–8." by Shaoxing Cui et al., abbreviated as MSRAN.
[0063] In order to evaluate the effect of the simulation of the present invention, the following performance evaluation index formula is used to evaluate the evaluation results in the simulation experiment of the present invention, and all the results are plotted in Table 1.
[0064]
[0065] The MAE in Table 1 is the sum of the absolute values of the differences between the target value and the predicted value. It only measures the average modulus of the prediction error, regardless of direction, and ranges from 0 to positive infinity. Low MAE values are often required for quantitative data.
[0066]
[0067] Among them, ∑ is the sum operation, n represents the total number of facial expression images in the test set, i is the sequence number of the facial expression image in the test set, and y i represents the target value, represents the predicted value, Indicates the absolute value of the difference between the target value and the predicted value.
[0068] The MSE in Table 1 is the sum of the squares of the distances between the predicted value and the true value, ranging from 0 to positive infinity. A low MSE value is often required for quantitative data.
[0069]
[0070] Among them, ∑ is the sum operation, n represents the total number of facial expression images in the test set, i is the sequence number of the facial expression image in the test set, and y i represents the target value, represents the predicted value, Represents the square of the distance between the target value and the predicted value.
[0071] Table 1 Comparison of simulation experiment evaluation results of four methods
[0072]
[0073] It can be seen from Table 1 that compared with the best existing technology, the method proposed by the present invention outperforms the other three existing technologies on the UNBC-McMaster shoulder pain archive database.
Claims
1. A pain expression assessment method based on Batchformer, characterized in that: The batchformer module is used to learn the local features of different samples in the same batch, and the weighted combination of features is used to train the neural network. The steps of this evaluation method include the following: Step 1: Generate training set: A sample set is formed by facial images of at least 25 subjects, which include four types of pain expressions: no pain, mild pain, moderate pain, and severe pain. Each image in the sample set is normalized, and all normalized samples and their corresponding pain expression labels form a training set. Step 2, build a neural network consisting of a feature extractor, a batchformer module, and a pain level classifier connected in series and set the parameters of the network; Step 3: Learn and train the neural network on the pain features of facial expression images: The sample images of the training set are input into the neural network in batches, and the feature extractor outputs the local features of each image block after segmentation; the batchformer module calculates the weighted combination features of different samples in each batch and outputs the sample relationship features of the batch; The pain level classifier outputs the pre-classification results and predicted labels of each sample image in each batch, calculates the cross entropy loss function and performs back propagation, uses the Adam optimizer to optimize the training network, and iteratively updates the network parameters until the network loss function converges to obtain a trained neural network; Step 4: Evaluate the pain expression of the facial image: The same method as step 1 is used to normalize the face images to be evaluated, and the normalized images are input into the trained neural network to output the pain level corresponding to each face image.
2. The pain expression assessment method based on Batchformer according to claim 1, characterized in that: The normalization process described in step 1 refers to using a sampling resolution of 224×224 to perform bilinear sampling on each image in the sample set and normalizing the sampled images.
3. The pain expression assessment method based on Batchformer according to claim 1, characterized in that: The structure of the feature extractor described in step 2 is composed of an input layer, a convolution layer, a maximum pooling layer, a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, a sixth convolution block, a seventh convolution block, an eighth convolution block, a ninth convolution block, a tenth convolution block, an eleventh convolution block, a twelfth convolution block, a thirteenth convolution block, a fourteenth convolution block, a fifteenth convolution block, and a sixteenth convolution block connected in series; the dimension size of the input layer is set to b×3×224×224, and the value of b is equal to the number of samples selected by the convolutional neural network input once; the number of convolution kernels of the convolution layer is set to 64, the convolution kernel size is set to 7×7, and the step size is set to 2; the pooling window of the maximum pooling layer is set to 3×3, and the step size is set to 2.
4. The pain expression assessment method based on Batchformer according to claim 1, characterized in that: The structures of the first to sixteenth convolution blocks in the feature extractor described in step 2 are the same, and are composed of a 1×1 convolution layer, a 3×3 convolution layer, a 1×1 convolution layer, a normalization layer, and an activation layer connected in series in sequence; the number of convolution kernels of the convolution layers in the first to third convolution blocks is set to 64, 64, and 256, respectively, the number of convolution kernels of the convolution layers in the fourth to seventh convolution blocks is set to 128, 128, and 512, respectively, the number of convolution kernels of the convolution layers in the eighth to thirteenth convolution layers is set to 256, 256, and 1024, respectively, and the number of convolution kernels of the convolution layers in the fourteenth to sixteenth convolution layers is set to 512, 512, and 2048, respectively; the step size of the convolution layers in the first to sixteenth convolution blocks is set to 2; the eps of the normalization layer is set to 1e-5, and the momentums is set to 0.1; the activation layer is implemented using the nonlinear activation function ReLu.
5. The pain expression assessment method based on Batchformer according to claim 1, characterized in that: The Batchformer module described in step 2 consists of a transformer unit and a feed-forward fully connected layer in series; The Transformer unit is composed of a self-attention layer, a first random dropout layer, a first fully connected layer, a second random dropout layer, a second fully connected layer, a third random dropout layer, a first normalization layer, and a second normalization layer in series; the self-attention layer includes three fully connected layers, and the number of attention heads in its multi-head attention mechanism is nhead, which is equal to the number of segmented images; the first to third random dropout layers all use the dropout function, and each neuron is set to 0 with a probability ρ, 0≤ρ<1; the input of the first and second fully connected layers are set to 2048 dimensions, and the output is set to 2048 dimensions; the first and second normalization layers use Formula implementation, where y represents the output of the normalization layer, x represents the input of the normalization layer, E[x] represents the expected value operation, Var[x] represents the variance operation, eps represents the stability coefficient, eps=1e-5; The feedforward fully connected layer is composed of a first fully connected layer and a second fully connected layer connected in series, and the input and output dimensions of the first and second fully connected layers are both set to 2048.
6. The pain expression assessment method based on Batchformer according to claim 1, characterized in that: The pain level classifier described in step 2 is composed of a fully connected layer and an output layer in series; the input of the fully connected layer is set to 1024 dimensions and the output is set to 4 dimensions; the output layer is implemented by a softmax function, which maps the output of the fully connected layer of the classifier to the probability of belonging to a category, and the sum of the probabilities of all categories is 1.
7. The pain expression assessment method based on Batchformer network according to claim 1, characterized in that: The local features of each image block in step 3 are obtained by the following steps: The first step is to divide the training set into multiple batches, with integer powers of 2 as one batch. Each batch has 32 sample images, and the last sample images less than 32 are combined into a sample batch. In the second step, the local features of each image block in each batch are calculated according to the following formula: f ij =E(x ij ) Among them, x ij represents the jth image block of the i-th sample in a batch, f ij represents the local features of the jth image block of the ith sample in a batch, and E(·) represents the feature extraction operation.
8. The pain expression assessment method based on Batchformer network according to claim 1, characterized in that: The sample relationship features of each batch described in step 3 are obtained by the following formula: Among them, f represents the sample relationship feature of a batch, n represents the total number of image blocks segmented by a batch of samples, softmax(·) represents the normalization operation, and q ij , k ij 、v ij represents the query features, key features, and value features of the jth image block of the ith sample extracted by the three fully connected layers of the self-attention layer in the Transformer unit, d k Represents the normalization coefficient of the stable gradient.
9. The pain expression assessment method based on Batchformer network according to claim 1, characterized in that: The pain level classifier output pre-classification result described in step 3 is obtained by the following formula: Among them, Softmax(z a ) represents the probability that the input image belongs to the ath category, a=1,2,3,4, a=1 means no pain, a=2 means mild pain, a=3 means moderate pain, a=4 means severe pain, ∑ represents the summation operation, e (·) represents the exponential operation with the natural constant e as the base, z a represents the ath output of the fully connected layer, N represents the total number of dimensions of the fully connected layer output, z c represents the output of the cth dimension of the fully connected layer, c∈[1,N].
10. The pain expression assessment method based on Batchformer network according to claim 1, characterized in that: The loss function described in step 3 is as follows: L=-log(ρ(y i |x i )) Where L represents the loss value between the predicted label output by the pain level classifier and the real pain expression label, log represents the logarithmic operation with base 10, and y i represents the i-th sample x i The corresponding real pain expression label, ρ(y i |x i ) represents the classifier for sample x i The predicted label of .
Citation Information
Patent Citations
Parkinson's disease intelligent data evaluation method and system based on facial expression
CN116052872A
Pain expression assessment method based on multi-task transformer
CN116246326A