Facial pain expression grading assessment method and device based on visual learning
By using RGB-D video streams and deep learning technology, we have achieved accurate classification of facial pain expressions, solving the problem of the influence of lighting and angle changes in traditional methods and improving the accuracy of pain expression recognition.
Patent Information
- Application Number
- CN202510346652.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Traditional methods cannot effectively cope with changes in ambient lighting and facial angles, ignore three-dimensional features, and have difficulty recognizing subtle facial changes and dynamic expressions, resulting in low accuracy in pain rating.
Facial expressions are captured using RGB-D video streams. Through 3D keypoint detection and micro-motion magnification, combined with convolutional neural networks and graph neural networks, muscle motion vector fields and motion maps are constructed to extract visual and motion features and classify pain expressions.
It improves the accuracy of facial expression analysis and the ability to capture dynamic changes, enhances the accuracy of pain expression recognition, and adapts to facial changes under different lighting and angles.
Smart Images

Figure CN120410972B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of pain classification, and particularly relates to a facial pain expression grading evaluation method and device based on visual learning. BACKGROUND
[0002] Traditional methods usually rely on static images or video streams for analysis, which cannot effectively cope with changes in environmental lighting and facial angle. Under different lighting conditions, the details of facial expressions may be obscured or distorted, resulting in reduced accuracy of expression recognition. Moreover, traditional methods usually rely only on two-dimensional images (RGB images) to extract facial expression features, ignoring the three-dimensional characteristics of the face. Many subtle changes in facial expressions, such as frowning and smiling, may be difficult to accurately capture in two-dimensional images, and the lack of depth information analysis also cannot effectively identify changes in facial expressions at different angles. Furthermore, traditional methods have weak recognition ability for small muscle movements and subtle facial changes. Facial expressions of complex emotions such as pain often involve small muscle movements. Moreover, traditional methods rely on static image analysis and ignore the dynamic changes of facial expressions. Facial expressions are usually dynamic, especially in the case of pain, facial expressions may change continuously with the change of pain intensity, and traditional static analysis methods are difficult to capture these dynamic changes, thereby affecting the accuracy of pain grading. In addition, traditional methods often cannot accurately model the coordinated movement of facial muscles. Facial expressions involve the coordinated action of multiple muscle groups, and many traditional methods only focus on a single muscle group or expression area, which cannot comprehensively model the complexity of facial expressions. SUMMARY
[0003] The technical problem to be solved by the present application is to overcome the shortcomings of the above-mentioned prior art and provide a facial pain expression grading evaluation method and device based on visual learning.
[0004] The technical solution adopted to solve the above technical problems is: a facial pain expression grading evaluation method based on visual learning, comprising:
[0005] obtaining an RGB-D video stream of a target patient's facial pain expression, and obtaining an RGB video frame sequence and a depth video frame sequence from the RGB-D video stream;
[0006] performing 3D key point detection on the RGB video frame sequence according to the depth video frame sequence to obtain a 3D key point sequence of the RGB video frame, and performing micro-motion amplification on the RGB video frame sequence to obtain an enhanced RGB video frame sequence;
[0007] performing visual feature extraction on each enhanced RGB video frame in the enhanced RGB video frame sequence according to a convolutional neural network to obtain a visual feature sequence of the target patient's face.
[0008] constructing each muscle movement vector field in the RGB video frame according to the 3D key point sequence of the RGB video frame, and constructing an action graph according to each muscle movement vector field and anatomical correlation between each muscle movement vector field;
[0009] extracting motion features of the RGB video frame sequence through the action graph according to the graph neural network, to obtain a motion feature sequence of the target patient face;
[0010] grading a pain expression of the target patient face according to the visual feature sequence and the motion feature sequence of the target patient face, to obtain a pain label of the target patient face.
[0011] Preferably, the 3D key point sequence of the RGB video frame is obtained by performing 3D key point extraction on the RGB video frame sequence according to the depth video frame sequence, comprising:
[0012] aligning the depth video frame and the RGB video frame through camera calibration parameters to generate a registered depth video frame and a registered RGB video frame;
[0013] detecting 3D key points of the registered depth video frame and the registered RGB video frame according to an improved HRNet-W64 model, to obtain a first 3D key point sequence of the RGB video frame;
[0014] performing key point regression on the first 3D key point sequence of the RGB video frame according to a preset regression loss function, to obtain a second 3D key point sequence of the RGB video frame, wherein the regression loss function is as follows:
[0015]
[0016] wherein L reg represents the regression loss function, N represents the total number of 3D key points in the first 3D key point sequence, φ(I t ,D t ) i represents the first 3D key point extracted by the improved HRNet-W64 model, P i gt represents the true 3D key point label;
[0017] constructing an anatomical loss function according to a preset muscle movement correlation matrix and gradient constraint, and performing anatomical constraint on the second 3D key point sequence of the RGB video frame according to the anatomical loss function, to obtain a final 3D key point sequence of the RGB video frame, wherein the anatomical loss function is as follows:
[0018]
[0019] wherein, L anatomy denotes an anatomical loss function, denotes the gradient of the second 3D keypoint motion of adjacent RGB video frames, and P t denotes the second 3D keypoint sequence of the t-th RGB video frame, M muscle denotes the muscle motion correlation matrix, and when the second 3D keypoint i and the second 3D keypoint j belong to the same muscle group, then M muscle [i,j] = 1, otherwise, M muscle [i,j] = 0.
[0020] Preferably, the improved part of the HRNet-W64 model comprises: in each transition layer of the HRNet-W64 model, mapping the depth video frame feature to the same channel number as the RGB video frame feature by 1x1 convolution, and performing element-wise addition, wherein the expression of the improved part of the HRNet-W64 model is as follows:
[0021]
[0022] wherein, denotes the fusion feature of the k-th transition layer in the HRNet-W64 model, denotes the RGB video frame feature of the k-th transition layer in the HRNet-W64 model, denotes the depth video frame feature of the k-th transition layer in the HRNet-W64 model, Conv 1×1 denotes a 1x1 convolution operation.
[0023] Preferably, the RGB video frame sequence is amplified for micro-movement to obtain an enhanced RGB video frame sequence, comprising:
[0024] constructing a Laplacian pyramid for each RGB video frame in the RGB video frame sequence, wherein each layer in the Laplacian pyramid is down-sampled by 2;
[0025] extracting a time series for each pixel point in each layer of the Laplacian pyramid to obtain a time signal, and performing Fourier transform on the time series to obtain a frequency domain signal;
[0026] filtering the frequency domain signal of each pixel point in each layer of the Laplacian pyramid according to a band-pass filter to obtain a motion signal;
[0027] determining an adaptive amplification factor for each pixel point in each RGB video frame in the RGB video frame sequence according to the actual depth and the reference depth of the pixel point;
[0028] performing inverse Fourier transform on the motion signal and taking real part to obtain a motion time signal, and amplifying the motion time signal according to the adaptive amplification factor to obtain an amplified motion time signal;
[0029] adding the amplified motion time signal back to the Laplacian pyramid to obtain an enhanced Laplacian pyramid;
[0030] performing upsampling operation layer by layer from the bottom layer of the enhanced Laplacian pyramid and fusing layer by layer to obtain an enhanced RGB video frame sequence.
[0031] Preferably, the calculation formula of the adaptive amplification factor is as follows:
[0032]
[0033] wherein, α(x, y) represents the adaptive amplification factor of the pixel point (x, y), γ represents an initial amplification factor, D ref (x, y) represents the reference depth of the pixel point (x, y), D t (x, y) represents the actual depth of the pixel point (x, y);
[0034] The expression of the enhanced Laplacian pyramid is as follows:
[0035]
[0036] wherein, represents the lth layer of the enhanced Laplacian pyramid, I t (l) represents the lth layer of the Laplacian pyramid, ΔI t (l) represents the amplified part of the lth layer of the Laplacian pyramid, and ΔI t (l) = α(x, y) · R(F -1 (F fit (l))), R represents a real part taking operation, F -1 represents inverse Fourier transform, F fit (l) represents the amplified motion time signal;
[0037] The expression of the enhanced RGB video frame is as follows:
[0038]
[0039] wherein, represents the enhanced RGB video frame, L represents the total number of layers of the enhanced Laplacian pyramid, and PyrUP represents the upsampling operation.
[0040] Preferably, constructing each muscle motion vector field of the target patient face in the RGB video frame according to the 3D key point sequence of the RGB video frame comprises:
[0041] Defining a muscle boundary set of each action unit in the target patient face, wherein the muscle boundary set contains indexes of pairs of 3D key points of adjacent RGB video frames;
[0042] Predicting a weight between each pair of 3D key points according to the motion of the pair of 3D key points between adjacent RGB video frames by a weight prediction network, determining the weight of each action unit according to the weight between the 3D key points, wherein the calculation formula of the weight between the 3D key points is as follows:
[0043]
[0044] wherein A jk represents the weight between the jth 3D key point in the RGB video frame and the kth 3D key point in the adjacent RGB video frame, Sigmoid represents an activation function, f represents an MLP network, P t j represents the jth 3D key point in the RGB video frame, represents the kth 3D key point in the adjacent RGB video frame, ΔP jk represents the relative displacement between the jth 3D key point in the RGB video frame and the kth 3D key point in the adjacent RGB video frame;
[0045] Constructing a muscle motion vector field according to the weight of each action unit and the motion of the pair of 3D key points in the action unit between adjacent RGB video frames, wherein the calculation formula of the muscle motion vector field is as follows:
[0046]
[0047] wherein, represents the muscle motion vector field of the ith action unit in the target patient face, E i represents the muscle boundary set of the ith action unit in the target patient face.
[0048] Preferably, the action graph comprises an action node set and a connection edge set, wherein the action nodes in the action node set correspond to muscle motion vector fields, and the connection edges in the connection edge set correspond to anatomical associations between action units.
[0049] Preferably, performing motion feature extraction on each RGB video frame in the sequence of RGB video frames according to the action graph by a graph neural network to obtain a sequence of motion features of the target patient face, comprising:
[0050] According to the graph neural network, the weight of the connected edge in the action graph is updated, wherein the update formula of the weight of the connected edge is as follows:
[0051]
[0052] wherein, denotes the updated weight of the connected edge between the i th action node and the j th action node in the action graph, W a denotes a first learnable weight matrix;
[0053] According to the graph neural network, the node feature propagation of the action graph is performed to obtain the motion feature, wherein the node feature propagation formula is as follows:
[0054]
[0055] wherein, denotes the node feature of the l+1 th layer in the graph neural network, denotes the adjacency matrix of the connected edge in the action graph, W g denotes a second learnable weight matrix.
[0056] Preferably, the target patient face is classified according to the visual feature sequence and the motion feature sequence of the target patient face to obtain a pain label of the target patient face, comprising:
[0057] The visual feature sequence and the motion feature sequence of the target patient face are spliced to obtain a visual motion feature sequence of the target patient face;
[0058] According to the GRU model, the target patient face is classified according to the visual motion feature sequence to obtain a pain label of the target patient face.
[0059] The technical solution adopted to solve the above technical problems is: a facial pain expression classification and evaluation device based on visual learning, which is suitable for the facial pain expression classification and evaluation method based on visual learning, comprising:
[0060] A video acquisition unit is configured to acquire an RGB-D video stream of a target patient's facial pain expression, and acquire an RGB video frame sequence and a depth video frame sequence according to the RGB-D video stream;
[0061] A key point detection unit is configured to perform 3D key point detection on the RGB video frame sequence according to the depth video frame sequence to obtain a 3D key point sequence of the RGB video frame sequence;
[0062] a motion enhancement unit, configured to perform micro-motion amplification on the sequence of RGB video frames to obtain a sequence of enhanced RGB video frames;
[0063] a visual extraction unit, configured to perform visual feature extraction on each enhanced RGB video frame in the sequence of enhanced RGB video frames according to a convolutional neural network to obtain a sequence of visual features of the target patient's face;
[0064] an action modeling unit, configured to construct each muscle movement vector field in the RGB video frames according to the sequence of 3D key points, and construct an action graph according to each muscle movement vector field and anatomical correlations between the muscle movement vector fields;
[0065] a motion extraction unit, configured to perform motion feature extraction on the sequence of RGB video frames through the action graph according to a graph neural network to obtain a sequence of motion features of the target patient's face;
[0066] a pain grading unit, configured to grade the target patient's face according to the sequence of visual features and the sequence of motion features to obtain a pain label of the target patient's face.
[0067] The present application has the following advantages: (1) The present application can provide more abundant facial expression data through the RGB-D video stream, and the depth information can help identify the three-dimensional features of the face, making the analysis of facial expressions more accurate and not affected by changes in light and facial angles. Moreover, by performing micro-motion amplification on the sequence of RGB video frames, the present application can enhance the subtle changes in facial expressions and capture some subtle muscle movements and expression changes, which are crucial for accurate identification of pain expressions; (2) The present application uses 3D key point detection, which can not only extract two-dimensional features of facial expressions but also capture spatial changes of each part of the face, providing more abundant contextual information for the analysis of facial muscle movements and enhancing the expressiveness of motion features. Moreover, by constructing muscle movement vector fields and action graphs combining anatomical correlations, the present application can model the coordinated movement of facial muscles more finely, thereby improving the recognition accuracy of facial actions, which makes the grading of pain expressions more scientific and can accurately identify facial actions of different pain levels; (3) The present application performs comprehensive analysis of visual features and motion features, which can not only identify pain through static facial expressions but also further verify and improve the accuracy of pain grading through dynamic facial action changes. Compared with single static image analysis, this method is more suitable for dynamic expression changes in actual applications. BRIEF DESCRIPTION OF DRAWINGS
[0068] Figure 1A schematic diagram of the step flow of the overall method in an embodiment of the present application is shown in the figure.
[0069] Figure 2 A schematic diagram of the device architecture of the overall device in an embodiment of the present application is shown in the figure.
[0070] Reference signs: 1, video acquisition unit; 2, key point detection unit; 3, motion enhancement unit; 4, visual extraction unit; 5, action modeling unit; 6, motion extraction unit; 7, pain grading unit. DETAILED DESCRIPTION
[0071] Embodiment one, as shown in the figure, the present application proposes a facial pain expression grading evaluation method based on visual learning, which comprises: Figure 1
[0072] S1, acquiring an RGB-D video stream of a target patient's facial pain expression, and acquiring an RGB video frame sequence and a depth video frame sequence according to the RGB-D video stream;
[0073] S2, performing 3D key point detection on the RGB video frame sequence according to the depth video frame sequence to obtain a 3D key point sequence of the RGB video frame, and performing micro-motion amplification on the RGB video frame sequence to obtain an enhanced RGB video frame sequence;
[0074] S3, performing visual feature extraction on each enhanced RGB video frame in the enhanced RGB video frame sequence according to a convolutional neural network to obtain a visual feature sequence of the target patient's face;
[0075] S4, constructing each muscle motion vector field in the RGB video frame according to the 3D key point sequence of the RGB video frame, and constructing an action graph according to each muscle motion vector field and the anatomical association between each muscle motion vector field;
[0076] S5, performing motion feature extraction on the RGB video frame sequence through the action graph according to a graph neural network to obtain a motion feature sequence of the target patient's face;
[0077] S6, performing pain expression grading on the target patient's face according to the visual feature sequence and the motion feature sequence of the target patient's face to obtain a pain label of the target patient's face.
[0078] In the present invention, RGB: represents color video frames (red, green, and blue color channels); D: represents depth information, usually obtained through depth sensors (such as Kinect), which can provide the distance from each pixel in the video frame to the camera; 3D key point detection refers to detecting and labeling key parts of the face in three-dimensional space, such as eyes, nose, mouth, etc., which can be determined through depth data and image data in the RGB-D video stream; micro-movement refers to the relatively subtle muscle movements in facial expressions, which are often difficult to capture through ordinary video frames, so they need to be amplified to enhance the visibility of these subtle actions; CNN is a deep learning network widely used in image and video processing, which automatically extracts features (such as edges, textures, etc.) in images through multiple convolutional layers, helping to identify objects, facial expressions, etc.; visual features refer to representative information extracted from images or videos, used to describe objects, scenes, or actions occurring in the scene; muscle movement vector field refers to the direction and amplitude of each facial muscle during facial expression changes, through a deep learning model, the movement vector of facial muscles can be calculated and a vector field is formed, representing the movement state of different muscles at different time points, these movement vectors can help analyze facial expression changes, especially in emotional states such as pain; action graph is a graph form that represents the relationship between facial muscle movement vector fields by mapping the anatomical relationship between them (i.e., which muscles interact with each other, work together), each node represents a facial muscle, and the edge represents the interaction between muscles, through the structure of the graph, the dynamic changes of facial expressions can be better understood; graph neural network is a deep learning model suitable for graph structure data, which can process graph data with nodes and edges, here, the graph neural network is used to model the movement of facial muscles through the action graph, so as to extract the movement feature sequence, the graph neural network can capture the complex interrelationships and dependencies between facial muscles, which is particularly important for facial expression analysis; movement features refer to dynamic change features in facial expressions, especially those related to muscle movement, these features help evaluate the changes in facial expressions over time, especially in emotional responses such as pain, facial muscle movements are often very obvious, through movement features, facial expressions in different emotional states can be distinguished; pain expression grading is an analysis of the target patient's facial expression, aiming to determine the patient's facial expression of pain through the analysis of visual and movement features, which is usually a classification problem, through the analysis of the patient's facial expression, a label representing the degree of pain can be obtained, such as no pain, mild pain, and severe pain, etc.
[0079] In the second embodiment, the facial pain expression grading evaluation method based on visual learning is further provided with the following steps: performing 3D key point extraction on the RGB video frame sequence according to the depth video frame sequence to obtain a 3D key point sequence of the RGB video frame, including:
[0080] A1, aligning the depth video frame and the RGB video frame by using the camera calibration parameters to generate the registered depth video frame and the registered RGB video frame;
[0081] A2, performing 3D key point detection on the registered depth video frame and the registered RGB video frame according to the improved HRNet-W64 model to obtain a first 3D key point sequence of the RGB video frame;
[0082] A3, performing key point regression on the first 3D key point sequence of the RGB video frame according to a preset regression loss function to obtain a second 3D key point sequence of the RGB video frame, wherein the regression loss function is as follows:
[0083]
[0084] wherein L reg represents the regression loss function, N represents the total number of 3D key points in the first 3D key point sequence, φ(I t ,D t ) i represents the first 3D key point extracted by the improved HRNet-W64 model, P i gt represents the real 3D key point label;
[0085] A4, constructing an anatomical loss function according to a preset muscle movement correlation matrix and gradient constraint, and performing anatomical constraint on the second 3D key point sequence of the RGB video frame according to the anatomical loss function to obtain a final 3D key point sequence of the RGB video frame, wherein the anatomical loss function is as follows:
[0086]
[0087] wherein L anatomy represents the anatomical loss function, represents the gradient of the second 3D key point motion of the adjacent RGB video frame, and P t represents the second 3D key point sequence of the tthRGB video frame, M muscle represents the muscle movement correlation matrix, and when the second 3D key point i and the second 3D key point j belong to the same muscle group, M muscle [i,j] = 1, otherwise, M muscle [i,j] = 0.
[0088] In this embodiment, camera calibration refers to measuring and calculating the internal and external parameters of the camera (such as focal length, distortion coefficient, position and pose of the camera, etc.), so that the camera can accurately capture images in the physical world. After calibration, the depth video frame (image obtained by the depth camera) and the RGB video frame (color image obtained by the ordinary camera) can be accurately aligned, so that the positions of the same object in both are consistent; HRNet (High-Resolution Network) is a deep neural network model widely used in human pose estimation, facial expression analysis and other tasks. The main feature of HRNet is to maintain high-resolution information flow in multiple network branches of different resolutions, which helps to capture more accurate details; gradient constraint is a technique that limits the rate of change of model output during training; anatomical constraint is to limit the output of the model through the anatomical understanding of facial muscles and skeletal structure.
[0089] In an optional embodiment, the improved part of the HRNet-W64 model includes: in each transition layer of the HRNet-W64 model, mapping the depth video frame features to the same number of channels as the RGB video frame features by 1x1 convolution, and performing element-wise addition, wherein the expression of the improved part of the HRNet-W64 model is as follows:
[0090]
[0091] wherein, represents the fusion features of the kth transition layer in the HRNet-W64 model, represents the RGB video frame features of the kth transition layer in the HRNet-W64 model, represents the depth video frame features of the kth transition layer in the HRNet-W64 model, Conv 1×1 represents the 1x1 convolution operation.
[0092] In an optional embodiment, the RGB video frame sequence is amplified for micro-movement to obtain an enhanced RGB video frame sequence, comprising:
[0093] B1, constructing a Laplacian pyramid for each RGB video frame in the RGB video frame sequence, wherein each layer in the Laplacian pyramid is down-sampled by 2;
[0094] B2, performing time series extraction on the pixel points of each layer in the Laplacian pyramid to obtain a time signal, and performing Fourier transform on the time series to obtain a frequency domain signal;
[0095] B3, filtering the frequency domain signal of each pixel in each layer of the Laplacian pyramid using a band-pass filter to obtain a motion signal;
[0096] B4, determining an adaptive magnification factor for each pixel in each RGB video frame in the sequence of RGB video frames based on the actual depth and the reference depth of the pixel;
[0097] B5, performing inverse Fourier transform on the motion signal and taking the real part to obtain a motion time signal, and magnifying the motion time signal based on the adaptive magnification factor to obtain a magnified motion time signal;
[0098] B6, adding the magnified motion time signal back to the Laplacian pyramid to obtain an enhanced Laplacian pyramid;
[0099] B7, performing upsampling operation on each layer of the enhanced Laplacian pyramid and fusing layer by layer to obtain an enhanced sequence of RGB video frames.
[0100] It should be noted that the Laplacian pyramid is a multi-resolution image representation method, which represents the original image by decomposing it into multiple resolution images layer by layer, in each layer, the resolution of the image is down-sampled (usually by blurring and sampling) to reduce the amount of information, each layer of the Laplacian pyramid represents the difference between the image of this layer and the last layer, which is commonly used in image compression, denoising and image enhancement applications; down-sampling is a process of reducing the resolution of an image, usually by reducing the number of pixels in the image; Fourier transform is a mathematical method for converting signals from time domain to frequency domain, in image processing, Fourier transform is often used to analyze the distribution of different frequency components in the image; frequency domain signal is a way to represent the frequency components of a signal, through Fourier transform, time domain signal (such as the sequence of brightness of each pixel in the RGB video frame changing over time) can be converted into frequency domain signal, revealing the frequency components in the signal; band-pass filter is a filter that allows certain range of frequencies to pass through, while suppressing frequencies above or below that range, band-pass filter is used to filter the frequency domain signal of each pixel to extract motion information in a certain frequency range, which is usually used to filter out background noise or irrelevant low-frequency, high-frequency signals, only relevant motion signals are retained; inverse Fourier transform is the inverse process of Fourier transform, which is used to convert frequency domain signal back to time domain; enhanced Laplacian pyramid refers to the original Laplacian pyramid structure, which is enhanced by adding magnified motion signal to enhance the motion performance in the image; up-sampling refers to the process of increasing the resolution of an image in image processing; layer-by-layer fusion refers to combining the processing results of different layers layer by layer in image processing.
[0101] In an optional embodiment, the formula for calculating the adaptive magnification factor is as follows:
[0102]
[0103] wherein, α(x, y) represents the adaptive magnification factor of the pixel point (x, y), γ represents the initial magnification factor, D ref represents the reference depth of the pixel point (x, y), D t (x, y) represents the actual depth of the pixel point (x, y);
[0104] The expression of the enhanced Laplacian pyramid is as follows:
[0105]
[0106] wherein, represents the l-th layer in the enhanced Laplacian pyramid, I t (l) represents the l-th layer in the Laplacian pyramid, ΔI t (l) represents the magnified part of the l-th layer in the Laplacian pyramid, and ΔI t (l) = α(x, y) · R(F -1 (F fit (l))), R represents the real part operation, F -1 represents the inverse Fourier transform, F fit (l) represents the magnified motion time signal;
[0107] The expression of the enhanced RGB video frame is as follows:
[0108]
[0109] wherein, represents the enhanced RGB video frame, L represents the total number of layers of the enhanced Laplacian pyramid, and PyrUP represents the up-sampling operation.
[0110] In an optional embodiment, each muscle motion vector field in the RGB video frame is constructed according to the 3D key point sequence of the RGB video frame, comprising:
[0111] C1, defining a muscle edge set of each action unit in the target patient's face, wherein the muscle edge set contains the indexes of the 3D key point pairs of adjacent RGB video frames;
[0112] C2, predicting the weight between each pair of 3D key points according to the motion of the 3D key point pairs between adjacent RGB video frames by a weight prediction network, and determining the weight of each action unit according to the edge weight between the 3D key points, wherein the calculation formula of the weight between the 3D key points is as follows:
[0113]
[0114] wherein, Ajk represents the weight between the jth 3D keypoint in the RGB video frame and the kth 3D keypoint in the adjacent RGB video frame, Sigmoid represents the activation function, f represents the MLP network, P t j represents the jth 3D keypoint in the RGB video frame, represents the kth 3D keypoint in the adjacent RGB video frame, ΔP jk represents the relative displacement between the jth 3D keypoint in the RGB video frame and the kth 3D keypoint in the adjacent RGB video frame;
[0115] C3, constructing a muscle motion vector field according to the weight of each action unit and the motion of the 3D keypoint pair in the action unit between adjacent RGB video frames, wherein the calculation formula of the muscle motion vector field is as follows:
[0116]
[0117] wherein v AUi represents the muscle motion vector field of the ith action unit in the target patient's face, E i represents the muscle edge set of the ith action unit in the target patient's face.
[0118] AU1 represents lifting the eyebrows, AU6 represents smiling, etc., and each action unit is composed of a plurality of 3D key point pairs (i.e., the active region of a facial muscle), which changes in space and time; the muscle edge set refers to the set of 3D key point pairs corresponding to the facial muscle movement, and the relative movement of each pair of 3D key points (such as the relative displacement of two key points in adjacent RGB video frames) can represent the activity of a certain facial muscle, and these edge sets connect different muscles of the face, reflecting the role of the muscles in different action units; the weight prediction network is a deep learning model used to predict the motion weight between each pair of 3D key points according to the motion of the 3D key points in adjacent RGB video frames; the MLP (Multi-Layer Perceptron) network is a basic feedforward neural network structure, which usually includes an input layer, multiple hidden layers and an output layer, each hidden layer is composed of a plurality of neurons, and the network fits the relationship between the input data and the output target by learning the weights between these layers, and the MLP network is usually used for function fitting and feature extraction, and here it is used to predict the corresponding weight according to the motion between the facial key points; the muscle motion vector field is a description of the movement of facial muscles in each action unit, which is composed of a plurality of vectors, each vector representing the movement direction and intensity of a muscle, and the muscle motion vector field of each action unit is constructed according to the motion between the 3D key points and their weights, for further facial expression analysis and synthesis.
[0119] In an optional embodiment, the action graph includes an action node set and a connection edge set, wherein the action nodes in the action node set correspond to the muscle motion vector fields, and the connection edges in the connection edge set correspond to the anatomical associations between the action units.
[0120] It should be noted that the anatomical association refers to the physiological and anatomical connection between different facial action units (AUs) or facial muscles, and different facial muscles and action units are connected to each other through anatomical structures, for example, lifting the eyebrows (AU1) is usually accompanied by changes in the muscles around the eyes, while smiling (AU6) is related to the movement of the muscles around the mouth.
[0121] In an optional embodiment, the motion feature of each RGB video frame in the sequence of RGB video frames is extracted according to the graph neural network through the action graph to obtain the motion feature sequence of the face of the target patient, including:
[0122] D1, updating the weight of the connection edge in the action graph according to the graph neural network, wherein the updating formula of the weight of the connection edge is as follows:
[0123]
[0124] wherein, denotes the updated weight of the connecting edge between the i th action node and the j th action node in the action graph, W a denotes the first learnable weight matrix;
[0125] D2, performing node feature propagation on the action graph according to the graph neural network to obtain motion features, wherein the node feature propagation formula is as follows:
[0126]
[0127] wherein, denotes the node feature of the l+1 th layer in the graph neural network, denotes the adjacency matrix of the connecting edge in the action graph, W g denotes the second learnable weight matrix.
[0128] In an optional embodiment, the target patient's face is classified according to the visual feature sequence and the motion feature sequence of the target patient's face to obtain a pain label of the target patient's face, comprising:
[0129] E1, splicing the visual feature sequence and the motion feature sequence of the target patient's face to obtain a visual motion feature sequence of the target patient's face;
[0130] E2, classifying the target patient's face according to the GRU model through the visual motion feature sequence to obtain a pain label of the target patient's face.
[0131] It should be noted that GRU (Gated Recurrent Unit) is a kind of recurrent neural network (RNN) model, which is used to process sequence data. Compared with traditional RNN, GRU solves the problem of gradient disappearance that RNN may encounter when processing long time sequence by introducing a gating mechanism. The GRU model has two main gates: update gate and reset gate. These two gates help the network to selectively update or reset the memory, so that the GRU performs more effectively when processing time sequence data.
[0132] Embodiment three, as Figure 2 shown, the application provides a facial pain expression classification and evaluation device based on visual learning, which is suitable for the facial pain expression classification and evaluation method based on visual learning, comprising:
[0133] A video acquisition unit 1, the video acquisition unit 1 is used for acquiring the RGB-D video stream of the target patient's face pain expression, and the RGB video frame sequence and the depth video frame sequence are acquired according to the RGB-D video stream;
[0134] The key point detection unit 2 is configured to perform 3D key point detection on the RGB video frame sequence according to the depth video frame sequence to obtain a 3D key point sequence of the RGB video frame.
[0135] The motion enhancement unit 3 is configured to perform micro-motion amplification on the RGB video frame sequence to obtain an enhanced RGB video frame sequence.
[0136] The visual extraction unit 4 is configured to perform visual feature extraction on each enhanced RGB video frame in the enhanced RGB video frame sequence according to a convolutional neural network to obtain a visual feature sequence of the target patient's face.
[0137] The action modeling unit 5 is configured to construct each muscle movement vector field in the RGB video frame according to the 3D key point sequence of the RGB video frame, and construct an action graph according to the each muscle movement vector field and the anatomical correlation between the each muscle movement vector fields.
[0138] The motion extraction unit 6 is configured to perform motion feature extraction on the RGB video frame sequence through the action graph according to a graph neural network to obtain a motion feature sequence of the target patient's face.
[0139] The pain grading unit 7 is configured to perform pain expression grading on the target patient's face according to the visual feature sequence and the motion feature sequence of the target patient's face to obtain a pain label of the target patient's face.
[0140] The embodiments of the present application are described in detail above in combination with the drawings, but the present application is not limited thereto, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.
Claims
1. A method for grading and assessing facial pain expressions based on visual learning, characterized in that, include: Acquire an RGB-D video stream of the target patient's facial pain expression, and obtain an RGB video frame sequence and a depth video frame sequence based on the RGB-D video stream; 3D key point detection is performed on the RGB video frame sequence based on the depth video frame sequence to obtain the 3D key point sequence of the RGB video frame, and micro-motion magnification is performed on the RGB video frame sequence to obtain the enhanced RGB video frame sequence. Visual features are extracted from each enhanced RGB video frame in the enhanced RGB video frame sequence using a convolutional neural network to obtain a visual feature sequence of the target patient's face. Based on the 3D key point sequence of the RGB video frame, construct the motion vector field of each muscle in the RGB video frame, and construct the motion map based on the anatomical relationship between each muscle motion vector field and the anatomical relationship between each muscle motion vector field. The motion features of the target patient's face are obtained by extracting motion features from the RGB video frame sequence using the motion graph through a graph neural network. Pain expression grading of the target patient's face is performed based on the visual feature sequence and motion feature sequence of the target patient's face to obtain the pain label of the target patient's face; Based on the depth video frame sequence, 3D keypoint extraction is performed on the RGB video frame sequence to obtain a 3D keypoint sequence of the RGB video frames, including: The depth video frame and the RGB video frame are image aligned using camera calibration parameters to generate registered depth video frames and RGB video frames. 3D keypoint detection is performed on the registered depth video frames and RGB video frames according to the improved HRNet-W64 model to obtain the first 3D keypoint sequence of the RGB video frames. Based on a preset regression loss function, keypoint regression is performed on the first 3D keypoint sequence of the RGB video frame to obtain the second 3D keypoint sequence of the RGB video frame. The regression loss function is as follows: Among them, L reg Let N represent the regression loss function, N represent the total number of 3D keypoints in the first 3D keypoint sequence, and φ(I) represent the regression loss function. t D t ) i This represents the first 3D keypoint extracted by the improved HRNet-W64 model. Represents true 3D keypoint annotations; An anatomical loss function is constructed based on a preset muscle motion correlation matrix and gradient constraints. This anatomical loss function is then used to apply anatomical constraints to the second 3D keypoint sequence of the RGB video frame, resulting in the final 3D keypoint sequence of the RGB video frame. The anatomical loss function is as follows: Among them, L anatomy Represents the anatomical loss function. This represents the gradient of the motion of the second 3D keypoint in adjacent RGB video frames, and P t M represents the sequence of the second 3D key points in the t-th RGB video frame. muscle Let M represent the muscle motion correlation matrix, and if the second 3D keypoint i and the second 3D keypoint j belong to the same muscle group, then M... muscle [i,j] = 1, otherwise, then M muscle [i,j]=0.
2. The method for assessing facial pain expression based on visual learning according to claim 1, characterized in that, The improved part of the HRNet-W64 model includes: in each transition layer of the HRNet-W64 model, mapping the depth video frame features to the same number of channels as the RGB video frame features through 1×1 convolution, and performing element-wise addition. The expression for the improved part of the HRNet-W64 model is as follows: in, This represents the fusion feature of the k-th transition layer in the HRNet-W64 model. This represents the RGB video frame features of the k-th transition layer in the HRNet-W64 model. Represents the deep video frame features of the k-th transition layer in the HRNet-W64 model, Conv 1×1 This represents a 1×1 convolution operation.
3. The method for assessing facial pain expression based on visual learning according to claim 2, characterized in that, Performing micro-motion amplification on the RGB video frame sequence to obtain an enhanced RGB video frame sequence includes: A Laplacian pyramid is constructed for each RGB video frame in the RGB video frame sequence, wherein each layer of the Laplacian pyramid is downsampled by 2 times; The time series of pixels in each layer of the Laplace pyramid is extracted to obtain a time signal, and the time series is subjected to Fourier transform to obtain a frequency domain signal. The frequency domain signals of each layer and each pixel in the Laplace pyramid are filtered using a bandpass filter to obtain the motion signal; The adaptive magnification factor of the pixel is determined based on the actual depth and reference depth of each pixel in each RGB video frame in the RGB video frame sequence; Perform an inverse Fourier transform on the motion signal and take the real part to obtain the motion time signal. Amplify the motion time signal according to the adaptive amplification factor to obtain an amplified motion time signal. The amplified motion time signal is added back into the Laplace pyramid to obtain the enhanced Laplace pyramid. Upsampling is performed layer by layer from the bottom layer of the enhanced Laplacian pyramid, and the layers are then fused to obtain an enhanced RGB video frame sequence.
4. The method for assessing facial pain expression based on visual learning according to claim 3, characterized in that, The formula for calculating the adaptive amplification factor is as follows: Where α(x,y) represents the adaptive magnification factor of pixel (x,y), γ represents the initial magnification factor, and D ref D represents the reference depth of pixel (x, y). t (x,y) represents the actual depth of the pixel (x,y); The expression for the enhanced Laplace pyramid is as follows: in, This represents the l-th layer in the enhanced Laplace's Pyramid. t (l) represents the l-th level in the Laplace pyramid, ΔI t (l) represents the magnified portion of the l-th layer in the Laplace pyramid, and ΔI t (l)=α(x,y)·R(F -1 (F fit (l))), R represents the real part operation, F -1 F represents the inverse Fourier transform. fit (l) indicates amplified motion time signal; The expression for the enhanced RGB video frame is as follows: in, This indicates an enhanced RGB video frame, where L represents the total number of layers in the enhanced Laplacian pyramid, and PyrUP represents the upsampling operation.
5. The method for assessing facial pain expression based on visual learning according to claim 4, characterized in that, Constructing the motion vector field of each muscle in the RGB video frame based on the 3D keypoint sequence of the RGB video frame, including: Define a muscle edge set for each action unit in the face of the target patient, wherein the muscle edge set contains indices of 3D keypoint pairs of adjacent RGB video frames; The weight prediction network predicts the weights between each pair of 3D keypoints based on their motion between adjacent RGB video frames. The weights of each action unit are then determined based on the edge weights between these 3D keypoints. The formula for calculating the weights between the 3D keypoints is as follows: Among them, A jk Let f represent the weight between the j-th 3D keypoint in an RGB video frame and the k-th 3D keypoint in an adjacent RGB video frame, where sigmoid represents the activation function and f represents the MLP network. This represents the j-th 3D keypoint in an RGB video frame. ΔP represents the k-th 3D keypoint in adjacent RGB video frames. jk This represents the relative displacement between the j-th 3D keypoint in an RGB video frame and the k-th 3D keypoint in an adjacent RGB video frame. A muscle motion vector field is constructed based on the weight of each action unit and the motion of 3D keypoint pairs in the action unit between adjacent RGB video frames. The calculation formula for the muscle motion vector field is as follows: in, E represents the muscle motion vector field of the i-th action unit in the face of the target patient. i This represents the muscle boundary set of the i-th action unit in the face of the target patient.
6. The method for assessing facial pain expression based on visual learning according to claim 5, characterized in that, The motion graph includes a set of motion nodes and a set of connecting edges. The motion nodes in the motion node set correspond to the muscle motion vector field, and the connecting edges in the connecting edge set correspond to the anatomical relationships between motion units.
7. The method for assessing facial pain expression based on visual learning according to claim 6, characterized in that, Based on the graph neural network, motion features are extracted from each RGB video frame in the RGB video frame sequence using the action graph to obtain a motion feature sequence of the target patient's face, including: The weights of the connecting edges in the action graph are updated using a graph neural network, and the update formula for the weights of the connecting edges is as follows: in, W represents the updated weight of the edge connecting the i-th and j-th action nodes in the action graph. a This represents the first learnable weight matrix; The motion graph is subjected to node feature propagation using a graph neural network to obtain motion features, wherein the node feature propagation formula is as follows: in, This represents the node features in the (l+1)th layer of a graph neural network. W represents the adjacency matrix of the connecting edges in the action graph. g This represents the second learnable weight matrix.
8. The method for assessing facial pain expression based on visual learning according to claim 7, characterized in that, Pain expression grading of the target patient's face is performed based on the visual and motion feature sequences of the target patient's face to obtain a pain label for the target patient's face, including: The visual feature sequence and motion feature sequence of the target patient's face are spliced together to obtain the visual motion feature sequence of the target patient's face. The pain expression of the target patient's face is graded based on the visual motion feature sequence using the GRU model to obtain the pain label of the target patient's face.
9. A facial pain expression grading and assessment device based on visual learning, applicable to the facial pain expression grading and assessment method based on visual learning as described in any one of claims 1 to 8, characterized in that, include: The video acquisition unit (1) is used to acquire an RGB-D video stream of the facial pain expression of the target patient, and to acquire an RGB video frame sequence and a depth video frame sequence based on the RGB-D video stream. Key point detection unit (2), the key point detection unit (2) is used to perform 3D key point detection on the RGB video frame sequence according to the depth video frame sequence, so as to obtain the 3D key point sequence of the RGB video frame; Motion enhancement unit (3) is used to perform micro-motion amplification on the RGB video frame sequence to obtain an enhanced RGB video frame sequence; Visual extraction unit (4), the visual extraction unit (4) is used to extract visual features from each enhanced RGB video frame in the enhanced RGB video frame sequence according to the convolutional neural network, so as to obtain the visual feature sequence of the target patient's face; Action modeling unit (5), the action modeling unit (5) is used to construct each muscle motion vector field in the RGB video frame according to the 3D key point sequence of the RGB video frame, and to construct an action map according to each muscle motion vector field and the anatomical relationship between each muscle motion vector field; Motion extraction unit (6), the motion extraction unit (6) is used to extract motion features from the RGB video frame sequence through the motion graph according to the graph neural network, so as to obtain the motion feature sequence of the target patient's face; Pain rating unit (7), the pain rating unit (7) is used to grade the pain expression of the target patient's face according to the visual feature sequence and motion feature sequence of the target patient's face, so as to obtain the pain label of the target patient's face; Based on the depth video frame sequence, 3D keypoint extraction is performed on the RGB video frame sequence to obtain a 3D keypoint sequence of the RGB video frames, including: The depth video frame and the RGB video frame are image aligned using camera calibration parameters to generate registered depth video frames and RGB video frames. 3D keypoint detection is performed on the registered depth video frames and RGB video frames according to the improved HRNet-W64 model to obtain the first 3D keypoint sequence of the RGB video frames. Based on a preset regression loss function, keypoint regression is performed on the first 3D keypoint sequence of the RGB video frame to obtain the second 3D keypoint sequence of the RGB video frame. The regression loss function is as follows: Among them, L reg Let N represent the regression loss function, N represent the total number of 3D keypoints in the first 3D keypoint sequence, and φ(I) represent the regression loss function. t D t ) i This represents the first 3D keypoint extracted by the improved HRNet-W64 model. Represents true 3D keypoint annotations; An anatomical loss function is constructed based on a preset muscle motion correlation matrix and gradient constraints. This anatomical loss function is then used to apply anatomical constraints to the second 3D keypoint sequence of the RGB video frame, resulting in the final 3D keypoint sequence of the RGB video frame. The anatomical loss function is as follows: Among them, L anatomy Represents the anatomical loss function. This represents the gradient of the motion of the second 3D keypoint in adjacent RGB video frames, and P t M represents the sequence of the second 3D key points in the t-th RGB video frame. muscle Let M represent the muscle motion correlation matrix, and if the second 3D keypoint i and the second 3D keypoint j belong to the same muscle group, then M... muscle [i,j] = 1, otherwise, then M muscle [i,j]=0.
Citation Information
Patent Citations
Facial paralysis detection method based on visual perception and audio information
CN112308037A
Child pain multi-modal data fusion evaluation method based on deep learning
CN118452821A