Facial pain grading methods, systems, devices, media, and products
By extracting the characteristic points of pain expressions in facial videos and building a model using optical flow method and long and short-term memory networks, the problem of single feature extraction and high computing resources in the prior art pain signal grading method is solved, and efficient and accurate pain grading is achieved.
Patent Information
- Application Number
- CN202510612622.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
AI Technical Summary
The existing pain signal grading methods have relatively single features extraction and processing methods, which leads to insufficient feature expression capabilities, making it difficult to capture subtle differences between different pain levels, and high computing resources demand, which affects grading efficiency.
By extracting the pain expression feature points in the facial video sample data, using the optical flow method to determine the relative motion feature samples, construct a sample training set and train a facial pain grading prediction model using long and short-term memory network to predict the pain level.
It improves the feature extraction and processing efficiency of pain grading, reduces the computational complexity, accurately captures the dynamic changes of pain microexpressions, and improves the accuracy and efficiency of pain grading.
Smart Images

Figure CN120472514A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a facial pain grading method, system, device, medium and product. Background Art
[0002] Currently, common pain signal grading technologies rely on physiological signals and hybrid deep learning networks for automated nociceptive pain assessment. These methods typically rely on automated feature learning and multimodal fusion for feature extraction and processing. These methods often require multimodal data or rely on single-dimensional feature extraction, ignoring the combined effects of directional movement and intensity changes in facial microexpressions. This results in insufficient feature expression and makes it difficult to capture subtle differences between pain levels.
[0003] Alternatively, facial feature extraction requires covering the entire face, resulting in high computational resource requirements, which may be limited in real-time applications or resource-constrained scenarios. Overall, current pain signal grading methods use relatively simple feature extraction and processing methods, and the generated features are complex, which affects feature expressiveness and reduces grading efficiency. Summary of the Invention
[0004] In view of this, the present invention provides a facial pain grading method, system, device, medium and product, which solves the technical problems that the feature extraction and processing methods of the current pain signal grading method are relatively simple, the generated features are complex, which affects the feature expression ability and reduces the grading efficiency.
[0005] A first aspect of the present invention provides a facial pain grading method, comprising:
[0006] Extracting pain expression feature point samples from each frame of facial image data in the facial video sample data according to the acquired facial video sample data;
[0007] Determining relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method based on the pain expression feature point samples;
[0008] constructing a sample training set according to the plurality of relative motion feature samples and facial pain level samples corresponding to the relative motion feature samples;
[0009] The long short-term memory network is trained using the sample training set to obtain a facial pain grading prediction model;
[0010] The relative motion feature samples of the pain expression feature points acquired in real time are input into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
[0011] Preferably, extracting pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data includes:
[0012] According to the acquired facial video sample data, the pain expression feature point samples in each frame of the facial image data are tracked and located based on a feature point detection algorithm to obtain the pain expression feature point samples in each frame of the facial image data; wherein the pain expression feature point samples are eyebrow feature points in the eyebrow area.
[0013] Preferably, the feature point detection algorithm is a 68-point facial feature point detector, and the eyebrow feature points include the eyebrow starting point, eyebrow ending point and eyebrow curvature change point captured by the 68-point facial feature point detector.
[0014] Preferably, determining the relative motion feature samples of the pain expression feature point samples between adjacent frames by an optical flow method based on the pain expression feature point samples includes:
[0015] Based on the optical flow method, under the condition of constant optical flow, the motion vector of the pain expression feature point sample in each adjacent frame is determined; wherein, the motion vector includes the pain expression feature point respectively in and Movement speed in the axial direction;
[0016] Determining, based on the motion vector, a first motion feature sample of the pain expression feature point sample between adjacent frames; wherein the first motion feature sample includes a motion angle and a motion amplitude;
[0017] Determine a second motion feature sample based on the first motion feature sample; wherein the second motion feature sample includes an overall facial motion angle and an overall facial motion amplitude;
[0018] Feature fusion is performed based on the second motion feature sample to obtain a relative motion feature sample; wherein the relative motion feature sample includes a comprehensive motion feature and a comprehensive motion change trend feature, wherein the comprehensive motion feature is used to characterize the comprehensive change amount of the motion angle and motion amplitude of the pain expression feature point sample, and the comprehensive motion change trend feature is used to characterize the comprehensive change frequency of the motion angle and motion amplitude of the pain expression feature point sample.
[0019] Preferably, the method further comprises:
[0020] A normalization operation is performed on the motion feature samples.
[0021] Preferably, the long short-term memory network is trained using the sample training set to obtain a facial pain grading prediction model, including:
[0022] The long short-term memory network is constructed, and based on the cross entropy loss function, the long short-term memory network is trained with the sample training set to obtain the facial pain grading prediction model.
[0023] In a second aspect, the present invention provides a facial pain grading system, comprising:
[0024] A feature point extraction module is used to extract pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data;
[0025] A motion feature extraction module, configured to determine, based on the pain expression feature point samples, relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method;
[0026] A training set construction module, configured to construct a sample training set based on a plurality of the relative motion feature samples and facial pain level samples corresponding to the relative motion feature samples;
[0027] A model training module, configured to train a long short-term memory network using the sample training set to obtain a facial pain grading prediction model;
[0028] The pain grading prediction module is used to input the relative motion feature samples of the pain expression feature points acquired in real time into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
[0029] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor performs the steps of the facial pain grading method as described in the first aspect.
[0030] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements the steps of the facial pain grading method as described in the first aspect.
[0031] In a fifth aspect, the present invention provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is caused to perform the steps of the facial pain grading method as described in the first aspect.
[0032] It can be seen from the above technical solutions that the present invention extracts pain expression feature point samples from facial image data of each frame in facial video sample data, thereby selecting feature points that effectively reflect the changes in pain signals, and performs feature extraction in a targeted manner. While reducing the amount of extracted data and reducing the computational complexity, the feature extraction and generation process as well as the classification efficiency are improved. The relative motion feature samples of the pain expression feature point samples between adjacent frames are determined by the optical flow method, thereby capturing the relative motion trend of the pain micro-expression evolving over time. A sample training set is constructed through the relative motion feature samples and the facial pain level samples, and the long short-term memory network is trained using the sample training set. The facial pain level is predicted by the trained facial pain grading prediction model, thereby improving the pain grading efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0034] Figure 1 A diagram illustrating an application environment of a facial pain grading method provided by an embodiment of the present invention;
[0035] Figure 2 A flowchart of a facial pain grading method provided by an embodiment of the present invention;
[0036] Figure 3 A layout diagram of key facial feature points provided by an embodiment of the present invention;
[0037] Figure 4 A schematic structural diagram of a facial pain grading system provided by an embodiment of the present invention;
[0038] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0040] The facial pain grading method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or it can be placed on the cloud or other network servers. The terminal 101 or the server 102 extracts the pain expression feature point samples in the facial image data of each frame in the facial video sample data based on the acquired facial video sample data; based on the pain expression feature point samples, the relative motion feature samples of the pain expression feature point samples between adjacent frames are determined by the optical flow method; based on the multiple relative motion feature samples and the facial pain level samples corresponding to the relative motion feature samples, a sample training set is constructed; the long short-term memory network is trained with the sample training set to obtain a facial pain grading prediction model; the relative motion feature samples of the pain expression feature points acquired in real time are input into the facial pain grading prediction model to obtain the facial pain level prediction result corresponding to the pain expression feature points.
[0041] The terminal 101 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and the like.
[0042] The server 102 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0043] like Figure 2 As shown, the embodiment of the present application provides a facial pain grading method, which is applied to Figure 1 The terminal 101 or the server 102 in the example is used to illustrate, including the following steps S1 to S5.
[0044] Step S1: extracting pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data.
[0045] The facial video sample data includes multiple frames of facial image data, and by extracting pain expression feature point samples from each frame of facial image data, the pain expression feature point samples are used to characterize feature points of pain signal changes.
[0046] It's understandable that this method doesn't directly extract feature signals from the entire face. Instead, it selects relatively stable features from a number of facial features that effectively reflect changes in pain signals, performing targeted feature extraction. This reduces the amount of extracted data and computational complexity, while also improving the efficiency of the feature extraction and generation process and the subsequent classification stage.
[0047] In some specific embodiments, extracting pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data includes:
[0048] According to the acquired facial video sample data, the pain expression feature point samples in each frame of facial image data are tracked and located based on the feature point detection algorithm to obtain the pain expression feature point samples in each frame of facial image data; wherein, the pain expression feature point samples are eyebrow feature points in the eyebrow area.
[0049] Among them, the feature point detection algorithm can be a 68-point facial feature point detector.
[0050] 68-point face landmark detector (shape_predictor_68_face_landmarks) is a face landmark detection algorithm based on machine learning, which can automatically identify and locate 68 key feature points from the input facial image, such as Figure 3 As shown in the figure, these feature points cover areas such as eyebrows, eyes, nose, mouth, and facial contours. By adopting this algorithm, accurate tracking and positioning of the feature point samples of pain expression can be achieved. In specific implementation, each frame of facial image can be pre-processed using a 68-point facial feature point detector. These 68 key points are labeled in a fixed order.
[0051] Among them, the eyebrow feature points include the eyebrow starting point, eyebrow ending point and eyebrow curvature change point captured by the 68-point facial feature point detector.
[0052] like Figure 3 As shown, the eyebrow landmarks are points 19-24 from the Dlib library, captured using a 68-point facial landmark detector. These points are located at the beginning and end of the eyebrows, as well as where the curvature of the eyebrows changes. These landmarks were selected because eyebrow movement is often the most noticeable in expressions of pain and best reflects changes in pain intensity. By tracking the subtle movements of these landmarks, subtle changes in micro-expressions of pain can be captured.
[0053] Dlib is an open-source library that includes various machine learning algorithms, including a 68-point facial landmark detector. This detector accurately captures landmarks within the eyebrow region, which often exhibit noticeable changes during expressions of pain, such as frowning or raising the eyebrows. By tracking the movement of these landmarks, changes in pain expressions can be more accurately analyzed, thereby improving the accuracy of pain grading.
[0054] Through experiments and comparisons, we found that eyebrows are the most stable points in a video. Their motion trajectory is least affected by external factors (such as head shaking), making them the most sensitive to micro-expression changes and, therefore, the easiest to extract. Even with interference from head shaking, eyebrows are less affected than other parts of the face. Therefore, the representation and extraction of micro-expression changes are easily affected by external shaking, while key points in other parts of the face are more susceptible to external interference, making micro-expression information more difficult to extract.
[0055] Furthermore, the feature point extraction in this method is targeted at specific scenarios (such as medical settings). Facial video data is acquired specifically, and in these scenarios, the face is usually captured as a whole. Therefore, cases without eyebrows are usually not considered.
[0056] Step S2: determining relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method based on the pain expression feature point samples.
[0057] Among them, the optical flow method is a technology for estimating the motion of objects between video frames. Its core assumption is that the brightness of pixels remains unchanged over a short period of time.
[0058] In some embodiments, determining relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method based on the pain expression feature point samples includes:
[0059] Step S201: Based on the optical flow method, under the condition of constant optical flow, determine the motion vector of the pain expression feature point sample in each adjacent frame; wherein the motion vector includes the pain expression feature point in each adjacent frame. and The speed of movement in the axis direction.
[0060] Let I(x,y,t) represent the brightness of pixel (x,y) at time t, then for adjacent frames, we have:
[0061]
[0062] in, Pixels in and The speed of movement in the axial direction, is the time interval between frames.
[0063] Performing a first-order Taylor expansion on the above equation at (x, y, t) and ignoring higher-order terms yields:
[0064]
[0065] Combined with the brightness invariance assumption, the left and right sides are equal, and after sorting, we can get the optical flow constant constraint equation:
[0066]
[0067] in:
[0068]
[0069]
[0070]
[0071] Where, 、 、 The images are The partial derivative in the direction Partial derivatives of the direction, the rate of change of the image over time.
[0072] The optical flow constrained equation forms the theoretical basis for calculating feature point motion vectors using the LK optical flow method. The resulting motion vectors are used for feature extraction (direction and magnitude). These features are further processed to support the five-class classification task of the LSTM model, ensuring that the classification process captures the dynamic changes in micro-expressions. Although the optical flow constraint equation is no longer used directly during feature processing and classification, the derived motion vectors are used throughout the entire technical process and serve as the core basis for feature generation and classification.
[0073] Since a single equation cannot be solved for two unknowns and This method uses the optical flow (Lucas-KanadeLK) method, assuming that the optical flow is constant in the local area. Define a window , solved by minimizing the weighted residual:
[0074]
[0075] in, is a weight function (such as Gaussian weight) that emphasizes the pixels in the center of the window. and Find the partial derivative and set it to zero, is a pixel within the local window Ω.
[0076] The window is defined around the target feature point (e.g., points 19-24 in the eyebrow region). By weighted summing the residuals of the optical flow constrained error for all pixels p within the window, the formula expands the optical flow constrained error equation into an optimization problem, providing sufficient information to estimate the optical flow vector (u, v). Accumulating constraint information across multiple pixels ensures accurate estimation of the optical flow vector, facilitating subsequent feature extraction and pain classification.
[0077] Solving the linear equations yields:
[0078]
[0079] The solution can be obtained by matrix inversion and , thus obtaining the motion vector of each pixel .
[0080] The formula for minimizing the weighted residual is an extension of the optical flow constraint equation. This equation provides theoretical constraints for individual pixels, but cannot be directly solved due to underdetermination. The LK optical flow method applies the optical flow constancy constraint equation to multiple pixels within a window through the local window assumption and least squares optimization, generating a solvable linear equation system.
[0081] In this embodiment, we focus on feature points 19-24 in the eyebrow area of the face. These points show significant and stable motion changes in the pain micro-expression. Using the LK optical flow method, the motion vectors of these feature points between adjacent frames are calculated as:
[0082] ,
[0083] Where, Indicates the feature point number (19 to 24).
[0084] Step S202: Determine first motion feature samples of the pain expression feature point samples between adjacent frames based on the motion vector; wherein the first motion feature samples include motion angle and motion amplitude.
[0085] To quantify motion characteristics, the following features are defined:
[0086] Movement angle (direction) :
[0087]
[0088] Where, Feature points angle of movement.
[0089] Range of motion :
[0090]
[0091] Where, Feature points range of motion.
[0092] Step S203: Determine a second motion feature sample based on the first motion feature sample; wherein the second motion feature sample includes an overall facial motion angle and an overall facial motion amplitude.
[0093] In order to eliminate the influence of overall facial shaking, the relative motion feature is calculated, that is, the second motion feature sample is:
[0094] Relative direction :
[0095]
[0096] Relative amplitude :
[0097]
[0098] in, and is the direction and magnitude of the overall facial motion, which can be estimated by the average value of all feature points:
[0099]
[0100] Here, (Corresponding to feature points 19-24).
[0101] Step S204: perform feature fusion based on the second motion feature sample to obtain a relative motion feature sample; wherein the relative motion feature sample includes a comprehensive motion feature and a comprehensive motion change trend feature, wherein the comprehensive motion feature is used to characterize the comprehensive change amount of the motion angle and motion amplitude of the pain expression feature point sample, and the comprehensive motion change trend feature is used to characterize the comprehensive change frequency of the motion angle and motion amplitude of the pain expression feature point sample.
[0102] In this embodiment, in order to capture the dynamic changes of micro-expressions, the present application performs feature fusion based on the second motion feature sample to generate two sets of features as relative motion feature samples, namely:
[0103] Comprehensive sports characteristics :
[0104]
[0105] Among them, the comprehensive motion feature combines direction and amplitude to reflect the comprehensive motion characteristics of the feature point.
[0106] Comprehensive sports change trend characteristics :
[0107]
[0108] It represents the time rate product of direction and magnitude, capturing the dynamic trend of motion. The rate of change is approximated by the difference between adjacent frames:
[0109]
[0110] in, Typically set to 1 (in frames).
[0111] This method multiplies the direction and amplitude of the optical flow to generate robust features. This feature generation method improves the expressiveness of the features, thereby improving classification accuracy. The product of the two effectively captures the dynamic characteristics of micro-expression changes. This fusion method enhances the expressiveness of the features, making them more capable of representing complex facial movement patterns related to pain.
[0112] This method not only generates static, comprehensive features but also calculates the rate of change of direction and amplitude, multiplying these to generate dynamic features. This method fully leverages the dynamic information of the time series, capturing subtle trends in the evolution of pain micro-expressions over time. Compared to existing technologies, this method enhances the robustness of features through the product of rate of change, enabling the model to more accurately identify transitions between pain levels.
[0113] This method focuses on six key feature points in the facial eyebrow region, significantly reducing computational complexity and improving processing efficiency. It also efficiently extracts features by combining motion direction and amplitude, achieving a balanced balance between resources and performance while maintaining high classification accuracy. This method enables real-time pain assessment using fewer features, making it suitable for resource-limited scenarios.
[0114] In some embodiments, to eliminate dimensional differences and improve model stability, a normalization operation is performed on the motion feature samples.
[0115] Specifically, standardization operations include:
[0116]
[0117] Where, is the standardized comprehensive motion characteristics, is the standardized comprehensive movement change trend characteristic, 、 Comprehensive sports characteristics The mean and standard deviation over all feature points and time steps, 、 Comprehensive movement change trend characteristics The mean and standard deviation over all feature points and time steps.
[0118] in,
[0119]
[0120] in, and Calculate similarly.
[0121] Step S3: construct a sample training set based on the multiple relative motion feature samples and the facial pain level samples corresponding to the relative motion feature samples.
[0122] The team used optical flow to extract relative motion feature samples from the eyebrow region, capturing the dynamic changes in micro-expressions. Different pain levels exhibit distinct characteristic patterns (feature point motion amplitude and frequency), and the statistical and dynamic characteristics of the features are distinguishable across different pain levels. By defining the facial pain levels corresponding to different relative motion feature samples, a mapping relationship between relative motion feature samples and facial pain level samples was obtained, and a sample training set was constructed.
[0123] The facial pain level sample includes multiple levels. For example, the facial pain level sample includes five pain levels: BL1 (no pain), PA1 (mild pain), PA2 (moderate pain), PA3 (severe pain), and PA4 (extreme pain).
[0124] Step S4: training the long short-term memory network using the sample training set to obtain a facial pain grading prediction model.
[0125] Among them, the Long Short-Term Memory (LSTM) network is a special type of recurrent neural network (RNN) suitable for processing and predicting important events with long intervals and delays in time series. The LSTM network solves the long-term dependency problem of traditional RNNs by introducing a clever self-loop design. The LSTM unit has three gating mechanisms: a forget gate, an input gate, and an output gate. These gating mechanisms control the flow of information through a sigmoid function, determining which information is retained, discarded, or updated. The forget gate determines which information in the cell state at the previous moment should be forgotten, the input gate determines which new information in the current input should be added to the cell state, and the output gate controls the output of the cell state. This design enables the LSTM to capture long-term dependencies in sequential data, making it ideal for processing time series data such as pain expressions, where the dynamic changes in micro-expressions are crucial for pain grading.
[0126] In the embodiment of the present application, the input sequence is (Here is the feature and ), the hidden state is , the cell state is , the update formula of LSTM is as follows:
[0127] Forget Gate:
[0128]
[0129] Control the cell state at the previous moment degree of forgetfulness.
[0130] Input Gate:
[0131]
[0132] Decide which information to update, is a candidate cell state.
[0133] Cell status update:
[0134]
[0135] Combine forgetting and input to update the current state.
[0136] Output gate:
[0137]
[0138] It is the hidden state at the current moment, which is used for subsequent classification.
[0139] in,
[0140]
[0141] is the sigmoid function,
[0142]
[0143] is the hyperbolic tangent function, and are trainable weights and biases.
[0144] Among them, the input layer: receives time series features, and the input dimension is the number of features for each time step (for example, 6 feature points and combination, dimension is 12).
[0145] STM layer: processes time series, outputs hidden states, and has adjustable hidden layer dimensions (such as hidden_size).
[0146] Fully connected layer: maps the hidden state of the last time step of LSTM to 5 categories (BL1, PA1, PA2, PA3, PA4).
[0147] In some embodiments, a long short-term memory network is constructed, and based on a cross-entropy loss function, the long short-term memory network is trained using a sample training set to obtain a facial pain grading prediction model.
[0148] Among them, the loss function adopts the cross entropy loss function, which is suitable for five-category tasks. The cross entropy loss function is:
[0149]
[0150] Where, is the loss value, is the true label (one-hot encoding), is the model prediction probability (normalized by softmax).
[0151] Optimizer: Use the Adam optimizer with an adjustable learning rate (e.g., 0.001).
[0152] Training parameters: batch size (such as 16), number of training rounds (such as 80), input sequence length is the number of video frames (such as 50 frames).
[0153] Step S5: inputting the relative motion feature samples of the pain expression feature points acquired in real time into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
[0154] Among them, the feature sequence of the test video is predicted through the trained facial pain grading prediction model, and five pain levels are output: BL1 (no pain), PA1 (mild pain), PA2 (moderate pain), PA3 (severe pain), and PA4 (extreme pain). This overcomes the limited multi-category classification capabilities of existing technologies and improves their practical value.
[0155] It should be noted that the embodiment of the present application extracts pain expression feature point samples from the facial image data of each frame in the facial video sample data, thereby selecting feature points that effectively reflect the changes in pain signals, and performs targeted feature extraction. While reducing the amount of extracted data and reducing the computational complexity, it improves the feature extraction and generation process and classification efficiency, and determines the relative motion feature samples of the pain expression feature point samples between adjacent frames through the optical flow method, thereby capturing the relative motion trend of the pain micro-expression evolving over time. A sample training set is constructed through the relative motion feature samples and the facial pain level samples, and the long short-term memory network is trained using the sample training set. The facial pain level is predicted through the trained facial pain grading prediction model, thereby improving the pain grading efficiency.
[0156] Based on the same inventive concept, an embodiment of the present application further provides a facial pain grading system for implementing the above-mentioned facial pain grading method.
[0157] The implementation solution provided by the system to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more facial pain grading system embodiments provided below can refer to the limitations on the facial pain grading method above and will not be repeated here.
[0158] like Figure 4 As shown, the embodiment of the present application provides a facial pain grading system, comprising:
[0159] The feature point extraction module 100 is used to extract the pain expression feature point samples in each frame of facial image data in the facial video sample data based on the acquired facial video sample data;
[0160] A motion feature extraction module 200 is used to determine relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method based on the pain expression feature point samples;
[0161] A training set construction module 300 is used to construct a sample training set based on a plurality of relative motion feature samples and facial pain level samples corresponding to the relative motion feature samples;
[0162] A model training module 400 is used to train a long short-term memory network using a sample training set to obtain a facial pain grading prediction model;
[0163] The pain grading prediction module 500 is used to input the relative motion feature samples of the pain expression feature points acquired in real time into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
[0164] In some embodiments, the feature point extraction module 100 is configured to:
[0165] According to the acquired facial video sample data, the pain expression feature point samples in each frame of facial image data are tracked and located based on the feature point detection algorithm to obtain the pain expression feature point samples in each frame of facial image data; wherein, the pain expression feature point samples are eyebrow feature points in the eyebrow area.
[0166] In some embodiments, the feature point detection algorithm is a 68-point facial feature point detector, and the eyebrow feature points include the eyebrow starting point, eyebrow ending point, and eyebrow curvature change point captured by the 68-point facial feature point detector.
[0167] In some embodiments, the motion feature extraction module 200 is configured to:
[0168] Based on the optical flow method, the motion vectors of the pain expression feature point samples in each adjacent frame are determined under the condition of constant optical flow; wherein the motion vectors include the pain expression feature points in each adjacent frame. and Movement speed in the axial direction;
[0169] Determine, based on the motion vector, a first motion feature sample of the pain expression feature point sample between adjacent frames; wherein the first motion feature sample includes a motion angle and a motion amplitude;
[0170] Determine a second motion feature sample based on the first motion feature sample; wherein the second motion feature sample includes an overall facial motion angle and an overall facial motion amplitude;
[0171] Feature fusion is performed based on the second motion feature sample to obtain a relative motion feature sample; wherein the relative motion feature sample includes a comprehensive motion feature and a comprehensive motion change trend feature, wherein the comprehensive motion feature is used to characterize the comprehensive change amount of the motion angle and motion amplitude of the pain expression feature point sample, and the comprehensive motion change trend feature is used to characterize the comprehensive change frequency of the motion angle and motion amplitude of the pain expression feature point sample.
[0172] In some embodiments, the system further includes: a standardization module for performing a standardization operation on the motion feature samples.
[0173] In some embodiments, the model training module 400 is used to:
[0174] A long short-term memory network was constructed and trained with a sample training set based on the cross-entropy loss function to obtain a facial pain grading prediction model.
[0175] like Figure 5 As shown, an embodiment of the present application provides an electronic device, the electronic device 10 includes a memory 20 and a processor 30, the memory 20 stores a computer program, and when the computer program is executed by the processor 30, the processor 30 performs the steps of the facial pain grading method in the above embodiment.
[0176] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the steps of the facial pain grading method in the above embodiment are implemented.
[0177] An embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer performs the steps of the facial pain grading method described in the above embodiment.
[0178] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, electronic devices, computer storage media, and computer program products can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0179] It should be noted that the user information (including but not limited to user facial images, user portrait information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0180] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0181] It should be understood that, although the various steps in the flowcharts involved in the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0182] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media, computer program products and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0183] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0184] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0185] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the method described in each embodiment of the present invention via a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0186] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A facial pain grading method, characterized in that: include: Extracting pain expression feature point samples from each frame of facial image data in the facial video sample data according to the acquired facial video sample data; Determining relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method based on the pain expression feature point samples; constructing a sample training set according to the plurality of relative motion feature samples and facial pain level samples corresponding to the relative motion feature samples; The long short-term memory network is trained using the sample training set to obtain a facial pain grading prediction model; The relative motion feature samples of the pain expression feature points acquired in real time are input into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
2. The facial pain grading method according to claim 1, characterized in that: The step of extracting pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data includes: According to the acquired facial video sample data, the pain expression feature point samples in each frame of the facial image data are tracked and located based on a feature point detection algorithm to obtain the pain expression feature point samples in each frame of the facial image data; wherein the pain expression feature point samples are eyebrow feature points in the eyebrow area.
3. The facial pain grading method according to claim 2, characterized in that: The feature point detection algorithm is a 68-point facial feature point detector, and the eyebrow feature points include the eyebrow starting point, eyebrow ending point and eyebrow curvature change point captured by the 68-point facial feature point detector.
4. The facial pain grading method according to claim 1, wherein: The step of determining relative motion feature samples of the pain expression feature point samples between adjacent frames by an optical flow method based on the pain expression feature point samples includes: Based on the optical flow method, under the condition of constant optical flow, the motion vector of the pain expression feature point sample in each adjacent frame is determined; wherein, the motion vector includes the pain expression feature point respectively in and The speed of movement in the axial direction; Determining, based on the motion vector, a first motion feature sample of the pain expression feature point sample between adjacent frames; wherein the first motion feature sample includes a motion angle and a motion amplitude; Determine a second motion feature sample based on the first motion feature sample; wherein the second motion feature sample includes an overall facial motion angle and an overall facial motion amplitude; Feature fusion is performed based on the second motion feature sample to obtain a relative motion feature sample; wherein the relative motion feature sample includes a comprehensive motion feature and a comprehensive motion change trend feature, wherein the comprehensive motion feature is used to characterize the comprehensive change amount of the motion angle and motion amplitude of the pain expression feature point sample, and the comprehensive motion change trend feature is used to characterize the comprehensive change frequency of the motion angle and motion amplitude of the pain expression feature point sample.
5. The facial pain grading method according to any one of claims 1 to 4, characterized in that: Also includes: A normalization operation is performed on the motion feature samples.
6. The facial pain grading method according to claim 1, wherein: The long short-term memory network is trained using the sample training set to obtain a facial pain grading prediction model, including: The long short-term memory network is constructed, and based on the cross entropy loss function, the long short-term memory network is trained with the sample training set to obtain the facial pain grading prediction model.
7. A facial pain grading system, characterized in that: include: A feature point extraction module is used to extract pain expression feature point samples from each frame of facial image data in the facial video sample data based on the acquired facial video sample data; A motion feature extraction module, configured to determine, based on the pain expression feature point samples, relative motion feature samples of the pain expression feature point samples between adjacent frames using an optical flow method; A training set construction module, configured to construct a sample training set based on a plurality of the relative motion feature samples and facial pain level samples corresponding to the relative motion feature samples; A model training module, configured to train a long short-term memory network using the sample training set to obtain a facial pain grading prediction model; The pain grading prediction module is used to input the relative motion feature samples of the pain expression feature points acquired in real time into the facial pain grading prediction model to obtain the facial pain level prediction results corresponding to the pain expression feature points.
8. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor performs the steps of the facial pain grading method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the steps of the facial pain grading method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer is caused to perform the steps of the facial pain grading method according to any one of claims 1 to 6.