A method, system, device and storage medium for identifying expression intensity

Through label probability distribution and neural network training model, the Gaussian distribution is constructed using the intensity distribution of multi-frame samples in the expression sequence, which solves the problem of inaccurate recognition of absolute value of expression and achieves higher accuracy and precision of expression intensity recognition.

CN115346086BActive Publication Date: 2025-09-05HUAZHONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210984712.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2025-09-05
Estimated Expiration
2042-08-17

AI Technical Summary

Technical Problem

Existing facial expression recognition methods perform poorly in estimating the absolute value of expressions and have difficulty in accurately identifying the intensity of expressions.

Method used

The intensity label is represented by label probability distribution. The neural network training model is used to determine the parameters of the label probability distribution using the original intensity distribution of multiple samples. The statistical information of the pseudo-intensity label and the true label in a short time window is combined to construct a Gaussian distribution, suppress label noise, and improve the accuracy of expression intensity estimation.

Benefits of technology

It improves the accuracy of expression intensity recognition, makes full use of the sequential information and intensity distribution characteristics in the sequence, can learn supervision information frame by frame without increasing the annotation burden, effectively suppresses label noise, and improves the accuracy of expression intensity estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346086B_ABST
    Figure CN115346086B_ABST
Patent Text Reader

Abstract

This application discloses a method, system, device, and storage medium for identifying facial expression intensity. The method includes: obtaining a data sample set, the data sample set comprising several facial expression sequences; collecting paired samples from the data sample set to construct a training sample set, wherein the intensity label of each sample in the training sample set is represented by a label probability distribution, the parameters of which are determined by the distribution of the original intensity labels of all samples within an observation window containing the target sample in the facial expression sequence; and using the training sample set to train a neural network-based facial expression intensity recognition model. This invention can improve the accuracy of facial expression intensity recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of expression recognition technology, and more specifically, to an expression intensity recognition method, system, device and storage medium. Background Art

[0002] Facial expressions play an important role in people's daily interactions, as they directly convey people's emotions. With the development of computer science and technology, automated facial expression analysis has become a research topic that has received increasing attention.

[0003] In recent years, numerous methods have been proposed in the field of facial expression recognition, achieving promising results. However, only a few have considered the emotional intensity of expressions. Variations in facial expression intensity provide temporal dynamics of facial behavior, which is crucial for interpreting the meaning of expressions. However, while existing ranking-based methods can effectively estimate the relative intensity of paired expressions in a continuous expression sequence, they perform poorly in estimating the absolute value of expressions. Summary of the Invention

[0004] In response to at least one shortcoming or improvement need in the prior art, the present invention provides a method, system, device and storage medium for expression intensity recognition, which can improve the accuracy of expression intensity recognition.

[0005] To achieve the above object, according to a first aspect of the present invention, a method for identifying expression intensity is provided, comprising:

[0006] Acquire a data sample set, wherein the data sample set includes a plurality of expression sequences;

[0007] Collecting paired samples from the data sample set to construct a training sample set, wherein the intensity label of each sample in the training sample set is represented by a label probability distribution, wherein the parameters of the label probability distribution are determined by the distribution of the original intensity labels of all samples within an observation window containing the target sample in the expression sequence where the target sample is located;

[0008] The training sample set is used to train an expression intensity recognition model based on a neural network.

[0009] Furthermore, the samples in the training sample set are recorded as The label probability distribution is recorded as Located at the center of the observation window, the observation window is expressed as [i-τ,i+τ], indicating The i-τ frame to i+τ frame samples in the expression sequence constitute the observation window, τ is a positive integer, and the distribution of the 2τ+1 frame samples in the observation window is determined. Parameters.

[0010] Furthermore, if we set Satisfies Gaussian distribution, Expressed as:

[0011]

[0012] Where y d yes The dth element represents The similarity between samples with intensity level d-1, D is The total number of elements, μ and σ 2 are the mean and variance of the Gaussian distribution, and Z is to ensure The normalization factor of .

[0013] Further,

[0014] y m Represents the original intensity label of the sample in the observation window.

[0015] Furthermore, each expression sequence includes multiple frames of expressions evolving from neutral expression frames to peak expression frames. If the original intensity labels of the target sample and at least some of the multiple samples adjacent to the target sample in the expression sequence to which the target sample belongs do not exist, a pseudo-intensity label is calculated based on the maximum intensity value and the minimum intensity value of the expression sequence to which the target sample belongs, and the generated pseudo-intensity label is used as the original intensity label of the sample.

[0016] Furthermore, if y m If it does not exist, a pseudo-intensity label is generated. The pseudo labels generated As the original intensity label of the sample, The calculation formula is:

[0017]

[0018] I h Indicates the maximum intensity value of the peak value of the expression sequence to which the target sample belongs, I l Indicates the lowest intensity value of the expression sequence to which the target sample belongs, δ is the representation function, If m<a k , indicating that the function output is 1, otherwise it is 0, If m≥a k , indicating that the function output is 1, otherwise it is 0, a k Indicates the frame number corresponding to the maximum intensity value in the expression sequence, m is y m The frame number of the corresponding sample in the expression sequence.

[0019] Furthermore, the loss function calculation formula for training is:

[0020]

[0021]

[0022]

[0023] is the regularization term, N is the minimum batch size, and The samples are and The label probability distribution of and Represents samples and The output after the convolutional layer and the fully connected layer, E(·) represents the mathematical expectation of the calculated predicted distribution.

[0024] According to a second aspect of the present invention, there is also provided an expression intensity recognition system, comprising:

[0025] An acquisition module, configured to acquire a data sample set, wherein the data sample set includes a plurality of expression sequences;

[0026] a label generation and training set construction module, configured to collect paired samples from the data sample set to construct a training sample set, wherein the intensity label of each sample in the training sample set is represented by a label probability distribution, wherein the parameters of the label probability distribution are determined by the distribution of the original intensity labels of all samples within an observation window containing the target sample in the expression sequence where the target sample is located;

[0027] The training module is used to train the expression intensity recognition model based on the neural network using the training sample set.

[0028] According to the third aspect of the present invention, an electronic device is also provided, which includes at least one processor and at least one storage module, wherein the storage module stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any one of the above methods.

[0029] According to a fourth aspect of the present invention, a storage medium is provided, which stores a computer program executable by a processor, and when the computer program runs on the processor, the processor executes the steps of any one of the above methods.

[0030] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0031] (1) The present invention uses label probability distribution to represent intensity labels and uses the intensity distribution of original labels of multiple samples to determine the parameters of label probability distribution. This can fully utilize the order information and intensity distribution characteristics in the sequence and improve the accuracy of the description of expressions with continuously changing intensity.

[0032] (2) This paper proposes a unified label distribution generation framework for datasets with and without intensity labels, which can learn supervision information frame by frame without increasing the annotation burden.

[0033] (3) The present invention uses the statistical information of pseudo labels and true labels in a short time window to construct a Gaussian distribution under semi-supervised settings and fully supervised settings, respectively, to effectively suppress label noise and improve the accuracy of expression intensity estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0035] Figure 1 is a flow chart of a method for estimating expression intensity according to an embodiment of the present invention;

[0036] Figure 2 This is a network structure diagram of the integrated label probability distribution learning and sequential regression model of an embodiment of the present invention;

[0037] Figure 3 is a schematic diagram of a sliding window used in a label probability distribution generation framework according to an embodiment of the present invention;

[0038] Figure 4 Schematic diagram of generating label probability distribution using real labels and pseudo labels under full supervision and semi-supervision, respectively, according to an embodiment of the present invention;

[0039] Figure 5 Schematic diagram of the Pearson correlation coefficient of the six basic expression intensity estimation results obtained from the ablation experiment performed on Extend CK according to an embodiment of the present invention;

[0040] Figure 6 Schematic diagram of the intra-group correlation coefficients of the six basic expression intensity estimation results obtained from the ablation experiment performed on Extend CK according to an embodiment of the present invention;

[0041] Figure 7 FIG. 4 is a schematic diagram of the mean absolute errors of the estimation results of six basic expression intensities obtained by performing an ablation experiment on Extend CK according to an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0043] The terms "including" and "having," and any variations thereof, in the specification and claims of this application and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to the process, method, product, or apparatus.

[0044] like Figure 1 As shown, a method for estimating facial expression intensity according to an embodiment of the present invention includes the following steps:

[0045] S101, obtaining a data sample set, where the data sample set includes several expression sequences.

[0046] Each expression sequence includes multiple frames of expression evolving from neutral expression frames to peak expression frames (apex frames).

[0047] Specifically, the first frame of each expression sequence is at its lowest intensity, gradually increases until it reaches its highest intensity in the apex frame, and then gradually decreases until it reaches its original intensity in the last frame. The change in facial expression intensity provides information about the temporal dynamics of facial behavior, which is crucial for interpreting the meaning of expressions.

[0048] Furthermore, before inputting into the next step for feature extraction, the sample is preprocessed. The steps of data preprocessing are as follows:

[0049] Crop the face area according to the facial feature points and normalize it;

[0050] Furthermore, the specific steps of face cropping and normalization are: determine the positions of the visible eye centers and mouth centers in the input image through facial feature points; calculate the transformation matrix T between the input image and the aligned image points. The calculation formula of the transformation matrix T is:

[0051] set up To align the horizontal center of the image, w is the horizontal length of the image, is the head pose angle. The position in the input and aligned images is calculated as follows:

[0052] s1=s mouth ,

[0053]

[0054] Among them, S1 is the coordinate of the mouth center of the input image, t1 is the coordinate of the mouth center of the aligned image, S2 is the coordinate of the eye center of the input image, t2 is the coordinate of the eye center of the aligned image, and s l.eye is the coordinate of the left eye center of the input image, s r.eye is the coordinate of the center of the right eye in the input image, and d represents the derivative operation. When one eye is in an invisible position, the coordinate of the visible eye is s v.eye In the case where the subject has only one eye visible, the visible eye coordinates s are used. v.eye Alternative That is, when the subject has only one eye visible, s2=s v.eye The transformation matrix T is obtained by solving the linear equations given by the two point correspondences. This alignment method is suitable for face alignment with a large range of head posture changes.

[0055] S102, collecting paired samples from the data sample set to construct a training sample set, wherein the label of each sample in the training sample set is represented by a label probability distribution, and the parameters of the label probability distribution are determined by the original labels of all samples in the observation window containing the target sample in the expression sequence where the target sample is located.

[0056] Given a dynamic facial expression dataset Where K is the total number of sequences, representing the sequence where |S k | is the sequence S k The total number of frames in represents the tth frame of the sequence. It is assumed that each expression sequence in the dataset evolves from a neutral frame to a peak frame. That is, the expression sequence is at its lowest intensity in the first frame, gradually increases until it reaches its highest intensity in the peak frame, and then gradually decreases until it reaches its original intensity in the last frame. The highest intensity is denoted as I h , the lowest intensity is expressed as I l The vertex frame is represented as where a k is the index into the vertex frame.

[0057] The training set is constructed by collecting paired data from the same monotonically changing segments, and the goal of ordinal regression is to learn the sequential information from these expression sequences.

[0058] In the prior art, the training set is represented as in yi and y j They are emoticon images and The intensity label of y i <y j In a fully supervised setting, all intensity labels are known. However, intensity labels are not available in most datasets. In this unsupervised setting, the ordinal regression model cannot estimate the absolute intensity of the expression.

[0059] Therefore, in an embodiment of the present invention, the intensity label of each sample in the training sample set is represented by a label probability distribution. Label probability distribution means that for a specific instance, the descriptive degree of all labels constitutes a data form similar to a probability distribution. The purpose of label distribution learning is to deal with subjective biases from manual labels, measurement noise, or inherent ambiguity of samples. By using label probability distribution to represent the intensity label and using the distribution of the original intensity labels of multiple samples to determine the parameters of the label probability distribution, the sequential information and intensity distribution characteristics in the sequence can be fully utilized to improve the accuracy of the description of expressions with continuously changing intensity.

[0060] In the embodiment of the present invention, the training set is represented as in and They are emoticon images and The probability distribution of intensity labels is such that

[0061] The parameters of the label probability distribution are determined by the raw intensity label distribution of all samples within the observation window containing the target sample in the expression sequence of the target sample. Raw intensity labels refer to conventional labels that use numerical values ​​to represent intensity, such as 0, 1, etc.

[0062] The following settings The parameter determination process is explained in detail using the Gaussian distribution as an example. If other probability distributions are assumed, the calculation can be performed according to the formula for each probability distribution.

[0063] In order to generate an appropriate label distribution, it is assumed that the sample Label distribution Satisfies discrete Gaussian distribution The details are as follows:

[0064]

[0065] Where y d is the label distribution The dth element represents the target sample and the similarity between samples with intensity level d-1. μ and σ 2 are the mean and variance of the Gaussian distribution respectively; Z is to ensure The closer d-1 is to the mean, the smaller the value of y d The larger the value, the farther the distance, d The smaller . The discrete Gaussian distribution is consistent with intuitive observations (i.e., samples with higher adjacent intensity levels are more similar to each other, while samples with lower levels are less similar).

[0066] The key to generating label distribution is to estimate the two parameters of discrete Gaussian distribution: mean and variance. Therefore, the target sample is located in the window τ refers to the radius of the sliding window, and i refers to the target sample. Therefore, the sliding window includes τ samples before and after the i-th sample in the entire expression sequence. And i-1 and i+1 are the samples adjacent to the i-th sample.

[0067] Compute the mean and variance of all samples within the window to estimate the mean and variance, respectively, of the following discrete Gaussian distribution:

[0068]

[0069] y m Represents the original intensity label of the sample in the mth frame in the observation window (the window size is 2τ+1).

[0070] The sliding window diagram is as follows Figure 2 shown.

[0071] In this way, the mean of all samples within the observation window is considered the ground truth value of the target sample. The greater the fluctuation in sample intensity within the window, the less reliable the sample label. In addition, a one-hot label (1, 0, ..., 0) and (0, 0, ..., 1) are assigned to the τ frames with the smallest intensity in the sequence (τ frames at the beginning and end of the sequence) and the τ frames with the largest intensity (τ frames near the peak), respectively.

[0072] Furthermore, if the original intensity labels of at least some of the target sample and the multiple samples adjacent to the target sample in the expression sequence to which the target sample belongs do not exist, a pseudo-intensity label is calculated based on the maximum intensity value and the minimum intensity value of the expression sequence to which the target sample belongs, and the generated pseudo-intensity label is used as the original intensity label of the sample.

[0073] In a fully supervised setting, all original intensity labels are known. Therefore, the label distribution can be generated using Equation 1, and the mean and variance can be estimated using Equation 2. However, most of the original labels are unknown in a semi-supervised setting. To generate a similar label distribution in a semi-supervised setting, pseudo-true values ​​are used as the original labels of the samples:

[0074]

[0075] where δ is the indicator function. If m<a k , the indicator function output is 1, otherwise it is 0 (indicating that the index of the target sample m in the sequence is less than a k , a k is the index of the peak frame of the expression sequence). If m≥a k , indicating that the function output is 1, otherwise it is 0. Figure 4 (a) Schematic diagram showing the label distribution generated using true labels under full supervision. Figure 4 (b) Schematic diagram showing the label distribution generated using pseudo labels in semi-supervised conditions.

[0076] S103: Using the training sample set, train the neural network-based expression intensity recognition model.

[0077] After obtaining the frame-by-frame label distribution, the expression intensity estimation model can be optimized through label distribution learning and cross-entropy loss function:

[0078]

[0079] Where f(·) represents the output of the backbone. The backbone adopts the VGG-16 model, whose parameters are represented as θ and w in the convolutional layer and the fully connected layer respectively. and Represents samples and Output after convolutional and fully connected layers. The number of neurons in the last layer is modified and labeled D to match the generated label distribution. N is the minimum batch size.

[0080] All elements of the label distribution have non-zero values. Therefore, label distribution learning not only learns the true intensity value of a sample, but in particular, it can, to a certain extent, learn other intensity levels from the sample, thereby enhancing the data. Unlike the hard decision made by supervising a hard label, label distribution learning can estimate expression intensity in a soft manner because it considers the similarity between samples of different intensity levels.

[0081] In order to improve the recognition ability of the soft decision method, an ordinal regression loss function is used to supervise the learning of relative intensity relationships. According to the literature, the rank support vector machine loss function is used for this purpose:

[0082]

[0083] where E(·) calculates the mathematical expectation of the forecast distribution to make the outputs comparable. Then there is a positive loss; otherwise, the loss is zero. By combining formula and formula, the objective function formula of the model is as follows:

[0084]

[0085] Among them, the first is a regularization to suppress overfitting of the fully connected layer. After optimization, the test sample is input into any path of the model to estimate the expression intensity.

[0086] In one embodiment, the Extend CK expression library was used. This library contains 593 facial expression sequences from 123 subjects aged 18 to 30, encompassing six expressions: anger, disgust, fear, happiness, sadness, and surprise. Each sequence gradually evolves from a neutral frame to a peak frame and then back to the starting state. No intensity annotation is provided.

[0087] For the Extend CK dataset, 309 sequences were selected from 118 subjects, totaling 5643 frames. For the BU-4DFE dataset, to keep consistent with other counting methods, 120 sequences were selected from 606 sequences for experiments, totaling 2415 frames. The first and last frames of each sequence are the starting frame and the peak frame, respectively. For these two datasets, the same settings as other methods were used, and the sequences were subsampled with a sampling interval of 3 and the The correlation pairs are constructed in this way as input to the two branches of the network.

[0088] The parameters in the backbone network were pre-trained using the VGG-Face dataset. The parameter D, which represents the number of neurons in the last fully connected layer, was set to 6. The parameter τ, which determines the window size for generating the label distribution, was set to 3. Then, face images normalized to a size of 224×224 were constructed as paired data to fine-tune the proposed model using the objective function defined in the formula. The ADAM optimizer was used to optimize the network parameters, with a learning rate set to 1e-5 and a minimum batch size of 32. Specifically, the network parameters are shown in Table 1.

[0089] Table 1

[0090]

[0091] During the testing phase, Pearson correlation coefficient (PCC), intraclass correlation (ICC), and mean absolute error (MAE) were used as evaluation criteria to compare the performance of our method with existing methods. PCC measures the predictive ability to capture intensity trends, and ICC measures the consistency of estimates within each intensity level. Both criteria range from [0,1], where larger values ​​are preferred. MAE measures the deviation between the predicted value and the true value. The smaller the value, the better the result. In a semi-supervised setting, these three metrics can compare the pseudo-true value defined in Equation 3 and the predicted value of the model, where I h and I l They are normalized to 1 and 0, respectively. In a fully supervised setting, the three metrics can compare the true values ​​provided by the dataset and the predicted values ​​of the model.

[0092] Results on ExtendCK. Our method outperforms other methods on all three evaluation metrics. Compared with OSVR-L2 using traditional handcrafted features, our method improves PCC, ICC, and MAE by 51.86%, 99.85%, and 85.88%, respectively. This reflects the superiority of deep learning methods in feature representation. Compared with SIE, ExtendCK improves PCC, ICC, and MAE by 9.34%, 14.18%, and 14.01%, respectively. It can be inferred that our method can learn the label distribution of pseudo-truth generation in a supervised manner, thereby improving the performance of the model.

[0093] To further demonstrate the effectiveness of this invention, an ablation study was conducted on the Extend CK dataset. The unsupervised approach trained the model while minimizing the support loss defined in Equation 5. The semi-supervised approach combined the SVM loss with the keyframe mean squared loss. The label distribution learning approach trained the model while minimizing the loss defined in Equation 4. The proposed method trained the model using the objective function defined in Equation 6. Figure 5 、 6 ,7 shows the experimental results.

[0094] For the estimation of the intensities of the six basic expressions, the average results of the two datasets were consistent. The semi-supervised method outperformed the unsupervised method, indicating that the absolute intensity labels of the keyframes can help to enhance or improve the performance. The label distribution learning method outperformed the semi-supervised method, which means that label distribution learning is robust to inaccurate annotation information and learns more models by using frame-by-frame supervision information. Finally, the method of the present invention achieved the best results, indicating that the combination of label distribution learning and regression learning can effectively improve the performance of each method.

[0095] An expression intensity recognition system according to an embodiment of the present invention includes:

[0096] An acquisition module is used to acquire a data sample set, wherein the data sample set includes a plurality of expression sequences;

[0097] The label generation and training set construction module is used to collect paired samples from the data sample set to construct a training sample set. The intensity label of each sample in the training sample set is represented by a label probability distribution. The parameters of the label probability distribution are determined by the distribution of the original intensity labels of all samples in the observation window containing the target sample in the expression sequence of the target sample.

[0098] The training module is used to train the expression intensity recognition model based on the neural network using the training sample set.

[0099] The implementation principle of the system is the same as the above method and will not be repeated here.

[0100] This embodiment also provides an electronic device, which includes at least one processor and at least one memory, wherein a computer program is stored in the memory. When the computer program is executed by the processor, the processor executes any step of the above-mentioned expression intensity recognition method. The specific steps are described in the method embodiment and are not repeated here. In this embodiment, the types of processor and memory are not specifically limited. For example, the processor can be a microprocessor, a digital information processor, an on-chip programmable logic system, etc.; the memory can be a volatile memory, a non-volatile memory, or a combination thereof.

[0101] The present application also provides a storage medium storing a computer program executable by a processor, which, when executed on the processor, causes the processor to perform any of the steps of the above-mentioned expression intensity recognition method. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive, a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0102] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0103] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0104] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of the system or module can be electrical or other forms.

[0105] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the purpose of this embodiment based on actual needs.

[0106] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0107] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0108] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0109] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the specification and practicing the disclosure herein, those skilled in the art will easily think of the implementation scheme of the present disclosure. This application is intended to cover any variation, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.

[0110] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for recognizing facial expression intensity, characterized in that: include: Acquire a data sample set, wherein the data sample set includes a plurality of expression sequences; Collecting paired samples from the data sample set to construct a training sample set, wherein the intensity label of each sample in the training sample set is represented by a label probability distribution, wherein the parameters of the label probability distribution are determined by the distribution of the original intensity labels of all samples within an observation window containing the target sample in the expression sequence where the target sample is located; Using the training sample set to train an expression intensity recognition model based on a neural network; The samples in the training sample set are recorded as The label probability distribution is recorded as Located at the center of the observation window, the observation window is expressed as [i-τ,i+τ], indicating The i-τ frame to i+τ frame samples in the expression sequence constitute the observation window, τ is a positive integer, and the distribution of the 2τ+1 frame samples in the observation window is determined. Parameters; If set Satisfies Gaussian distribution, Expressed as: Where y d yes The dth element represents The similarity between samples with intensity level d-1, D is The total number of elements, μ and σ 2 are the mean and variance of the Gaussian distribution, and Z is to ensure The normalization factor of y m represents the original intensity label of the sample in the observation window; If y m If it does not exist, a pseudo-intensity label is generated. The pseudo labels generated As the original intensity label of the sample, The calculation formula is: I h Indicates the maximum intensity value of the peak value of the expression sequence to which the target sample belongs, I l Indicates the lowest intensity value of the expression sequence to which the target sample belongs, δ is the representation function, If m<a k , indicating that the function output is 1, otherwise it is 0, If m≥a k , indicating that the function output is 1, otherwise it is 0, a k Indicates the frame number corresponding to the maximum intensity value in the expression sequence, m is y m The frame number of the corresponding sample in the expression sequence.

2. The method for recognizing facial expression intensity according to claim 1, wherein: Each expression sequence includes multiple frames of expressions evolving from neutral expression frames to peak expression frames. If the original intensity labels of the target sample and at least some of the multiple samples adjacent to the target sample in the expression sequence to which the target sample belongs do not exist, a pseudo-intensity label is calculated based on the maximum intensity value and the minimum intensity value of the expression sequence to which the target sample belongs, and the generated pseudo-intensity label is used as the original intensity label of the sample.

3. The method for recognizing facial expression intensity according to claim 1, wherein: The loss function calculation formula for training is: is the regularization term, N is the minimum batch size, and The samples are and The label probability distribution of and Represents samples and The output after the convolutional layer and the fully connected layer, E(·) represents the mathematical expectation of the calculated predicted distribution.

4. An expression intensity recognition system, characterized in that: Implementing the expression intensity recognition method according to any one of claims 1 to 3, comprising: An acquisition module, configured to acquire a data sample set, wherein the data sample set includes a plurality of expression sequences; a label generation and training set construction module, configured to collect paired samples from the data sample set to construct a training sample set, wherein the intensity label of each sample in the training sample set is represented by a label probability distribution, wherein the parameters of the label probability distribution are determined by the distribution of the original intensity labels of all samples within an observation window containing the target sample in the expression sequence where the target sample is located; The training module is used to train the expression intensity recognition model based on the neural network using the training sample set.

5. An electronic device, characterized in that: The method comprises at least one processor and at least one storage module, wherein the storage module stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 3.

6. A storage medium, characterized in that The device stores a computer program, which, when executed on a processor, enables the processor to execute the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Expression synthesis method fused with attention mechanism

    CN111369646A

  • AI Platform For Pixel Spacing, Distance, And Volumetric Predictions From Dental Images

    US20210358123A1