Facial expression recognition method and apparatus
By performing feature extraction and denoising on short videos, and utilizing causal 3D convolution and spatiotemporal attention denoising modules, the problem of poor facial expression recognition accuracy in short videos is solved, achieving efficient facial expression recognition even in low-quality environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM HANGZHOU CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-06-26
AI Technical Summary
Existing facial expression analysis technologies cannot effectively recognize facial expressions in short videos. Due to limitations in low quality, complex environments, and high dynamic characteristics, the recognition accuracy is poor.
By extracting features and denoising short video frames, facial expression features and expression change features of facial images are obtained using an expression recognition model. Then, clear expression information is reconstructed using a causal 3D convolution and spatiotemporal attention denoising module.
In low-quality and noisy environments, it improves the accuracy of electronic devices in recognizing facial expressions in short videos and can extract complete facial expression feature information.
Smart Images

Figure CN122290190A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an expression recognition method and apparatus. Background Technology
[0002] Currently, because facial expressions are a type of facial muscle movement that lasts for a very short time and is not entirely controlled by the subject's subjective consciousness, facial expressions can reveal an individual's true emotional state and have important application value in fields such as affective computing, psychological assessment, and security review.
[0003] In related technologies, electronic devices can train models using training data collected in a high-quality, controlled environment, such as training data containing high frame rates, high resolutions, stable lighting, and frontal facial poses, so that the models can recognize micro-expressions contained in the training data.
[0004] However, in the above methods, model training usually uses high-quality facial data, while short videos are generally characterized by extremely short duration, low resolution, and high compression rate. Furthermore, the duration of facial expressions is extremely short, making it difficult to capture them completely in a limited number of video frames. In addition, low-quality video data causes severe degradation of facial details. Therefore, the model trained by the above methods has poor accuracy in recognizing facial expressions in short videos. Summary of the Invention
[0005] The purpose of this application is to provide an expression recognition method and apparatus that can improve the accuracy of electronic devices in recognizing expressions in short videos.
[0006] In a first aspect, embodiments of this application provide an expression recognition method, which includes: extracting features from a first set of video frames corresponding to a first video to obtain a first image feature set corresponding to the first set of video frames, wherein all video frames in the first set of video frames include facial images; inputting the first image feature set into a first expression recognition model, and performing denoising processing on the image features in the image feature set through a denoising module in the first expression recognition model to obtain a first expression feature set, wherein the first expression feature set includes expression features of facial images in video frames and expression change features of facial images between every two video frames; and reconstructing expressions based on the first expression feature set through the first expression recognition model to recognize facial expressions in the first video.
[0007] Secondly, embodiments of this application provide an expression recognition device, which includes an extraction module, a processing module, and a reconstruction module. The extraction module is used to extract features from a first set of video frames corresponding to a first video, obtaining a first image feature set corresponding to the first set of video frames, wherein each video frame in the first set of video frames includes a face image. The processing module is used to input the first image feature set obtained by the extraction module into a first expression recognition model, and to perform denoising processing on the image features in the image feature set through a denoising module in the first expression recognition model, obtaining a first expression feature set, which includes expression features of face images in video frames and expression change features of face images between every two video frames. The reconstruction module is used to reconstruct expressions based on the first expression feature set obtained by the processing module using the first expression recognition model, and to recognize facial expressions in the first video.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0011] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.
[0012] In this embodiment, features are extracted from the first set of video frames corresponding to the first video to obtain a first set of image features. Each video frame in the first set includes a face image. Then, the first set of image features is input into a first expression recognition model. The image features in the first expression recognition model are denoised using a denoising module to obtain a first set of expression features. This first set of expression features includes expression features of the face images in the video frames and expression change features of the face images between every two video frames. Finally, the first expression recognition model reconstructs the expression based on the first set of expression features to recognize the facial expressions in the first video. In this solution, by acquiring expression features of the face images and expression change features of the face images between every two video frames through the first expression recognition model, the first expression recognition model in the electronic device can extract complete expression feature information from low-quality, noisy short video image features. Therefore, a clear and complete facial expression can be reconstructed using the complete expression feature information. In other words, the first expression recognition model in the electronic device can obtain effective expression information even when video frame information is severely degraded, thus improving the accuracy of expression recognition by the electronic device. Attached Figure Description
[0013] Figure 1 This is one of the flowcharts of an expression recognition method provided in the embodiments of this application;
[0014] Figure 2 This is a second flowchart of an expression recognition method provided in an embodiment of this application;
[0015] Figure 3 This is the third flowchart of an expression recognition method provided in the embodiments of this application;
[0016] Figure 4 This is the fourth flowchart of an expression recognition method provided in the embodiments of this application;
[0017] Figure 5 This is the fifth flowchart of an expression recognition method provided in the embodiments of this application;
[0018] Figure 6 This is the sixth flowchart of an expression recognition method provided in the embodiments of this application;
[0019] Figure 7 This is the seventh flowchart of an expression recognition method provided in the embodiments of this application;
[0020] Figure 8 This is a schematic diagram of the structure of an expression recognition device provided in an embodiment of this application;
[0021] Figure 9This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;
[0022] Figure 10 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0024] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects. For example, a first object can be one or more, where "more" means at least two. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The terms "at least one" and "at least one of" in this application's specification refer to any one, any two, or a combination of two or more of the included objects. For example, "at least one of a, b, and c" can mean "a", "b", "c", "a and b", "a and c", "b and c", and "a, b, and c", where a, b, and c can be single or multiple, and multiple means at least two. Similarly, "at least two" means two or more, and its meaning is similar to "at least one". The identifiers in this application are text, symbols, images, etc., used to indicate information, and can use controls or other containers as carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.
[0026] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.
[0027] Short videos: Short videos refer to frequently pushed video content played on various new media platforms, suitable for viewing in mobile and short leisure time, and generally less than 5 minutes in length, ranging from a few seconds to a few minutes. Short videos have the following characteristics: Short length and rich content: Generally ranging from 15 seconds to 5 minutes, the content covers a variety of themes such as skill sharing, humor and entertainment, and fashion trends, catering to users' fragmented viewing habits. Rapid dissemination and strong interactivity: Low barriers to dissemination, abundant channels, and easy to achieve viral spread; users can interact through sharing, commenting, sending bullet comments, and liking. Creative and highly personalized: Diverse content and presentation formats meet the personalized and diverse aesthetic needs of the younger generation; users can create short videos using personalized production and editing techniques. Precise targeting and triggering marketing effects: Short video marketing can accurately target users and achieve precise marketing.
[0028] Spatiotemporal alignment: Spatiotemporal alignment refers to the process of precisely matching and calibrating data in both time and space dimensions. In data analysis and processing, spatiotemporal alignment typically involves synchronizing and integrating data from different sources or from the same source but at different points in time to facilitate comparison, analysis, or fusion. For example, in video processing, it may be necessary to align video frames captured by different cameras according to their timestamps for multi-view analysis. In geographic information systems, spatiotemporal alignment may mean aligning geographic data from different points in time according to their geographic locations to observe trends.
[0029] Model quantization: Model quantization is the process of converting floating-point parameters and activation values in a deep learning model into low-bit representations, such as converting 32-bit floating-point numbers to 8-bit integers. The main goal of model quantization is to reduce the model's storage space, lower computational complexity, and improve inference speed, while maintaining model accuracy as much as possible. For example, a model represented by 32-bit floating-point numbers can be quantized to 8-bit integers, reducing its size to one-quarter and increasing inference speed by 2-4 times. This is significant for deploying deep learning models on mobile or edge devices.
[0030] Multi-granularity knowledge distillation: Multi-granularity knowledge distillation is a technique that improves the performance of student models by integrating knowledge from different granularity levels. It has applications in multiple fields, such as text embedding, image generation, graph classification, and quantum relation knowledge distillation.
[0031] The facial expression recognition method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0032] The facial expression recognition method provided in this application can be applied to facial expression recognition scenarios in short videos.
[0033] Currently, because facial expressions are a type of facial muscle movement that lasts for a very short time and is not entirely controlled by the subject's subjective consciousness, facial expressions can reveal an individual's true emotional state and have important application value in fields such as affective computing, psychological assessment, and security review.
[0034] Traditional facial expression analysis methods primarily rely on high-quality video data collected under controlled conditions, such as high frame rates, high resolution, stable lighting, and frontal facial poses. However, when applying these methods to the increasingly prevalent short video scenarios, they face the following insurmountable technical bottlenecks:
[0035] (1) The contradiction between low data quality and scarcity of effective information.
[0036] Videos on short video platforms are generally extremely short, ranging from a few seconds to tens of seconds, with low resolution and high compression rates, resulting in blurry images or loss of detail. Moreover, facial expressions themselves are extremely short-lived, making them difficult to capture completely within the limited number of video frames. In addition, the low-quality video data causes severe degradation of facial details, especially subtle muscle movements, leading to a sharp decline in the performance of traditional analysis models that rely on high-definition video frame sequences, or even their failure.
[0037] (2) Robustness challenges in complex uncontrolled environments.
[0038] The highly uncontrollable shooting environment of short videos introduces numerous strong interfering factors to facial analysis, including but not limited to: drastic lighting changes, such as overexposure, backlighting, and localized shadows; frequent and significant changes in head posture, such as turning the face to the side, tilting, or rapid rotation; partial facial occlusion, such as hair, glasses, gestures, or props; and cluttered backgrounds and rapid camera transitions. These factors severely interfere with the accuracy of face detection and facial landmark alignment, leading to insufficient stability and reliability in subsequent feature extraction.
[0039] (3) Facial expressions are “submerged” in complex facial movements.
[0040] Short videos feature vivid and varied facial expressions, including a large number of macro-expressions that are wide-ranging and long-lasting, such as laughing and frowning in anger, as well as a large number of facial movements unrelated to emotion, such as changes in mouth shape while speaking, natural blinking, and coughing. As a weak and transient signal, facial expressions are easily masked, confused, or interrupted by these dominant and high-energy facial activities, making it extremely difficult to accurately separate and identify true facial expressions.
[0041] In summary, existing facial expression analysis technologies are limited by their strong dependence on data quality and controlled environment, making them unable to adapt to the low-quality, highly interfering, and highly dynamic characteristics of short video scenarios, and even more difficult to achieve the high-efficiency, low-latency real-time analysis required on mobile devices.
[0042] Based on the above-mentioned scenarios applied in the embodiments of this application, the expression recognition method provided in the embodiments of this application obtains the expression features of a face image and the expression change features of the face image between every two video frames through a first expression recognition model. The first expression recognition model in the electronic device can extract complete expression feature information from the features of low-quality, noisy short video images, thereby reconstructing a clear and complete expression through the complete expression feature information. In other words, the first expression recognition model in the electronic device can obtain effective expression information even when the video frame information is severely degraded, thus improving the accuracy of expression recognition by the electronic device.
[0043] The facial expression recognition method provided in this application is executed by an facial expression recognition device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this device. The following will use an electronic device as an example to illustrate the facial expression recognition method provided in this application.
[0044] This application provides an expression recognition method. Figure 1 A flowchart of an expression recognition method provided in an embodiment of this application is shown. Figure 1 As shown, the facial expression recognition method provided in this application embodiment may include the following steps 201 to 203.
[0045] Step 201: The electronic device extracts features from the first video frame set corresponding to the first video to obtain the first image feature set corresponding to the first video frame set.
[0046] In this embodiment of the application, all video frames in the first video frame set include human face images.
[0047] For example, the first video mentioned above can be a short video.
[0048] Optionally, in this embodiment of the application, the electronic device can perform face detection on each video frame in the first video, thereby determining the video frames containing faces as the first video frame set.
[0049] For example, an electronic device can employ a lightweight face detector, such as 224×224, to detect face regions in video frames in real time. The detector uses the MobileFaceNet architecture and has a detection frequency baseline of 15 FPS, achieving a balance between detection accuracy and speed.
[0050] Then, the electronic device performs image rotation compensation on the video frames in the first video frame set to compensate for changes in head tilt, pitch, and other angles. Then, the electronic device can perform spatiotemporal alignment based on 3D facial key points, extract facial features through a multi-task cascaded convolutional neural network and key points, and convert them into homogeneous coordinates to obtain a normalized facial feature sequence.
[0051] Next, the electronic device divides the normalized facial sequence into temporal blocks of fixed length, such as 16 frames, and extracts spatiotemporal features from each temporal block through causal 3D convolution.
[0052] For example, electronic devices can construct temporally-space-normalized facial image sequences to ensure consistency in the scale and position of facial regions across frames. This causal 3D convolution serves as the causal constraint for the generated data. Causal convolution is used to implement a causal mask in the temporal dimension. The causal 3D convolution is causal in both the temporal and spatial dimensions; forward frames can only see historical and current frames, not future frames.
[0053] The temporal stride of the causal 3D convolution is 1: all temporal information is preserved, i.e., no downsampling is performed; the spatial stride is 4: the spatial resolution is reduced by 4 times. The causal mask is implemented by _padding = (2, 2, 1, 1, 2, 0) to achieve (2, 2); the spatial y-dimensional padding is 1, 1 on both sides; the spatial x-dimensional padding is 2, 0 on both sides; the temporal dimension padding is before (historical frames) and after (implementing causality). The caching mechanism supports the cache_x parameter, which allows the reuse of the previous frame information during block-level generation.
[0054] For example, such as Figure 2 As shown, the specific implementation process of step 201 above will be explained in detail below. Specifically, it can be implemented through steps 1 to 7 below.
[0055] Step 1: The electronic device inputs the short video into the face detector.
[0056] For example, the face detector described above could be MobileVid.
[0057] Step 2: If the electronic device does not detect a face in any frame of the short video using a face detector, it skips that frame.
[0058] Step 3: When the electronic device detects a face in any frame of the short video using a face detector, it extracts the region of interest of the face.
[0059] Step 4: The electronic device performs head pose control and spatial transformation alignment on the region of interest of the face to obtain aligned facial feature information.
[0060] Step 5: The electronic device generates a normalized facial sequence based on the aligned facial feature information.
[0061] Step 6: The electronic device divides the normalized facial sequence into 16-frame temporal blocks.
[0062] Step 7: The electronic device performs multi-scale feature extraction on the 16-frame time block to obtain the first image feature set.
[0063] Step 202: The electronic device inputs the first image feature set into the first expression recognition model, and performs noise reduction processing on the image features in the image feature set through the denoising module in the first expression recognition model to obtain the first expression feature set.
[0064] In this embodiment of the application, the first set of facial expression features includes facial expression features of face images in video frames and facial expression change features of face images between every two video frames.
[0065] Optionally, in the embodiments of this application, the above-mentioned facial expression features may include at least one of the following: facial expression location information, facial expression shape information, and facial expression amplitude information.
[0066] It can be understood that the facial expression features in the aforementioned video frames can be spatial feature information corresponding to the spatial dimension, and the facial expression change features between each pair of video frames can be temporal feature information corresponding to the temporal dimension.
[0067] For example, the above-mentioned denoising module can be a spatiotemporal attention denoising module.
[0068] For example, the aforementioned spatiotemporal attention denoising module can use a complex-domain rotation matrix to encode spatiotemporal positions, converting absolute position information into relative positional relationships. This encoding method maintains the linear complexity of the sequence length while ensuring the relative invariance of positional information.
[0069] Then, the spatiotemporal attention denoising module adopts a separate structure of one-dimensional temporal convolution kernels and two-dimensional spatial convolution kernels. Causal convolution is used in the temporal dimension to ensure temporal consistency, while dilated convolution is used in the spatial dimension to increase the receptive field.
[0070] Next, the spatiotemporal attention denoising module extends standard self-attention to spatiotemporal joint attention, extracting features from the query, key, and value vectors simultaneously from both spatial and temporal dimensions. The calculation of attention weights takes into account the spatial proximity and temporal continuity between pixels.
[0071] Next, the spatiotemporal attention denoising module uses gated linear units and depthwise separable convolutions to build an efficient feedforward network, which significantly reduces computational complexity while maintaining expressive power.
[0072] Finally, the spatiotemporal attention denoising module adopts a "time-first, space-later" processing strategy: first, convolution operations are performed in the temporal dimension to capture inter-frame motion patterns; then, self-attention is applied in the spatial dimension to model intra-frame structural relationships; finally, spatiotemporal features are fused through residual connections. In particular, the spatiotemporal attention denoising module introduces a relative position bias term to explicitly encode the spatiotemporal distance information between pixels.
[0073] Step 203: The electronic device uses the first expression recognition model to reconstruct expressions based on the first expression feature set and recognizes the facial expressions in the first video.
[0074] It should be noted that the specific implementation process of step 203 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0075] Optionally, in this embodiment of the application, the first set of facial expression features includes facial expression features under different facial postures.
[0076] For example, combined Figure 1 ,like Figure 3 As shown, step 203 can be implemented through steps 203a and 203b below.
[0077] Step 203a: The electronic device inputs the first expression feature set into the expression transient focusing module in the first expression recognition model through the first expression recognition model, identifies the expression feature information in the first expression feature set, and performs feature enhancement on the expression feature information to obtain the enhanced first expression feature set.
[0078] In this embodiment of the application, the transient focus module of the first facial expression recognition model of the electronic device can locate the brief and weak bursts of muscle movement in the time and space dimensions, and adaptively enhance its features.
[0079] Step 203b: The electronic device reconstructs facial expressions based on the enhanced first set of facial expression features to obtain facial expressions.
[0080] It should be noted that the specific implementation of step 203b above can be found in the description in the relevant technology. To avoid repetition, it will not be repeated here.
[0081] In the expression recognition method provided in this application embodiment, features are extracted from a first set of video frames corresponding to a first video to obtain a first set of image features corresponding to the first set of video frames. Each video frame in the first set of video frames includes a face image. Then, the first set of image features is input into a first expression recognition model. The image features in the image feature set are denoised using a denoising module in the first expression recognition model to obtain a first set of expression features. The first set of expression features includes expression features of face images in the video frames and expression change features of face images between every two video frames. Finally, expression reconstruction is performed based on the first set of expression features using the first expression recognition model to obtain expression information. In this solution, by acquiring expression features of face images and expression change features of face images between every two video frames through the first expression recognition model, the first expression recognition model in the electronic device can extract complete expression feature information from low-quality, noisy short video image features. Therefore, a clear and complete expression can be reconstructed using the complete expression feature information. In other words, the first expression recognition model in the electronic device can obtain effective expression information even when video frame information is severely degraded, thus improving the accuracy of expression recognition by the electronic device.
[0082] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 4 As shown, prior to step 201 above, the facial expression recognition method provided in this application embodiment further includes steps 301 and 302 as described below.
[0083] Step 301: The electronic device trains the first initial model based on the second expression recognition model to obtain the trained first initial model.
[0084] In this embodiment of the application, the second expression recognition model has the same model structure as the first expression recognition model. The storage space occupied by the first expression recognition model is smaller than that occupied by the second expression recognition model. The second expression recognition model is a teacher model, and the first expression recognition model is a student model.
[0085] Optionally, in this embodiment of the application, the electronic device can perform multi-granularity knowledge distillation on the first initial model through the second expression recognition model to obtain the trained first initial model.
[0086] For example, an electronic device can perform single-step denoising distillation on a first initial model using a second expression recognition model to obtain a distilled first initial model, and then perform multi-granularity distillation supervision on the distilled first initial model to obtain a trained first initial model.
[0087] For example, the single-step denoising distillation described above can be: the student model processes the noisy input... Perform single-step denoising to directly predict clean features. Then, the prediction results... Add random Gaussian noise to obtain intermediate samples Next, for the real intermediate samples Adding the same noise yields the noisy real data. Finally, the student model was tested on noisy real data. Perform model inference to obtain expression classification results. .
[0088] Then, the teacher model performs feature extraction, that is, the noisy input. Input the pre-trained teacher model and obtain Then, Input the initialized teacher model and obtain Next, Input the pre-trained teacher model to obtain feature representations of real data. Finally Input the auxiliary classifier and obtain the expression classification results. .
[0089] Next, the electronic device uses the recognition results and The distribution matching loss is calculated to ensure that the student model is consistent with the teacher model on the overall distribution trajectory. Finally, the electronic device calculates the feature loss based on the intermediate latent space features CA and DA, enabling the student model to approximate the true data distribution.
[0090] Step 302: The electronic device quantizes the model parameters of the first initial model after training to obtain the first expression recognition model.
[0091] It should be noted that the specific implementation process of step 302 above can be found in the following embodiments, and will not be repeated here to avoid repetition.
[0092] Optionally, in the embodiments of this application, combined with Figure 4 ,like Figure 5 As shown, step 302 can be implemented through steps 302a and 302b below.
[0093] Step 302a: The electronic device quantizes the model parameters corresponding to each model layer based on the quantization strategy corresponding to each model layer of the trained first initial model to obtain the quantized first initial model.
[0094] Optionally, in the embodiments of this application, the first initial model may include a spatiotemporal attention denoising module, a 3D convolution module, and a fully connected classification module.
[0095] For example, the spatiotemporal attention denoising module, as a core component of the first initial model, is responsible for modeling intra-frame spatial relationships and inter-frame temporal dependencies. This application adopts a hybrid strategy of retaining full precision or reducing it to INT8 precision. In specific implementation, attention calculation is decomposed into multiple quantizable sub-operations: the attention, key, and value projection matrices are quantized to INT8 using a channel-wise symmetric quantization scheme, while the attention score calculation and Softmax operation retain FP16 or FP32 precision.
[0096] The 3D convolution module is responsible for extracting the spatiotemporal features of the video. This application adopts the INT8 per-channel symmetric quantization scheme. The convolution kernel weights are quantized per channel, and independent quantization parameters are calculated for each output channel. Input activation uses dynamic range quantization. Considering the continuity of the time dimension, a time-aware quantization mechanism is implemented: an independent scaling factor and zero-point parameter are assigned to each time slice along the time dimension.
[0097] Fully connected classification module: The parameters of the fully connected classification module account for the main part of the model's time consumption. This application uses INT4 quantization, through fake quantization + straight-through estimator (STE); the original BF16 parameters are retained as learnable parameters; the forward propagation simulates the INT4 quantization process, i.e., quantization → dequantization, introducing quantization noise; the gradient is still passed back to the BF16 parameters during backpropagation to avoid the discrete gradient problem.
[0098] The following section provides a detailed explanation of the multi-stage quantization calibration for each module in the first initial model described above.
[0099] Among them, the quantization calibration of the spatiotemporal attention denoising module adopts a hierarchical calibration strategy: weight calibration: attention projection matrix ( , , ) Symmetrical quantization is employed for each output channel. During calibration, the maximum absolute value of the weights for each channel is collected, and the scaling factor is calculated. The INT8 symmetric quantization range is -127 to 127. For multi-head attention, each head is calibrated independently to avoid interference between quantization parameters between heads.
[0100] Activation Calibration: The input activation of the spatiotemporal attention denoising module employs dynamic statistical calibration. During forward propagation, the activation value distribution of each attention head is recorded in real time, constructing a 256-bin histogram. The optimal quantization range is determined using the KL divergence minimization criterion: different cutoff thresholds are searched, the KL divergence of the distribution before and after quantization is calculated, and the threshold that minimizes the KL divergence is selected as the quantization boundary. In practice, the 99.9th percentile method is used for pre-screening, followed by fine-tuning using KL divergence.
[0101] Softmax temperature factor quantization: attention scaling factor Static quantization is used to quantize it into a fixed-point representation, avoiding division operations. This is applied to different head dimensions. The quantization table is pre-calculated and retrieved during runtime by looking up the table.
[0102] Quantization calibration of the 3D convolutional module takes into account the special characteristics of the temporal dimension: Spatiotemporal separation calibration: The 3D convolution is decomposed into spatial and temporal dimensions for separate calibration. The spatial convolution kernel uses traditional channel-by-channel quantization, while the temporal convolution kernel uses time-slice-independent quantization parameters. During calibration, the statistical distributions of spatial and temporal feature maps are collected separately. Special handling for causal convolution: For causal convolution, a strategy combining online and offline calibration is adopted. Online exponential moving average statistics are used during the training phase, and pre-calibrated fixed parameters are used during the inference phase. This ensures that the quantization error of the previous frame does not accumulate and propagate in the temporal dimension. Grouped convolution optimization: For grouped convolution, each group is quantized independently. The calibration samples must cover all group combinations to ensure the balance of quantization parameters across groups.
[0103] The fully connected classification module INT4 quantization calibration faces the challenge of maintaining accuracy due to its extremely low bit width. The calibration process employs several methods: Grouped quantization calibration: The weight matrix is divided into groups of 32 elements, with each group calibrated independently. The calibration process uses a min-max method to determine the range, but introduces a boundary expansion mechanism: the minimum and maximum values are expanded outwards by 1% to avoid truncation error. Cluster-assisted calibration: k-means clustering is applied to divide the weight values into 16 clusters, with the cluster centers serving as the quantization level. During calibration, the joint objective function of the cluster centers and quantization range is iteratively optimized to minimize the reconstruction error. Non-uniform quantization option: For highly non-uniform weight distributions, a non-uniform quantization option is provided. The Lloyd-Max algorithm is used to iteratively optimize the quantization level, minimizing the quantization error.
[0104] Step 302b: The electronic device performs quantization perception adjustment processing on the quantized first initial model to obtain the first expression recognition model.
[0105] In some embodiments of this application, the electronic device can be retrained for 2-3 epochs on the quantized student model.
[0106] For example, an electronic device can use a pass-through estimator to bypass the gradient interruption problem of quantization operation, as shown in Equation (1), which is as follows:
[0107] (1)
[0108] Where X is the original value to be quantized, S is the quantization step size, and b is the bit width. To quantify the results.
[0109] Implementation mechanism: Pseudo-quantization nodes are inserted into facial expression features. During forward propagation, the true quantization-dequantization operation is performed, and during backpropagation, the gradient is directly passed to the input. Customized STE variants are implemented for different quantization types.
[0110] Add a quantization error regularization term:
[0111]
[0112]
[0113] Then, the quantization error of the electronic device is regularized, and a quantization error regularization term is added to the fine-tuning loss function, along with regularization weights assigned based on the sensitivity of each layer to quantization. The sensitivity is obtained through pre-analysis, and the formula is: .
[0114] For example, the progressive quantization fine-tuning process provided in the embodiments of this application will be described below.
[0115] Phase 1: Weight freeze fine-tuning (1 epoch), only fine-tuning the quantization parameters (scaling factor, zero point) while keeping the original weights unchanged; learning rate: 0.1 times the original learning rate; objective: minimize the initial error introduced by quantization.
[0116] Phase 2: Partial unfreezing and fine-tuning (1-2 epochs), unfreezing the weights of the sensitive layer (based on sensitivity analysis) and optimizing the weights and quantization parameters at the same time; learning rate: 0.05 times the original learning rate; objective: to recover the accuracy loss of the sensitive layer.
[0117] Phase 3: Global fine-tuning (0-1 epochs), unfreezing all trainable parameters, focusing on optimizing the balance between task loss and quantization loss; learning rate: 0.01 times the original learning rate; goal: to reach the final accuracy-efficiency tradeoff point.
[0118] Then, the electronic device can calibrate the data selection strategy, and fine-tuning the selection of calibration data affects the final performance: Representative sample selection: Select the samples that best represent the data distribution from the training set. Use clustering methods to select center samples from the feature space to ensure coverage of all major patterns.
[0119] Hard sample augmentation: This specifically includes hard samples whose performance degrades significantly after quantization. These samples are identified through quantization simulation, and their weights are increased in the fine-tuned dataset.
[0120] Temporal continuity considerations: For video models, select segments with temporal continuity rather than isolated frames. Ensure that the calibration data includes typical temporal dynamic patterns.
[0121] For example, such as Figure 6 As shown, steps 301 to 302a above will be explained in detail below. Specifically, they can be implemented through steps 10 to 12 below.
[0122] Step 10: The electronic device adopts different quantization strategies according to different data layers of the model.
[0123] Step 11: The electronic device obtains a quantized miniature model.
[0124] Step 12: The electronic device uses online data to fine-tune the model parameters of the quantized small model.
[0125] In some embodiments of this application, after obtaining a small model with finely tuned parameters, the electronic device can deploy the small model in the electronic device.
[0126] For example, the electronic device converts the quantized small model into TFLite or MNN format, and then performs operator fusion optimization, such as Conv+BN+ReLU fusion.
[0127] For example, this application proposes a deep operator fusion technique based on mathematical equivalence transformation, specifically targeting the common Conv+BN+ReLU pattern. The principle of mathematical equivalence transformation is as follows: For the cascaded combination of convolutional layers, batch normalization layers, and ReLU activation functions, complete fusion is achieved through mathematical derivation. Let the original computation be:
[0128]
[0129]
[0130] .
[0131] The equivalent calculation after fusion is as follows:
[0132]
[0133]
[0134] .
[0135] Thus, the emergence of fusion operators primarily addresses the issue of the amount of data read in during model training, while reducing write-back operations of intermediate results and decreasing memory access operations, thereby reducing processing time.
[0136] Optionally, in the embodiments of this application, the facial expression recognition method provided in the embodiments of this application further includes the following steps 401 to 406.
[0137] Step 401: The electronic device performs noise addition processing on the second image feature set through the second initial model to obtain the noise-added second image feature set.
[0138] Optionally, in the embodiments of this application, the above-mentioned noise addition processing may include at least one of the following: applying a random direction motion blur kernel, varying the illumination: randomly adjusting brightness and contrast, adding shadows, and adding compression artifacts: simulating the block effect and ringing effect of JPEG compression. The specific method can be determined according to actual usage requirements, and the embodiments of this application do not impose limitations.
[0139] Step 402: The electronic device inputs the noise-added second image feature set into the face action detection module in the second initial model through the second initial model to perform face action detection and obtain the first face feature set.
[0140] In this embodiment of the application, the facial features included in the first facial feature set correspond to an action amplitude that is less than or equal to a preset amplitude threshold.
[0141] For example, the second initial model can automatically identify and suppress continuous, large-scale motion areas in video frames through a face motion detection module, such as the mouth when speaking or the cheeks when laughing.
[0142] Step 403: The electronic device inputs the first face feature set into the multi-scale feature injection module of the second initial model through the second initial model to extract features and obtain the second face feature set.
[0143] In this embodiment of the application, the second facial feature set includes facial feature information under different image sizes.
[0144] For example, the multi-scale feature injection module described above constructs a pyramid-shaped multi-scale feature extraction and fusion network, which includes the following key components:
[0145] Hierarchical encoder network: It is constructed using cascaded residual convolutional blocks. Each layer achieves spatial downsampling through stride convolution, and the number of channels grows exponentially, forming a multi-scale feature representation from fine-grained to coarse-grained.
[0146] Cross-scale attention mechanism: Each scale is equipped with an independent cross-attention unit, allowing information exchange between feature maps of different resolutions. In practice, high-resolution features are aligned to the low-resolution space through bilinear interpolation, and then weighted and fused after calculating attention weights.
[0147] Adaptive Gated Fusion: A temporally embedded gated network is designed to dynamically calculate the fusion weights of features at each scale. The gated network receives temporally encoded vectors and outputs a weight matrix with dimensions equal to the square of the scale, which is then normalized using Softmax to guide feature fusion.
[0148] Feature pyramid output: The final output contains feature pyramids at multiple scales, providing rich contextual information for the subsequent denoising module.
[0149] For example, the input image feature set is first processed by a hierarchical encoder to extract multi-scale features. Each scale feature, in addition to being processed by its own layer, also interacts with features from other scales through attention. A cross-scale attention mechanism calculates the similarity matrix between the current scale feature and all other scale features, and based on this, features are recombined. Finally, an adaptive fusion module dynamically adjusts the contribution of each scale feature according to the importance of the current denoising stage, generating a fused multi-scale feature representation.
[0150] Step 404: The electronic device inputs the second facial feature set into the spatiotemporal attention denoising module of the second initial model through the second initial model, and performs denoising processing on the second facial feature set to obtain the second expression feature set.
[0151] In this embodiment of the application, the second set of facial expression features includes facial expression features of face images in video frames and facial expression change features of face images between every two video frames.
[0152] It should be noted that the specific implementation process of the above-mentioned spatiotemporal attention denoising module can be found in the above embodiments, and will not be repeated here to avoid repetition.
[0153] Step 405: The electronic device inputs the second expression feature set into the expression transient focusing module in the second initial model through the second initial model, identifies the expression feature information, and performs feature enhancement on the expression feature information to obtain the enhanced second expression feature set.
[0154] It should be noted that the specific implementation process of the above-mentioned transient expression focusing module can be found in the above embodiments, and will not be repeated here to avoid repetition.
[0155] Step 406: The electronic device trains a second initial model based on the enhanced second expression feature set to obtain a second expression recognition model.
[0156] Optionally, in this embodiment of the application, the electronic device can calculate the cross-entropy loss value between the enhanced second expression feature set and the reference expression feature set, and train a second initial model based on the cross-entropy loss value through backpropagation to obtain a second expression recognition model.
[0157] In this embodiment, the second expression recognition model does not directly process the original pixels through the network, but instead simulates a "denoising diffusion" process in the feature domain. A lightweight conditional diffusion model is used to learn and reconstruct clear and complete temporal feature maps of expressions from low-quality, noisy short video facial features. It is specifically trained for the sparsity (few frames) and noise characteristics of short videos, enabling it to "complete" effective facial dynamic information even with severely degraded data. During model training, pose and illumination estimators are introduced as adversarial networks to force the expression features learned by the backbone feature extraction network to decouple from head pose and illumination conditions. Simultaneously, a 3D facial deformation model is used for geometric normalization, mapping facial movements under different poses to a standard, frontal, neutral face model, thereby generating geometrically and illumination-invariant expression motion features. A temporal saliency detection module, namely the aforementioned face action detection module, is employed to automatically identify and suppress continuous, large-amplitude motion regions in the video. The second path (enhanced path) employs a transient facial expression focusing module to locate brief, subtle bursts of muscle movement in both time and space, and adaptively enhances their features. The two paths are dynamically fused to "filter out interference and amplify the signal" from complex facial dynamics, thereby providing a high-precision benchmark on the server side.
[0158] Optionally, in this embodiment of the application, before step 401 above, the facial expression recognition method provided in this embodiment of the application further includes the following steps 501 and 502.
[0159] Step 501: The electronic device obtains the noise value corresponding to the second image feature set through the noise scheduling module in the second initial model.
[0160] In this embodiment, the noise value is used to characterize the noise level of the image feature information contained in the second image feature set.
[0161] It should be noted that the specific implementation process for obtaining the noise value corresponding to the second image feature set can be found in the description in the relevant technology. To avoid repetition, it will not be repeated here.
[0162] Step 502: The electronic device calculates the initial module parameters of each module in the second initial model based on the noise value through the noise scheduling module.
[0163] For example, the noise scheduling module described above abandons the traditional fixed-parameter noise scheduling strategy and adopts a parameterized adaptive scheduling mechanism. The core innovation lies in the learnable β parameter sequence: the noise scheduling parameter β is transformed from a fixed value into a learnable logarithmic space parameter, enabling the model to autonomously optimize the noise injection strategy based on the characteristics of the training data.
[0164] Temperature regulation mechanism: Introduce temperature parameters to control the sharpness of the distribution of scheduling weights, so as to achieve a smooth transition from uniform sampling to key sampling.
[0165] Dynamic weighted fusion: Learnable weight coefficients are assigned to each time step, and the importance of each time step is weighted through a soft maximization function.
[0166] For example, during the forward propagation phase of the diffusion process, the noise scheduling module calculates the weighted fused noise coefficients based on the current time step t. Specifically, it multiplies the standard model parameters by their cumulative product and sums the learned weights to obtain an adaptively adjusted noise scaling factor.
[0167] This design allows the model to allocate more attention to the critical denoising stage and quickly pass through the simpler stages, significantly improving training efficiency and generation quality.
[0168] This application provides an expression recognition method. Figure 7 A flowchart of an expression recognition method provided in an embodiment of this application is shown. Figure 7 As shown, the facial expression recognition method provided in this application embodiment may include the following steps 20 to 25.
[0169] Step 20: The user uploads a short video.
[0170] Step 21: The electronic device performs data preprocessing and feature enhancement on the short video.
[0171] Step 22: The electronic device trains the teacher model to obtain the trained teacher model.
[0172] Step 23: The electronic device performs multi-granularity knowledge distillation on the student model using the trained teacher model to obtain the trained student model.
[0173] Step 24: The electronic device performs adaptive quantization and fine-tuning on the trained student model to obtain the quantized student model.
[0174] Step 25: Deploy the quantized student model on the electronic device and perform facial expression reasoning to obtain facial expression information.
[0175] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0176] It should be noted that the facial expression recognition method provided in this application embodiment can be executed by an facial expression recognition device. This application embodiment uses an facial expression recognition device executing the facial expression recognition method as an example to illustrate the facial expression recognition device provided in this application embodiment.
[0177] Figure 8 A schematic diagram of a possible structure of the facial expression recognition device involved in an embodiment of this application is shown. For example... Figure 8 As shown, the facial expression recognition device 70 may include an extraction module 71, a processing module 72, and a reconstruction module 73.
[0178] The extraction module 71 is used to extract features from the first set of video frames corresponding to the first video to obtain a first image feature set corresponding to the first set of video frames. All video frames in the first set of video frames include facial images. The processing module 72 is used to input the first image feature set obtained by the extraction module into a first expression recognition model. The first expression recognition model uses a denoising module to denoise the image features in the image feature set to obtain a first expression feature set. The first expression feature set includes expression features of facial images in the video frames and expression change features of facial images between every two video frames. The reconstruction module 73 is used to reconstruct expressions based on the first expression feature set obtained by the processing module using the first expression recognition model to recognize facial expressions in the first video.
[0179] In the facial expression recognition device provided in this application embodiment, the facial expression features of the face image and the facial expression change features between every two video frames are obtained through the first facial expression recognition model. The first facial expression recognition model in the electronic device can extract complete facial expression feature information from the features of low-quality, noisy short video images, thereby reconstructing a clear and complete facial expression through the complete facial expression feature information. In other words, the first facial expression recognition model in the facial expression recognition device can obtain effective facial expression information even when the video frame information is severely degraded, thus improving the accuracy of facial expression recognition device in recognizing facial expressions.
[0180] In one possible implementation, the aforementioned first expression feature set includes expression features under different facial poses. The processing module 72 is specifically used to input the first expression feature set into the expression transient focusing module of the first expression recognition model, identify the expression feature information in the first expression feature set, and perform feature enhancement on the expression feature information to obtain an enhanced first expression feature set. The reconstruction module 73 is specifically used to perform expression reconstruction based on the enhanced first expression feature set obtained by the processing module to obtain a facial expression.
[0181] In one possible implementation, the processing module 72 is further configured to: extract features from the first video frame set corresponding to the first video to obtain the image feature set corresponding to the first video frame set; train the first initial model based on the second expression recognition model to obtain the trained first initial model; and quantize the model parameters of the trained first initial model to obtain the first expression recognition model; wherein the second expression recognition model has the same model structure as the first expression recognition model, the storage space occupied by the first expression recognition model is smaller than that occupied by the second expression recognition model, the second expression recognition model is a teacher model, and the first expression recognition model is a student model.
[0182] In one possible implementation, the processing module 72 is specifically used to quantize the model parameters corresponding to each model layer based on the quantization strategy corresponding to each model layer of the trained first initial model, to obtain the quantized first initial model; and to perform quantization perception adjustment processing on the quantized first initial model to obtain the first expression recognition model.
[0183] In one possible implementation, the processing module 72 is further configured to: add noise to the second image feature set using a second initial model to obtain a noisy second image feature set; input the noisy second image feature set into a face action detection module within the second initial model for face action detection to obtain a first face feature set, wherein the face features included in the first face feature set correspond to action amplitudes less than or equal to a preset amplitude threshold; and input the first face feature set into a multi-scale feature injection module within the second initial model for feature extraction to obtain a second face feature set, wherein the second face feature set includes different images. The system first obtains facial feature information at a specific size. Then, it inputs the second facial feature set into the spatiotemporal attention denoising module of the second initial model to denoise the facial feature set, resulting in a second expression feature set. This second expression feature set includes the expression features of the facial images in the video frames and the expression change features between every two video frames. Next, it inputs the second expression feature set into the expression transient focusing module of the second initial model to identify expression feature information and enhance it, resulting in an enhanced second expression feature set. Based on the enhanced second expression feature set, the second initial model is trained to obtain the second expression recognition model.
[0184] The facial expression recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0185] The facial expression recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0186] The facial expression recognition device provided in this application embodiment can implement all the processes implemented in the above method embodiments, and will not be described again here to avoid repetition.
[0187] Optionally, such as Figure 9 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described expression recognition method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0188] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0189] Figure 10 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0190] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0191] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 10 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0192] The processor 110 is configured to extract features from a set of first video frames corresponding to a first video to obtain a first image feature set corresponding to the first video frame set, wherein all video frames in the first video frame set include face images; input the first image feature set into a first expression recognition model, and perform denoising processing on the image features in the image feature set through the spatiotemporal attention denoising module in the first expression recognition model to obtain a first expression feature set, wherein the first expression feature set includes expression features of face images in video frames and expression change features of face images between every two video frames; and perform expression reconstruction based on the first expression feature set through the first expression recognition model to recognize the face expressions in the first video.
[0193] In the electronic device provided in this application embodiment, facial expression features of a face image and facial expression change features between every two video frames are obtained through a first expression recognition model. The first expression recognition model in the electronic device can extract complete expression feature information from the features of low-quality, noisy short video images, thereby reconstructing a clear and complete expression through the complete expression feature information. In other words, the first expression recognition model in the electronic device can obtain effective expression information even when the video frame information is severely degraded, thus improving the accuracy of the electronic device in recognizing expressions.
[0194] Optionally, in this embodiment of the application, the first expression feature set includes expression features under different facial poses. Specifically, the processor 110 is used to input the first expression feature set into the expression transient focusing module of the first expression recognition model, identify the expression feature information in the first expression feature set, and perform feature enhancement on the expression feature information to obtain an enhanced first expression feature set; based on the enhanced first expression feature set, perform expression reconstruction to obtain a facial expression.
[0195] Optionally, in this embodiment of the application, the processor 110 is further configured to perform feature extraction on the first video frame set corresponding to the first video to obtain the image feature set corresponding to the first video frame set, and then train the first initial model based on the second expression recognition model to obtain the trained first initial model; quantize the model parameters of the trained first initial model to obtain the first expression recognition model; wherein the second expression recognition model has the same model structure as the first expression recognition model, the storage space occupied by the first expression recognition model is smaller than that occupied by the second expression recognition model, the second expression recognition model is a teacher model, and the first expression recognition model is a student model.
[0196] Optionally, in this embodiment of the application, the processor 110 is specifically used to quantize the model parameters corresponding to each model layer based on the quantization strategy corresponding to each model layer of the trained first initial model, to obtain the quantized first initial model; and to perform quantization perception adjustment processing on the quantized first initial model to obtain the first expression recognition model.
[0197] Optionally, in this embodiment, the processor 110 is further configured to: add noise to the second image feature set using a second initial model to obtain a noisy second image feature set; input the noisy second image feature set into a face action detection module in the second initial model for face action detection to obtain a first face feature set, wherein the face features included in the first face feature set correspond to action amplitudes less than or equal to a preset amplitude threshold; and input the first face feature set into a multi-scale feature injection module in the second initial model for feature extraction to obtain a second face feature set, wherein the second face feature set includes different images. The system first obtains facial feature information at a specific size. Then, it inputs the second facial feature set into the spatiotemporal attention denoising module of the second initial model to denoise the facial feature set, resulting in a second expression feature set. This second expression feature set includes the expression features of the facial images in the video frames and the expression change features between every two video frames. Next, it inputs the second expression feature set into the expression transient focusing module of the second initial model to identify expression feature information and enhance it, resulting in an enhanced second expression feature set. Based on the enhanced second expression feature set, the second initial model is trained to obtain the second expression recognition model.
[0198] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0199] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.
[0200] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0201] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0202] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0203] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0204] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0205] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0206] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0207] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.
[0208] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0209] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0210] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A facial expression recognition method, characterized in that, The method includes: Feature extraction is performed on the first video frame set corresponding to the first video to obtain the first image feature set corresponding to the first video frame set, wherein all video frames in the first video frame set include face images. The first image feature set is input into the first expression recognition model. The image features in the image feature set are denoised by the denoising module in the first expression recognition model to obtain the first expression feature set. The first expression feature set includes the expression features of the face images in the video frames and the expression change features of the face images between every two video frames. The first expression recognition model is used to reconstruct expressions based on the first expression feature set and identify facial expressions in the first video.
2. The method according to claim 1, characterized in that, The first set of facial expression features includes facial expression features under different facial postures; The step of reconstructing facial expressions in the first video using the first expression recognition model based on the first expression feature set and identifying facial expressions includes: The first expression recognition model inputs the first expression feature set into the expression transient focusing module in the first expression recognition model to identify the expression feature information in the first expression feature set and enhances the expression feature information to obtain the enhanced first expression feature set. Based on the enhanced first set of facial expression features, facial expression reconstruction is performed to obtain the facial expression.
3. The method according to claim 1, characterized in that, Before performing feature extraction on the first video frame set corresponding to the first video to obtain the image feature set corresponding to the first video frame set, the method further includes: The first initial model is trained based on the second facial expression recognition model to obtain the trained first initial model. The model parameters of the first initial model after training are quantized to obtain the first expression recognition model; The second expression recognition model has the same model structure as the first expression recognition model, the first expression recognition model occupies less storage space than the second expression recognition model, the second expression recognition model is a teacher model, and the first expression recognition model is a student model.
4. The method according to claim 3, characterized in that, The step of quantizing the model parameters of the trained first initial model to obtain the first expression recognition model includes: Based on the quantization strategy corresponding to each model layer of the first initial model after training, the model parameters corresponding to each model layer are quantized respectively to obtain the quantized first initial model. The quantized first initial model is subjected to quantization perception adjustment processing to obtain the first expression recognition model.
5. The method according to claim 3, characterized in that, The method further includes: The second image feature set is noise-added using the second initial model to obtain the noise-added second image feature set. Through the second initial model, the noisy second image feature set is input into the face action detection module in the second initial model to perform face action detection and obtain a first face feature set. The action amplitude corresponding to the face features included in the first face feature set is less than or equal to a preset amplitude threshold. The first face feature set is input into the multi-scale feature injection module of the second initial model through the second initial model to extract features and obtain the second face feature set, which includes face feature information under different image sizes. The second face feature set is input into the spatiotemporal attention denoising module of the second initial model through the second initial model to perform denoising processing on the second face feature set, and a second expression feature set is obtained. The second expression feature set includes the expression features of the face images in the video frames and the expression change features of the face images between every two video frames. The second expression feature set is input into the expression transient focusing module in the second initial model through the second initial model to identify expression feature information and perform feature enhancement on the expression feature information to obtain the enhanced second expression feature set. Based on the enhanced second expression feature set, the second initial model is trained to obtain the second expression recognition model.
6. An expression recognition device, characterized in that, The facial expression recognition device includes: an extraction module, a processing module, and a reconstruction module; The extraction module is used to extract features from the first video frame set corresponding to the first video to obtain a first image feature set corresponding to the first video frame set, wherein all video frames in the first video frame set include face images. The processing module is used to input the first image feature set obtained by the extraction module into the first expression recognition model, and to perform noise reduction processing on the image features in the image feature set through the noise reduction module in the first expression recognition model to obtain the first expression feature set. The first expression feature set includes the expression features of the face images in the video frames and the expression change features of the face images between every two video frames. The reconstruction module is used to reconstruct facial expressions in the first video by using the first expression recognition model and the first expression feature set obtained by the processing module.
7. The apparatus according to claim 6, characterized in that, The first set of facial expression features includes facial expression features under different facial postures; The processing module is specifically used to input the first expression feature set into the expression transient focusing module in the first expression recognition model through the first expression recognition model, identify the expression feature information in the first expression feature set, and perform feature enhancement on the expression feature information to obtain the enhanced first expression feature set. The reconstruction module is specifically used to reconstruct facial expressions based on the enhanced first set of facial expression features obtained by the processing module, thereby obtaining the facial expression.
8. The apparatus according to claim 6, characterized in that, The processing module is further configured to perform feature extraction on the first video frame set corresponding to the first video before the extraction module obtains the image feature set corresponding to the first video frame set, and then train the first initial model based on the second expression recognition model to obtain the trained first initial model. The model parameters of the first initial model after training are quantized to obtain the first expression recognition model; The second expression recognition model has the same model structure as the first expression recognition model, the first expression recognition model occupies less storage space than the second expression recognition model, the second expression recognition model is a teacher model, and the first expression recognition model is a student model.
9. The apparatus according to claim 8, characterized in that, The processing module is specifically used to quantize the model parameters corresponding to each model layer based on the quantization strategy corresponding to each model layer of the trained first initial model, so as to obtain the quantized first initial model. The quantized first initial model is then subjected to quantization perception adjustment processing to obtain the first expression recognition model.
10. The apparatus according to claim 8, characterized in that, The processing module is further configured to add noise to the second image feature set using the second initial model to obtain a noisy second image feature set. Through the second initial model, the noisy second image feature set is input into the face action detection module in the second initial model to perform face action detection and obtain a first face feature set. The action amplitude corresponding to the face features included in the first face feature set is less than or equal to a preset amplitude threshold. The first face feature set is input into the multi-scale feature injection module of the second initial model through the second initial model to extract features and obtain the second face feature set, which includes face feature information under different image sizes. The second face feature set is input into the spatiotemporal attention denoising module of the second initial model through the second initial model to perform denoising processing on the second face feature set, and a second expression feature set is obtained. The second expression feature set includes the expression features of the face images in the video frames and the expression change features of the face images between every two video frames. The second expression feature set is input into the expression transient focusing module in the second initial model through the second initial model to identify expression feature information and perform feature enhancement on the expression feature information to obtain the enhanced second expression feature set. Based on the enhanced second expression feature set, the second initial model is trained to obtain the second expression recognition model.