Micro-expression Analysis Method and System

Through the multi-scale residual channel attention network model, the problem of insufficient feature extraction capability of micro-expression analysis methods in the prior art is solved, and a higher positioning and recognition accuracy is achieved.

CN119942620BActive Publication Date: 2025-06-27NANCHANG UNIV

Patent Information

Application Number
CN202510424098.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-27
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing micro-expression analysis methods lack the feature extraction ability, resulting in low localization and recognition accuracy.

Method used

A multi-scale residual channel attention network model is adopted, including multi-scale shared subnet, positioning subnet and identification subnet. The key features of micro-expressions are extracted through channel attention mechanism, convolutional layer processing and residual connection strategy.

Benefits of technology

The feature extraction ability of micro-expressions is improved, the model's performance in complex emotion recognition tasks is enhanced, and the accuracy of positioning and recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942620B_ABST
    Figure CN119942620B_ABST
Patent Text Reader

Abstract

The present application discloses a micro-expression analysis method and system, belonging to the field of computer vision technology. The method includes: obtaining a micro-expression data set and preprocessing it to obtain a target data set; constructing a multi-scale residual channel attention network model, and training and testing the multi-scale residual channel attention network model according to the target data set; inputting the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputting the localization data and recognition data of the image stream to be analyzed; processing the localization data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval. By comprehensively using the design of the multi-scale shared sub-network and the residual connection channel attention module, the multi-scale residual channel attention network model can provide higher accuracy in the localization and recognition of micro-expressions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision technology, and particularly relates to a micro-expression analysis method and system. Background Art

[0002] Micro-expression is a very short-lived (less than 0.5 seconds) and hardly noticeable facial expression change. The generation of micro-expression is usually unconscious and spontaneous, often occurring before people are aware of it. The muscle movement on the face is very small and difficult to observe with the naked eye. It is a natural expression of people's inner emotions, usually reflecting people's true emotions and feelings. Therefore, it is very difficult to avoid micro-expression from being unnoticed through control or concealment. Since micro-expression is undeceiving and can reflect people's true emotions, micro-expression recognition has great practical significance in real life.

[0003] Existing micro-expression analysis methods generally adopt the MEAN framework and a simple three-layer convolutional network, but their feature extraction ability is insufficient. In addition, although attention mechanisms such as CBAM and ViT perform well in other visual tasks, when dealing with micro-expressions, due to their short and subtle changes, traditional attention mechanisms are difficult to adjust the focus in time to capture these key features, resulting in low localization and recognition accuracy. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a micro-expression analysis method and system, which can solve the technical problem that the existing technology has insufficient extraction of key features of micro-expressions, resulting in low localization and recognition accuracy.

[0005] To solve the above technical problem, this application is implemented as follows:

[0006] In a first aspect, the embodiments of this application provide a micro-expression analysis method, and the method includes:

[0007] Obtain a micro-expression data set, and preprocess the micro-expression data set to obtain a target data set;

[0008] Construct a multi-scale residual channel attention network model, and the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a localization sub-network, and an identification sub-network;

[0009] Train and test the multi-scale residual channel attention network model according to the target data set;

[0010] Input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the localization data and identification data of the image stream to be analyzed;

[0011] Process the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval;

[0012] Among them, the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0013] As an optional implementation manner of the first aspect of the present application, obtain a micro-expression data set, and preprocess the micro-expression data set to obtain a target data set; specifically:

[0014] Obtain multiple initial image streams containing micro-expressions, and perform face cropping processing on each initial image stream to obtain each face image stream corresponding to each initial image stream;

[0015] Perform optical flow extraction, ROI selection, and resampling processing on each face image stream to obtain each target image stream corresponding to each face image stream;

[0016] Construct the target data set according to each target image stream, where each target image stream includes three types of data: horizontal optical flow, vertical optical flow, and optical strain.

[0017] As an optional implementation manner of the first aspect of the present application, train and test the multi-scale residual channel attention network model according to the target data set; specifically:

[0018] Divide the target data set into a target training set and a target test set according to a preset ratio;

[0019] Construct a loss function, and train the multi-scale residual channel attention network model according to the target training set and the loss function;

[0020] Test the trained multi-scale residual channel attention network model according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

[0021] As an optional implementation manner of the first aspect of the present application, the multi-scale shared sub-network; includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers, and one splicing layer; the process of the multi-scale residual channel attention network model processing the image stream to be analyzed is:

[0022] According to the three multi-scale convolutional layers, perform multi-scale convolutional processing on the horizontal optical flow data, vertical optical flow data, and optical strain data of the image stream to be analyzed respectively to obtain a horizontal optical flow feature map, a vertical optical flow feature map, and an optical strain feature map;

[0023] Perform attention weighting processing on the horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map respectively according to the three residual connection channel attention modules to obtain a horizontal weighted feature map, a vertical weighted feature map, and an optical strain weighted feature map;

[0024] Perform convolution processing on the horizontal weighted feature map, vertical weighted feature map, and the optical strain weighted feature map respectively according to the three convolutional layers to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map, and an optical strain weighted convolutional feature map;

[0025] Perform normalization processing, max-pooling processing, and stochastic depth dropout processing on the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional feature map in sequence, and splice the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional feature map after the stochastic depth dropout processing according to the splicing layer to obtain a spliced feature map;

[0026] The positioning sub-network performs positioning processing on the spliced feature map to obtain the positioning data of the image flow to be analyzed, and the recognition sub-network processes the positioning data and the spliced feature map to obtain recognition data.

[0027] As an optional implementation manner of the first aspect of the present application, the process of the multi-scale convolutional layer performing multi-scale convolution on the horizontal optical flow data is as follows:

[0028] Perform two-dimensional convolution processing on the horizontal optical flow data to obtain multiple convolution feature maps with different scales, and perform activation processing on the multiple convolution feature maps with different scales to obtain multiple activation feature maps;

[0029] Perform batch normalization processing on the multiple activation feature maps, and splice the multiple activation feature maps after the batch normalization processing along the channel dimension to obtain the horizontal optical flow feature map;

[0030] According to the process of the multi-scale convolutional layer performing multi-scale convolution processing on the horizontal optical flow data, perform multi-scale convolution processing on the vertical optical flow data and optical strain data to obtain the vertical optical flow feature map and the optical strain feature map.

[0031] As an optional implementation manner of the first aspect of the present application, the process of the residual connection channel attention module performing attention weighting processing on the horizontal optical flow feature map; specifically:

[0032] Perform average pooling and max-pooling processing on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a max-pooling feature map;

[0033] Perform channel compression and channel restoration processing on both the average pooling feature map and the max pooling feature map to obtain an average pooling weighted feature map and a max pooling weighted feature map;

[0034] Perform residual connection on the average pooling weighted feature map and the max pooling weighted feature map to obtain a weighted concatenated feature map, perform 1×1 convolution processing on the weighted concatenated feature map to obtain a convolution weighted feature map, and perform normalization processing on the convolution weighted feature map to obtain a weighted matrix;

[0035] Multiply the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and perform residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain the horizontal weighted feature map;

[0036] According to the process of performing attention weighting processing on the horizontal optical flow feature map by the residual connection channel attention module, perform attention weighting processing on the vertical optical flow feature map and the optical strain feature map to obtain the vertical weighted feature map and the optical strain weighted feature map.

[0037] As an optional implementation manner of the first aspect of the present application, process the positioning data and the recognition data to obtain the micro-expression interval of the to-be-analyzed image stream and the micro-expression category of the micro-expression interval; specifically:

[0038] If only one micro-expression interval is included in the to-be-analyzed image stream, obtain the maximum peak position in the positioning data, identify the micro-expression interval of the to-be-analyzed image stream according to the maximum peak position and the front and back interval positions of the maximum peak position, and obtain the micro-expression category corresponding to the micro-expression interval according to the recognition data;

[0039] If multiple micro-expression intervals are included in the to-be-analyzed image stream, perform smoothing processing on the positioning data to obtain smoothed positioning data, calculate the average value and the maximum value of the smoothed positioning data, calculate an interval threshold according to the average value and the maximum value, obtain multiple micro-expression intervals in the to-be-analyzed image stream according to the interval threshold, and obtain each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals according to the recognition data.

[0040] In a second aspect, an embodiment of the present application provides a micro-expression analysis system, and the system includes:

[0041] An acquisition module: acquire a micro-expression data set, and perform preprocessing on the micro-expression data set to obtain a target data set;

[0042] Building module: Build a multi-scale residual channel attention network model, where the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a localization sub-network, and an identification sub-network;

[0043] Training and testing module: Train and test the multi-scale residual channel attention network model according to the target data set;

[0044] Analysis module: Input the image stream to be analyzed into the multi-scale residual channel attention network model after testing for analysis, and output the localization data and identification data of the image stream to be analyzed; wherein, the localization data includes the position information of the micro-expression interval in the image stream to be analyzed, and the identification data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed;

[0045] Processing module: Process the localization data and the identification data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

[0046] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0047] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0048] In the embodiment of the present application, compared with the prior art, the following beneficial effects are achieved:

[0049] (1) Improve feature extraction ability: Through the residual connection channel attention module that integrates the channel attention mechanism, convolutional layer processing, and residual connection strategy, the multi-scale residual channel attention network model can more effectively extract key features in micro-expressions; this design not only enhances the performance of the model in complex emotion recognition tasks, but also improves its robustness and generalization performance.

[0050] (2) Enhance the ability to capture features of different scales: The multi-scale convolutional layers introduced in the multi-scale shared sub-network enable the network to capture important information at different scales in the image, which is particularly important for micro-expression analysis because the changes in micro-expressions often involve details at multiple scales, which helps to improve the accuracy of localization and recognition.

[0051] (3) Improve the attention to important information: The feature map processed by the residual connection channel attention module enables the model to focus more on the key regions in the image rather than being disturbed by irrelevant background information. This method ensures that even if the micro-expression changes are both short-lived and subtle, the model can timely adjust its focus to effectively capture the information at these critical moments.

[0052] (4) Improve the localization and recognition accuracy: By comprehensively using the design of the multi-scale shared sub-network and the residual connection channel attention module, the multi-scale residual channel attention network model can provide higher accuracy in the localization and recognition of micro-expressions. Brief Description of the Drawings

[0053] Figure 1 is a flowchart of a micro-expression analysis method provided by some embodiments of the present application;

[0054] Figure 2 is a structural diagram of a multi-scale residual channel attention network model of a micro-expression analysis method provided by some embodiments of the present application;

[0055] Figure 3 is a structural diagram of a residual connection channel attention module of a micro-expression analysis method provided by some embodiments of the present application. Detailed Embodiments

[0056] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0057] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0058] Next, a micro-expression analysis method and system provided by embodiments of the present application will be described in detail in conjunction with the accompanying drawings through specific embodiments and their application scenarios.

[0059] Embodiment

[0060] Refer to Figure 1 As shown, a micro-expression analysis method includes the following steps:

[0061] S100: Obtain a micro-expression dataset, preprocess the micro-expression dataset to obtain a target dataset;

[0062] It should be noted that S100 specifically is:

[0063] S110: Obtain multiple initial image streams containing micro-expressions, perform face cropping on each initial image stream to obtain each face image stream corresponding to each initial image stream;

[0064] S120: Perform optical flow extraction, ROI selection, and resampling on each face image stream to obtain each target image stream corresponding to each face image stream;

[0065] S130: Construct a target dataset based on each target image stream, where each target image stream contains three types of data: horizontal optical flow, vertical optical flow, and optical strain.

[0066] Furthermore, after face detection, it is cropped to 128×128 based on the first frame of the initial image stream, and then 68 key points of the face are marked to assist in ROI selection and resampling. The TV-L1 algorithm is used to extract the optical flow. This algorithm has strong robustness to noise and can accurately capture the dynamic changes of micro-expressions to obtain the horizontal optical flow (u) and the vertical optical flow (v). At the same time, calculate the optical strain to approximate the facial deformation intensity, represented by the Hessian matrix, and calculate the optical strain amplitude value as one of the input features. Combine the optical flow and the optical strain to form a triple as the input of the micro-expression analysis network. In addition, select three regions rich in micro-expression information: the left eye and the left eyebrow, the right eye and the right eyebrow, and the mouth for ROI selection and image resampling, adjust the region sizes to 21×21 and 21×42, and finally combine them into a target image stream of 42×42.

[0067] S200: Construct a multi-scale residual channel attention network model, and the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a localization sub-network, and an identification sub-network.

[0068] Refer to Figure 2 as shown, Figure 2 is the structure diagram of the multi-scale residual channel attention network model. Furthermore, the multi-scale shared sub-network includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers, and one splicing layer.

[0069] S300: Train and test the multi-scale residual channel attention network model according to the target dataset;

[0070] It should be noted that S300 specifically is:

[0071] S310: Divide the target data set into a target training set and a target test set according to a preset ratio;

[0072] S320: Construct a loss function, and train the multi-scale residual channel attention network model according to the target training set and the loss function;

[0073] S330: Test the trained multi-scale residual channel attention network model according to the target test set, and complete the training and testing of the multi-scale residual channel attention network model.

[0074] Further, in this embodiment, the target data set is divided into a target training set and a target test set according to a ratio of 8:2; construct a loss function, and train the multi-scale residual channel attention network model according to the target training set and the loss function; the loss function is used to optimize the parameters in the training process of the multi-scale residual channel attention network model; test the trained multi-scale residual channel attention network model according to the target test set, and complete the training and testing of the multi-scale residual channel attention network model.

[0075] Still further, in this embodiment, the target data set may be a first target data set or a second target data set. Each target image stream in the first target data set only contains one micro-expression interval; and each target image stream in the second target data set contains multiple micro-expression intervals; however, whether it is for the first target data set or the second target data set, the target training set and the target test set are divided according to a ratio of 8:2, and the training and testing processes of the first target data set and the second target data set are independent. That is to say, the multi-scale residual channel attention network model can learn the target image stream containing only one micro-expression interval, and can also learn the target image stream containing multiple micro-expression intervals, so that the multi-scale residual channel attention network model can identify the micro-expression interval and the category of the micro-expression interval in the target image stream containing only one micro-expression interval. At the same time, it can also identify multiple micro-expression intervals in the target image stream containing multiple micro-expression intervals, and the micro-expression category corresponding to each micro-expression interval among the multiple micro-expression intervals.

[0076] Furthermore, in this embodiment, a pseudo-labeling technique is also used to perform pseudo-labeling operations on each target image stream in the target data set, obtaining each pseudo-label set corresponding to each target image stream. Pseudo-labeling the target image stream before training the target data set can improve the generalization ability and robustness of the model. When putting the target training set into the multi-scale residual channel attention network model for training, each pseudo-label set corresponding to each target image stream in the target training set is also put into the multi-scale residual channel attention network model for training at the same time. And the pseudo-label set is only used for the training of the localization sub-network. That is, after the multi-scale shared sub-network in the multi-scale residual channel attention network model trains a target image stream in the target data set and outputs the output feature corresponding to the target image stream, when the output feature enters the localization sub-network for training, the pseudo-label set corresponding to the target image stream and the output feature enter the localization sub-network for training together.

[0077] S400: Input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the localization data and recognition data of the image stream to be analyzed;

[0078] It should be noted that S400 specifically is:

[0079] S410: According to three multi-scale convolutional layers, perform multi-scale convolutional processing on the horizontal optical flow data, vertical optical flow data, and optical strain data of the image stream to be analyzed respectively, obtaining a horizontal optical flow feature map, a vertical optical flow feature map, and an optical strain feature map;

[0080] S420: According to three residual connection channel attention modules, perform attention weighting processing on the horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map respectively, obtaining a horizontal weighted feature map, a vertical weighted feature map, and an optical strain weighted feature map;

[0081] S430: According to three convolutional layers, perform convolutional processing on the horizontal weighted feature map, vertical weighted feature map, and optical strain weighted feature map respectively, obtaining a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map, and an optical strain weighted convolutional feature map;

[0082] S440: Perform normalization processing, max-pooling processing, and random path dropout processing on the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional feature map in sequence, and splice the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional after random path dropout processing according to the splicing layer, obtaining a spliced feature map;

[0083] S450: The localization sub-network performs localization processing on the spliced feature map to obtain the localization data of the image stream to be analyzed, and the recognition sub-network processes the localization data and the spliced feature map to obtain the recognition data.

[0084] Further, according to the three multi-scale convolutional layers in the multi-scale shared network (MSSN), multi-scale convolutional processing is respectively performed on the horizontal optical flow data, vertical optical flow data, and optical strain data of the image stream to be analyzed, obtaining a horizontal optical flow feature map, a vertical optical flow feature map, and an optical strain feature map; this can enhance the network model's ability to capture features of different scales in the image; the horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map respectively output by the three multi-scale convolutional layers are processed through three residual connection channel attention modules (RCCAM) to obtain a horizontal weighted feature map, a vertical weighted feature map, and an optical strain weighted feature map respectively. This operation can enhance the network model's attention to important information; convolutional processing is respectively performed on the horizontal weighted feature map, vertical weighted feature map, and optical strain weighted feature map again to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map, and an optical strain weighted convolutional feature map respectively to further enhance the feature representation ability; batch normalization (BatchNormalization) and max pooling (MaxPooling2D) operations are performed on the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional feature map to further improve the distinguishability and robustness of the features, and a random path dropout layer is used to reduce the overfitting risk of the model; finally, according to the concatenation layer, the horizontal weighted convolutional feature map, vertical weighted convolutional feature map, and optical strain weighted convolutional feature map after random path dropout processing are concatenated to obtain a concatenated feature map; finally, the concatenated feature map is processed by the localization sub-network to obtain the localization data of the image stream to be analyzed, and the recognition sub-network processes the localization data and the concatenated feature map to obtain the recognition data; among them, the localization data contains the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data contains the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0085] It should be noted that S410 specifically is:

[0086] S411: Perform two-dimensional convolutional processing on the horizontal optical flow data to obtain multiple convolutional feature maps of different scales, and perform activation processing on the multiple convolutional feature maps of different scales to obtain multiple activation feature maps;

[0087] S412: Perform batch normalization processing on the multiple activation feature maps, and concatenate the multiple activation feature maps after batch normalization processing along the channel dimension to obtain a horizontal optical flow feature map;

[0088] S413: According to the processing process of the horizontal optical flow data in S411 - S412, perform multi-scale convolutional processing on the vertical optical flow data and the optical strain data to obtain a vertical optical flow feature map and an optical strain feature map.

[0089] Furthermore, for the multi-scale shared network (MSSN) of this embodiment, in order to enhance the network's ability to capture features of different scales in images, a multi-scale convolutional layer is added to the MSSN; the design of this layer aims to extract multi-level feature representations of the input data through convolutional kernels of different sizes. Specifically, given the input tensor inputs (horizontal optical flow data, vertical optical flow data, and optical strain data), whose number of channels is filters, a list kernel_sizes containing various sizes of convolutional kernels is defined. The inputs of the horizontal optical flow data and the vertical optical flow data are set to [1, 3], and the input of the optical strain data is set to [3, 5]. For each convolutional kernel size in the list, the following operations are performed: Apply a two-dimensional convolutional operation (Conv2D), where the size of the convolutional kernel is equal to the currently selected kernel_size, and the 'same' type of padding is used to maintain the consistency of the input and output sizes. To prevent overfitting and introduce a regularization effect, an L2 regularization term is added during the convolution process. The activation function is selected as RelU to increase the non-linear expression ability of the model. To further improve the generalization ability and stability of the model, a batch normalization operation is added after each convolutional layer. This helps to accelerate the training process and reduce problems caused by internal covariate shift. To integrate the feature information at different scales, all the feature maps processed as above are concatenated together along the channel dimension to form an output of multi-scale feature fusion (horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map). This structure allows the network to simultaneously obtain local details and broader context information from the same input, thus improving the expressiveness of the final model.

[0090] Specifically, the multi-scale convolutional layer is represented by the following formula:

[0091] ,

[0092] where, represents the input of the multi-scale convolutional layer, represents the two-dimensional convolutional processing, represents the activation function, represents batch normalization, represents feature concatenation, represents the output of the multi-scale convolutional layer.

[0093] Refer to Figure 3 as shown, Figure 3 is the structural diagram of the residual connection channel attention module. It should be noted that S420 is specifically:

[0094] S421: Perform average pooling and max pooling on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a max pooling feature map;

[0095] S422: Perform channel compression and channel restoration on both the average pooling feature map and the max pooling feature map to obtain an average pooling weighted feature map and a max pooling weighted feature map;

[0096] S423: Perform residual connection on the average pooling weighted feature map and the max pooling weighted feature map to obtain a weighted concatenated feature map, perform 1×1 convolution on the weighted concatenated feature map to obtain a convolution weighted feature map, and perform normalization on the convolution weighted feature map to obtain a weighted matrix;

[0097] S424: Multiply the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and perform residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain a horizontal weighted feature map;

[0098] S425: According to the process of performing attention weighting on the horizontal optical flow feature map in S421 - S424, perform attention weighting on the vertical optical flow feature map and the optical strain feature map to obtain a vertical weighted feature map and an optical strain weighted feature map.

[0099] Furthermore, global max pooling and global average pooling are performed on the input feature maps (horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map) to obtain the average pooling feature map and the max pooling feature map respectively, both with a shape of [C]. Then, the average pooling feature map and the max pooling feature map are reshaped into the shape [1, 1, C] for convenient subsequent operations. The reshaped average pooling feature map and max pooling feature map are input into a network containing two fully connected layers RD (each fully connected layer is equipped with a ReLU activation function) to obtain the average pooling weighted feature map and the max pooling weighted feature map. The first fully connected layer reduces the number of channels to r times the original number of channels (r is a hyperparameter used to control the ratio of channel compression) to learn the complex relationships between different channels. The second fully connected layer then restores the number of channels to the original number. This learning process can reveal which channels are more critical for micro-expression recognition and then generate a weight vector that matches the number of input channels. In this way, fine-grained weighting of each channel of the input feature map can be achieved, thereby enhancing the model's attention to important features. The average pooling weighted feature map and the max pooling weighted feature map are connected residually to obtain a weighted concatenated feature. Subsequently, the weighted concatenated feature is fed into a carefully designed 1x1 convolutional layer for further processing to obtain the convolutional weighted feature map. The main purpose of this convolutional layer is to further extract and refine features to better capture the subtle changes in micro-expressions. The configuration of the convolutional layer (i.e., the number of filters is set to 1, the kernel size is 1x1, the stride is 1, and the padding method is'same') is to keep the size of the feature map unchanged while learning cross-channel feature combinations. Subsequently, the Sigmoid function is used to normalize the weights of the convolutional weighted feature map to obtain a weight matrix ranging from [0, 1], and the input feature map is multiplied by the weight matrix to obtain the channel weighted feature map corresponding to the input feature map (horizontal channel weighted feature map, vertical channel weighted feature map, optical strain channel weighted feature map). Through this step, the model can learn richer feature representations, thereby improving its accuracy in recognizing micro-expressions. To further improve the training efficiency and generalization ability of the model, a residual connection is established between the channel weighted feature map and the input feature map to obtain the final output feature map (horizontal weighted feature map, vertical weighted feature map, and optical strain weighted feature map). The core idea of this strategy is that by directly adding the output feature map of the convolutional layer to the input feature map, the key information in the input features can be retained, and the problem of gradient disappearance that may occur during the training of deep networks can be alleviated.

[0100] Specifically, the data processing process of the residual connection channel attention module (RCCAM) is represented by the following formula

[0101] ,

[0102] ,

[0103] ,

[0104] ,

[0105] ,

[0106] Among them, represents the input feature map, represents two-dimensional max pooling processing, represents two-dimensional average pooling processing, represents reshaping the shape of the feature to , represents the max pooling feature map, represents the average pooling feature map; represents converting the number of channels to , represents a constant greater than 0 and less than 1, represents converting the number of channels to , represents the max pooling weighted feature map, represents the average pooling weighted feature map, represents the residual connection, represents the weighted concatenated feature map, represents performing two-dimensional convolution and activation processing on in sequence. The convolution kernel size of the two-dimensional convolution is , and the stride is , represents the normalization function, represents element-wise multiplication, represents the channel weighted feature map, represents the output feature map.

[0107] S500: Process the positioning data and recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

[0108] It should be noted that the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0109] It should be noted that in S500, if the image stream to be analyzed only contains one micro-expression interval, the maximum peak position in the positioning data is obtained, and based on the maximum peak position and the interval positions before and after the maximum peak position, the micro-expression interval of the image stream to be analyzed is identified, and the micro-expression category corresponding to the micro-expression interval is obtained according to the recognition data;

[0110] If the image stream to be analyzed contains multiple micro-expression intervals, the positioning data is smoothed to obtain smoothed positioning data, the average value and the maximum value of the smoothed positioning data are calculated, so as to calculate the interval threshold according to the average value and the maximum value, obtain multiple micro-expression intervals in the image stream to be analyzed according to the interval threshold, and obtain each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals according to the recognition data.

[0111] Furthermore, different micro-expression positioning strategies are adopted for image streams to be analyzed with different lengths. For short videos (the case where the image stream only contains one micro-expression interval), this embodiment adopts a simple peak discovery technique, and accurately identifies the micro-expression interval of the short video by positioning to obtain the maximum local peak and its front and back segments, and obtains the micro-expression category corresponding to the micro-expression interval of the short video according to the recognition data. For long videos (the case where the image stream contains multiple micro-expression intervals), a threshold-based technique is introduced to determine the micro-expression interval. Specifically, first, the positioning data corresponding to the long video is simply smoothed to avoid misidentifying sudden spikes from the positioning data. Subsequently, the average value and the maximum value of the smoothed positioning data are calculated, so as to calculate the interval threshold according to the average value and the maximum value in combination with preset parameters. This interval threshold is used to divide multiple micro-expression intervals from the smoothed positioning data, and obtain each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals according to the recognition data. In terms of micro-expression recognition, the research is based on the experimental results of previous work, that is, when predicting the emotion category in the recognition task, selecting the segment from the start frame to the vertex frame of the micro-expression for emotion recognition can obtain the best effect.

[0112] A micro-expression analysis method according to this embodiment, through a residual connection channel attention module that integrates a channel attention mechanism, convolutional layer processing, and a residual connection strategy, the multi-scale residual channel attention network model can extract key features in micro-expressions more effectively; this design not only enhances the performance of the model in complex emotion recognition tasks, but also improves its robustness and generalization performance. The multi-scale convolutional layers introduced in the multi-scale shared sub-network enable the network to capture important information at different scales in the image, which is particularly important for micro-expression analysis because the changes in micro-expressions often involve details at multiple scales. This helps to improve the accuracy of localization and recognition. The feature map processed by the residual connection channel attention module can make the model focus more on the key regions in the image rather than being disturbed by irrelevant background information. This method ensures that even if the micro-expression changes are both short-lived and subtle, the model can adjust its focus in a timely manner to effectively capture the information at these critical moments. By comprehensively using the design of the multi-scale shared sub-network and the residual connection channel attention module, the multi-scale residual channel attention network model can provide higher accuracy in the localization and recognition of micro-expressions. Experimental results show that compared with existing methods, the method proposed in the present invention has a significant improvement in overall performance, demonstrating its superiority in the field of micro-expression analysis.

[0113] It should be noted that for a micro-expression analysis method provided in an embodiment of the present application, the execution subject can be a micro-expression analysis system, or a control module in the micro-expression analysis system for executing and loading a micro-expression analysis method. In an embodiment of the present application, taking a micro-expression analysis system executing and loading a micro-expression analysis method as an example, a micro-expression analysis method provided in an embodiment of the present application is described.

[0114] A micro-expression analysis system includes:

[0115] An acquisition module: acquires a micro-expression data set, preprocesses the micro-expression data set to obtain a target data set;

[0116] A construction module: constructs a multi-scale residual channel attention network model, and the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a localization sub-network, and an identification sub-network;

[0117] A training and testing module: trains and tests the multi-scale residual channel attention network model according to the target data set;

[0118] An analysis module: inputs the image stream to be analyzed into the multi-scale residual channel attention network model after testing for analysis, and outputs the localization data and identification data of the image stream to be analyzed; wherein, the localization data includes the position information of the micro-expression interval in the image stream to be analyzed, and the identification data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed;

[0119] Processing module: processes the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

[0120] A micro-expression analysis system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.

[0121] A micro-expression analysis system in an embodiment of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0122] According to a micro-expression analysis system of this embodiment, the acquisition module can acquire the micro-expression data set required for training and testing, and preprocess the micro-expression data set to obtain a target data set; the construction module is used to construct a multi-scale residual channel attention network model, where the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a positioning sub-network, and a recognition sub-network; then the target data set processed by the acquisition module is used to train and test the multi-scale residual channel attention network model; then the data to be analyzed is input into the analysis module, and the analysis module inputs the image stream to be analyzed into the multi-scale residual channel attention network model after testing for analysis, and outputs the positioning data and recognition data of the image stream to be analyzed; finally, the processing module processes the positioning data and recognition data output by the analysis module to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval; through the mutual cooperation between the modules, the efficient positioning of the micro-expression interval in the image stream to be analyzed is realized, and the micro-expression category in the micro-expression interval is accurately recognized.

[0123] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-described embodiment of a micro-expression analysis method and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0124] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-described embodiment of a micro-expression analysis method and can achieve the same technical effects. To avoid repetition, details are not described herein again.

[0125] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0126] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element. In addition, it should be pointed out that the methods and devices in the embodiments of the present application are not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0128] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A micro-expression analysis method, characterized in that: The method comprises: Acquire a micro-expression dataset, and pre-process the micro-expression dataset to obtain a target dataset; A multi-scale residual channel attention network model is constructed, wherein the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a positioning sub-network and an identification sub-network, wherein the multi-scale shared sub-network includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers and a splicing layer; the process of the multi-scale residual channel attention network model processing the image stream to be analyzed is: According to the three multi-scale convolution layers, multi-scale convolution processing is performed on the horizontal optical flow data, the vertical optical flow data and the optical strain data of the image stream to be analyzed, respectively, to obtain a horizontal optical flow feature map, a vertical optical flow feature map and an optical strain feature map; According to the three residual connection channel attention modules, the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map are respectively subjected to attention weighting processing to obtain a horizontal weighted feature map, a vertical weighted feature map and an optical strain weighted feature map; According to the three convolutional layers, the horizontal weighted feature map, the vertical weighted feature map and the optical strain weighted feature map are respectively convolved to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map and an optical strain weighted convolutional feature map; The horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map are sequentially subjected to normalization processing, maximum pooling processing and random path discarding processing, and the horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map after the random path discarding processing are spliced ​​according to the splicing layer to obtain a spliced ​​feature map; The positioning sub-network performs positioning processing on the splicing feature map to obtain positioning data of the image stream to be analyzed, and the recognition sub-network processes the positioning data and the splicing feature map to obtain recognition data; Training and testing the multi-scale residual channel attention network model according to the target data set; Inputting the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputting the positioning data and recognition data of the image stream to be analyzed; Processing the positioning data and the recognition data to obtain a micro-expression interval of the image stream to be analyzed and a micro-expression category of the micro-expression interval; The positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

2. A micro-expression analysis method according to claim 1, characterized in that: The micro-expression dataset is obtained, and the micro-expression dataset is preprocessed to obtain a target dataset; specifically: Acquire multiple segments of initial image streams containing micro-expressions, perform face cropping processing on each of the initial image streams, and obtain each face image stream corresponding to each of the initial image streams; Performing optical flow extraction, ROI selection and resampling processing on each of the face image streams to obtain each target image stream corresponding to each of the face image streams; The target data set is constructed according to each of the target image streams, wherein each of the target image streams includes three types of data: horizontal optical flow, vertical optical flow and optical strain.

3. A micro-expression analysis method according to claim 1, characterized in that: The multi-scale residual channel attention network model is trained and tested according to the target data set; specifically: Dividing the target data set into a target training set and a target test set according to a preset ratio; Constructing a loss function, and training the multi-scale residual channel attention network model according to the target training set and the loss function; The trained multi-scale residual channel attention network model is tested according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

4. A micro-expression analysis method according to claim 1, characterized in that: The process of the multi-scale convolution layer performing multi-scale convolution on the horizontal optical flow data is as follows: Performing two-dimensional convolution processing on the horizontal optical flow data to obtain a plurality of convolution feature maps of different scales, and performing activation processing on the plurality of convolution feature maps of different scales to obtain a plurality of activation feature maps; Performing batch normalization processing on the multiple activation feature maps, and splicing the multiple activation feature maps after the batch normalization processing along the channel dimension to obtain the horizontal optical flow feature map; According to the process of performing multi-scale convolution processing on the horizontal optical flow data by the multi-scale convolution layer, multi-scale convolution processing is performed on the vertical optical flow data and the optical strain data to obtain the vertical optical flow feature map and the optical strain feature map.

5. A micro-expression analysis method according to claim 1, characterized in that: The residual connection channel attention module performs attention weighted processing on the horizontal optical flow feature map; specifically: Performing average pooling and maximum pooling processing on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a maximum pooling feature map; Performing channel compression and channel recovery processing on the average pooling feature map and the maximum pooling feature map to obtain an average pooling weighted feature map and a maximum pooling weighted feature map; Performing residual connection on the average pooling weighted feature map and the maximum pooling weighted feature map to obtain a weighted splicing feature map, performing 1×1 convolution processing on the weighted splicing feature map to obtain a convolution weighted feature map, and normalizing the convolution weighted feature map to obtain a weighted matrix; Multiplying the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and performing a residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain the horizontal weighted feature map; According to the process of performing attention weighted processing on the horizontal optical flow feature map by the residual connection channel attention module, attention weighted processing is performed on the vertical optical flow feature map and the optical strain feature map to obtain the vertical weighted feature map and the optical strain weighted feature map.

6. A micro-expression analysis method according to claim 1, characterized in that: The positioning data and the recognition data are processed to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval; specifically: If the image stream to be analyzed only includes a micro-expression interval, the maximum peak position in the positioning data is obtained, the micro-expression interval of the image stream to be analyzed is identified according to the maximum peak position and the interval positions before and after the maximum peak position, and the micro-expression category corresponding to the micro-expression interval is obtained according to the identification data; If the image stream to be analyzed contains multiple micro-expression intervals, the positioning data is smoothed to obtain smoothed positioning data, the average value and the maximum value of the smoothed positioning data are calculated, and the interval threshold is calculated based on the average value and the maximum value. The multiple micro-expression intervals in the image stream to be analyzed are obtained according to the interval threshold, and each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals is obtained according to the recognition data.

7. A micro-expression analysis system, characterized in that: The system comprises: Acquisition module: acquiring a micro-expression data set, preprocessing the micro-expression data set, and obtaining a target data set; Construction module: construct a multi-scale residual channel attention network model, the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a positioning sub-network and an identification sub-network, wherein the multi-scale shared sub-network includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers and a splicing layer; the process of the multi-scale residual channel attention network model processing the image stream to be analyzed is: According to the three multi-scale convolution layers, multi-scale convolution processing is performed on the horizontal optical flow data, the vertical optical flow data and the optical strain data of the image stream to be analyzed, respectively, to obtain a horizontal optical flow feature map, a vertical optical flow feature map and an optical strain feature map; According to the three residual connection channel attention modules, the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map are respectively subjected to attention weighting processing to obtain a horizontal weighted feature map, a vertical weighted feature map and an optical strain weighted feature map; According to the three convolutional layers, the horizontal weighted feature map, the vertical weighted feature map and the optical strain weighted feature map are respectively convolved to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map and an optical strain weighted convolutional feature map; The horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map are sequentially subjected to normalization processing, maximum pooling processing and random path discarding processing, and the horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map after the random path discarding processing are spliced ​​according to the splicing layer to obtain a spliced ​​feature map; The positioning sub-network performs positioning processing on the splicing feature map to obtain positioning data of the image stream to be analyzed, and the recognition sub-network processes the positioning data and the splicing feature map to obtain recognition data; Training and testing module: training and testing the multi-scale residual channel attention network model according to the target data set; Analysis module: input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the positioning data and identification data of the image stream to be analyzed; wherein the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the identification data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed; Processing module: Processing the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

8. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of a micro-expression analysis method as described in any one of claims 1 to 6.

9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a micro-expression analysis method as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Micro-expression recognition method based on space-time appearance movement attention network

    CN112307958A

  • Face micro-expression recognition method in video image sequence

    CN113496217A

Cited By

  • Personnel facial expression intelligent analysis method and system based on convolutional neural network

    CN121392927A