Micro-expression analysis method and system

Through the multi-scale residual channel attention network model, the problem of insufficient feature extraction in micro-expression analysis is solved, the accuracy of positioning and recognition is improved, and the robustness and generalization performance of the model are enhanced.

CN119942620AActive Publication Date: 2025-05-06NANCHANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510424098.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing micro-expression analysis methods lack the ability to extract features, resulting in low accuracy in positioning and recognition, and it is difficult for traditional attention mechanisms to capture the key features of micro-expression.

Method used

A multi-scale residual channel attention network model is adopted, including a multi-scale shared subnet, a positioning subnet and an identification subnet. The channel attention module is connected through a multi-scale convolutional layer and residual, and feature extraction and key information capture capabilities are enhanced.

Benefits of technology

提高了微表情分析的定位与识别准确率,增强了模型的鲁棒性和泛化性能,能够有效捕捉图像中的关键特征。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942620A_ABST
    Figure CN119942620A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-expression analysis method and system, and belongs to the technical field of computer vision. The method comprises the following steps: acquiring and preprocessing a micro-expression data set to obtain a target data set; constructing a multi-scale residual channel attention network model, and training and testing the multi-scale residual channel attention network model according to the target data set; inputting a to-be-analyzed image stream into the tested multi-scale residual channel attention network model for analysis, and outputting positioning data and identification data of the to-be-analyzed image stream; and processing the positioning data and the identification data to obtain a micro-expression interval of the to-be-analyzed image stream and a micro-expression category of the micro-expression interval. The design of comprehensively using a multi-scale shared sub-network and a residual connection channel attention module is adopted, and a multi-scale residual channel attention network model can provide higher accuracy in the aspect of micro-expression positioning and recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer vision technology, and specifically relates to a micro-expression analysis method and system. Background Art

[0002] Micro-expressions are a type of facial expression that lasts for a very short time (less than 0.5 seconds) and is difficult to detect. Micro-expressions are usually unconscious and spontaneous, often appearing before people notice them. The facial muscles are very small and difficult to detect with the naked eye. They are the natural expression of people's inner feelings, usually reflecting people's true feelings and emotions. Therefore, it is difficult to avoid micro-expressions by controlling or concealing them. Since micro-expressions are not deceptive and can reflect people's true emotions, micro-expression recognition has great practical significance in real life.

[0003] Existing micro-expression analysis methods generally use the MEAN framework and a simple three-layer convolutional network, but their feature extraction capabilities are insufficient. In addition, although attention mechanisms such as CBAM and ViT perform well in other visual tasks, when processing micro-expressions, due to their short-lived and subtle changes, traditional attention mechanisms have difficulty adjusting their focus in time to capture these key features, resulting in low positioning and recognition accuracy. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a micro-expression analysis method and system, which can solve the technical problem that the prior art is insufficient in extracting key features of micro-expressions, resulting in low positioning and recognition accuracy.

[0005] In order to solve the above technical problems, this application is implemented as follows: In a first aspect, an embodiment of the present application provides a micro-expression analysis method, the method comprising: Acquire a micro-expression dataset, and pre-process the micro-expression dataset to obtain a target dataset; Constructing a multi-scale residual channel attention network model, wherein the multi-scale residual channel attention network model includes a multi-scale sharing sub-network, a positioning sub-network and a recognition sub-network; Training and testing the multi-scale residual channel attention network model according to the target data set; Inputting the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputting the positioning data and recognition data of the image stream to be analyzed; Processing the positioning data and the recognition data to obtain a micro-expression interval of the image stream to be analyzed and a micro-expression category of the micro-expression interval; The positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0006] As an optional implementation of the first aspect of the present application, the micro-expression dataset is obtained, and the micro-expression dataset is preprocessed to obtain a target dataset; specifically: Acquire multiple segments of initial image streams containing micro-expressions, perform face cropping processing on each of the initial image streams, and obtain each face image stream corresponding to each of the initial image streams; Performing optical flow extraction, ROI selection and resampling processing on each of the face image streams to obtain each target image stream corresponding to each of the face image streams; The target data set is constructed according to each of the target image streams, wherein each of the target image streams includes three types of data: horizontal optical flow, vertical optical flow and optical strain.

[0007] As an optional implementation of the first aspect of the present application, the multi-scale residual channel attention network model is trained and tested according to the target data set; specifically: Dividing the target data set into a target training set and a target test set according to a preset ratio; Constructing a loss function, and training the multi-scale residual channel attention network model according to the target training set and the loss function; The trained multi-scale residual channel attention network model is tested according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

[0008] As an optional implementation of the first aspect of the present application, the multi-scale shared subnetwork comprises: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers and a splicing layer; the process of the multi-scale residual channel attention network model processing the image stream to be analyzed is: According to the three multi-scale convolution layers, multi-scale convolution processing is performed on the horizontal optical flow data, the vertical optical flow data and the optical strain data of the image stream to be analyzed, respectively, to obtain a horizontal optical flow feature map, a vertical optical flow feature map and an optical strain feature map; According to the three residual connection channel attention modules, the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map are respectively subjected to attention weighting processing to obtain a horizontal weighted feature map, a vertical weighted feature map and an optical strain weighted feature map; According to the three convolutional layers, the horizontal weighted feature map, the vertical weighted feature map and the optical strain weighted feature map are respectively convolved to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map and an optical strain weighted convolutional feature map; The horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map are sequentially subjected to normalization processing, maximum pooling processing and random path discarding processing, and the horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map after the random path discarding processing are spliced ​​according to the splicing layer to obtain a spliced ​​feature map; The positioning sub-network performs positioning processing on the splicing feature map to obtain positioning data of the image stream to be analyzed, and the recognition sub-network processes the positioning data and the splicing feature map to obtain recognition data.

[0009] As an optional implementation manner of the first aspect of the present application, the process of the multi-scale convolution layer performing multi-scale convolution on the horizontal optical flow data is as follows: Performing two-dimensional convolution processing on the horizontal optical flow data to obtain a plurality of convolution feature maps of different scales, and performing activation processing on the plurality of convolution feature maps of different scales to obtain a plurality of activation feature maps; Performing batch normalization processing on the multiple activation feature maps, and splicing the multiple activation feature maps after the batch normalization processing along the channel dimension to obtain the horizontal optical flow feature map; According to the process of performing multi-scale convolution processing on the horizontal optical flow data by the multi-scale convolution layer, multi-scale convolution processing is performed on the vertical optical flow data and the optical strain data to obtain the vertical optical flow feature map and the optical strain feature map.

[0010] As an optional implementation of the first aspect of the present application, the residual connection channel attention module performs attention weighted processing on the horizontal optical flow feature map; specifically: Performing average pooling and maximum pooling processing on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a maximum pooling feature map; Performing channel compression and channel recovery processing on the average pooling feature map and the maximum pooling feature map to obtain an average pooling weighted feature map and a maximum pooling weighted feature map; Performing residual connection on the average pooling weighted feature map and the maximum pooling weighted feature map to obtain a weighted splicing feature map, performing 1×1 convolution processing on the weighted splicing feature map to obtain a convolution weighted feature map, and normalizing the convolution weighted feature map to obtain a weighted matrix; Multiplying the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and performing a residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain the horizontal weighted feature map; According to the process of performing attention weighted processing on the horizontal optical flow feature map by the residual connection channel attention module, attention weighted processing is performed on the vertical optical flow feature map and the optical strain feature map to obtain the vertical weighted feature map and the optical strain weighted feature map.

[0011] As an optional implementation of the first aspect of the present application, the positioning data and the recognition data are processed to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval; specifically: If the image stream to be analyzed only includes a micro-expression interval, the maximum peak position in the positioning data is obtained, the micro-expression interval of the image stream to be analyzed is identified according to the maximum peak position and the interval positions before and after the maximum peak position, and the micro-expression category corresponding to the micro-expression interval is obtained according to the identification data; If the image stream to be analyzed contains multiple micro-expression intervals, the positioning data is smoothed to obtain smoothed positioning data, the average value and the maximum value of the smoothed positioning data are calculated, and the interval threshold is calculated based on the average value and the maximum value. The multiple micro-expression intervals in the image stream to be analyzed are obtained according to the interval threshold, and each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals is obtained according to the recognition data.

[0012] In a second aspect, an embodiment of the present application provides a micro-expression analysis system, the system comprising: Acquisition module: acquiring a micro-expression data set, preprocessing the micro-expression data set, and obtaining a target data set; Construction module: construct a multi-scale residual channel attention network model, wherein the multi-scale residual channel attention network model includes a multi-scale sharing sub-network, a positioning sub-network and a recognition sub-network; Training and testing module: training and testing the multi-scale residual channel attention network model according to the target data set; Analysis module: input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the positioning data and identification data of the image stream to be analyzed; wherein the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the identification data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed; Processing module: Processing the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In the embodiments of the present application, compared with the prior art, the following beneficial effects are achieved: (1) Improving feature extraction capabilities: By integrating the channel attention mechanism, convolutional layer processing, and residual connection strategy into the residual connection channel attention module, the multi-scale residual channel attention network model can more effectively extract the key features of micro-expressions; this design not only enhances the performance of the model in complex emotion recognition tasks, but also improves its robustness and generalization performance.

[0016] (2) Enhanced ability to capture features of different scales: The multi-scale convolutional layer introduced in the multi-scale shared subnetwork enables the network to capture important information of different scales in the image. This is particularly important for micro-expression analysis because changes in micro-expressions often involve details at multiple scales, which helps to improve the accuracy of positioning and recognition.

[0017] (3) Improve attention to important information: The feature maps processed by the residual connection channel attention module enable the model to focus more on the key areas in the image instead of being distracted by irrelevant background information. This approach ensures that even if the micro-expression changes are brief and subtle, the model can adjust its focus in time, thereby effectively capturing the information at these critical moments.

[0018] (4) Improved positioning and recognition accuracy: By combining the design of a multi-scale shared sub-network and a residual connection channel attention module, the multi-scale residual channel attention network model can provide higher accuracy in the positioning and recognition of micro-expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of a micro-expression analysis method provided by some embodiments of the present application; Figure 2 It is a structural diagram of a multi-scale residual channel attention network model of a micro-expression analysis method provided by some embodiments of the present application; Figure 3 This is a structural diagram of a residual connection channel attention module of a micro-expression analysis method provided by some embodiments of the present application. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.

[0022] In the following, in conjunction with the accompanying drawings, a micro-expression analysis method and system provided by an embodiment of the present application are described in detail through specific embodiments and their application scenarios.

[0023] Example See also Figure 1 As shown, a micro-expression analysis method comprises the following steps: S100: Acquire a micro-expression dataset, pre-process the micro-expression dataset, and obtain a target dataset; It should be noted that S100 is specifically: S110: acquiring multiple segments of initial image streams containing micro-expressions, performing face cropping processing on each initial image stream, and obtaining each face image stream corresponding to each initial image stream; S120: performing optical flow extraction, ROI selection and resampling processing on each face image stream to obtain each target image stream corresponding to each face image stream; S130: constructing a target data set according to each target image stream, wherein each target image stream includes three types of data: horizontal optical flow, vertical optical flow and optical strain.

[0024] Furthermore, after face detection, the first frame of the initial image stream is cropped to 128×128, and then 68 key points of the face are marked to assist in ROI selection and resampling. The TV-L1 algorithm is used to extract optical flow, which is robust to noise and can accurately capture the dynamic changes of micro-expressions to obtain horizontal optical flow (u) and vertical optical flow (v). At the same time, the optical strain is calculated The facial deformation intensity is approximated by the Hessian matrix and the optical strain amplitude is calculated. As one of the input features. Combined with optical flow and optical strain, a triplet is formed As the input of the micro-expression analysis network. In addition, three regions rich in micro-expression information, the left eye and left eyebrow, the right eye and right eyebrow, and the mouth, are selected for ROI selection and image resampling, and the region sizes are adjusted to 21×21 and 21×42, and finally combined into a 42×42 target image stream.

[0025] S200: Construct a multi-scale residual channel attention network model, which includes a multi-scale shared sub-network, a positioning sub-network and a recognition sub-network.

[0026] See also Figure 2 As shown, Figure 2 This is a structural diagram of the multi-scale residual channel attention network model. Furthermore, the multi-scale shared sub-network includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers and one splicing layer.

[0027] S300: Train and test the multi-scale residual channel attention network model according to the target dataset; It should be noted that S300 is specifically: S310: Dividing the target data set into a target training set and a target test set according to a preset ratio; S320: construct a loss function, and train the multi-scale residual channel attention network model according to the target training set and the loss function; S330: Testing the trained multi-scale residual channel attention network model according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

[0028] Furthermore, this embodiment divides the target data set into a target training set and a target test set in a ratio of 8:2; constructs a loss function, and trains the multi-scale residual channel attention network model according to the target training set and the loss function; the loss function is used to perform parameter tuning on the training process of the multi-scale residual channel attention network model; the trained multi-scale residual channel attention network model is tested according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

[0029] Furthermore, in the present embodiment, the target data set may be a first target data set or a second target data set, and each target image stream in the first target data set contains only one micro-expression interval; and each target image stream in the second target data set contains multiple micro-expression intervals; but whether it is the first target data set or the second target data set, the target training set and the target test set are divided into a ratio of 8:2, and the training and testing processes of the first target data set and the second target data set are independent, which means that the multi-scale residual channel attention network model can learn a target image stream containing only one micro-expression interval, or a target image stream containing multiple micro-expression intervals, so that the multi-scale residual channel attention network model can identify the micro-expression interval and the category of the micro-expression interval in the target image stream containing only one micro-expression interval, and at the same time, it can also identify multiple micro-expression intervals in the target image stream containing multiple micro-expression intervals, and the micro-expression category corresponding to each of the multiple micro-expression intervals.

[0030] Furthermore, the present embodiment also uses a pseudo-labeling technology to perform a pseudo-labeling operation on each target image stream in the target data set to obtain each pseudo-label set corresponding to each target image stream. Before training the target data set, the target image stream is pseudo-labeled to improve the generalization ability and robustness of the model. When the target training set is put into the multi-scale residual channel attention network model training, each pseudo-label set corresponding to each target image stream in the target training set is also put into the multi-scale residual channel attention network model training, and the pseudo-label set is only used for the training of the positioning subnetwork, that is, after the multi-scale shared subnetwork in the multi-scale residual channel attention network model trains a target image stream in the target data set, it outputs the output feature corresponding to the target image stream. When the output feature enters the positioning subnetwork for training, the pseudo-label set corresponding to the target image stream enters the positioning subnetwork training together with the output feature.

[0031] S400: inputting the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputting positioning data and recognition data of the image stream to be analyzed; It should be noted that S400 is specifically: S410: performing multi-scale convolution processing on the horizontal optical flow data, the vertical optical flow data, and the optical strain data of the image stream to be analyzed according to the three multi-scale convolution layers, respectively, to obtain a horizontal optical flow feature map, a vertical optical flow feature map, and an optical strain feature map; S420: performing attention weighting processing on the horizontal optical flow feature map, the vertical optical flow feature map, and the optical strain feature map according to the three residual connection channel attention modules, respectively, to obtain a horizontal weighted feature map, a vertical weighted feature map, and an optical strain weighted feature map; S430: performing convolution processing on the horizontal weighted feature map, the vertical weighted feature map, and the optical strain weighted feature map according to three convolution layers, respectively, to obtain a horizontal weighted convolution feature map, a vertical weighted convolution feature map, and an optical strain weighted convolution feature map; S440: performing normalization processing, maximum pooling processing, and random path discarding processing on the horizontal weighted convolution feature map, the vertical weighted convolution feature map, and the optical strain weighted convolution feature map in sequence, and splicing the horizontal weighted convolution feature map, the vertical weighted convolution feature map, and the optical strain weighted convolution feature map after the random path discarding processing according to the splicing layer to obtain a spliced ​​feature map; S450: The positioning sub-network performs positioning processing on the spliced ​​feature map to obtain positioning data of the image stream to be analyzed, and the recognition sub-network processes the positioning data and the spliced ​​feature map to obtain recognition data.

[0032] Furthermore, according to the three multi-scale convolutional layers in the multi-scale shared network (MSSN), the horizontal optical flow data, vertical optical flow data and optical strain data of the image stream to be analyzed are subjected to multi-scale convolution processing to obtain the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map; the network model's ability to capture features of different scales in the image can be enhanced; the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map output by the three multi-scale convolutional layers are processed by three residual connection channel attention modules (RCCAM) to obtain the horizontal and vertical optical flow feature maps and the optical strain feature maps. The weighted feature map, vertical weighted feature map and optical strain weighted feature map are convolved again to obtain horizontal weighted convolution feature map, vertical weighted convolution feature map and optical strain weighted convolution feature map, so as to further enhance the feature representation capability; the horizontal weighted convolution feature map, vertical weighted convolution feature map and optical strain weighted convolution feature map are batch normalized. And maximum pooling (MaxPooling2D) operations are performed to further improve the discrimination and robustness of the features, and a random path dropout layer is used to reduce the overfitting risk of the model; finally, the horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution after the random path dropout processing are spliced ​​according to the splicing layer to obtain a spliced ​​feature map; finally, the spliced ​​feature map is positioned by the positioning subnetwork to obtain the positioning data of the image stream to be analyzed, and the recognition subnetwork processes the positioning data and the spliced ​​feature map to obtain the recognition data; wherein, the positioning data contains the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data contains the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0033] It should be noted that S410 specifically refers to: S411: performing two-dimensional convolution processing on the horizontal optical flow data to obtain a plurality of convolution feature maps of different scales, and performing activation processing on the plurality of convolution feature maps of different scales to obtain a plurality of activation feature maps; S412: performing batch normalization processing on the multiple activation feature maps, and splicing the multiple activation feature maps after the batch normalization processing along the channel dimension to obtain a horizontal optical flow feature map; S413: According to the processing of the horizontal optical flow data in S411-S412, multi-scale convolution processing is performed on the vertical optical flow data and the optical strain data to obtain a vertical optical flow feature map and an optical strain feature map.

[0034] Furthermore, in order to enhance the network's ability to capture features of different scales in the image, the multi-scale shared network (MSSN) of this embodiment adds a multi-scale convolution layer to the MSSN; the design of this layer is to extract multi-level feature representations of the input data through convolution kernels of different sizes. Specifically, given the input tensor inputs (horizontal optical flow data, vertical optical flow data, and optical strain data), the number of channels is filters, and a list kernel_sizes containing convolution kernels of various sizes is defined. The inputs of horizontal optical flow data and vertical optical flow data are set to [1, 3], and the input of optical strain data is set to [3,5]. For each convolution kernel size in the list, the following operations are performed: a two-dimensional convolution operation (Conv2D) is applied, where the size of the convolution kernel is equal to the currently selected kernel_size, and the padding of the 'same' type is used to maintain the consistency of the input and output sizes. In order to prevent overfitting and introduce regularization effects, we add an L2 regularization term to the convolution process. The activation function is selected as RelU to increase the nonlinear expression ability of the model. To further improve the generalization and stability of the model, a batch normalization operation is added after each convolutional layer. This helps speed up the training process and reduce problems caused by internal covariate shift. In order to integrate feature information at different scales, all feature maps processed above are spliced ​​together along the channel dimension to form a multi-scale feature fusion output (horizontal optical flow feature map, vertical optical flow feature map, and optical strain feature map). This structure allows the network to obtain local details and broader contextual information from the same input at the same time, thereby improving the expressiveness of the final model.

[0035] Specifically, the multi-scale convolutional layer is expressed by the following formula: , in, represents the input of the multi-scale convolutional layer, represents two-dimensional convolution processing, represents the activation function, represents batch normalization, represents feature concatenation, Represents the output of a multi-scale convolutional layer.

[0036] See also Figure 3 As shown, Figure 3 This is a structural diagram of the residual connection channel attention module. It should be noted that S420 is specifically: S421: performing average pooling and maximum pooling processing on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a maximum pooling feature map; S422: performing channel compression and channel recovery processing on both the average pooling feature map and the maximum pooling feature map to obtain an average pooling weighted feature map and a maximum pooling weighted feature map; S423: performing residual connection on the average pooling weighted feature map and the maximum pooling weighted feature map to obtain a weighted splicing feature map, performing 1×1 convolution processing on the weighted splicing feature map to obtain a convolution weighted feature map, and normalizing the convolution weighted feature map to obtain a weighted matrix; S424: multiplying the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and performing a residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain a horizontal weighted feature map; S425: According to the process of performing attention weighted processing on the horizontal optical flow feature map in S421-S424, attention weighted processing is performed on the vertical optical flow feature map and the optical strain feature map to obtain a vertical weighted feature map and an optical strain weighted feature map.

[0037] Furthermore, the input feature maps (horizontal optical flow feature map, vertical optical flow feature map and optical strain feature map) are subjected to global maximum pooling and global average pooling, respectively, to obtain average pooling feature maps and maximum pooling feature maps, both of which have the shape of [C], and the average pooling feature maps and maximum pooling feature maps are reshaped into the shape of [1,1,C] for the convenience of subsequent operations; the reshaped average pooling feature maps and maximum pooling feature maps are input into a network containing two fully connected layers RD (each fully connected layer is equipped with a ReLU activation function) to obtain average pooling weighted feature maps and maximum pooling weighted feature maps. The first fully connected layer reduces the number of channels to r times the original number of channels (r is a hyperparameter used to control the ratio of channel compression) to learn the complex relationship between different channels. The second fully connected layer restores the number of channels to the original number. This learning process can reveal which channels are more critical for recognizing micro-expressions, and then generate a weight vector that matches the number of input channels. In this way, it is possible to achieve fine weighting of each channel of the input feature map, thereby enhancing the model's attention to important features. The average pooled weighted feature map and the maximum pooled weighted feature map are residually connected to obtain a weighted splicing feature; then, the weighted splicing feature is sent to a carefully designed 1x1 convolution layer for further processing to obtain a convolution weighted feature map. The main purpose of this convolution layer is to further extract and refine features in order to better capture subtle changes in micro-expressions. The configuration of the convolution layer (i.e., the number of filters is set to 1, the kernel size is 1x1, the stride is 1, and the padding method is 'same') is to keep the size of the feature map unchanged while learning cross-channel feature combinations. The weights of the convolution weighted feature map are then normalized using the Sigmoid function to obtain a weight matrix in the range of [0,1], and the input feature map is multiplied by the weight matrix to obtain the channel weighted feature map corresponding to the input feature map (horizontal channel weighted feature map, vertical channel weighted feature map, optical strain channel weighted feature map). Through this step, the model can learn richer feature representations, thereby improving its accuracy in identifying micro-expressions. In order to further improve the training efficiency and generalization ability of the model, a residual connection is established between the channel weighted feature map and the input feature map to obtain the final output feature map (horizontal weighted feature map, vertical weighted feature map, and optical strain weighted feature map). The core idea of ​​this strategy is that by directly adding the output feature map of the convolutional layer to the input feature map, the key information in the input features can be retained and the gradient vanishing problem that may occur during deep network training can be alleviated.

[0038] Specifically, the data processing process of the residual connection channel attention module (RCCAM) is expressed by the following formula: , , , , , in, represents the input feature map, represents the two-dimensional maximum pooling process, represents two-dimensional average pooling processing, Reshape the feature into , represents the maximum pooling feature map, represents the average pooling feature map; Indicates converting the number of channels to , Represents a constant greater than 0 and less than 1. Indicates converting the number of channels into , represents the maximum pooling weighted feature map, represents the average pooled weighted feature map, represents the residual connection, represents the weighted concatenated feature map, Express Two-dimensional convolution and activation processing are performed in sequence. The convolution kernel size of the two-dimensional convolution is , the step length is , represents the normalization function, represents element-wise multiplication, represents the channel weighted feature map, Represents the output feature map.

[0039] S500: Processing the positioning data and the recognition data to obtain the micro-expression intervals of the image stream to be analyzed and the micro-expression categories of the micro-expression intervals.

[0040] It should be noted that the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

[0041] It should be noted that in S500, if the image stream to be analyzed contains only one micro-expression interval, the maximum peak position in the positioning data is obtained, and the micro-expression interval of the image stream to be analyzed is identified according to the maximum peak position and the interval positions before and after the maximum peak position, and the micro-expression category corresponding to the micro-expression interval is obtained according to the identification data; If the image stream to be analyzed contains multiple micro-expression intervals, the positioning data is smoothed to obtain smoothed positioning data, and the average value and maximum value of the smoothed positioning data are calculated to calculate the interval threshold based on the average value and the maximum value. According to the interval threshold, multiple micro-expression intervals in the image stream to be analyzed are obtained, and each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals is obtained according to the recognition data.

[0042] Furthermore, different micro-expression positioning strategies are adopted for image streams to be analyzed of different lengths. For short videos (when the image stream contains only one micro-expression interval), this embodiment adopts a simple peak finding technology to accurately identify the micro-expression interval of the short video by locating and obtaining the maximum local peak and its preceding and following segments, and obtain the micro-expression category corresponding to the micro-expression interval of the short video according to the recognition data. For long videos (when the image stream contains multiple micro-expression intervals), a threshold-based technology is introduced to determine the micro-expression interval. Specifically, the positioning data corresponding to the long video is first obtained by simple smoothing to avoid erroneously identifying sudden spikes from the positioning data. Subsequently, the average and maximum values ​​of the smoothed positioning data are calculated, and the interval threshold is calculated based on the average and maximum values ​​combined with preset parameters. The interval threshold is used to divide multiple micro-expression intervals from the smoothed positioning data, and each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals is obtained according to the recognition data. In terms of micro-expression recognition, the research is based on the experimental results of previous work, that is, when predicting the emotion category in the recognition task, selecting the segment from the micro-expression start frame to the vertex frame for emotion recognition can obtain the best effect.

[0043] According to a micro-expression analysis method of this embodiment, through the residual connection channel attention module integrating the channel attention mechanism, convolution layer processing and residual connection strategy, the multi-scale residual channel attention network model can more effectively extract the key features in the micro-expression; this design not only enhances the performance of the model in complex emotion recognition tasks, but also improves its robustness and generalization performance. The multi-scale convolution layer introduced in the multi-scale shared sub-network enables the network to capture important information of different scales in the image, which is particularly important for micro-expression analysis, because the changes in micro-expressions often involve details at multiple scales. This helps to improve the accuracy of positioning and recognition. The feature map processed by the residual connection channel attention module enables the model to focus more on the key areas in the image instead of being disturbed by irrelevant background information. This method ensures that even if the micro-expression changes are both short-lived and subtle, the model can adjust the focus in time, thereby effectively capturing the information of these key moments. By combining the design of the multi-scale shared sub-network and the residual connection channel attention module, the multi-scale residual channel attention network model can provide higher accuracy in the positioning and recognition of micro-expressions. The experimental results show that compared with the existing methods, the method proposed in this invention has significantly improved the overall performance, demonstrating its superiority in the field of micro-expression analysis.

[0044] It should be noted that the micro-expression analysis method provided in the embodiment of the present application can be executed by a micro-expression analysis system, or a control module in the micro-expression analysis system for executing and loading a micro-expression analysis method. In the embodiment of the present application, a micro-expression analysis system is used to execute and load a micro-expression analysis method as an example to illustrate that the embodiment of the present application provides a micro-expression analysis method.

[0045] A micro-expression analysis system, comprising: Acquisition module: acquires micro-expression dataset, pre-processes the micro-expression dataset, and obtains the target dataset; Construction module: Construct a multi-scale residual channel attention network model, which includes a multi-scale sharing sub-network, a positioning sub-network and a recognition sub-network; Training and testing module: train and test the multi-scale residual channel attention network model according to the target dataset; Analysis module: input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the positioning data and recognition data of the image stream to be analyzed; wherein the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed; Processing module: processes the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

[0046] A micro-expression analysis system in the embodiment of the present application may be a device, or a component, integrated circuit, or chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.

[0047] A micro-expression analysis system in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0048] According to a micro-expression analysis system of the present embodiment, the micro-expression data set required for training and testing can be obtained through the acquisition module, and the micro-expression data set is preprocessed to obtain the target data set; and the construction module is used to construct a multi-scale residual channel attention network model, wherein the multi-scale residual channel attention network model includes a multi-scale shared sub-network, a positioning sub-network and an identification sub-network; then the multi-scale residual channel attention network model is trained and tested by the target data set processed by the acquisition module; then the data to be analyzed is input into the analysis module, and the analysis module inputs the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputs the positioning data and identification data of the image stream to be analyzed; finally, the positioning data and the identification data output by the analysis module are processed according to the processing module to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval; through the mutual cooperation between the various modules, efficient positioning of the micro-expression interval in the image stream to be analyzed and accurate identification of the micro-expression category in the micro-expression interval are achieved.

[0049] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned micro-expression analysis method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0050] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned micro-expression analysis method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0051] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0052] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0053] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0054] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A micro-expression analysis method, characterized in that: The method comprises: Acquire a micro-expression dataset, and pre-process the micro-expression dataset to obtain a target dataset; Constructing a multi-scale residual channel attention network model, wherein the multi-scale residual channel attention network model includes a multi-scale sharing sub-network, a positioning sub-network and a recognition sub-network; Training and testing the multi-scale residual channel attention network model according to the target data set; Inputting the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and outputting the positioning data and recognition data of the image stream to be analyzed; Processing the positioning data and the recognition data to obtain a micro-expression interval of the image stream to be analyzed and a micro-expression category of the micro-expression interval; The positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the recognition data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed.

2. A micro-expression analysis method according to claim 1, characterized in that: The micro-expression dataset is obtained, and the micro-expression dataset is preprocessed to obtain a target dataset; specifically: Acquire multiple segments of initial image streams containing micro-expressions, perform face cropping processing on each of the initial image streams, and obtain each face image stream corresponding to each of the initial image streams; Performing optical flow extraction, ROI selection and resampling processing on each of the face image streams to obtain each target image stream corresponding to each of the face image streams; The target data set is constructed according to each of the target image streams, wherein each of the target image streams includes three types of data: horizontal optical flow, vertical optical flow and optical strain.

3. A micro-expression analysis method according to claim 1, characterized in that: The multi-scale residual channel attention network model is trained and tested according to the target data set; specifically: Dividing the target data set into a target training set and a target test set according to a preset ratio; Constructing a loss function, and training the multi-scale residual channel attention network model according to the target training set and the loss function; The trained multi-scale residual channel attention network model is tested according to the target test set to complete the training and testing of the multi-scale residual channel attention network model.

4. A micro-expression analysis method according to claim 1, characterized in that: The multi-scale shared sub-network includes: three multi-scale convolutional layers, three residual connection channel attention modules, three convolutional layers and a splicing layer; the process of the multi-scale residual channel attention network model processing the image stream to be analyzed is: According to the three multi-scale convolution layers, multi-scale convolution processing is performed on the horizontal optical flow data, the vertical optical flow data and the optical strain data of the image stream to be analyzed, respectively, to obtain a horizontal optical flow feature map, a vertical optical flow feature map and an optical strain feature map; According to the three residual connection channel attention modules, the horizontal optical flow feature map, the vertical optical flow feature map and the optical strain feature map are respectively subjected to attention weighting processing to obtain a horizontal weighted feature map, a vertical weighted feature map and an optical strain weighted feature map; According to the three convolutional layers, the horizontal weighted feature map, the vertical weighted feature map and the optical strain weighted feature map are respectively convolved to obtain a horizontal weighted convolutional feature map, a vertical weighted convolutional feature map and an optical strain weighted convolutional feature map; The horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map are sequentially subjected to normalization processing, maximum pooling processing and random path discarding processing, and the horizontal weighted convolution feature map, the vertical weighted convolution feature map and the optical strain weighted convolution feature map after the random path discarding processing are spliced ​​according to the splicing layer to obtain a spliced ​​feature map; The positioning sub-network performs positioning processing on the splicing feature map to obtain positioning data of the image stream to be analyzed, and the recognition sub-network processes the positioning data and the splicing feature map to obtain recognition data.

5. A micro-expression analysis method according to claim 4, characterized in that: The process of the multi-scale convolution layer performing multi-scale convolution on the horizontal optical flow data is as follows: Performing two-dimensional convolution processing on the horizontal optical flow data to obtain a plurality of convolution feature maps of different scales, and performing activation processing on the plurality of convolution feature maps of different scales to obtain a plurality of activation feature maps; Performing batch normalization processing on the multiple activation feature maps, and splicing the multiple activation feature maps after the batch normalization processing along the channel dimension to obtain the horizontal optical flow feature map; According to the process of performing multi-scale convolution processing on the horizontal optical flow data by the multi-scale convolution layer, multi-scale convolution processing is performed on the vertical optical flow data and the optical strain data to obtain the vertical optical flow feature map and the optical strain feature map.

6. A micro-expression analysis method according to claim 4, characterized in that: The residual connection channel attention module performs attention weighted processing on the horizontal optical flow feature map; specifically: Performing average pooling and maximum pooling processing on the horizontal optical flow feature map respectively to obtain an average pooling feature map and a maximum pooling feature map; Performing channel compression and channel recovery processing on the average pooling feature map and the maximum pooling feature map to obtain an average pooling weighted feature map and a maximum pooling weighted feature map; Performing residual connection on the average pooling weighted feature map and the maximum pooling weighted feature map to obtain a weighted splicing feature map, performing 1×1 convolution processing on the weighted splicing feature map to obtain a convolution weighted feature map, and normalizing the convolution weighted feature map to obtain a weighted matrix; Multiplying the weighted matrix by the horizontal optical flow feature map to obtain a horizontal channel weighted feature map, and performing a residual connection on the horizontal channel weighted feature map and the horizontal optical flow feature map to obtain the horizontal weighted feature map; According to the process of performing attention weighted processing on the horizontal optical flow feature map by the residual connection channel attention module, attention weighted processing is performed on the vertical optical flow feature map and the optical strain feature map to obtain the vertical weighted feature map and the optical strain weighted feature map.

7. A micro-expression analysis method according to claim 1, characterized in that: The positioning data and the recognition data are processed to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval; specifically: If the image stream to be analyzed only includes a micro-expression interval, the maximum peak position in the positioning data is obtained, the micro-expression interval of the image stream to be analyzed is identified according to the maximum peak position and the interval positions before and after the maximum peak position, and the micro-expression category corresponding to the micro-expression interval is obtained according to the identification data; If the image stream to be analyzed contains multiple micro-expression intervals, the positioning data is smoothed to obtain smoothed positioning data, the average value and the maximum value of the smoothed positioning data are calculated, and the interval threshold is calculated based on the average value and the maximum value. The multiple micro-expression intervals in the image stream to be analyzed are obtained according to the interval threshold, and each micro-expression category corresponding to each micro-expression interval in the multiple micro-expression intervals is obtained according to the recognition data.

8. A micro-expression analysis system, characterized in that: The system comprises: Acquisition module: acquiring a micro-expression data set, preprocessing the micro-expression data set, and obtaining a target data set; Construction module: construct a multi-scale residual channel attention network model, wherein the multi-scale residual channel attention network model includes a multi-scale sharing sub-network, a positioning sub-network and a recognition sub-network; Training and testing module: training and testing the multi-scale residual channel attention network model according to the target data set; Analysis module: input the image stream to be analyzed into the tested multi-scale residual channel attention network model for analysis, and output the positioning data and identification data of the image stream to be analyzed; wherein the positioning data includes the position information of the micro-expression interval in the image stream to be analyzed, and the identification data includes the micro-expression category information of the micro-expression interval in the image stream to be analyzed; Processing module: Processing the positioning data and the recognition data to obtain the micro-expression interval of the image stream to be analyzed and the micro-expression category of the micro-expression interval.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of a micro-expression analysis method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of a micro-expression analysis method as described in any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Micro-expression recognition method based on space-time appearance movement attention network

    CN112307958A

  • Face micro-expression recognition method in video image sequence

    CN113496217A

  • Micro-expression recognition method based on double attention mechanism

    CN114550270A

  • Micro-expression recognition method and system based on Inception-CBAM +

    CN116052245A

  • Micro-expression recognition method and device based on multi-level graph convolutional network

    CN116311472A