A micro-expression recognition method and system based on multi-receptive field visual feature extraction

By adopting a multi-receptive field visual feature extraction network in micro-expression recognition, using asymmetric multiple scanning strategies and space-channel attention mechanisms, the problem of insufficient spatial dependency capture in micro-expression recognition is solved, and more efficient micro-expression feature extraction and recognition performance is achieved.

CN119863830BActive Publication Date: 2025-05-23UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336419.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-23
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the spatial dependence between different facial regions in micro-expression recognition, resulting in insufficient learning of micro-expression features, and the commonly used two-way or symmetric scanning strategies increase computational overhead and redundancy.

Method used

A method based on multi-receptive field visual feature extraction is adopted to extract micro-expression feature through a multi-receptive field visual feature extraction network (including the fast feature extraction stage and multiple local-global feature integration stages). The network utilizes asymmetric multiple scanning strategies and space-channel attention mechanisms to reduce redundancy and enhance spatial perception.

Benefits of technology

Effectively capture the fine-grained motion characteristics and global dependencies of micro-expressions, reduces computational overhead and redundancy, and improves the performance of micro-expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863830B_ABST
    Figure CN119863830B_ABST
Patent Text Reader

Abstract

The present invention discloses a micro-expression recognition method and system based on multi-receptive field visual feature extraction, the method comprising: pre-processing the original image frames of the micro-expression data set, including face detection, alignment and cropping, and calculating the TV-L1 optical flow features based on the starting frame and the peak frame to obtain input features; converting the input features into overlapping patches and inputting them into a multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network comprises: a plurality of continuous local-global feature integration stages, each stage comprising a combination layer of several local extractors and multi-layer perceptrons, followed by a combination layer of global self-attention and multi-layer perceptrons. By combining the local feature extractor with the global self-attention mechanism, the subtle facial features and spatial long-range dependencies of micro-expressions can be effectively captured. At the same time, the asymmetric multiple scanning strategy reduces redundancy while enhancing the spatial perception ability of the model, thereby improving the micro-expression recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of micro-expression recognition, and specifically to a micro-expression recognition method and system based on multi-receptive field visual feature extraction. Background Art

[0002] Micro-expressions are spontaneous facial movements that can reveal an individual's true emotions and intentions, and are of great value in psychological research. Unlike regular facial expressions, micro-expressions are short-lasting (usually less than 0.5 seconds), low-intensity, and occur in local facial areas. In recent years, micro-expression recognition has shown good application prospects in fields such as criminal investigation, interrogation, and psychological diagnosis.

[0003] With the rapid development of artificial intelligence, automatic micro-expression recognition has received increasing attention in the field of emotional computing. In order to extract effective features from micro-expressions, researchers have developed a variety of feature extractors based on convolutional neural networks or visual transformers. Although convolutional neural networks perform well in capturing local features, they have certain limitations in modeling global dependencies; while visual transformers can effectively process global information, they have high computational overhead and require a large amount of data, which often leads to serious overfitting problems in micro-expression recognition.

[0004] Recently, the Mamba architecture based on the state-space model has shown excellent performance in multiple computer vision tasks, and is particularly favored for its efficient linear time complexity, strong context-awareness, and simplified network structure. Although there have been some attempts to apply Mamba to micro-expression recognition-related tasks, the key requirement of capturing the spatial dependencies between different facial regions, which is essential for accurate micro-expression feature learning, has been generally overlooked. This is because micro-expressions are usually presented as action units in local facial regions, while the combination of emotion-related action units spans multiple local facial regions.

[0005] It is worth noting that a reasonable scanning strategy can help improve the performance of "Mamba" in visual tasks. Existing visual feature extractors based on "Mamba" usually adopt bidirectional or symmetrical scanning directions to better capture spatial adjacency. However, for micro-expressions, the fixed spatial structure of the face after preprocessing and the inherent left-right symmetry of facial expressions lead to the redundancy of bidirectional or symmetrical scanning, and also increase the computational overhead. Therefore, it is necessary to design a scanning strategy suitable for micro-expression recognition to improve spatial perception ability while reducing redundancy. Summary of the invention

[0006] In this embodiment, a micro-expression recognition method, system, electronic device and storage medium based on multi-receptive field visual feature extraction are provided to solve the problem of insufficient fine-grained motion feature learning and global dependency modeling of micro-expressions in related technologies.

[0007] In a first aspect, an embodiment of the present invention provides a micro-expression recognition method based on multi-receptive field visual feature extraction, and the micro-expression recognition method based on multi-receptive field visual feature extraction includes:

[0008] The original image frames of the micro-expression dataset are preprocessed, including face detection, alignment and cropping, and the TV-L1 optical flow features are calculated based on the starting frame and the peak frame to obtain the input features;

[0009] The input features are converted into overlapping patches and input into a multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network comprises:

[0010] The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling;

[0011] Multiple consecutive local-global feature integration stages, each stage includes several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; wherein the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, and after being processed by a mixer block based on the Mamba architecture, the localized feature representation is generated by fusion through a spatial-channel attention mechanism;

[0012] Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution;

[0013] The final output features are flattened and classified through a linear classifier to obtain the micro-expression recognition results.

[0014] In an optional embodiment, the multi-receptive field visual feature extraction network includes three local-global feature integration stages, the number of combined layers of the local extractor and the multi-layer perceptron in each stage is 2, 6, and 3 respectively, and the output features of the last stage are input into the linear classifier after two-dimensional batch normalization and average pooling.

[0015] In an optional embodiment, in the local-global feature integration stage, the window size of the local extractor is 7×7, the window size of the global self-attention layer is consistent with the spatial resolution of the input features, and in the fourth stage, the window size of the local extractor is the same as that of the global self-attention layer.

[0016] In an optional embodiment, the asymmetric multiple scanning strategy includes four directions: horizontal raster scanning, vertical raster scanning, horizontal zigzag scanning and vertical zigzag scanning, and the scanning directions are not bidirectional or symmetrical with each other.

[0017] In an optional embodiment, the spatial-channel attention mechanism adaptively assigns contribution weights of four scanning directions to perform weighted fusion on the reconstructed two-dimensional feature maps to generate a final feature representation of the local window.

[0018] In an optional embodiment, the TV-L1 optical flow feature has three dimensions, which are a horizontal component, a vertical component and a modulus.

[0019] In an optional embodiment, cross entropy classification loss is used as the loss function for training the multi-receptive field visual feature extraction network.

[0020] Compared with the prior art, the micro-expression recognition method based on multi-receptive field visual feature extraction of the present invention has the following beneficial effects:

[0021] The local-global feature integration stage is composed of several local extractors and multi-layer perceptron combination layers followed by two global self-attention and multi-layer perceptron combination layers; among them, the local extractor applies an asymmetric multi-scanning strategy, and the introduced scanning directions do not present a bidirectional or symmetrical relationship between each other, so as to reduce redundancy while enhancing spatial perception ability; for the local-global feature integration stage, the receptive field of the local feature extractor is gradually increased. By combining the local feature extractor with the global self-attention mechanism, the subtle facial features and spatial long-range dependencies of micro-expressions can be effectively captured. At the same time, the asymmetric multi-scanning strategy reduces redundancy while enhancing the spatial perception ability of the model, improving the micro-expression recognition performance.

[0022] In a second aspect, an embodiment of the present invention provides a micro-expression recognition system based on multi-receptive field visual feature extraction, comprising:

[0023] The preprocessing module is used to preprocess the original image frames of the micro-expression dataset, including face detection, alignment and cropping, and calculate the TV-L1 optical flow features based on the starting frame and the peak frame to obtain the input features;

[0024] The multi-receptive field visual feature extraction module is used to convert the input features into overlapping patches and input them into the multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network includes:

[0025] The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling;

[0026] Multiple consecutive local-global feature integration stages, each stage includes several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; wherein the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, and after being processed by a mixer block based on the Mamba architecture, the localized feature representation is generated by fusion through a spatial-channel attention mechanism;

[0027] Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution;

[0028] The classification and recognition module is used to flatten the final output features and classify them through a linear classifier to obtain micro-expression recognition results.

[0029] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor, a communication interface, a memory and a bus, wherein the processor, the communication interface and the memory communicate with each other through the bus, and the processor can call logic instructions in the memory to execute the steps of the method provided in the first aspect.

[0030] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the micro-expression recognition method based on multi-receptive field visual feature extraction as described in the first aspect.

[0031] Compared with the prior art, the beneficial effects of the micro-expression recognition system, electronic device and storage medium based on multi-receptive field visual feature extraction of the present invention are the same as the micro-expression recognition method based on multi-receptive field visual feature extraction described in the first aspect, so they will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0033] Figure 1 Flow chart of a micro-expression recognition method based on multi-receptive field visual feature extraction in an embodiment of the present invention;

[0034] Figure 2 1 is a processing flow chart of a multi-receptive field visual feature extraction network in an embodiment of the present invention;

[0035] Figure 3Schematic diagram of a micro-expression recognition method based on multi-receptive field visual feature extraction in an embodiment of the present invention;

[0036] Figure 4 Schematic diagram of an asymmetric multi-scanning strategy in an embodiment of the present invention;

[0037] Figure 5 It is a structural block diagram of a micro-expression recognition system based on multi-receptive field visual feature extraction in an embodiment of the present invention;

[0038] Figure 6 1 is a structural block diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0040] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0041] In an embodiment of the present invention, a micro-expression recognition method based on multi-receptive field visual feature extraction is provided. Figure 1 is a flow chart of the micro-expression recognition method based on multi-receptive field visual feature extraction of the present invention, such as Figure 1 , Figure 2 , Figure 3 As shown, the process includes the following steps:

[0042] S100, preprocessing the original image frames of the micro-expression dataset, including face detection, alignment and cropping, and calculating the TV-L1 optical flow features based on the start frame and the peak frame to obtain input features;

[0043] Specifically, in terms of data preprocessing, firstly, all the original image frames of the micro-expression dataset are uniformly subjected to face detection, alignment, and cropping. The TV-L1 optical flow features of the samples are calculated using the calibrated start frame and peak frame in the public dataset: It should be noted that the TV-L1 optical flow feature has three dimensions, namely the horizontal component, the vertical component and the modulus, which are used as the input of the end-to-end multi-receptive field visual feature extraction network, and the resolution of the TV-L1 optical flow feature as an input feature is 224×224.

[0044] S200, converting the input features into overlapping patches and inputting them into a multi-receptive field visual feature extraction network;

[0045] Specifically, the input (TV-L1 optical flow features) is first converted to The overlapping patches are then passed to the first feature extraction stage of the multi-receptive field visual feature extraction network.

[0046] More specifically, the multi-receptive field visual feature extraction network includes:

[0047] S210, a fast feature extraction stage, which is composed of multiple residual connected convolutional neural network layers and outputs features through downsampling;

[0048] It should be noted that the first stage of the multi-receptive field visual feature extraction network is a fast feature extraction stage, which is composed of three convolutional neural network layers with nonlinear activation in the form of residual connections. After the downsampling layer processing, the spatial resolution of the feature map is halved and the number of feature channels is doubled. The output features of the first stage are .

[0049] S220, multiple consecutive local-global feature integration stages, each stage comprising a combination layer of several local extractors and multi-layer perceptrons, followed by a combination layer of global self-attention and multi-layer perceptrons; wherein the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, and after being processed by a mixer block based on the Mamba architecture, generates a localized feature representation through a spatial-channel attention mechanism;

[0050] S230, downsampling operation, performing downsampling at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution;

[0051] It should be noted that the multi-receptive field visual feature extraction network contains three local-global feature integration stages. The number of combined layers of local extractors and multi-layer perceptrons in each stage is 2, 6, and 3 respectively, and the output features of the last stage are input into the linear classifier after two-dimensional batch normalization and average pooling.

[0052] The output features of the first stage Directly used as input features for the next stage The second stage is a local-global feature integration stage. Each local-global feature integration stage k is composed of several ( ) local extractor and multi-layer perceptron combination layer followed by 2 global self-attention and multi-layer perceptron combination layers, for the second stage ( )have .

[0053] It should be further explained that in the local-global feature integration stage, the window size of the local extractor is 7×7, the window size of the global self-attention layer is consistent with the spatial resolution of the input features, and in the fourth stage, the window size of the local extractor is the same as that of the global self-attention layer.

[0054] In particular, the local extractor and the global self-attention layer have different spatial receptive fields, and the feature extraction of the local extractor is fixed at The operation is performed in non-overlapping local windows, while the window size of the global self-attention layer is consistent with the spatial resolution of the input feature, that is, the operation is performed on the scale of the entire feature map. Similar downsampling layer processing is performed in the first stage (fast feature extraction stage) to obtain the output features of the second stage .

[0055] Further, Directly used as input features for the next stage The third stage is still a local-global feature integration stage. )have , the window size of the local extractor is still Similar to the first stage (fast feature extraction stage) and the second stage (local-global feature integration stage), similar downsampling layer processing is performed to obtain the output features of the third stage ;

[0056] The output features of the third stage Directly used as input features for the next stage The fourth stage is also a local-global feature integration stage. )have Since the spatial resolution of the input feature map has been reduced to , so the window size of the local extractor and the global self-attention layer is consistent in this stage. After the fourth stage feature extraction, there is no downsampling, but it is replaced by two-dimensional batch normalization and two-dimensional average pooling to obtain the output features of the fourth stage .

[0057] Output features of the fourth stage It is flattened to a one-dimensional vector of length 1024 and a linear classifier is used to perform C classification tasks, where C represents the number of categories of emotional labels. The cross entropy classification loss is used as the loss function for training the multi-receptive field visual feature extraction network.

[0058] Special explanation of the structure of the combined layer of the local extractor and the multi-layer perceptron in the local-global feature integration stage: for each local window , first adopt an asymmetric multi-scanning strategy, specifically, Figure 4 As shown in the figure, the asymmetric multi-scanning strategy includes four directions: horizontal raster scanning, vertical raster scanning, horizontal zigzag scanning, and vertical zigzag scanning, and there is no bidirectional or symmetric relationship between the scanning directions. The spatial-channel attention mechanism adaptively assigns the contribution weights of the four scanning directions, performs weighted fusion on the reconstructed two-dimensional feature maps, and generates the final feature representation of the local window.

[0059] Rearrange it into 4 one-dimensional sequences with different orders , where s (superscript) refers to Figure 4 One of the four scanning directions shown: a (horizontal raster scanning), b (vertical raster scanning), c (horizontal zigzag scanning), and d (vertical zigzag scanning).

[0060] Each sequence is then processed by a Mamba-based mixer block to extract local spatial features, e.g. Figure 3 As shown in the upper right part. This operation can be expressed as: ,in represents the intermediate representation of the sequence after processing. Subsequently, the sequence is reversed back to its original spatial arrangement to reconstruct the corresponding two-dimensional feature map Next, the four reconstructed feature maps are fused through a spatial-channel attention mechanism that adaptively assigns contribution weights to each scanning direction. , thus obtaining a fused two-dimensional feature map . It is further refined through a linear layer and a multi-layer perceptron block to generate the final localized feature representation for each local window .

[0061] S300, flattening the final output features and classifying them through a linear classifier to obtain micro-expression recognition results.

[0062] The present invention constructs an end-to-end multi-receptive field visual feature extraction network, including a fast feature extraction stage and three consecutive local-global feature integration stages: the fast feature extraction stage is composed of a plurality of convolutional neural network layers, which are used for initialization feature extraction of TV-L1 optical flow features calculated from the starting frame and the peak frame of micro-expression samples; the local-global feature integration stage is composed of a combination layer of a plurality of local extractors and a multi-layer perceptron followed by two combination layers of global self-attention and a multi-layer perceptron; wherein the local extractor applies an asymmetric multiple scanning strategy, and the introduced scanning directions do not present a bidirectional or symmetrical relationship between each other, so as to reduce redundancy while enhancing spatial perception capability; for the local-global feature integration stage, the receptive field of the local feature extractor is gradually increased, which is achieved by a downsampling operation at the end of each stage;

[0063] Afterwards, a linear classification head is connected to the feature extractor as the classifier of the entire neural network; the micro-expression samples to be identified are classified by the multi-receptive field visual feature extraction network obtained through training.

[0064] It can be seen from the technical solution provided by the present invention that by combining the local feature extractor with the global self-attention mechanism, the subtle facial features and spatial long-range dependencies of micro-expressions can be effectively captured. At the same time, the asymmetric multi-scanning strategy reduces redundancy while enhancing the spatial perception ability of the model, thereby improving the micro-expression recognition performance.

[0065] In order to intuitively demonstrate the recognition effect of the above scheme of the present invention, three-classification experiments were carried out on the public data sets CASME II, SAMM, SMIC-HS data sets and the joint data set of the three data sets. The experimental results are shown in Table 1. The recognition UF1 value and UAR value both reach the level of the current optimal recognition scheme.

[0066] Table 1 Three-classification experiment

[0067]

[0068] In addition, seven-category experiments were conducted on the A and B test sets of the public dataset DFME-public. The experimental results are shown in Table 2. The recognition accuracy, UF1 value and UAR value also reach the level of the current optimal recognition solution.

[0069] Table 2 DFME-public test set performance

[0070]

[0071] The embodiment of the present invention also provides a micro-expression recognition system based on multi-receptive field visual feature extraction, which is used to implement the above method embodiment, which has been described and will not be repeated. The terms "module", "unit", "sub-unit", etc. used below can implement a combination of software and / or hardware for predetermined functions. Although the system described in the following embodiments is preferably implemented in software, the implementation of hardware or a combination of software and hardware is also possible and conceivable.

[0072] like Figure 5 As shown, Figure 5 It is a structural block diagram of a micro-expression recognition system based on multi-receptive field visual feature extraction in the present invention, and the system includes:

[0073] A preprocessing module 101 is used to preprocess the original image frames of the micro-expression dataset, including face detection, alignment and cropping, and calculate the TV-L1 optical flow features based on the start frame and the peak frame to obtain input features;

[0074] The multi-receptive field visual feature extraction module 102 is used to convert the input features into overlapping patches and input them into the multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network includes:

[0075] The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling;

[0076] Multiple consecutive local-global feature integration stages, each of which contains several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, which are processed by the mixer block based on the Mamba architecture and then fused to generate localized feature representations through the spatial-channel attention mechanism;

[0077] Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution;

[0078] The classification and recognition module 103 is used to flatten the final output features and classify them through a linear classifier to obtain micro-expression recognition results.

[0079] Figure 6 A structural block diagram of an electronic device provided by an embodiment of the present invention, such as Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute the following method:

[0080] The original image frames of the micro-expression dataset are preprocessed, including face detection, alignment and cropping, and the TV-L1 optical flow features are calculated based on the starting frame and the peak frame to obtain the input features;

[0081] The input features are converted into overlapping patches and input into the multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network includes:

[0082] The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling;

[0083] Multiple consecutive local-global feature integration stages, each of which contains several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, which are processed by the mixer block based on the Mamba architecture and then fused to generate localized feature representations through the spatial-channel attention mechanism;

[0084] Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution;

[0085] The final output features are flattened and classified through a linear classifier to obtain the micro-expression recognition results.

[0086] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0087] An embodiment of the present invention further provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.

[0088] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A micro-expression recognition method based on multi-receptive field visual feature extraction, characterized in that: The micro-expression recognition method based on multi-receptive field visual feature extraction includes: The original image frames of the micro-expression dataset are preprocessed, including face detection, alignment and cropping, and the TV-L1 optical flow features are calculated based on the starting frame and the peak frame to obtain the input features; The input features are converted into overlapping patches and input into a multi-receptive field visual feature extraction network; the multi-receptive field visual feature extraction network comprises: The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling; Multiple consecutive local-global feature integration stages, each stage includes several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; wherein the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, and after being processed by a mixer block based on the Mamba architecture, the localized feature representation is generated by fusion through a spatial-channel attention mechanism; Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution; The final output features are flattened and classified through a linear classifier to obtain the micro-expression recognition results.

2. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1 is characterized in that: The multi-receptive field visual feature extraction network includes three local-global feature integration stages, the number of combined layers of the local extractor and the multi-layer perceptron in each stage is 2, 6, and 3 respectively, and the output features of the last stage are input into the linear classifier after two-dimensional batch normalization and average pooling.

3. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1 is characterized in that: In the local-global feature integration stage, the window size of the local extractor is 7×7, the window size of the global self-attention layer is consistent with the spatial resolution of the input features, and in the fourth stage, the window size of the local extractor is the same as that of the global self-attention layer.

4. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1 is characterized in that: The asymmetric multiple scanning strategy includes four directions: horizontal raster scanning, vertical raster scanning, horizontal zigzag scanning and vertical zigzag scanning, and there is no bidirectional or symmetrical relationship between the scanning directions.

5. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1 is characterized in that: The spatial-channel attention mechanism adaptively assigns contribution weights of four scanning directions, performs weighted fusion on the reconstructed two-dimensional feature maps, and generates the final feature representation of the local window.

6. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1, characterized in that: The TV-L1 optical flow feature has three dimensions, which are a horizontal component, a vertical component, and a modulus.

7. The micro-expression recognition method based on multi-receptive field visual feature extraction according to claim 1 is characterized in that: Cross entropy classification loss is used as the loss function for training multi-receptive field visual feature extraction network.

8. A micro-expression recognition system based on multi-receptive field visual feature extraction, characterized in that: include: The preprocessing module is used to preprocess the original image frames of the micro-expression dataset, including face detection, alignment and cropping, and calculate the TV-L1 optical flow features based on the starting frame and the peak frame to obtain the input features; A multi-receptive field visual feature extraction module, used to convert input features into overlapping patches and input them into a multi-receptive field visual feature extraction network; The multi-receptive field visual feature extraction network includes: The fast feature extraction stage consists of multiple residual-connected convolutional neural network layers that output features through downsampling; Multiple consecutive local-global feature integration stages, each stage includes several local extractor and multi-layer perceptron combination layers, followed by global self-attention and multi-layer perceptron combination layers; wherein the local extractor adopts an asymmetric multi-scanning strategy to rearrange the local window features into four one-dimensional sequences with different scanning directions, and after being processed by a mixer block based on the Mamba architecture, the localized feature representation is generated by fusion through a spatial-channel attention mechanism; Downsampling operation, performed at the end of each stage, gradually increasing the receptive field and reducing the spatial resolution; The classification and recognition module is used to flatten the final output features and classify them through a linear classifier to obtain micro-expression recognition results.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the micro-expression recognition method based on multi-receptive field visual feature extraction as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the micro-expression recognition method based on multi-receptive field visual feature extraction as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Micro-expression recognition method based on space-time appearance movement attention network

    CN112307958A

  • Micro-expression recognition method based on dynamic graph convolutional network

    CN114140846A