Micro-expression recognition method and system based on appearance, motion and geometric information fusion

By extracting appearance, motion, and geometric information from micro-expression videos, a multi-information feature fusion module is constructed, which solves the problems of low feature discrimination and robustness in existing technologies and achieves higher accuracy and stability in micro-expression recognition.

CN120877347APending Publication Date: 2025-10-31NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510971248.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing micro-expression recognition methods have low feature discrimination and robustness, and lack effective fusion of appearance, motion and geometric information, resulting in insufficient recognition accuracy and stability.

Method used

By extracting appearance, motion, and geometric information from micro-expression videos, a multi-information feature extraction module is constructed, which is divided into local features and fused locally. Global fusion of information is achieved through a self-attention mechanism and a Transformer layer to improve feature robustness.

Benefits of technology

It improves the accuracy and stability of micro-expression recognition, effectively integrates multiple information features, and enhances the performance of the recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877347A_ABST
    Figure CN120877347A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-expression recognition method and system based on appearance, motion and geometric information fusion. The method mainly comprises the following steps: (1) extracting micro-expression appearance, motion and geometric information data from a video sample of a micro-expression data set; (2) constructing a micro-expression recognition model based on appearance, motion and geometric information fusion; (3) training a micro-expression recognition model based on appearance, motion and geometric information fusion by adopting micro-expression appearance, motion and geometric information data of the video samples in the micro-expression data set; and (4) extracting appearance, motion and geometric information data from the micro-expression test video, and inputting the appearance, motion and geometric information data into the micro-expression recognition model based on appearance, motion and geometric information fusion to obtain a micro-expression category. According to the method, the identifiable information in the micro-expression video is fully mined, and the micro-expression recognition accuracy can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing target detection, specifically to a micro-expression recognition method and system based on the fusion of appearance, motion, and geometric information. Background Technology

[0002] Humans express their emotions through facial expressions, and facial expression recognition technology is crucial for communication and interaction between people. However, when people attempt to suppress or hide their true emotions, their true feelings cannot be determined through macro-expressions alone. Instead, they rely on micro-expressions. Micro-expressions are very brief and spontaneous facial expressions that arise when humans try to suppress or hide their true emotions, and they can reveal a person's true feelings. Because micro-expressions can effectively infer people's true emotions, micro-expression recognition technology has significant application value and potential prospects in fields such as lie detection, public safety, criminal investigation and interrogation, negotiation, and medicine. For example, in criminal investigations, the micro-expressions of suspects can be used to determine their true emotions, thereby guiding the investigation and accelerating case processing. In military negotiations, the micro-expressions of opposing personnel can be used to infer their true intentions and make correct negotiation decisions.

[0003] Compared to macro-expressions, micro-expressions occur for a shorter duration (only 1 / 25 to 1 / 5 of a second) and cause weaker facial movements. Therefore, micro-expressions are difficult for the human eye to detect, making manual recognition extremely challenging. According to existing literature, even trained personnel achieve a recognition rate of only about 47%. With the development and advancement of artificial intelligence technology, automatic micro-expression recognition has become a research hotspot. Employing computer and AI algorithms for micro-expression recognition can advance the process of machines understanding genuine human emotions and further promote the high-quality development of intelligent human-computer interaction. However, given the current state of research, the accuracy of micro-expression recognition remains low, and many problems and challenges persist.

[0004] Extracting micro-expression features from micro-expression videos is a core step in micro-expression recognition systems. The effectiveness and robustness of these features directly impact the accuracy and stability of micro-expression recognition in real-world applications. Because the facial muscle movements caused by micro-expressions are subtle and inconspicuous, micro-expression features are easily masked by irrelevant features, making the extraction of effective and robust micro-expression features extremely challenging. Micro-expression videos contain various feature information related to micro-expressions, such as appearance, motion, and geometric information. Fully utilizing this feature information is crucial for extracting effective and robust micro-expression features. Existing literature indicates that appearance and motion information are effective for micro-expression recognition, and the inventors of this patent recently discovered that geometric information can also be effectively applied to micro-expression recognition. The fusion of different information can fully exploit the micro-expression-related feature information in micro-expression videos, extracting richer, more stable, and robust micro-expression features, thereby improving the accuracy and stability of the micro-expression recognition system. However, on the one hand, there is currently a lack of micro-expression recognition methods that integrate appearance, motion, and geometric information; on the other hand, micro-expression recognition technologies based on the fusion of these three types of information have many unresolved issues, such as how to effectively extract different information from micro-expression videos; how to build models to extract and align features from multi-information data with different structures; and how to effectively fuse multi-information features. Summary of the Invention

[0005] The purpose of this invention is to address the problems of low feature discrimination and robustness in existing micro-expression recognition methods. This invention aims to provide a micro-expression recognition method and system based on the fusion of appearance, motion, and geometric information. It mines appearance, motion, and geometric information from micro-expression videos, extracts appearance, motion, and geometric features, and divides them into local appearance, motion, and geometric features. Then, it fuses the appearance, motion, and geometric local features within local facial regions to obtain multi-information local fusion features. Finally, it fuses the multi-information local fusion features from different regions to obtain multi-information global fusion features, achieving effective fusion of appearance, motion, and geometric information, thereby improving the robustness of micro-expression features and the accuracy of micro-expression recognition.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] A micro-expression recognition method based on the fusion of appearance, motion, and geometric information, the method comprising:

[0008] S100. Extract appearance, motion, and geometric information data from video samples in the micro-expression dataset;

[0009] S200. Construct a micro-expression recognition model based on the fusion of appearance, motion, and geometric information;

[0010] S300: Use the appearance, motion, and geometric information data of video samples in the micro-expression dataset to train a micro-expression recognition model based on the fusion of appearance, motion, and geometric information;

[0011] S400: Extract appearance, motion, and geometric data from the micro-expression test video, and input the appearance, motion, and geometric data into a micro-expression recognition model based on the fusion of appearance, motion, and geometric information to obtain the micro-expression category.

[0012] Preferably, in S100, appearance, motion, and geometric information data are extracted from video samples in the micro-expression dataset, including:

[0013] S101. Locate the starting frame image of the micro-expression occurrence from the micro-expression video. and peak frame image Peak frame image This is the appearance information data; its length and width are respectively... and ;

[0014] S102. Extract the starting frame image using an optical flow extraction algorithm. and peak frame image Optical flow images between As motion information data, its length and width are respectively and ;

[0015] S103. Extract the starting frame image using a face calibration point detection algorithm. and peak frame image Face landmarks and geometric information data ;

[0016] in, and These are the first and second frames in the starting frame and peak frame, respectively. The coordinates of the individual's face marker. The number of face markers.

[0017] Preferably, the micro-expression recognition model in S200 includes: a multi-information feature extraction module, a local feature segmentation module, a multi-information local fusion module, a multi-information global fusion module, and a micro-expression classification module;

[0018] S201, The multi-information feature extraction module is used to extract appearance feature maps, motion feature maps and geometric features from micro-expression appearance, motion and geometric information data respectively;

[0019] S202, The local feature segmentation module is used to segment the appearance feature map, motion feature map and geometric features into appearance, motion and geometric local features;

[0020] S203. The multi-information local fusion module is used to fuse appearance, motion and geometric local features within a local area to obtain multi-information local fusion features.

[0021] S204. The multi-information global fusion module is used to fuse multi-information local fusion features in different regions to obtain micro-expression multi-information global fusion features.

[0022] S205. The micro-expression classification module is used to classify multi-information global fusion features to obtain micro-expression categories.

[0023] Preferably, the multi-information feature extraction module in S201 employs a convolutional neural network. From peak frame image Extracting appearance feature maps Using a convolutional neural network model From optical flow images Extracting motion feature maps Using graph convolutional neural networks From facial landmarks Extracting geometric features ;

[0024] in, and These are the length and width of the feature map, respectively. The number of feature channels, The number of feature channels, , It is the dimension of the feature of each graph node.

[0025] Preferably, the local feature segmentation module in S202 includes: dividing the appearance feature map, motion feature map, and geometric features into... Local features of appearance Local features of motion and geometric local features ,in, , The partitioning method includes the following steps:

[0026] S202-1 Calculate the appearance feature maps respectively and motion feature map The Middle Coordinates of feature points and ;in, , , For rounding functions, you can use round-up functions, round-to-the-floor functions, etc.

[0027] S202-2, with and Centered on, from the perspective of appearance characteristics and motion characteristics Extracting local feature blocks of appearance from the middle and motion local feature blocks ;in and These are the length and width of the feature block, respectively;

[0028] S202-3, respectively and Flattening to obtain local appearance features and local features of motion ;in, = ;

[0029] S202-4, No. Geometric local features ;in, ;

[0030] S202-5, To For each value, repeat S202-1 to S202-4 to obtain Local features of appearance , Local features of motion and Geometric local features .

[0031] Preferably, the multi-information local fusion module in S203 employs a single-output self-attention mechanism to fuse appearance, motion, and geometric local features to obtain... Multi-information local fusion features , The single-output self-attention mechanism includes the following steps:

[0032] S203-1, the first A multi-information local feature group is composed of appearance, motion, and geometric local features. ;in, ;

[0033] S203-2. Calculate the query according to the formula. ,key Sum : ;

[0034] in, , and All are learnable parameter matrices. = ; for Transpose of;

[0035] S203-3. Select one feature from appearance, motion, and geometric local features as the core feature, and the other two features as supplementary features. Calculate the result using the following formula to obtain the... Multiple information local fusion features:

[0036] ;

[0037] in, The OK, =1, 2, or 3, corresponding to appearance, motion, and geometric local features, respectively.

[0038] Preferably, the multi-information global fusion module in S204 will... Multi-information local fusion features Encoded as A token vector, and A token vector is input to In the Transformer layer, global fusion features of micro-expression information are obtained. .

[0039] Preferably, a micro-expression recognition method based on the fusion of appearance, motion, and geometric information, by mining appearance, motion, and geometric information in micro-expression videos, and by first fusing multi-information features of local regions, and then fusing multi-information fusion features of different regions, achieves the fusion of appearance, motion, and geometric information, specifically including:

[0040] Peak frame image, optical flow image between peak frame and start frame, and face calibration point coordinates between peak frame and start frame are extracted from micro-expression video, respectively, and used as appearance, motion and geometric information data;

[0041] Appearance, motion, and geometric features are extracted from peak frame images, optical flow images, and face calibration point coordinates. The appearance, motion, and geometric features are divided into appearance, motion, and geometric local features, with each local facial region as a unit.

[0042] The appearance, motion, and geometric local features of each facial region are fused to obtain multi-information local fusion features; the correlation between the local fusion features of different facial regions is modeled, and the multi-information local fusion features of different facial regions are fused to obtain the final micro-expression multi-information global fusion features.

[0043] Preferably, a micro-expression recognition system is characterized in that: the system comprises:

[0044] The multi-information data extraction module is used to extract data containing appearance, motion, and geometric information from micro-expression videos;

[0045] The multi-information feature extraction module is used to extract appearance feature maps, motion feature maps, and geometric features from appearance, motion, and geometric data, respectively.

[0046] The local feature segmentation module is used to segment the appearance feature map, motion feature map, and geometric features into appearance, motion, and geometric local features, respectively.

[0047] The multi-information local fusion module is used to fuse the appearance, motion, and geometric local features of each facial region and learn multi-information local fusion features.

[0048] The multi-information global fusion module is used to model the correlation between multi-information local fusion features in different facial regions, fuse multi-information local fusion features in different regions, and learn the multi-information global fusion features of micro-expressions.

[0049] The micro-expression classification module is used to classify the global fusion features of multiple micro-expression information to obtain micro-expression recognition results.

[0050] A micro-expression recognition system includes at least one computing device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the aforementioned micro-expression recognition method based on the fusion of appearance, motion, and geometric information.

[0051] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0052] (1) In view of the problem of low discriminativeness and robustness of micro-expression features in micro-expression recognition, this invention proposes a micro-expression recognition method based on the fusion of appearance, motion and geometric information. The appearance, motion and geometric features are extracted from micro-expression videos and divided into appearance, motion and geometric local features. Then, the multi-information features in the local area are fused to learn the multi-information local fusion features and construct a multi-information fusion module to fuse the multi-information fusion features of different facial local areas. Finally, the appearance, motion and geometric information are effectively fused from local to global. The experimental results verify that the micro-expression recognition method based on the fusion of appearance, motion and geometric information proposed in this invention can effectively fuse multiple information in micro-expression videos and improve the accuracy of micro-expression recognition.

[0053] (2) In the micro-expression recognition method based on the fusion of appearance, motion and geometric information proposed in this invention, different deep learning models are used to extract features for information of different data types. Specifically, for information based on image data, such as appearance and motion information, convolutional neural networks are used to extract appearance and motion features, while for information based on coordinate data, such as geometric information, graph convolutional networks are used to extract geometric features.

[0054] (3) In the multi-information local fusion module constructed in this invention, a single-output self-attention mechanism is designed to divide local regions based on face calibration points, and to fuse multi-information local features in each local region with one type of information as the core and the other two types of information as supplements.

[0055] (4) In the global fusion module constructed in this invention, multiple Transformer layers are stacked to model the correlation between multi-information local fusion features in different facial regions, effectively fuse multi-information local fusion features in different facial regions, and learn robust and effective micro-expression multi-information global fusion features. Attached Figure Description

[0056] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0057] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.

[0058] Figure 2 This is a framework diagram of a micro-expression recognition model based on the fusion of appearance, motion, and geometric information in an embodiment of the present invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Please see Figures 1-2 The present invention provides the following technical solution:

[0061] Example 1: The micro-expression recognition method based on the fusion of appearance, motion, and geometric information disclosed in this embodiment of the invention first obtains peak frame images, optical flow images, and face calibration point data from micro-expression video dataset samples, respectively; then, a multi-information feature extraction module is constructed to extract appearance, motion, and geometric features from the peak frames, optical flow images, and face calibration point data, respectively; next, a multi-information local fusion module is constructed to fuse multi-information features within local facial regions, learning multi-information local fusion features from appearance, motion, and geometric features; and a global feature fusion module is constructed to fuse multi-information local fusion features from different local facial regions, learning global multi-information fusion features for micro-expressions; finally, a classifier is used to classify the global fusion features of micro-expressions to obtain micro-expression categories.

[0062] Before introducing the method of the embodiments of the present invention, a brief description of the micro-expression video dataset used in the present invention will be given first. Those skilled in the art will understand that the scope of protection of the present invention is not limited to the specific micro-expression video dataset used in this embodiment. The micro-expression video dataset used in the embodiments of the present invention is the CASME II micro-expression video dataset, containing 255 micro-expression video samples, divided into 7 emotion categories. Due to the small sample size of some categories, categories with fewer than 10 samples were deleted. Therefore, this embodiment retains 246 micro-expression video samples, divided into 5 emotion categories: happiness, disgust, surprise, repression, and others. This dataset contains peak frame and start frame labels, which can be used to locate the start frame and peak frame in the micro-expression video samples.

[0063] Specifically, such as Figure 1 As shown in the figure, the micro-expression recognition method based on the fusion of appearance, motion, and geometric information provided by this invention mainly includes the following steps:

[0064] (1) The starting frame image of the micro-expression was located from the CASME II micro-expression video dataset samples using the label information of the starting frame and the peak frame. and peak frame image The starting frame image is extracted using an optical flow extraction algorithm. and peak frame image Optical flow images between The starting frame image is extracted using a face landmark detection algorithm. and peak frame image Set of face calibration point coordinates Among them, the starting frame image Peak frame image and optical flow images The length and width are respectively and 224; set of face calibration point coordinates ,in, and These are the first and second frames in the starting frame and peak frame, respectively. The coordinates of individual face landmarks. Those skilled in the art will understand that the scope of protection of this invention is not limited to the specific start frame image and peak frame image localization algorithm, optical flow extraction algorithm, and face landmark detection algorithm used in this embodiment.

[0065] (2) Construct a micro-expression recognition model based on the fusion of appearance, motion, and geometric information. The model framework is as follows: Figure 2 As shown, the model includes a multi-information feature extraction module, a multi-information local feature segmentation module, a multi-information local fusion module, a multi-information global fusion module, and a micro-expression classification module.

[0066] Multi-information feature extraction module: This module extracts peak frame images The input is fed into the ResNet50 model, and the output feature map of the first layer is taken as the appearance feature map. ,Right now, , =64; Input the optical flow image T into the ResNet50 model, and take the output feature map of the first layer as the motion feature map. ,Right now, A graph convolutional neural network containing one graph convolutional layer is used to extract the coordinates of face calibration points from the set of face calibration points. Extracting geometric features That is, the dimension of each graph node feature is 576.

[0067] Multi-information local feature segmentation module: This module segments the appearance feature map Motion feature map and geometric features Divided into Local features of appearance 68 local motion features and 68 geometric local features ,in, 68.

[0068] The specific division method includes the following steps:

[0069] S202-1 Calculate the appearance feature maps respectively and motion feature map The Middle Coordinates of feature points and ,in , ;

[0070] S202-2, with and Centered on the external feature diagrams and motion feature map Extracting local feature blocks of appearance from the middle and local movement ;

[0071] S202-3, respectively and Flattening to obtain local appearance features and local features of motion ,in, = ;

[0072] S202-4, No. Geometric local features ;

[0073] S202-5, To For each value, repeat steps (3.1.1)-(3.1.4) to obtain 68 local appearance features. 68 local motion features and 68 geometric local features .

[0074] Multi-information local fusion module: Employs a single-output self-attention mechanism to fuse appearance, motion, and geometric local features to obtain... Multi-information local fusion features , The single-output self-attention mechanism method includes the following steps:

[0075] S203-1, the first A multi-information local feature group is composed of appearance, motion, and geometric local features. ;

[0076] S203-2, Calculation and Query ,key Sum The formula is as follows: ;

[0077] in, , and For learnable parameter matrix, for Transpose of;

[0078] S203-3. Calculate the following formula to obtain the... Multiple information local fusion features:

[0079] ;

[0080] in, The OK, =1, 2, or 3, corresponding to appearance, motion, and geometric local features, respectively, in the embodiments of the invention. The value is 2.

[0081] Construct a global feature fusion module: Multi-information local fusion features Encoded as A token vector, and A token vector is input to In the Transformer layer, multi-information fusion features of micro-expressions are obtained. .

[0082] (3) Use the appearance, motion and geometric information data of video samples in the CASME II micro-expression dataset to train a micro-expression recognition model based on the fusion of appearance, motion and geometric information;

[0083] (4) Extract appearance, motion and geometry data from the micro-expression test video, and input the appearance, motion and geometry data into the micro-expression recognition model based on the fusion of appearance, motion and geometry information to obtain the micro-expression category.

[0084] The beneficial effects of the present invention can be further illustrated by the comparative experimental results in Table 1.

[0085] Table 1 Ablation experimental results on CASME II dataset

[0086] In Table 1, UF1 and UAR are the unweighted F1 score and the unweighted average recall, respectively.

[0087] As can be seen from the results in Table 1, fusing the three types of information achieved the best performance.

[0088] Based on the same inventive concept, this invention provides a micro-expression recognition system based on the fusion of appearance, motion, and geometric information, comprising:

[0089] Multi-information data extraction module: used to extract data containing appearance, motion, and geometric information from micro-expression videos;

[0090] Multi-information feature extraction module: used to extract appearance feature maps, motion feature maps, and geometric features from appearance, motion, and geometric data, respectively;

[0091] Local feature segmentation module: used to segment appearance feature map, motion feature map and geometric features into appearance, motion and geometric local features respectively;

[0092] Multi-information local fusion module: used to fuse appearance, motion and geometric local features within each facial region to learn multi-information local fusion features;

[0093] Multi-information global fusion module: used to model the correlation between multi-information local fusion features of different facial regions, fuse multi-information local fusion features of different regions, and learn the multi-information global fusion features of micro-expressions;

[0094] The micro-expression classification module is used to classify the global fusion features of multiple micro-expression information to obtain micro-expression recognition results.

[0095] The specific implementation of each module is described in the above method embodiments and will not be repeated here. Those skilled in the art will understand that the modules in the embodiments can be adaptively changed and placed in one or more systems different from this embodiment. The modules, units, or components in the embodiments can be combined into a single module, unit, or component, and they can also be divided into multiple sub-modules, sub-units, or sub-components.

[0096] Based on the same inventive concept, the present invention provides a micro-expression recognition system based on the fusion of appearance, motion and geometric information, including at least one computing device. The computing device includes a memory, a processor and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the aforementioned micro-expression recognition method based on the fusion of appearance, motion and geometric information.

[0097] The technical solutions disclosed in this invention include not only the technical methods involved in the above-described embodiments, but also technical solutions formed by any combination of the above technical methods. Those skilled in the art can make certain improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered to be within the scope of protection of this invention.

[0098] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A micro-expression recognition method based on the fusion of appearance, motion, and geometric information, characterized in that: The method includes: S100. Extract appearance, motion, and geometric information data from video samples in the micro-expression dataset; S200. Construct a micro-expression recognition model based on the fusion of appearance, motion, and geometric information; S300: Use the appearance, motion, and geometric information data of video samples in the micro-expression dataset to train a micro-expression recognition model based on the fusion of appearance, motion, and geometric information; S400: Extract appearance, motion, and geometric data from the micro-expression test video, and input the appearance, motion, and geometric data into a micro-expression recognition model based on the fusion of appearance, motion, and geometric information to obtain the micro-expression category.

2. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 1, characterized in that, The step S100 involves extracting appearance, motion, and geometric information data from video samples in the micro-expression dataset, including: S101. Locate the starting frame image of the micro-expression occurrence from the micro-expression video. and peak frame image Peak frame image This is the appearance information data; its length and width are respectively... and ; S102. Extract the starting frame image using an optical flow extraction algorithm. and peak frame image Optical flow images between As motion information data, its length and width are respectively and ; S103. Extract the starting frame image using a face calibration point detection algorithm. and peak frame image Face landmarks and geometric information data ; in, and These are the first and second frames in the starting frame and peak frame, respectively. The coordinates of the individual's face marker. The number of face markers.

3. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 1, characterized in that, The micro-expression recognition model in S200 includes: a multi-information feature extraction module, a local feature segmentation module, a multi-information local fusion module, a multi-information global fusion module, and a micro-expression classification module; S201, The multi-information feature extraction module is used to extract appearance feature maps, motion feature maps and geometric features from micro-expression appearance, motion and geometric information data respectively; S202, The local feature segmentation module is used to segment the appearance feature map, motion feature map and geometric features into appearance, motion and geometric local features; S203. The multi-information local fusion module is used to fuse appearance, motion and geometric local features within a local area to obtain multi-information local fusion features. S204. The multi-information global fusion module is used to fuse multi-information local fusion features in different regions to obtain micro-expression multi-information global fusion features. S205. The micro-expression classification module is used to classify multi-information global fusion features to obtain micro-expression categories.

4. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 3, characterized in that, The multi-information feature extraction module in S201 employs a convolutional neural network. From peak frame image Extracting appearance feature maps Using a convolutional neural network model From optical flow images Extracting motion feature maps Using graph convolutional neural networks From facial landmarks Extracting geometric features ; in, and These are the length and width of the feature map, respectively. The number of feature channels, The number of feature channels, , It is the dimension of the feature of each graph node.

5. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 3, characterized in that, The local feature segmentation module in S202 includes: dividing the appearance feature map, motion feature map, and geometric features into... Local features of appearance Local features of motion and geometric local features ,in, , The partitioning method includes the following steps: S202-1 Calculate the appearance feature maps respectively and motion feature map The Middle Coordinates of feature points and ;in, , , It is a rounding function; S202-2, with and Centered on, from the perspective of appearance characteristics and motion characteristics Extracting local feature blocks of appearance from the middle and motion local feature blocks ;in and These are the length and width of the feature block, respectively; S202-3, respectively and Flattening to obtain local appearance features and local features of motion ;in, = ; S202-4, No. Geometric local features ;in, ; S202-5, To For each value, repeat S202-1 to S202-4 to obtain Local features of appearance , Local features of motion and Geometric local features .

6. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 3, characterized in that, The multi-information local fusion module in S203 uses a single-output self-attention mechanism to fuse appearance, motion, and geometric local features to obtain... Multi-information local fusion features , The single-output self-attention mechanism includes the following steps: S203-1, the first A multi-information local feature group is composed of appearance, motion, and geometric local features. ;in, ; S203-2. Calculate the query according to the formula. ,key Sum : ; in, , and All are learnable parameter matrices. = ; for Transpose of; S203-3. Select one feature from appearance, motion, and geometric local features as the core feature, and the other two features as supplementary features. Calculate the result using the following formula to obtain the... Multiple information local fusion features: ; in, The OK, =1, 2, or 3, corresponding to appearance, motion, and geometric local features, respectively.

7. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 3, characterized in that, The multi-information global fusion module in S204 will... Multi-information local fusion features Encoded as A token vector, and A token vector is input to In the Transformer layer, global fusion features of micro-expression information are obtained. .

8. The micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in claim 1, characterized in that, The method extracts appearance, motion, and geometric information from micro-expression videos, and achieves the fusion of appearance, motion, and geometric information by first fusing multi-information features of local regions and then fusing multi-information fusion features of different regions. Specifically, it includes: Peak frame image, optical flow image between peak frame and start frame, and face calibration point coordinates between peak frame and start frame are extracted from micro-expression video, respectively, and used as appearance, motion and geometric information data; Appearance, motion, and geometric features are extracted from peak frame images, optical flow images, and face calibration point coordinates. The appearance, motion, and geometric features are divided into appearance, motion, and geometric local features, with each local facial region as a unit. By fusing appearance, motion, and geometric local features within each facial region, multi-information local fusion features are obtained. The correlation between local fusion features of different facial regions is modeled, and multi-information local fusion features of different facial regions are fused to obtain the final micro-expression multi-information global fusion features.

9. A micro-expression recognition system for implementing the micro-expression recognition method based on the fusion of appearance, motion, and geometric information as described in any one of claims 1-8, characterized in that: The system includes: The multi-information data extraction module is used to extract data containing appearance, motion, and geometric information from micro-expression videos; The multi-information feature extraction module is used to extract appearance feature maps, motion feature maps, and geometric features from appearance, motion, and geometric data, respectively. The local feature segmentation module is used to segment the appearance feature map, motion feature map, and geometric features into appearance, motion, and geometric local features, respectively. The multi-information local fusion module is used to fuse the appearance, motion, and geometric local features of each facial region and learn multi-information local fusion features. The multi-information global fusion module is used to model the correlation between multi-information local fusion features in different facial regions, fuse multi-information local fusion features in different regions, and learn the multi-information global fusion features of micro-expressions. The micro-expression classification module is used to classify the global fusion features of multiple micro-expression information to obtain micro-expression recognition results.

10. The micro-expression recognition system as described in claim 9, characterized in that, The device includes at least one computing device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements a micro-expression recognition method based on the fusion of appearance, motion, and geometric information according to any one of claims 1-8.