Micro-expression recognition method and system fusing action unit trend analysis and sequence enhancement technology

CN122821602APending Publication Date: 2026-09-25BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004206.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这些噪声若被输入模型,将严重干扰训练过程,降低模型的鲁棒性和跨场景泛化能力

Benefits of technology

[0020]本申请针对微表情识别任务对开源大模型进行领域适配微调,显著提升模型对微表情面部动作单元(Action Unit,AU)特征的提取精度,有效增强了特征精确度以及判别能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821602A_ABST
    Figure CN122821602A_ABST
Patent Text Reader

Abstract

The application discloses a micro-expression recognition method and system fusing action unit trend analysis and sequence enhancement technology, and relates to the technical field of biometric identification. The method comprises the following steps: acquiring a micro-expression video segment to be detected; using a trained facial action unit feature extraction model to frame by frame extract intensity value time sequence of multiple facial action units, and obtaining a unit set composed of the multiple facial action units; intercepting effective sub-sequences from each intensity value time sequence, and expanding all the effective sub-sequences to a preset length by using a linear interpolation method; according to a preset unstable facial action unit list, eliminating unstable facial action units from the unit set; composing one-dimensional time sequence features of the effective sub-sequences of the facial action units remaining after the elimination into a trained classification network to obtain a recognition result. The application can significantly improve the accuracy and robustness of micro-expression recognition under the condition of using only a lightweight network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biometric recognition technology, specifically to a micro-expression recognition method and system that integrates action unit trend analysis and sequence enhancement technology. Background Technology

[0002] The statements in this section are merely to provide background information in relation to this application to aid in understanding it, and such background information does not necessarily constitute prior art.

[0003] Microexpressions are brief, unconscious, and usually unintentional facial expressions characterized by their rapid and localized nature. They often reveal an individual's true emotions that they are trying to conceal. Due to their extremely short duration and subtle amplitude, microexpressions are difficult to directly perceive and analyze with the naked eye. Because they are not under conscious control, microexpressions can be considered an unconscious revelation of true emotional states, and they have significant value in scenarios such as lie detection, psychological counseling, and clinical diagnosis, providing crucial clues for identifying hidden human emotions. Therefore, they have attracted widespread attention from the academic and engineering application fields.

[0004] In recent years, action unit analysis based on facial action coding systems has become a mainstream research direction in micro-expression recognition. This system deconstructs complex expressions into multiple independent facial muscle movement units, achieving a mapping from a high-dimensional image space to a low-dimensional action semantic space. This not only significantly compresses data dimensionality and improves model processing efficiency but also provides physiologically interpretable semantic features, enabling the recognition model to transcend individual appearance differences and focus on the essential dynamic changes in emotional expression. This lays the theoretical and methodological foundation for building efficient and robust micro-expression recognition systems.

[0005] However, the above methods still have significant drawbacks, mainly including: First, individual differences in facial structure and muscle activity baselines significantly increase the within-class variance of similar emotions, causing the model to tend to learn static intensity rather than dynamic changes, thus introducing discrimination bias. Simply using global normalization to eliminate such differences weakens the discriminative power of key motor unit change patterns corresponding to different emotions, making it even more difficult to effectively extract the already weak and scarce effective signals from limited samples. Furthermore, in terms of model training, the currently available micro-expression datasets have limited sample size and high annotation costs, resulting in scarce training data and limiting the performance ceiling and generalization ability of data-driven models. Moreover, during feature extraction, fluctuations in motor units caused by non-emotional factors such as head rotation and sudden changes in lighting can be mistakenly identified as effective emotional signals, forming noisy features. If this noise is input into the model, it will severely interfere with the training process, reducing the model's robustness and cross-scene generalization ability. Summary of the Invention

[0006] This application aims to provide a micro-expression recognition scheme that integrates action unit trend analysis and sequence enhancement technology for micro-expression recognition tasks in complex real-world scenarios. By combining the proposed data augmentation strategy, temporal standardization mechanism and key feature selection method, the accuracy and robustness of micro-expression recognition are significantly improved using only a lightweight classification model.

[0007] To achieve the above objectives, this application provides the following technical solution:

[0008] According to the first aspect of this application, a micro-expression recognition method integrating action unit trend analysis and sequence enhancement technology is provided, comprising: Step 1, acquiring a micro-expression video segment to be detected, using a trained facial action unit feature extraction model to extract the intensity value temporal sequence of multiple facial action units frame by frame, and obtaining a unit set composed of multiple facial action units corresponding to the micro-expression video segment to be detected; Step 2, extracting effective sub-sequences from each intensity value temporal sequence, and using linear interpolation to extend all effective sub-sequences to a preset length; Step 3, removing unstable facial action units from the unit set according to a preset list of unstable facial action units; Step 4, forming a one-dimensional temporal feature from the effective sub-sequences corresponding to the facial action units in the unit set obtained after removal, and inputting it into a trained classification network to obtain the emotion category recognition result.

[0009] Preferably, step 2, which involves extracting a valid subsequence from each intensity value time sequence, includes: extracting a valid subsequence from the start frame to the peak frame from each intensity value time sequence, or extracting a falling segment sequence from the peak frame to the end frame and then reversing the time of the falling segment sequence to obtain a valid subsequence; wherein, the start frame is the image frame at the beginning of the micro-expression, the peak frame is the image frame at the peak of the micro-expression intensity, and the end frame is the image frame at the end of the micro-expression.

[0010] Preferably, step 2 further includes: performing differential scaling normalization on the effective subsequence according to the following steps to eliminate intensity differences of the same facial motion unit among different individuals: obtaining the intensity values ​​of a preset number of frames before the starting frame from the intensity value temporal sequence, and calculating the baseline mean of each facial motion unit based on the intensity values ​​of the preset number of frames; constructing a differential sequence based on the effective subsequence and the baseline mean, and performing min-max normalization on the differential sequence; multiplying the normalized result element-wise with the facial motion unit intensity values ​​in the effective subsequence to obtain an effective subsequence with fused temporal information.

[0011] Preferably, the facial action unit feature extraction model in step 1 is trained according to the following process: acquiring a video sample set, which includes videos of multiple facial images, with each image frame in each video labeled with multiple facial action units and their intensity values; and fine-tuning a preset pre-trained model using the video sample set to obtain the facial action unit feature extraction model.

[0012] Preferably, the classification network in step 4 is trained according to the following process: using a fine-tuned facial action unit feature extraction model, the intensity value temporal sequence of multiple facial action units is extracted frame by frame from the video sample set, and a unit set composed of multiple facial action units corresponding to the video sample is obtained; effective sub-sequences from the start frame to the end frame are extracted from each intensity value temporal sequence; data augmentation processing is performed on the effective sub-sequences based on the dynamic symmetry of micro-expressions to obtain extended training samples; length unification processing and differential scaling standardization processing are performed on the extended training samples to obtain processed extended training samples; unstable facial action units are removed from the unit set based on a preset trend consistency screening rule; effective sub-sequences of facial action units remaining in the unit set are obtained from the extended training samples and formed into a one-dimensional temporal training feature; the emotion category temporal sequence corresponding to the video sample is obtained; the classification network is trained according to the one-dimensional temporal training feature and the emotion category temporal sequence until the network converges to obtain a trained classification network.

[0013] Preferably, the extended training samples are obtained by performing data augmentation processing on the effective sub-sequences based on the dynamic symmetry of micro-expressions, including: extracting the rising segment sequence and falling segment sequence of each effective sub-sequence, wherein the rising segment sequence is the effective sub-sequence from the start frame to the peak frame, and the falling segment sequence is the effective sub-sequence from the peak frame to the end frame, and the peak frame is the image frame at the peak moment of micro-expression intensity; reversing the falling segment sequence in time to obtain the reversed sequence; and using the reversed sequence and the rising segment sequence as new effective sub-sequences for each facial action unit to obtain extended training samples composed of new effective sub-sequences.

[0014] Preferably, based on a preset trend consistency screening rule, unstable facial action units are removed from the unit set, including: for each facial action unit under each emotion category in the extended training samples, determining the trend direction of its intensity value based on a sign function to obtain a trend feature vector, wherein the trend feature vector includes four trend patterns: increase-increase, increase-decrease, decrease-increase, and decrease-decrease; determining the dominant trend pattern of each facial action unit under each emotion category based on the trend pattern of the trend feature vector; comparing the trend feature vector corresponding to each facial action unit under each emotion category with the dominant trend pattern, identifying samples with completely opposite trends as noise samples, and calculating the global noise sensitivity score of each facial action unit based on the number of noise samples; removing facial action units with a global noise sensitivity score exceeding a preset threshold from the unit set.

[0015] Preferably, a list of eliminated facial motion units is generated based on the global noise sensitivity scores exceeding a preset threshold, and the list of eliminated facial motion units is used as the preset list of unstable facial motion units in step 3.

[0016] According to a second aspect of this application, a micro-expression recognition system integrating action unit trend analysis and sequence enhancement technology is provided, comprising: a feature extraction module, an effective frame extraction module, a feature filtering module, and a recognition module; wherein: the feature extraction module is configured to acquire a micro-expression video segment to be detected, and use a trained facial action unit feature extraction model to extract the intensity value temporal sequence of multiple facial action units frame by frame, and obtain a unit set composed of multiple facial action units corresponding to the micro-expression video segment to be detected; the effective frame extraction module is configured to extract effective sub-sequences from each intensity value temporal sequence, and use linear interpolation to extend all effective sub-sequences to a preset length; the feature filtering module is configured to remove unstable facial action units from the unit set according to a preset list of unstable facial action units; the recognition module is configured to form a one-dimensional temporal feature from the effective sub-sequences corresponding to the facial action units in the unit set obtained after removal, and input it into a trained classification network to obtain an emotion category recognition result.

[0017] According to a third aspect of this application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of the first aspect of this application.

[0018] According to a fourth aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method of the first aspect of this application.

[0019] Compared with the prior art, this application has the following advantages:

[0020] This application performs domain-adaptation fine-tuning on a large open-source model for micro-expression recognition tasks, significantly improving the model's extraction accuracy of facial action unit (AU) features for micro-expressions, and effectively enhancing feature accuracy and discrimination ability.

[0021] We propose a time-series data augmentation method based on a bimodal distribution to effectively expand the sample space. Combined with differential scaling and normalization, this method preserves key temporal information while amplifying the samples, reduces intra-class differences and noise interference, and provides high-quality time-series input for lightweight classification models.

[0022] A key facial action unit screening mechanism based on trend consistency is constructed to eliminate facial action units that are prone to misjudgment, thereby further improving the discriminative power of the feature set and enabling high-precision micro-expression recognition using a lightweight classification network. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0024] Figure 1 This is a feature image of a video sample converted into one-dimensional data according to an embodiment of this application;

[0025] Figure 2 This is a flowchart illustrating a micro-expression recognition method according to an embodiment of this application;

[0026] Figure 3 This is a flowchart illustrating a micro-expression recognition method according to another embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. Furthermore, the technical features involved in the various embodiments described below can be combined with each other as long as they do not conflict with each other.

[0028] The solutions of this application embodiment are described in detail below from the training phase and the application phase, respectively.

[0029] I. Training Phase:

[0030] 1. Fine-tune the facial motion unit feature extraction model;

[0031] First, a video sample set is obtained, such as the open-source DISFA (Denver Intensity of Spontaneous Facial Action) dataset, a widely used public dataset for facial expression analysis, facial action unit detection, and intensity estimation. This video sample set includes videos of multiple facial images, with each person's video representing a single sample across all emotion categories. Each frame in each video sample is labeled with multiple facial action units and their corresponding intensity values. Then, this video sample set is used to fine-tune a pre-trained model to obtain a facial action unit feature extraction model. The pre-trained model can be the open-source GPT2 large model.

[0032] 2. Train the classification network;

[0033] Using a fine-tuned facial action unit feature extraction model, the intensity value temporal sequence of multiple facial action units is extracted frame by frame from the video sample set, resulting in a unit set consisting of multiple facial action units corresponding to the video sample. Specifically, for the 1st... Each action unit, whose intensity values ​​are arranged in chronological order, forms a one-dimensional temporal sequence. :

[0034] ;

[0035] in, Indicates the first The action unit in the first Intensity values ​​of each image frame, , This indicates the number of frames in the video sample.

[0036] Since the time interval from the peak expression frame to the action termination frame contains complete dynamic information of the gradual decline of emotion from its strongest state, it is a key temporal interval for distinguishing micro-expression categories. Therefore, the time sequence of each intensity value is used to determine the micro-expression category. Extract the valid subsequence from the start frame to the end frame. :

[0037] ;

[0038] in, Indicates the first The intensity value of each action unit in the initial frame. Indicates the first The intensity value of each action unit at the peak frame. Indicates the first The intensity value of each action unit at the termination frame. The start frame is the image frame at the moment when the micro-expression begins, the peak frame is the image frame at the moment when the micro-expression intensity reaches its peak, and the termination frame is the image frame at the moment when the micro-expression ends.

[0039] Due to the limited size of currently available micro-expression datasets and the need to use temporal data of facial action units extracted from videos for modeling, the number of training samples is relatively scarce. To alleviate the problem of sample scarcity, one embodiment of this application proposes data augmentation processing of effective sub-sequences based on the dynamic symmetry of micro-expressions to obtain expanded training samples. Specifically, this includes: extracting each effective sub-sequence separately. rising segment sequence With descending segment sequence The rising segment sequence is the valid subsequence from the start frame to the peak frame, and the falling segment sequence is the valid subsequence from the peak frame to the end frame. The falling segment sequence... Time reversal yields the reversed sequence Reverse the sequence With rising segment sequence As new effective subsequences for each facial action unit, an expanded training sample consisting of these new effective subsequences is obtained, meaning that each video sample in the original video sample set corresponds to two new effective subsequences. This embodiment of the application, without introducing external noise, fully utilizes the inherent symmetry of micro-expression dynamic changes to double the amount of training samples, without requiring additional computational burden.

[0040] Since the number of frames from the start to the peak and from the end to the peak differs among different samples, in some embodiments, linear interpolation is used to unify all sequences in the obtained extended training samples to a preset standard length. . The value of is determined based on the maximum number of frames in all sequences in the original extended training samples, that is, the maximum number of frames in all rising and falling sequences, to ensure that all input sequences have a consistent time dimension during subsequent processing.

[0041] To eliminate differences in facial motion unit intensity between individuals and to standardize the range of temporal dynamic changes, in some embodiments, differential scaling normalization is performed on the expanded training samples after standardization, including:

[0042] (1) Obtain the preset number of frames before the start frame from the intensity value time sequence. The intensity value is calculated, and the baseline mean of each facial motion unit is calculated based on the intensity value of a preset number of frames. For the first... Each action unit was used to calculate its baseline mean during the pre-emotional neutral phase. :

[0043] .

[0044] (2) Construct a difference sequence based on the effective subsequences in the expanded training samples and the baseline mean:

[0045] , ;

[0046] in, Indicates the first The difference sequence corresponding to the effective subsequence of each action unit Indicates the first The effective subsequence corresponding to the action unit is the first... The difference values ​​corresponding to each image frame Indicates the first The effective subsequence corresponding to the action unit is the first... Intensity values ​​of each image frame, .

[0047] (3) Perform min-max normalization on the obtained difference sequence to obtain the normalized difference ratio:

[0048] ;

[0049] in, express The corresponding normalized difference ratio, Represents the difference sequence The minimum value in, Represents the difference sequence The maximum value in, is the numerical stability constant.

[0050] (4) In order to preserve the original facial motion unit intensity information, the normalized difference ratio is multiplied element-wise with the original facial motion unit intensity value to obtain the final standardized sequence:

[0051] , ;

[0052] in, Indicates the first The final standardized sequence of each action unit, i.e., the effective subsequence that incorporates temporal information; express The normalized value.

[0053] This application's embodiments remove individual baseline offsets through differential calculation, effectively highlighting the temporal changes of facial motion feature units. Simultaneously, it maintains the relative weights of the original intensity distribution using normalized difference ratios, thereby suppressing intra-class variance while enhancing the ability to discriminate dynamic patterns related to micro-expressions.

[0054] During micro-expression analysis, environmental interference or non-expressional facial movements may cause the temporal changes of some facial action units to deviate from their typical patterns, thus introducing noise. To effectively identify and eliminate such easily disturbed facial action units, some embodiments propose a reliability assessment of facial action units based on the consistency of dominant trends. This includes: for each facial action unit in each emotion category in the extended training samples, determining the trend direction of its intensity value based on a sign function to obtain a trend feature vector, where the trend feature vector includes four trend patterns: increasing-increasing, increasing-decreasing, decreasing-increasing, and decreasing-decreasing. Based on the trend pattern of the trend feature vector, determining the dominant trend pattern of each facial action unit in each emotion category. Comparing the trend feature vector corresponding to each facial action unit in each emotion category with the dominant trend pattern, identifying samples with completely opposite patterns as noise samples, and calculating the global noise sensitivity score of each facial action unit based on the number of noise samples. Facial action units with a global noise sensitivity score exceeding a preset threshold are removed from the unit set. The specific implementation is as follows:

[0055] For the The first under the emotion category Action unit The intensity value is determined by comparing the relative changes in the average intensity over a set time period with the average intensity of its preceding and following action units, thus identifying its temporal trend. The calculation is performed on the 1st... In each sample Trend feature vector :

[0056] ;

[0057] in, Indicates the first In the nth sample The average intensity of each action unit within a set time period Indicates the first In the nth sample The average intensity of each action unit within the set time period. Indicates the first In the nth sample The average intensity of each action unit within the set time period. For symbolic functions, output Indicates enhancement (Intensify). This indicates a decrease (Diminish). From this, four trend patterns can be derived: Increase-Increase (II), Increase-Decrease (ID), Decrease-Increase (DI), and Decrease-Decrease (DD).

[0058] In the In the samples of emotion categories, the trend feature vector of each facial action unit is statistically analyzed, and the trend pattern that appears most frequently is taken as the first trend. Dominant trend patterns of facial movement units under emotion categories Samples whose trend patterns in the trend feature vectors are completely different from the dominant trend pattern are classified as noise samples. For example, If the dominant trend pattern is increasing (II), then samples with a decreasing (DD) trend pattern are noise samples. Samples with increasing (ID) or decreasing (DI) trend patterns are not completely different from the dominant trend pattern, and therefore are not noise samples.

[0059] According to the Number of noise samples under the emotion category Calculate the first Each facial movement unit in emotion category noise sensitivity ,in Indicates the first The number of all samples under the emotion category. Noise sensitivity reflects the facial movement unit's ability to express emotion. The proportion of severe trend anomalies is indicated by a higher value, which suggests that the facial movement unit is more susceptible to interference.

[0060] To comprehensively evaluate the stability of this facial action unit throughout the entire recognition task, its global noise sensitivity score is calculated:

[0061] ;

[0062] in, Indicates the first Global noise sensitivity score of each facial motion unit, total number of samples , This indicates the number of emotion categories. If... If the facial motion unit is deemed too sensitive to external interference and exhibits excessive inconsistency in its trend, it will be removed from the list. This indicates a preset threshold.

[0063] This application's embodiments, by focusing on anomalous samples that completely deviate from the dominant trend, can effectively capture drastic fluctuations in facial action units caused by sudden non-expression movements, thereby excluding unstable facial action units in feature selection and improving the robustness of subsequent micro-expression recognition.

[0064] After the above preprocessing and filtering, effective subsequences of facial action units retained in the unit set are obtained from the expanded training samples and used to form a one-dimensional temporal training feature, thus obtaining a concise and high-quality facial action unit feature set. Therefore, a micro-expression video sample is ultimately represented as two serial sequences, each with a length of... One-dimensional sequence set , , This represents the number of facial motion units remaining in the unit set after removing unstable facial motion units. This one-dimensional sequence can be represented as follows: Figure 1 As shown, this representation method maps the recognition task from a high-dimensional image space to a low-dimensional temporal signal analysis space with explicit physiological semantics. Furthermore, there is no overall temporal relationship; only the values ​​within each facial action unit have temporal relationships, which have already been processed and integrated into the data.

[0065] Because the data is one-dimensional, this embodiment can use a simple CNN (Convolutional Neural Network) as the classification network. For example, the classification network includes multiple cascaded one-dimensional convolutional layers for layer-by-layer feature extraction and mapping of the one-dimensional temporal training features. After global feature aggregation and dimensionality reduction via pooling layers, a linear mapping layer outputs a classification feature vector corresponding to the preset number of emotion categories to obtain the emotion category recognition result. The classification network is trained based on the one-dimensional temporal training features and the temporal sequence of the emotion categories corresponding to the acquired video samples until the network converges, resulting in a trained classification network.

[0066] 3. Loss function;

[0067] In this embodiment, cross-entropy loss is used as the final classification loss function during training. Cross-entropy is compared between the model's predictions and the true labels of the data. As the predictions become more accurate, the cross-entropy value decreases; if the predictions are completely correct, the cross-entropy value is 0. Cross-entropy loss It can be written as:

[0068] ;

[0069] in, The number of samples input to the classification network for training. Indicates the first The true label corresponding to each sample Indicates the first The predicted label for each sample.

[0070] II. Application Phase:

[0071] Figure 2 This is a schematic flowchart of a micro-expression recognition method according to an embodiment of this application. Figure 2As shown, the method includes: Step S201, acquiring the micro-expression video segment to be detected, using a trained facial action unit feature extraction model to extract the intensity value temporal sequence of multiple facial action units frame by frame, and obtaining a unit set composed of multiple facial action units corresponding to the micro-expression video segment to be detected; Step S202, extracting effective sub-sequences from each intensity value temporal sequence, and using linear interpolation to extend all effective sub-sequences to a preset length, such as the standard length preset during training. Step S203: Based on the preset list of unstable facial action units, remove unstable facial action units from the unit set; Step S204: Combine the effective subsequences corresponding to the facial action units in the unit set obtained after removal into a one-dimensional temporal feature, and input it into the trained classification network to obtain the emotion category recognition result.

[0072] In some embodiments, step S202, which involves extracting a valid subsequence from each intensity value time sequence, includes: extracting a valid subsequence from the start frame to the peak frame from each intensity value time sequence, or extracting a falling segment sequence from the peak frame to the end frame and then reversing the time of the falling segment sequence as a valid subsequence; wherein, the start frame is the image frame at the beginning of the micro-expression, the peak frame is the image frame at the peak of the micro-expression intensity, and the end frame is the image frame at the end of the micro-expression.

[0073] In some embodiments, step S202 further includes: performing differential scaling normalization on the effective subsequence to eliminate intensity differences of the same facial action unit among different individuals, including: obtaining the intensity values ​​of a preset number of frames before the starting frame from the intensity value temporal sequence, and calculating the baseline mean of each facial action unit based on the intensity values ​​of the preset number of frames; constructing a differential sequence based on the effective subsequence and the baseline mean, and performing min-max normalization on the differential sequence; multiplying the normalized result element-wise with the facial action unit intensity values ​​in the effective subsequence to obtain the effective subsequence with fused temporal information.

[0074] In some embodiments, the preset list of unstable facial motion units in step S203 is generated based on facial motion units whose global noise sensitivity scores exceed a preset threshold during the training phase.

[0075] Figure 3 This is a flowchart illustrating a micro-expression recognition method according to another embodiment of this application. Figure 3As shown, Dataset 1 is input into the open-source large model for fine-tuning. Then, Dataset 1 is input into the fine-tuned open-source large model for feature value extraction. The open-source large model outputs AU feature values, from which calm expression feature values ​​and micro-expression feature values ​​are extracted. The micro-expression feature values ​​are processed by data expansion to obtain bimodal feature values. After interpolation expansion, these are then interpolated and subtracted from the calm expression feature values ​​to generate a difference feature value sequence. This difference feature value sequence is then fused with the bimodal feature values ​​to obtain a standardized sequence with fused temporal information. The standardized sequence with fused temporal information is divided into category 1 to category 2 according to emotion category. For each category, calculate each facial motion unit separately. , … noise sensitivity, The number of facial motion units is represented. The noise sensitivity of the same facial motion unit under each emotion category is summarized across categories to obtain the corresponding global noise sensitivity score. The global noise sensitivity score of each facial motion unit is compared with a preset threshold. Based on the comparison result, facial motion unit features are selected and composed into one-dimensional feature data. The one-dimensional feature data is input into a lightweight temporal network for processing, and finally, the emotion category recognition result is output.

[0076] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.

[0077] Based on the same inventive concept, one embodiment of this application provides a micro-expression recognition system that integrates action unit trend analysis and sequence enhancement technology, including: a feature extraction module, an effective frame extraction module, a feature filtering module, and a recognition module. The feature extraction module is configured to acquire a micro-expression video segment to be detected, and using a trained facial action unit feature extraction model, extract the temporal sequence of intensity values ​​of multiple facial action units frame by frame, obtaining a unit set corresponding to the micro-expression video segment to be detected, composed of multiple facial action units. The effective frame extraction module is configured to extract effective sub-sequences from each intensity value temporal sequence, and use linear interpolation to extend all effective sub-sequences to a preset length. The feature filtering module is configured to remove unstable facial action units from the unit set according to a preset list of unstable facial action units. The recognition module is configured to form a one-dimensional temporal feature from the effective sub-sequences corresponding to the facial action units in the unit set obtained after removal, and input it into a trained classification network to obtain an emotion category recognition result. Since the system embodiment is basically similar to the method embodiment, the details of the relevant technical features and the effects of implementation can be found in the corresponding descriptions of the method embodiment provided above. In the embodiments provided in this application, it should be understood that the division of modules (or units) is merely a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection between modules can be through some interfaces, or indirect coupling or communication connection between modules, or it can be an electrical, mechanical, or other form of connection. Additionally, the functional modules in the various embodiments of this application can be integrated into one module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above modules can be implemented in hardware or as software functions.

[0078] The embodiments provided in this application can also be computer program products. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this application. A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.

[0079] This application uses specific embodiments to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the solution and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A micro-expression recognition method integrating action unit trend analysis and sequence enhancement technology, characterized in that, include: Step 1: Obtain the micro-expression video segment to be detected, and use the trained facial action unit feature extraction model to extract the intensity value temporal sequence of multiple facial action units frame by frame, and obtain a unit set consisting of multiple facial action units corresponding to the micro-expression video segment to be detected. Step 2: Extract effective subsequences from the time series of each intensity value, and use linear interpolation to extend all effective subsequences to a preset length; Step 3: Remove unstable facial motion units from the set of units according to the preset list of unstable facial motion units; Step 4: Combine the effective subsequences corresponding to the facial action units in the unit set obtained after removal into a one-dimensional temporal feature, and input it into the trained classification network to obtain the emotion category recognition result.

2. The method according to claim 1, characterized in that, Step 2 involves extracting valid subsequences from the time series of each intensity value, including: Extract the effective subsequence from the start frame to the peak frame from the time sequence of each intensity value, or extract the falling segment sequence from the peak frame to the end frame and reverse the time of the falling segment sequence as the effective subsequence; The start frame is the image frame at the moment the micro-expression begins, the peak frame is the image frame at the moment the micro-expression intensity reaches its peak, and the end frame is the image frame at the moment the micro-expression ends.

3. The method according to claim 1, characterized in that, Step 2 further includes: performing differential scaling normalization on the effective subsequence according to the following steps to eliminate intensity differences of the same facial action unit between different individuals: The intensity values ​​of a preset number of frames prior to the starting frame are obtained from the intensity value time sequence, and the baseline mean of each facial motion unit is calculated based on the intensity values ​​of the preset number of frames. The starting frame is the image frame at the moment when the micro-expression begins. A difference sequence is constructed based on the effective subsequence and the baseline mean, and the difference sequence is subjected to min-max normalization. The normalized result is multiplied element-wise with the facial motion unit intensity value in the effective subsequence to obtain the effective subsequence with fused temporal information.

4. The method according to claim 1, characterized in that, The facial action unit feature extraction model described in step 1 is trained according to the following process: Obtain a video sample set, which includes videos of multiple facial images, and each image frame in each video is labeled with multiple facial action units and their intensity values; The pre-trained model is fine-tuned using the video sample set to obtain a facial motion unit feature extraction model.

5. The method according to claim 4, characterized in that, The classification network described in step 4 is trained according to the following process: Using a finely tuned facial action unit feature extraction model, the intensity value temporal sequence of multiple facial action units is extracted frame by frame from the video sample set, and a unit set consisting of multiple facial action units corresponding to the video sample is obtained. Extract the valid subsequence from the start frame to the end frame from the temporal sequence of each intensity value; Data augmentation processing is performed on the effective subsequences based on the dynamic symmetry of micro-expressions to obtain expanded training samples; The extended training samples are subjected to length unification and differential scaling standardization to obtain the processed extended training samples. Based on a preset trend consistency screening rule, unstable facial motion units are removed from the unit set; Effective subsequences of facial action units retained in the unit set are obtained from the expanded training samples and formed into one-dimensional temporal training features; Obtain the temporal sequence of the emotion category corresponding to the video sample; The classification network is trained based on the one-dimensional temporal training features and the temporal sequence of emotion categories until the network converges, resulting in a well-trained classification network.

6. The method according to claim 5, characterized in that, Based on the dynamic symmetry of micro-expressions, the effective subsequences are augmented to obtain expanded training samples, including: The rising segment sequence and falling segment sequence of each effective subsequence are extracted respectively. The rising segment sequence is the effective subsequence from the start frame to the peak frame, and the falling segment sequence is the effective subsequence from the peak frame to the end frame. The peak frame is the image frame at the peak of the micro-expression intensity. The descending segment sequence is time-reversed to obtain the inverted sequence; The inverted sequence and the rising sequence are used as new effective sub-sequences for each facial action unit to obtain an expanded training sample composed of the new effective sub-sequences.

7. The method according to claim 5, characterized in that, Based on a preset trend consistency screening rule, unstable facial motion units are removed from the unit set, including: For each facial action unit under each emotion category in the extended training samples, the trend direction of its intensity value is determined based on the sign function to obtain the trend feature vector, which includes four trend patterns: increase-increase, increase-decrease, decrease-increase, and decrease-decrease. Based on the trend patterns of the trend feature vectors, determine the dominant trend patterns of each facial action unit under each emotion category. The trend feature vector corresponding to each facial action unit under each emotion category is compared with the dominant trend pattern. Samples with completely opposite trends are identified as noise samples, and the global noise sensitivity score of each facial action unit is calculated based on the number of noise samples. Facial motion units whose global noise sensitivity scores exceed a preset threshold are removed from the unit set.

8. The method according to claim 7, characterized in that, A list of eliminated facial motion units is generated based on the global noise sensitivity score exceeding a preset threshold, and this list of eliminated facial motion units is used as the preset list of unstable facial motion units in step 3.

9. A micro-expression recognition system integrating action unit trend analysis and sequence enhancement technology, characterized in that, include: The module comprises a feature extraction module, a valid frame extraction module, a feature filtering module, and a recognition module; among which: The feature extraction module is configured to acquire the micro-expression video segment to be detected, and use the trained facial action unit feature extraction model to extract the intensity value temporal sequence of multiple facial action units frame by frame, and obtain a unit set consisting of multiple facial action units corresponding to the micro-expression video segment to be detected. The effective frame extraction module is configured to extract effective subsequences from the time sequence of each intensity value and use linear interpolation to extend all effective subsequences to a preset length; The feature filtering module is configured to remove unstable facial motion units from the unit set based on a preset list of unstable facial motion units; The recognition module is configured to form a one-dimensional temporal feature from the effective subsequences corresponding to facial action units in the unit set obtained after removal, and input it into the trained classification network to obtain the emotion category recognition result.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.