Method, system, device and medium for magnifying micro-expressions to macro-expressions

Through the image migration model FOMM and dual attention mechanism, the image distortion problem in the nonlinear amplification task of micro-expression is solved, and the approximate visual effect from micro-expression to macro-expression is achieved, and the feature extraction effect of micro-expression recognition is improved.

CN115116112BActive Publication Date: 2025-08-19SOUTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210740101.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-08-19
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

In the prior art, linear amplification method is not suitable for nonlinear amplification tasks of micro-expressions, and artificial adjustment of the magnification ratio can easily lead to image distortion and distortion. The existing deep learning methods have limited effect on micro-expressions.

Method used

The image migration model FOMM is adopted, combining the motion estimation module, image generation module and motion amplification module, and by introducing a dual attention mechanism, nonlinear amplification from micro-expression to macro-expression is achieved, and a macro-expression data set is used for training, and a motion amplification module is added to realize image migration.

Benefits of technology

Nonlinear amplification of micro-expressions is achieved, and the generated image is approaching macro-expressions, with good visual effects and no distortion, which improves the feature extraction effect of micro-expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116112B_ABST
    Figure CN115116112B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of micro-expressions and discloses a method, system, device and medium for magnifying micro-expressions into macro-expressions. The method comprises extracting an initial frame and an intermediate frame from a macro-expression sequence to simulate the facial movement of the micro-expression, and using the intermediate frame of the macro-expression as the top frame of the micro-expression; selecting an excellent image migration model FOMM, the image migration model FOMM comprising a motion estimation module and an image generation module; the FOMM network takes a source image and a driving frame as input, so that an object in the source image generates a new image according to the action in the driving frame; the FOMM network is trained according to a given macro-expression sequence data set, so that the network grasps the characteristics of macro-expression changes; and the motion magnification module is added between the motion estimation module and the image generation module to realize the micro-expression magnification function based on image migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of micro-expression technology, and in particular relates to a method, system, device and medium for amplifying micro-expressions into macro-expressions. Background Art

[0002] Microexpression recognition, a crucial task in computer vision, has reached a bottleneck in recent years. Researchers have found that microexpression upscaling can aid microexpression recognition by magnifying facial features, facilitating feature extraction. In recent years, linear Euler video upscaling algorithms and phase-based video upscaling methods have been commonly applied to microexpression upscaling. Since the introduction of deep learning upscaling methods, they have also been applied to microexpression upscaling, achieving superior results compared to the previous two methods. However, these methods mostly rely on preprocessing of microexpression images, and the upscaling results are limited in terms of both visual quality and accuracy. Current deep learning-based video upscaling methods are not specifically designed for microexpression upscaling. Instead, they adjust the upscaling factor and then multiply the difference between two frames to perform upscaling. However, this linear upscaling approach is not suitable for nonlinear upscaling tasks like microexpressions, and artificially adjusting the upscaling factor often results in significant distortion of the upscaled image. Therefore, these methods can only perform small-scale upscaling preprocessing on microexpressions, resulting in limited upscaling results. In summary, how to achieve nonlinear amplification of micro-expressions based on deep learning while ensuring that the image is not distorted has become an important issue in the field of micro-expressions.

[0003] Through the above analysis, the problems and defects of the existing technology are: the linear magnification in the existing technology is not suitable for the nonlinear magnification task of micro-expressions, and artificially adjusting the magnification factor for the micro-expression magnification task often causes serious distortion of the magnified image. Summary of the Invention

[0004] In response to the problems existing in the prior art, the present invention provides a method, system, device and medium for amplifying micro-expressions to macro-expressions.

[0005] The present invention is implemented as follows: a method for amplifying a micro-expression to a macro-expression, the method comprising:

[0006] The initial frame and intermediate frames are extracted from the macro-expression sequence to simulate the facial movements of micro-expressions, and the intermediate frames of the macro-expressions are used as the top frames of the micro-expressions; an excellent image migration model FOMM is selected, which includes a motion estimation module and an image generation module; the FOMM network takes the source image and the driving frame as input, so that the object in the source image generates a new image according to the action in the driving frame; the FOMM network is trained based on a given macro-expression sequence dataset so that the network can grasp the characteristics of macro-expression changes; the initial frame in the given macro-expression sequence is used as the source image, and the other frames are used as driving frames; the dataset includes MMI and CK+ macro-expression datasets; the motion amplification module is added between the motion estimation module and the image generation module to realize the micro-expression amplification function based on image migration.

[0007] Furthermore, the specific process of the method of amplifying micro-expressions to macro-expressions is as follows:

[0008] Step 1: Based on a given pre-trained motion estimation module, the initial frame and intermediate frame as well as the initial frame and top frame of the macro expression sequence are taken as input, and the micro expression feature map and the macro expression feature map are output respectively;

[0009] Step 2: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into a macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image;

[0010] Step three: Add a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to realize the micro-expression amplification function based on image transfer.

[0011] Furthermore, in step 3, a motion amplification module is added between the pre-trained motion estimation module and the pre-trained image generation module to implement the micro-expression amplification function based on image migration. The specific process is as follows:

[0012] The motion amplification module is an encoder-decoder structure. The first half is feature extraction, and the second half is upsampling. The motion amplification module is used for skip connections for multi-scale feature fusion. The span of the amplification process from micro-expressions to macro-expressions is too large, so a dual attention mechanism is introduced.

[0013] Furthermore, the amplification process from micro-expressions to macro-expressions spans too long, so the dual attention mechanism is introduced. The specific process is as follows:

[0014] The attention module calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain better amplification effect.

[0015] Furthermore, the specific process of obtaining a better amplification effect is as follows:

[0016] The feature map in the attention module is fed into three different convolutional layers as input and generates feature maps of new dimensions; the three different convolutional layers are value_conv, query_conv and key_conv; the two feature maps output by the query_conv and key_conv convolutional layers are reshaped, multiplied and weighted normalized using softmax to obtain the attention map; the attention map and the feature map output by the value_conv convolutional layer are finally output as a new feature map through a series of mathematical operations.

[0017] Furthermore, the upsampling is provided with two upsampling layers with scales of 128×128 and 64×64 respectively, and an attention module is added after the two upsampling layers.

[0018] Another object of the present invention is to provide a system for magnifying micro-expressions into macro-expressions for implementing the method for magnifying micro-expressions into macro-expressions, wherein the system comprises:

[0019] The pre-trained motion estimation module takes the initial and intermediate frames, as well as the initial and top frames, of the macro-expression sequence as input, and outputs micro-expression feature maps and macro-expression feature maps, respectively;

[0020] The pre-trained image generation module transforms the generated macro expression feature map into the final enlarged image;

[0021] The motion magnification module performs feature extraction and upsampling through an encoder-decoder structure, and uses skip connections for multi-scale feature fusion.

[0022] Furthermore, the motion amplification module is provided with an attention module, which calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain a better amplification effect.

[0023] Another object of the present invention is to provide a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps:

[0024] Step 1: Based on a given pre-trained motion estimation module, the initial frame and intermediate frame as well as the initial frame and top frame of the macro expression sequence are taken as input, and the micro expression feature map and the macro expression feature map are output respectively;

[0025] Step 2: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into a macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image;

[0026] Step three: Add a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to realize the micro-expression amplification function based on image transfer.

[0027] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0028] Step 1: Based on a given pre-trained motion estimation module, the initial frame and intermediate frame as well as the initial frame and top frame of the macro expression sequence are taken as input, and the micro expression feature map and the macro expression feature map are output respectively;

[0029] Step 2: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into a macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image;

[0030] Step three: Add a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to realize the micro-expression amplification function based on image transfer.

[0031] In combination with the above technical solutions and the technical problems solved, please analyze the advantages and positive effects of the technical solutions to be protected by the present invention from the following aspects:

[0032] In view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, this paper closely combines the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyzes in detail and in depth how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:

[0033] Micro-expressions are low-amplitude, incomplete macro-expressions with similar facial movement tendencies as macro-expressions. At the same time, since the micro-expression dataset does not have enough samples, for the above reasons, the present invention applies the macro-expression dataset to the micro-expression amplification task. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a flow chart of a method for magnifying micro-expressions into macro-expressions provided by an embodiment of the present invention;

[0035] Figure 2 2. It is a schematic diagram of the system structure for amplifying micro-expressions to macro-expressions provided by an embodiment of the present invention;

[0036] Figure 3 This is a schematic diagram of the principle of amplifying micro-expressions to macro-expressions using deep transfer learning provided by an embodiment of the present invention;

[0037] Figure 4 The embodiment of the present invention provides an input of the initial frame and the top frame of the micro-expression, and outputs a schematic diagram of an enlarged image;

[0038] Figure 5 This is a schematic diagram showing some samples after testing on three micro-expression datasets, CASME II, SAMM, and SMIC, provided by an embodiment of the present invention;

[0039] Figure 5 Middle: Figure a, top frame of micro-expression; Figure b, enlarged image;

[0040] In the figure: 1. Pre-trained motion estimation module; 2. Pre-trained image generation module; 3. Motion amplification module; 4. Attention module. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] 1. Explanatory Examples In order to enable those skilled in the art to fully understand how to implement the present invention, this section provides an illustrative example that expands upon the technical solutions of the claims.

[0043] like Figure 1 As shown, the method for magnifying a micro-expression to a macro-expression provided by an embodiment of the present invention includes:

[0044] S101: Based on a given pre-trained motion estimation module, the initial frame and the middle frame as well as the initial frame and the top frame of the macro expression sequence are taken as input, and a micro expression feature map and a macro expression feature map are output respectively.

[0045] S102: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into the macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image.

[0046] S103: adding a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to implement a micro-expression amplification function based on image migration.

[0047] In S103 provided in the embodiment of the present invention, the motion amplification module is similar to U-Net, which is an encoder-decoder structure. The first half is feature extraction and the second half is upsampling. The module also has skip connections for multi-scale feature fusion. Considering that the span of the amplification process from micro-expressions to macro-expressions is too large, a dual attention mechanism is introduced. The attention module can calculate the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain a better amplification effect. The attention module is added after the two upsampling layers with scales of 128×128 and 64×64 respectively.

[0048] The feature map in the attention module is fed into three different convolutional layers (value_conv, query_conv, and key_conv) to generate a new feature map. The two feature maps output by the query_conv and key_conv convolutional layers are reshaped, multiplied, and weighted normalized using softmax to create the attention map. A series of mathematical operations are performed on the attention map and the feature map output by the value_conv convolutional layer to ultimately output a new feature map.

[0049] like Figure 2 As shown, the system for magnifying micro-expressions to macro-expressions provided by an embodiment of the present invention includes:

[0050] The pre-trained motion estimation module 1 takes the initial frame and the middle frame as well as the initial frame and the top frame of the macro expression sequence as input, and outputs a micro expression feature map and a macro expression feature map respectively.

[0051] The pre-trained image generation module 2 transforms the generated macro expression feature map into the final enlarged image.

[0052] The motion magnification module 3 performs feature extraction and upsampling through an encoder-decoder structure and uses skip connections for multi-scale feature fusion.

[0053] Attention module 4 calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain better amplification effect.

[0054] 2. Evidence of the Effects of the Embodiments The embodiments of the present invention have achieved some positive effects during the development or use process, and indeed have great advantages over the prior art. The following content describes them with reference to data, charts, etc. during the test process.

[0055] In terms of visual effects, the amplified effect of micro-expressions is close to that of macro-expressions. Figure 5As shown in the figure, some samples of the present invention after testing on three micro-expression datasets: CASME II, SAMM and SMIC. Figure 5 The top frame of Figure a is the top frame of the micro-expression. Figure 5 The middle image (b) is an enlarged image. From a visual perspective, the enlarged image is close to the macro expression.

[0056] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0057] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A method for magnifying micro-expressions to macro-expressions, characterized in that: The method of amplifying a micro-expression to a macro-expression includes: Extract the initial frame and intermediate frames from the macro-expression sequence to simulate the facial movement of micro-expressions, and use the intermediate frame of the macro-expression as the top frame of the micro-expression; select an excellent image migration model FOMM, which includes a motion estimation module and an image generation module; the FOMM network takes the source image and the driving frame as input, so that the object in the source image generates a new image based on the action in the driving frame; the FOMM network is trained based on a given macro-expression sequence dataset so that the network can grasp the characteristics of macro-expression changes; the initial frame in the given macro-expression sequence is used as the source image, and the other frames are used as driving frames; the dataset includes the MMI and CK+ macro-expression datasets; the motion amplification module is added between the motion estimation module and the image generation module to realize the micro-expression amplification function based on image migration; The specific process of the method of amplifying micro-expressions to macro-expressions is as follows: Step 1: Based on a given pre-trained motion estimation module, the initial frame and intermediate frame as well as the initial frame and top frame of the macro expression sequence are taken as input, and the micro expression feature map and the macro expression feature map are output respectively; Step 2: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into a macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image; Step 3: Add a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to achieve the micro-expression amplification function based on image transfer; In step 3, a motion amplification module is added between the pre-trained motion estimation module and the pre-trained image generation module to implement the micro-expression amplification function based on image migration. The specific process is as follows: The motion amplification module uses an encoder-decoder structure, with the first half performing feature extraction and the second half performing upsampling. The motion amplification module uses skip connections for multi-scale feature fusion. The amplification process from micro-expressions to macro-expressions spans a large range, so a dual attention mechanism is introduced. The process of amplifying micro-expressions to macro-expressions is too long, so the dual attention mechanism is introduced as follows: The attention module calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain better amplification effect; The specific process of obtaining a better magnification effect is as follows: The feature map in the attention module is fed into three different convolutional layers as input and generates feature maps of new dimensions; the three different convolutional layers are value_conv, query_conv and key_conv; the two feature maps output by the query_conv and key_conv convolutional layers are reshaped, multiplied and weighted normalized using softmax to obtain the attention map; the attention map and the feature map output by the value_conv convolutional layer are finally output as a new feature map through a series of mathematical operations.

2. The method for magnifying micro-expressions to macro-expressions according to claim 1, wherein: The upsampling is set up with two upsampling layers with scales of 128×128 and 64×64 respectively, and an attention module is added after the two upsampling layers.

3. A system for amplifying micro-expressions to macro-expressions, characterized in that: The system for amplifying micro-expressions to macro-expressions includes: The pre-trained motion estimation module takes the initial and intermediate frames, as well as the initial and top frames, of the macro-expression sequence as input, and outputs micro-expression feature maps and macro-expression feature maps, respectively; The pre-trained image generation module transforms the generated macro expression feature map into the final enlarged image; Motion amplification module, which performs feature extraction and upsampling through an encoder-decoder structure and skip connections for multi-scale feature fusion; The method of amplifying a micro-expression to a macro-expression includes: Extract the initial frame and intermediate frames from the macro-expression sequence to simulate the facial movement of micro-expressions, and use the intermediate frame of the macro-expression as the top frame of the micro-expression; select an excellent image migration model FOMM, which includes a motion estimation module and an image generation module; the FOMM network takes the source image and the driving frame as input, so that the object in the source image generates a new image based on the action in the driving frame; the FOMM network is trained based on a given macro-expression sequence dataset so that the network can grasp the characteristics of macro-expression changes; the initial frame in the given macro-expression sequence is used as the source image, and the other frames are used as driving frames; the dataset includes the MMI and CK+ macro-expression datasets; the motion amplification module is added between the motion estimation module and the image generation module to realize the micro-expression amplification function based on image migration; The specific process of the method of amplifying micro-expressions to macro-expressions is as follows: Step 1: Based on a given pre-trained motion estimation module, the initial frame and intermediate frame as well as the initial frame and top frame of the macro expression sequence are taken as input, and the micro expression feature map and the macro expression feature map are output respectively; Step 2: Under the guidance of the macro expression feature map, the micro expression feature map is input into the motion amplification module to train how to transform into a macro expression feature map; the generated macro expression feature map is input into the pre-trained image generation module to generate the final amplified image; Step 3: Add a motion amplification module between the pre-trained motion estimation module and the pre-trained image generation module to achieve the micro-expression amplification function based on image transfer; In step 3, a motion amplification module is added between the pre-trained motion estimation module and the pre-trained image generation module to implement the micro-expression amplification function based on image migration. The specific process is as follows: The motion amplification module uses an encoder-decoder structure, with the first half performing feature extraction and the second half performing upsampling. The motion amplification module uses skip connections for multi-scale feature fusion. The amplification process from micro-expressions to macro-expressions spans a large range, so a dual attention mechanism is introduced. The process of amplifying micro-expressions to macro-expressions is too long, so the dual attention mechanism is introduced as follows: The attention module calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain better amplification effect; The specific process of obtaining a better magnification effect is as follows: The feature map in the attention module is fed into three different convolutional layers as input and generates feature maps of new dimensions; the three different convolutional layers are value_conv, query_conv and key_conv; the two feature maps output by the query_conv and key_conv convolutional layers are reshaped, multiplied and weighted normalized using softmax to obtain the attention map; the attention map and the feature map output by the value_conv convolutional layer are finally output as a new feature map through a series of mathematical operations.

4. The system for magnifying micro-expressions to macro-expressions as claimed in claim 3, characterized in that: The motion amplification module is provided with an attention module, which calculates the response of a certain position as the weighted sum of all features from different spatial positions, thereby connecting the long-term dependency and nonlinear transformation information of any two positions in the feature map to obtain a better amplification effect.

5. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the method of magnifying micro-expressions to macro-expressions according to any one of claims 1 to 2. 6 . A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the method for magnifying a micro-expression to a macro-expression according to claim 1 .

Citation Information

Patent Citations

  • Micro-expression recognition method based on adaptive motion amplification and convolutional neural network

    CN113537008A

  • Micro-expression feature extraction and recognition method based on deep learning

    CN114220154A