Method and device for realizing micro-expression recognition based on dynamic weak texture information adaptive amplification, processor and readable storage medium thereof
The method addresses the challenge of weak dynamic texture information in micro-expression recognition by using adaptive enhancement techniques, improving accuracy and reliability in identifying subtle facial movements.
Patent Information
- Application Number
- CN202510448470.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-15
AI Technical Summary
Existing micro-expression recognition methods are difficult to accurately extract facial tiny movement information, resulting in insufficient recognition accuracy, especially when facial texture characteristics change are weak.
Adaptive amplification method based on dynamic weak texture information is adopted, dynamic areas are identified through differential algorithms, dynamic masks are constructed for amplification, and combined with global enhancement networks and local attention feature extraction networks, and image features are captured and fused for classification.
It improves the accuracy and flexibility of micro-expression recognition, reduces noise interference, and can adaptively select appropriate amplification factors in different video sequences, enhancing the extraction ability of global and local features.
Smart Images

Figure CN120318882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and pattern recognition, in particular to the technical field of micro-expression recognition, and specifically refers to a method, device, processor and computer-readable storage medium for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information. Background Technique
[0002] Facial micro-expressions are an important way for humans to express emotions and are prevalent in various social interactions. At the same time, facial micro-expressions are closely related to various diseases. They can not only serve as an external manifestation of diseases but also may affect the disease process and the psychological state of patients. Subtle facial expression changes have important reference value, which helps in the early identification of diseases, thus providing support for timely treatment. In recent years, with the development of artificial intelligence technology and the continuous breakthroughs in computer vision and pattern recognition technology in the field of micro-expression analysis, intelligent real-time and fast micro-expression recognition has become a research field that is both challenging and valuable.
[0003] In the prior art, currently, the relatively mature micro-expression recognition methods are mainly divided into two categories: micro-expression recognition methods based on facial feature points and micro-expression recognition methods based on facial texture features. The former mainly uses local feature points on the face for recognition, and the latter mainly recognizes according to the changing trend of facial texture features. The micro-expression recognition method based on feature points provides a relatively complete implementation framework for micro-expression recognition research. However, it is easily affected by the selection of feature points and has strong limitations. The micro-expression recognition method based on facial texture features can capture the dynamic changes of the face through the global features of the face and can recognize micro-expressions through the dynamic change information of the face. However, it is difficult for the prior art to accurately extract the changes in the minute motion information of the face, which affects the accuracy of micro-expression recognition. Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide a method, device, processor and computer-readable storage medium for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information, which has high accuracy, small error and a relatively wide application range.
[0005] In order to achieve the above purpose, the method, device, processor and computer-readable storage medium for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention are as follows:
[0006] The method for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information is mainly characterized in that the method includes the following steps:
[0007] (1) Collect the information of the facial micro-expression video segment of the tester by the visual system and store it in the facial micro-expression data storage memory bank;
[0008] (2) Extract the micro-expression emotional video clips in the face micro-expression data storage repository, and amplify the micro-expression video information by a weak signal amplification method based on the amplification of dynamic weak texture information; and compare the high-frequency information of the amplified video with the high-frequency information in the original video to obtain the high-frequency increment ratio value; feedback-adjust the amplification factor through the high-frequency increment ratio value;
[0009] (3) Capture and amplify image features by constructing two parallel branches of a global enhancement network and a local attention feature extraction network;
[0010] (4) Fuse the global enhancement features and the local attention features, and classify the micro-expressions through the fused features.
[0011] Preferably, in the step (2), when extracting the micro-expression emotional video clips in the face micro-expression data storage repository and amplifying the micro-expression video information by a weak signal amplification method based on the amplification of dynamic weak texture information, specifically:
[0012] Identify the dynamic part in the image through a differential algorithm, construct a dynamic region mask, and amplify the weak texture information in the dynamic mask region; calculate the brightness value of the pixel point (x, y) in the enhanced image at time t, specifically:
[0013] Calculate the brightness value of the pixel point (x, y) in the enhanced image at time t according to the following formula:
[0014]
[0015] where, I(x, y, t) represents the brightness value of the pixel point (x, y) at time t, k represents the level of the pyramid. I k (x, y, t) represents the pixel point of the k-th layer image at time t. α is the amplification factor, F and F -1 represent the Fourier transform and the inverse transform respectively, H(f) is a band-pass filter. M(x, y, t) is the mask of the dynamic region. I en (x, y, t) represents the brightness value of the pixel point (x, y) in the enhanced image at time t.
[0016] Preferably, in the step (2), when comparing the high-frequency information of the amplified video with the high-frequency information in the original video to obtain the high-frequency increment ratio value, specifically:
[0017] Obtain the high-frequency increment ratio value according to the following formula:
[0018]
[0019] where u and v are respectively spatial frequency components, w is frequency information, and H space is a spatial high-pass filter, and H time (w) is a temporal high-frequency filter, F(u, v, w) is the frequency-domain information of the original video after Fourier transform, and F amplified (u, v, w) is the frequency-domain information of the magnified video after Fourier transform, ∈ is a constant with a value of 10 -8 , and α is the magnification factor.
[0020] Preferably, in step (2), the magnification factor is feedback-regulated by the high-frequency increment ratio value, which specifically includes:
[0021] The updated value of the magnification factor is calculated iteratively to gradually converge the loss function to the optimal solution, specifically as follows:
[0022] The updated value of the magnification factor is calculated according to the following formula:
[0023]
[0024] where α t is the magnification factor at time t, α t+1 is the magnification factor at time t + 1, α t-1 is the magnification factor at time t - 1, E th is the increment ratio threshold constant, β is the learning rate, R(α t ) is the increment ratio at time t, and R(α t-1 ) is the increment ratio at time t - 1.
[0025] Preferably, the global enhancement network in step (3) is used to extract the global enhancement features of micro-expressions, which specifically includes the following steps:
[0026] The receptive field of the channel sub-block is gradually expanded in two directions using a symmetric multi-scale strategy, and the global representation semantics is used to guide the channel sub-block to obtain multi-scale channel enhancement features; the two multi-scale channel features are concatenated along the channel dimension and dimension-reduced to obtain enhanced global channel flow semantics to enhance the extraction ability of the global network features.
[0027] Preferably, the obtaining of the global channel flow semantic information is specifically as follows:
[0028] The global channel flow semantic information is obtained according to the following formula:
[0029]
[0030] where is the multi-scale feature of the i-th channel sub-block calculated from left to right, The multi-scale feature of the i-th channel sub-block calculated from right to left, G is the global representation semantics, and F i ′ is the shuffled feature of the i-th channel, and f 3×3 (·) is a convolution operation with a convolution kernel of 3, and Concat C (·) is a feature concatenation operation in the channel dimension, and f 1×1 (·) is a 1×1 convolution, M is to group the channel shuffled features into M sub-blocks by channel, and ω k is each feed-forward feature and of represents the weighted sum of the cumulative left-side feature information from level 1 to i - 1, represents the weighted sum of the cumulative right-side feature information from i + 1 to M levels.
[0031] Preferably, the local attention feature extraction network in step (3) is used to obtain the local attention features of micro-expressions, and specifically includes the following steps:
[0032] Divide the micro-expression image into four non-overlapping local feature maps F ∈ R H×W×C , perform two 3×3 convolution operations on each local feature map, use the attention module as the attention feature for extracting local information, perform an element-wise multiplication operation on the attention map and the input feature map, and adaptively refine the features.
[0033] Preferably, the local attention feature extraction network in step (3) is specifically:
[0034] Obtain the output of the local attention feature extraction network according to the following formula:
[0035]
[0036] where F is the local feature map after division, and f 3×3 (·) represents a convolution operation with a convolution kernel of 3, and f 1×1 (·) represents a 1×1 convolution, represents an element-wise multiplication operation, represents an element-wise addition operation, and M c represents ChannelAttention calculation, and M s represents Spatial Attention calculation, and f down (·) represents a downsampling operation, and F r is the obtained local attention feature.
[0037] Preferably, step (4) specifically includes the following steps:
[0038] The global enhanced features and local attention features are concatenated and fused, and the fused feature vector information is used as the input of the fully connected neural network. The fully connected neural network minimizes the objective function by iteratively updating the model parameters, making the value of the objective function continuously decrease. The fully connected neural network learns the relationship between the input data and the output labels by continuously adjusting the weights and biases, and classifies the newly input facial microexpressions using the weight relationship between the input data and the labels.
[0039] The device for micro-expression recognition based on adaptive amplification of dynamic weak texture information is mainly characterized in that the device includes:
[0040] A processor configured to execute computer-executable instructions;
[0041] A memory storing one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, each step of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information as described above is implemented.
[0042] The processor for micro-expression recognition based on adaptive amplification of dynamic weak texture information is mainly characterized in that the processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, each step of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information as described above is implemented.
[0043] The computer-readable storage medium is mainly characterized in that a computer program is stored thereon, and the computer program can be executed by a processor to implement each step of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information as described above.
[0044] By adopting the method, device, processor and computer-readable storage medium for micro-expression recognition based on adaptive amplification of dynamic weak texture information, the dynamic part in the image is recognized through differential technology, a dynamic region mask is constructed, and the amplification of signals outside the mask region is suppressed, effectively reducing the noise in the amplified image. This method has a flexible structure and stronger learning ability and higher accuracy compared with the original Eulerian video magnification algorithm. The present invention can accurately and quickly recognize the facial micro-expressions of the tester. It can be used as a clue for diagnosing certain diseases of the tester and helping the tester recognize some psychological and mental diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic flowchart of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention.
[0046] Figure 2 It is a schematic framework diagram of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention.
[0047] Figure 3 Flowchart of the learning method of the gradient descent method for the method of realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention.
[0048] Figure 4 Schematic diagram of the local attention extraction module for the method of realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention.
[0049] Figure 5 Schematic diagram of amplification based on dynamic weak texture information for the method of realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention.
[0050] Figure 6 Schematic diagram of the fixed threshold method and the method of the present invention for amplification for the method of realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention. Detailed implementation manner
[0051] In order to be able to more clearly describe the technical content of the present invention, the following will be further described in conjunction with specific embodiments. The method of the present invention for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information includes the following steps:
[0052] (1) Collect the facial micro-expression video segment information of the tester according to the visual system and store it in the facial micro-expression data storage memory bank;
[0053] (2) Extract the micro-expression emotion video segment in the facial micro-expression data storage memory bank, and amplify the micro-expression video information by the weak signal amplification method based on amplification of dynamic weak texture information; and compare the high-frequency information of the amplified video with the high-frequency information in the original video to obtain the high-frequency increment ratio value; feedback-adjust the amplification factor through the high-frequency increment ratio value;
[0054] (3) Capture and amplify the image features by constructing two parallel branches of a global enhancement network and a local attention feature extraction network;
[0055] (4) Fuse the global enhancement feature and the local attention feature, and classify the micro-expression through the fused feature.
[0056] As a preferred implementation manner of the present invention, in the step (2), the micro-expression emotion video segment in the facial micro-expression data storage memory bank is extracted, and the micro-expression video information is amplified by the weak signal amplification method based on amplification of dynamic weak texture information, specifically as follows:
[0057] Identify the dynamic part in the image through the differential algorithm, construct a dynamic region mask, and amplify the weak texture information within the dynamic mask region; calculate the brightness value of the pixel point (x, y) in the image at time t after enhancement, specifically as follows:
[0058] Calculate the brightness value of the pixel point (x, y) in the image at time t after enhancement according to the following formula:
[0059]
[0060] Among them, I(x, y, t) represents the brightness value of the pixel point (x, y) at time t, and k represents the level of the pyramid. I k (x, y, t) represents the pixel point of the k-th layer image at time t. α is the amplification factor, F and F -1 respectively represent the Fourier transform and the inverse transform, and H(f) is the band-pass filter. M(x, y, t) is the mask of the dynamic region. I en (x, y, t) represents the brightness value of the pixel point (x, y) in the image at time t after enhancement.
[0061] As a preferred embodiment of the present invention, in the step (2), the high-frequency information of the amplified video is compared with the high-frequency information in the original video to obtain the high-frequency increment ratio value, specifically as follows:
[0062] Obtain the high-frequency increment ratio value according to the following formula:
[0063]
[0064] Among them, u and v are respectively the spatial frequency components, w is the frequency information, H space is the spatial high-pass filter, H time (w) is the time high-frequency filter, F(u, v, w) is the frequency domain information of the original video after Fourier transform, F amplified (u, v, w) is the frequency domain information of the amplified video after Fourier transform, ε is a constant with a value of 10 -8 , and α is the amplification factor.
[0065] As a preferred embodiment of the present invention, in the step (2), the amplification factor is feedback-regulated through the high-frequency increment ratio value, specifically including:
[0066] Calculate the updated value of the amplification factor through iterative calculation to gradually converge the loss function to the optimal solution, specifically as follows:
[0067] Calculate the updated value of the amplification factor according to the following formula:
[0068]
[0069] Among them, αt is the amplification factor at time t, α t+1 is the amplification factor at time t + 1, α t-1 is the amplification factor at time t - 1, E th is the increment ratio threshold constant, β is the learning rate, R(α t ) is the increment ratio at time t, R(α t-1 ) is the increment ratio at time t - 1.
[0070] As a preferred embodiment of the present invention, the global enhancement network in step (3) is used to extract the global enhancement features of micro - expressions, specifically including the following steps:
[0071] Adopt a symmetric multi - scale strategy to gradually expand the receptive field of the channel sub - block in two directions respectively, and use the global representation semantics to guide the channel sub - block to obtain multi - scale channel enhancement features; splice and dimension - reduce the two multi - scale channel features along the channel dimension to obtain enhanced global channel flow semantics, so as to enhance the extraction ability of the global network features.
[0072] As a preferred embodiment of the present invention, the obtaining of the global channel flow semantic information is specifically as follows:
[0073] Obtain the global channel flow semantic information according to the following formula:
[0074]
[0075] Among them, is the multi - scale feature of the i - th channel sub - block calculated from left to right, is the multi - scale feature of the i - th channel sub - block calculated from right to left, G is the global representation semantics, F i ′ is the shuffled feature of the i - th channel, f 3×3 (·) is a convolution operation with a convolution kernel of 3, Concat C (·) is a feature splicing operation on the channel dimension, f 1×1 (·) is a 1×1 convolution, M is to group the channel shuffled features into M sub - blocks by channel, ω k is each feed - forward feature and 's weight, represents the weighted sum of the cumulative left - hand side feature information from level 1 to i - 1, represents the weighted sum of the cumulative right - hand side feature information from level i + 1 to M.
[0076] As a preferred embodiment of the present invention, the local attention feature extraction network in step (3) is used to obtain the local attention features of micro - expressions, specifically including the following steps:
[0077] Divide the micro-expression image into four non-overlapping local feature maps \(F\in R\) H×W×C , perform two 3×3 convolution operations on each local feature map, use the attention module as the attention feature for extracting local information, perform an element-wise multiplication operation on the attention map and the input feature map, and adaptively refine the features.
[0078] As a preferred embodiment of the present invention, the local attention feature extraction network in step (3) is specifically:[[]]
[0079] Obtain the output of the local attention feature extraction network according to the following formula:[[]]
[0080]
[0081] where \(F\) is the local feature map after division, \(f\) 3×3 (·) represents a convolution operation with a convolution kernel of 3, \(f\) 1×1 (·) represents a 1×1 convolution, represents an element-wise multiplication operation, represents an element-wise addition operation, \(M\) c represents ChannelAttention calculation, \(M\) s represents Spatial Attention calculation, \(f\) down (·) represents a downsampling operation, \(F\) r is the obtained local attention feature.
[0082] As a preferred embodiment of the present invention, step (4) specifically includes the following steps:[[]]
[0083] Concatenate and fuse the obtained global enhanced feature and local attention feature, use the fused feature vector information as the input of the fully connected neural network. The fully connected neural network iteratively updates the model parameters to minimize the objective function, making the objective function value continuously decrease. The fully connected neural network learns the relationship between the input data and the output label by continuously adjusting the weights and biases, and classifies the newly input facial micro-expression using the weight relationship between the input data and the label.
[0084] The device for micro-expression recognition based on adaptive amplification of dynamic weak texture information of the present invention, wherein the device includes:[[]]
[0085] A processor configured to execute computer-executable instructions;
[0086] A memory storing one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above method for micro-expression recognition based on adaptive amplification of dynamic weak texture information are implemented.
[0087] The processor for micro-expression recognition based on adaptive amplification of dynamic weak texture information according to the present invention, wherein the processor is configured to execute computer-executable instructions, and when the computer-executable instructions are executed by the processor, each step of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information as described above is implemented.
[0088] The computer-readable storage medium of the present invention, on which a computer program is stored, and the computer program can be executed by a processor to implement each step of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information as described above.
[0089] The purpose of the present invention is to provide a micro-expression recognition method, device and medium based on adaptive amplification of dynamic weak texture information, so as to solve technical problems such as the weak movement intensity of micro-expressions, the locality of movement, and the difficulty of existing algorithms to quickly and accurately recognize. The framework of the micro-expression recognition method based on adaptive amplification of dynamic weak texture information of the present invention is as Figure 2 shown. This method first amplifies the weak operation signal through the Eulerian video magnification algorithm designed based on dynamic weak texture information of the present invention to eliminate the influence of static area noise. Then, an adaptive optimization method of the amplification factor based on feedback regulation is used to adjust the amplification factor of the Eulerian video magnification algorithm. Finally, the amplified image is used as the input of a fusion global enhancement-local attention feature network to classify micro-expressions. This network uses two parallel branches to capture global enhancement features and local attention features respectively, and finally fuses the global information and local information to classify micro-expressions through the fused features.
[0090] In a specific embodiment of the present invention, a micro-expression recognition method based on adaptive amplification of dynamic weak texture information is involved, which specifically includes:
[0091] Step S1, collecting information of a micro-expression video clip of a tester's face by a vision system and storing it in a face micro-expression data storage memory bank;
[0092] Step S2, extracting the micro-expression emotion video clip in the face micro-expression data storage memory bank, and amplifying the micro-expression video information by a weak signal amplification method based on amplification of dynamic weak texture information. And comparing the high-frequency information of the amplified video with the high-frequency information in the original video to obtain a high-frequency increment ratio value. The amplification factor is feedback-regulated through the high-frequency increment ratio value to suppress the generation of noise during the amplification process.
[0093] Step S3: By constructing two parallel branches, i.e., a global enhancement network and a local attention feature extraction network, the image features are effectively captured and amplified. Among them, the global enhancement network is used to extract the global enhancement features of micro-expressions. And the local attention feature extraction network focuses on obtaining the local attention features of micro-expressions.
[0094] Step S4: Finally, the global enhancement features and the local attention features are fused, and the micro-expressions are classified based on the fused features.
[0095] Preferably, in step S2, the micro-expression emotion video segments in the face micro-expression data storage repository are extracted, and the micro-expression video information is amplified by a weak signal amplification method based on dynamic weak texture information amplification. The implementation steps of this method are as follows. Specifically as follows:
[0096] The motion texture information of micro-expressions is weak, and it is difficult for traditional algorithms to extract its texture motion information. The present invention uses the Eulerian video magnification algorithm to amplify the weak information in micro-expressions. However, when the traditional Eulerian video magnification algorithm amplifies weak information, it will misamplify the noise in the static area as weak motion information, affecting the recognition accuracy of micro-expressions. For this reason, the present invention proposes an Eulerian video magnification algorithm based on dynamic weak texture information. This method identifies the dynamic part in the image through the differential algorithm, constructs a dynamic region mask, and only amplifies the weak texture information within the dynamic mask region, effectively removing the amplification of the noise information in the static area. The implementation steps of this method are as follows.
[0097] Step 1: Since the micro-expression video is a time series, the present invention uses I(x, y, t) to represent the time series function of the micro-expression video segment, where I(x, y, t) represents the luminance value of the pixel point (x, y) at time t, and this function is combined with dynamic signals and static signals.
[0098] I(x, y, t) = S(x, y, t) + D(x, y, t)
[0099] Among them, S(x, y, t) is the static component, representing the background or static area. D(x, y, t) is the dynamic component, representing the area that changes over time.
[0100] Since the dynamic component and the static component in the video are unknown, the present invention identifies the dynamic component D(x, y, t) of the image through the differential method and amplifies the dynamic component. Given two adjacent frames I(x, y, t) and I(x, y, t + Δt), their difference is defined as:
[0101] ΔI(x, y, t) = I(x, y, t + Δt) - I(x, y, t)
[0102] If ΔI(x, y, t) ≈ 0, it indicates that the pixel at (x, y) has not changed, and this part is the static region. If ΔI(x, y, t) is not equal to 0, it means the pixel at (x, y) has changed, and this part is the dynamic region.
[0103] Step 2: Determine the mask M(x, y, t) of the dynamic region through the threshold τ. When the change value of the pixel is greater than τ, set the information at the mask position of this dynamic region to 1, otherwise set it to 0. Among them Figure 1 is a schematic diagram of the dynamic region mask established by the difference method. It can be seen from the image that the constructed dynamic mask is mainly distributed in the mouth, eyes, and face contour parts.
[0104]
[0105] The multiplication of the face sorting information and the dynamic mask information is the face dynamic signal, and the expression of the face dynamic signal is:
[0106] D(x, y, t) = ΔI(x, y, t) · M(x, y, t)
[0107] Step 3: Obtain the position of the weak motion information through the dynamic mask, and only amplify the weak texture information within the dynamic mask area.
[0108]
[0109] Among them, I(x, y, t) represents the brightness value of the pixel at (x, y) at time t, and k represents the level of the pyramid. I k (x, y, t) represents the pixel at time t of the k-th layer image. α is the amplification factor, F and F -1 respectively represent the Fourier transform and the inverse transform, and H(f) is the band-pass filter. M(x, y, t) is the mask of the dynamic region. I en (x, y, t) represents the brightness value of the pixel at (x, y) of the enhanced image at time t.
[0110] When the Eulerian video magnification algorithm based on suppressing the static part constructed by the present invention is used to magnify micro-expressions, the signals in the static region will not be affected, and the signals in the dynamic region will be significantly enhanced.
[0111] In the step S2, compare the high-frequency information of the magnified video with the high-frequency information in the original video to obtain the high-frequency increment ratio value. Specifically as follows:
[0112] The Eulerian video magnification algorithm mainly enhances subtle motions by magnifying the time-domain signals in specific frequency bands. However, when the magnification factor is set too high, not only the target signal is magnified, but also the quantization error during the discretization of continuous signals and the interpolation error during the generation of continuous signals from discretized data are magnified. The error superposition effect in these two processes will introduce additional high-frequency artifact signals, resulting in the generation of artifact noise. This phenomenon is manifested in the frequency domain as a significant increase in high-frequency components, making the intensity of high-frequency signals in the magnified video no longer proportional to the original signal, so that the proportion of high-frequency increment signals in the magnified video is significantly higher than that in the original signal. Therefore, the present invention detects artifact noise through high-frequency increment information.
[0113] The present invention converts the original video and the magnified video into frequency-domain information through Fourier transform, and the conversion method is as follows:
[0114]
[0115] Among them, I orig (x, y) is the pixel value of the original image at the position (x, y). I ampl (x, y) is the pixel value of the magnified image at the position (x, y). F orig (u, v) is the complex value in the frequency domain of the original image, and F ampl (u, v) is the complex value in the frequency domain of the magnified image. u and v are spatial frequency components respectively. F orig (u, v) is the frequency-domain information of the original video signal. F ampl (u, v) is the frequency-domain information of the magnified video signal. w is the frequency information.
[0116] The spatial high-frequency information is related to the magnitudes of u and v, and the present invention extracts it through the spatial high-pass filter H space . The expression of the spatial high-frequency filter is:
[0117]
[0118] Among them, f c is the threshold of the spatial high-pass filter. The temporal high-frequency information is related to the magnitude of w, and the present invention extracts the temporal high-frequency information through the temporal high-pass filter H time (w).
[0119]
[0120] Among them, f is the threshold of the temporal high-pass filter. In this patent, the high-frequency information in the micro-expression video is extracted through the spatial high-pass filter and the temporal high-pass filter, and the high-frequency energy is characterized by the sum of the squared amplitudes of the high-frequency components. The frequency-domain information of the amplified video is compared with the frequency-domain information before amplification, and the high-frequency increment ratio is solved. The calculation method is as follows.
[0121]
[0122] Among them, u and v are the spatial frequency components respectively, w is the frequency information, and H space is the spatial high-pass filter, and H time (w) is the temporal high-frequency filter, F(u, v, w) is the frequency-domain information of the original video after Fourier transform, and F amplified (u, v, w) is the frequency-domain information of the amplified video after Fourier transform, ∈ is a constant with a value of 10 -8 , and α is the amplification factor.
[0123] Preferably, in step S2, the amplification factor is feedback-regulated by the high-frequency increment ratio to suppress the generation of noise during the amplification process. Specifically as follows:
[0124] When the high-frequency increment ratio is greater than the set threshold or no convergence occurs, the gradient descent method will be used to learn the next amplification factor to ensure that the system can adaptively adjust the amplification strategy, thereby optimizing the distribution of high-frequency energy to make it more in line with the target requirements. The learning method of the gradient descent method is as follows. This method iteratively calculates the updated value of the amplification factor, making the loss function gradually converge to the optimal solution. The flowchart of the learning method of the gradient descent method is as Figure 3 shown, clearly demonstrating the complete learning process from calculating the loss function to updating the amplification factor. The update method is as follows.
[0125]
[0126] Among them, α t is the amplification factor at time t, α t+1 is the amplification factor at time t + 1, α t-1 is the amplification factor at time t - 1, E th is the increment ratio threshold constant, β is the learning rate, R(α t ) is the increment ratio at time t, and R(α t-1 ) is the increment ratio at time t - 1.
[0127] This iterative process will continue until the loss function converges, that is, R and the set threshold E thThe error between them is small enough to ensure that the amplification process of high-frequency energy reaches an optimal state. This adaptive learning method based on gradient descent can effectively improve the control accuracy of the high-frequency increment ratio, enabling the system to maintain excellent performance under different input conditions. This method can adaptively select appropriate amplification coefficients according to different video sequences. It has a flexible structure and stronger learning ability and higher accuracy compared to the original Euler video magnification algorithm.
[0128] Preferably, in step S3, by constructing two parallel branches, namely a global enhancement network and a local attention feature extraction network, image features are effectively captured and amplified. Among them, the global enhancement network is used to extract the global enhancement features of micro-expressions. Specifically as follows:
[0129] To further improve the accuracy and effectiveness of global feature extraction, the present invention innovatively proposes a multi-scale feature extraction network based on channel-spatial global feature enhancement. This module adopts a symmetric multi-scale strategy to gradually expand the receptive field of the channel sub-blocks in two directions respectively, and uses global representation semantics to guide the channel sub-blocks to obtain multi-scale channel enhancement features. Finally, the two multi-scale channel features are concatenated along the channel dimension and dimension-reduced to obtain enhanced global channel flow semantics to enhance the extraction ability of global network features.
[0130] This network is composed of a channel feature module and a spatial feature module. In the channel feature module, to fully capture the mutual relationship between different channels and cross-channel interaction information, the feature map F ∈ R output by ResNet18 is used C×H×W as the input of the global feature enhancement branch, and on this basis, a channel shuffle operation is carried out.
[0131] F′ = Shuffle(F)
[0132] where Shuffle(·) represents the channel shuffle operation, and F′ is the channel shuffle feature. To further refine the channel flow information, F′ is grouped into M sub-blocks by channel:
[0133]
[0134] where represents the feature of the i-th channel sub-block. Split C (·) represents the channel division operation. F ∈ R C×H×W means that F is a three-dimensional tensor, and its shape is C×H×W, belonging to the real number set R. C represents the number of channels, that is, the depth of the feature map; H represents the height of the feature map; W represents the width of the feature map. In addition, to measure the channel shuffle feature F′ and the channel sub-block feature Regarding the relationship and expanding the receptive field of the channel sub-blocks, the present invention adopts a symmetric multi-scale strategy to obtain the interaction information between different channel sub-blocks. First, a global representation of F′ is obtained by using a convolution operation:
[0135] G = f 3×3 (F′)
[0136] where f 3×3 (·) represents a convolution operation with a convolution kernel of 3. To obtain multi-scale channel interaction semantics, a symmetric multi-scale strategy is adopted to gradually expand the receptive field of the channel sub-blocks in two directions respectively, and the global representation semantics G is used to guide the channel sub-blocks to obtain multi-scale channel enhanced features. The way to obtain the global channel flow semantic information is as shown in the formula.
[0137]
[0138] where is the multi-scale feature of the i-th channel sub-block calculated from left to right, is the multi-scale feature of the i-th channel sub-block calculated from right to left, G is the global representation semantics, F i ′ is the shuffled feature of the i-th channel, f 3×3 (·) is a convolution operation with a convolution kernel of 3, Concat C (·) is a feature concatenation operation in the channel dimension, f 1×1 (·) is a 1×1 convolution, M is to group the channel shuffled features into M sub-blocks by channel, ω k is the weight of each feed-forward feature and , represents the weighted sum of the cumulative left-side feature information from level 1 to i - 1, represents the weighted sum of the cumulative right-side feature information from i + 1 to M levels.
[0139] The constructed symmetric multi-scale structure gradually expands the receptive field of the channel sub-blocks in two directions respectively, so as to better capture the multi-scale channel interaction semantics and realize the enhancement of the global channel features. Although the symmetric multi-scale channel flow module pays attention to the information interaction between channels, it ignores the capture of spatial information. To obtain the global context information containing spatial semantics, the present invention adopts a maximum operation to obtain the pixel-level space of the global representation semantics G.
[0140]
[0141] where G pixel represents the maximum value feature of G, W max and H maxrespectively represent the maximum functions in the width and height of the feature map. On the other hand, the mean operation is used to obtain the mean representation of the global representation semantics G:
[0142] G c = C Mean (G)
[0143] where G c represents the mean feature of G, and C Mean represents the mean function of the channel dimension. Then, to highlight the significant semantics of the spatial flow, according to the pixel-level spatial representation G pixel and the mean representation G c calculate the spatial position weight:
[0144] λ G = δ(Conv 1×1 (Concat C (G pixel + G c )))
[0145] where δ(·) represents the Sigmoid activation function. Concat C (·) represents the feature concatenation operation on the channel dimension. Finally, weight the global representation semantics with the spatial position weight λ G to obtain the enhanced global spatial flow semantics:
[0146]
[0147] where B represents the spatial flow semantics. The spatial flow module fully represents the spatial information of the global features and realizes the enhancement of the global features on the spatial flow. Fuse the obtained channel flow semantics A and spatial flow semantics B and use them as the global enhanced features of the feature map F:
[0148] F G = A + B
[0149] From the calculation process of the global enhanced feature F G it can be seen that the constructed channel-spatial global feature enhancement structure, on the one hand, realizes the extraction of channel semantics through the symmetric multi-scale strategy and the guidance of the global representation semantics; on the other hand, it uses the spatial position weight to weight the spatial semantics. Therefore, it enhances the global features from two perspectives of the channel flow and the spatial flow, improving the discriminability of the global context features.
[0150] Preferably, in step (3), the local attention feature extraction network focuses on obtaining the local attention features of the micro-expression. Specifically as follows:
[0151] The global enhanced features obtained by the global feature enhancement branch can effectively represent the overall information of the feature map, but they pay insufficient attention to local details. In the task of facial micro-expression recognition, the expression details in local regions are crucial, and there are significant differences in the importance of different regions. For example, regions such as the mouth, eyebrows, and eyes contribute more to expression recognition than regions such as the nose, chin, and forehead. To obtain effective local features, this patent precisely divides the feature map output by ResNet18 into four non-overlapping local feature maps F ∈ R H×W×C . After performing two 3×3 convolutional operations on each local feature map, this invention patent uses an attention module to extract the attention features of local information. Finally, these attention maps are multiplied element-wise with the input feature map to achieve adaptive refinement of the features. The local attention extraction module of the present invention is as shown in Figure 4 . The output of the local attention feature extraction network is as shown in the formula.
[0152]
[0153] Among them, F is the local feature map after division. f 3×3 (·) represents a convolutional operation with a convolution kernel of 3. f 1×1 (·) represents a 1×1 convolution. represents an element-wise multiplication operation. represents an element-wise addition operation. M c represents ChannelAttention calculation, and M s represents Spatial Attention calculation. f down (·) represents a downsampling operation. The local attention module consists of four parallel local attention networks, and each local attention network contains four attention blocks. This module takes four local feature maps with a size of 14×14×128 as inputs, and each local feature map is respectively input into the corresponding local attention network for processing. After the operation of the local attention module, four local feature maps with a size of 7×7×512 are obtained. Subsequently, these feature maps are concatenated along the spatial axis to form a feature map with a size of 14×14×512, and a global average pooling layer is applied to it, finally generating a feature vector with a dimension of 512.
[0154] Preferably, in step S4, finally, the global enhanced features and the local attention features are fused, and the micro-expression is classified through the fused features. Specifically:
[0155] The present invention concatenates and fuses the obtained global enhancement features and local attention features, and uses the fused feature vector information as the input of a fully connected neural network. The fully connected neural network minimizes the objective function by iteratively updating the model parameters, thereby continuously reducing the value of the objective function. To solve the problem of severe imbalance in the number of samples in the task, this patent uses the cross-entropy loss function to adjust the weights of different categories of data, making the model pay more attention to small samples or samples that are difficult to classify. This method improves the overall classification accuracy by increasing the weights of misclassified samples. The specific form of the cross-entropy loss function is as follows:
[0156]
[0157] Among them, p represents the class probability predicted by the model, y is the true label, and L is the output result of the cross-loss function. The focal loss function introduces a modulation factor γ (γ > 0) on the basis of the cross-entropy loss to enhance the model's attention to difficult-to-classify and misclassified samples. μ is a balance factor used to balance the problem of class imbalance between positive and negative samples. The fully connected neural network incorporating the cross-entropy loss function learns the relationship between the input data and the output label by continuously adjusting the weights and biases, and classifies the newly input micro-expression using the weight relationship between the input data and the label.
[0158] Considering that the facial movement amplitude of micro-expressions is small and the movements are mostly local movements, it is difficult for existing mainstream expression recognition algorithms to extract the features of micro-expressions, resulting in low recognition accuracy. The present invention proposes a micro-expression recognition method based on the adaptive amplification of dynamic weak texture information. This method collects the facial micro-expression video segment information of the test subject according to the visual system and stores it in the facial micro-expression data storage memory bank. The micro-expression emotion video segment in the facial micro-expression data storage memory bank is extracted, and the weak movement information of the micro-expression is amplified through a designed weak texture information amplification algorithm based on the adaptive amplification of dynamic weak texture information. The global enhancement features and local attention features of the amplified image are captured through two parallel branches of a global enhancement network and a local attention feature extraction network. Finally, the global enhancement features and local attention features are fused, and the micro-expression is classified through the fused features.
[0159] The present invention includes multiple embodiments as follows:
[0160] Embodiment 1:
[0161] As Figures 1 to 6 shown, a micro-expression recognition method based on the adaptive amplification of dynamic weak texture information is proposed, and the specific steps are as follows:
[0162] Step S1: According to the facial micro-expression video clip information of the tester collected by the vision system, store the micro-expression video information collected by the device into the face data storage memory bank;
[0163] Specifically, it includes: according to the facial micro-expression video clip information of the tester collected by the vision acquisition system, convert the collected video clip information into a continuous image sequence, and use a Gaussian smoothing filter to perform distortion removal processing on the continuous images obtained in Step S1 to exclude the graphic distortion caused by changes in the external environment. According to the working principle of the vision acquisition system, the vision acquisition system will store the collected facial micro-expression video clip information into the emotion memory bank after processing.
[0164] In Step S2, extract the micro-expression emotion video clip from the face micro-expression data storage memory bank, and amplify the micro-expression video information by a weak signal amplification method based on dynamic weak texture information amplification. The specific method is as follows:
[0165] The motion texture information of micro-expressions is weak, and it is difficult for traditional algorithms to extract its texture motion information. The present invention uses the Eulerian video magnification algorithm to amplify the weak information in micro-expressions. However, when the traditional Eulerian video magnification algorithm amplifies the weak information, it will misinterpret the noise in the static area as weak motion information for amplification, affecting the recognition accuracy of micro-expressions. For this reason, the present invention proposes an Eulerian video magnification algorithm based on dynamic weak texture information. This method identifies the dynamic part in the image through the difference algorithm, constructs a dynamic region mask, and only amplifies the weak texture information within the dynamic mask region, effectively removing the amplification of the noise information in the static area. The implementation steps of this method are as follows.
[0166] Step 1: Since the micro-expression video is a time series, the present invention uses I(x, y, t) to represent the time series function of the micro-expression video clip, where I(x, y, t) represents the brightness value of the pixel point (x, y) at time t, and this function is combined with a dynamic signal and a static signal.
[0167] I(x, y, t) = S(x, y, t) + D(x, y, t)
[0168] Among them, S(x, y, t) is the static component, representing the background or the stationary area. D(x, y, t) is the dynamic component, representing the area that changes over time.
[0169] Since the dynamic component and the static component in the video are unknown, the present invention identifies the dynamic component D(x, y, t) of the image through the difference method and amplifies the dynamic component. Given two adjacent frames I(x, y, t) and I(x, y, t + Δt), their difference is defined as:
[0170] ΔI(x,y,t) = I(x,y,t + Δt) - I(x,y,t)
[0171] If ΔI(x,y,t) ≈ 0, it indicates that the pixel at (x,y) has not changed, and this part is a static region. If ΔI(x,y,t) is not equal to 0, it indicates that the pixel at (x,y) has changed, and this part is a dynamic region.
[0172] Step 2: Determine the mask M(x,y,t) of the dynamic region through the threshold τ. When the pixel change value is greater than τ, set the information at the mask position of the dynamic region to 1, otherwise set it to 0. Among them Figure 1 is a schematic diagram of the dynamic region mask established by the difference method. It can be seen from the image that the constructed dynamic mask is mainly distributed in the mouth, eyes, and face contour parts.
[0173]
[0174] The multiplication of the face sorting information and the dynamic mask information is the face dynamic signal, and the expression of the face dynamic signal is:
[0175] D(x,y,t) = ΔI(x,y,t) · M(x,y,t)
[0176] Step 3: Obtain the position of the weak motion information through the dynamic mask, and only amplify the weak texture information within the dynamic mask area.
[0177]
[0178] Among them, I(x,y,t) represents the brightness value of the pixel at (x,y) at time t, and k represents the level of the pyramid. I k (x,y,t) represents the pixel at time t of the k-th layer image. α is the amplification factor, F and F -1 represent the Fourier transform and the inverse transform respectively, and H(f) is a band-pass filter. M(x,y,t) is the mask of the dynamic region. I en (x,y,t) represents the brightness value of the pixel at (x,y) at time t of the enhanced image.
[0179] When magnifying micro-expressions by the Eulerian video magnification algorithm based on suppressing static parts constructed by the present invention, the signals in the static region will not be affected, and the signals in the dynamic region will be significantly enhanced.
[0180] Preferably, in the step S2, the high-frequency information of the magnified video is compared with the high-frequency information in the original video to obtain the high-frequency increment ratio value. Specifically as follows:
[0181] The Eulerian video magnification algorithm mainly enhances subtle motions by magnifying the time-domain signals in specific frequency bands. However, when the magnification factor is set too high, not only the target signal is magnified, but also the quantization error during the discretization of continuous signals and the interpolation error during the generation of continuous signals from discretized data are magnified. The error superposition effect in these two processes will introduce additional high-frequency artifact signals, resulting in the generation of artifact noise. This phenomenon is manifested in the frequency domain as a significant increase in high-frequency components, making the intensity of high-frequency signals in the magnified video no longer proportional to the original signal, so that the proportion of high-frequency increment signals in the magnified video is significantly higher than that in the original signal. Therefore, the present invention detects artifact noise through high-frequency increment information.
[0182] The present invention converts the original video and the magnified video into frequency-domain information through Fourier transform, and the conversion method is as follows:
[0183]
[0184] Among them, I orig (x, y) is the pixel value of the original image at the position (x, y). I ampl (x, y) is the pixel value of the magnified image at the position (x, y). F orig (u, v) is the complex value of the original image in the frequency domain, and F ampl (u, v) is the complex value of the magnified image in the frequency domain. u and v are spatial frequency components respectively. F orig (u, v) is the frequency-domain information of the original video signal. F ampl (u, v) is the frequency-domain information of the magnified video signal. w is the frequency information.
[0185] The spatial high-frequency information is related to the magnitudes of u and v, and the present invention extracts it through the spatial high-pass filter H space . The expression of the spatial high-frequency filter is:
[0186]
[0187] Among them, f c is the threshold of the spatial high-pass filter. The temporal high-frequency information is related to the magnitude of w, and the present invention extracts the temporal high-frequency information through the temporal high-pass filter H time (w).
[0188]
[0189] Among them, f is the threshold of the temporal high-pass filter. In this patent, the high-frequency information in the micro-expression video is extracted through the spatial high-pass filter and the temporal high-pass filter, and the high-frequency energy is characterized by the sum of the squared amplitudes of the high-frequency components. The frequency-domain information of the amplified video is compared with the frequency-domain information before amplification, and the high-frequency increment ratio is solved. The calculation method is as follows.
[0190]
[0191] Among them, u and v are the spatial frequency components respectively. w is the frequency information. H space is the spatial high-pass filter. H time (w) is the temporal high-frequency filter. F(u, v, w) is the frequency-domain information of the original video after Fourier transform. F amplified (u, v, w) is the frequency-domain information of the amplified video after Fourier transform. ∈ is a constant with a value of 10 -8 , used to avoid the denominator being zero. α is the amplification factor.
[0192] Preferably, in the step S2, the amplification factor is feedback-regulated by the high-frequency increment ratio value to suppress the generation of noise during the amplification process. The specific steps are as follows:
[0193] When the high-frequency increment ratio is greater than the set threshold or no convergence occurs, the gradient descent method will be used to learn the next amplification factor to ensure that the system can adaptively adjust the amplification strategy, thereby optimizing the distribution of high-frequency energy to make it more in line with the target requirements. The learning method of the gradient descent method is as follows. This method iteratively calculates the updated value of the amplification factor, making the loss function gradually converge to the optimal solution. The flow chart of the learning method of the gradient descent method is as Figure 3 shown, clearly showing the complete learning process from calculating the loss function to updating the amplification factor. The update method is as follows.
[0194]
[0195] Among them, α t is the amplification factor at time t, α t+1 is the amplification factor at time t + 1, α t-1 is the amplification factor at time t - 1. E th is the increment ratio threshold constant. β is the learning rate, used to control the step size of each update. R(α t ) represents the increment ratio at time t, and R(α t-1 ) represents the increment ratio at time t - 1.
[0196] This iterative process will continue until the loss function converges, that is, R and the set threshold E thThe error between them is small enough to ensure that the amplification process of high-frequency energy reaches an optimal state. This adaptive learning method based on gradient descent can effectively improve the control accuracy of the high-frequency increment ratio, enabling the system to maintain excellent performance under different input conditions. This method can adaptively select appropriate amplification coefficients according to different video sequences. This method has a flexible structure and stronger learning ability and higher accuracy compared to the original Eulerian video magnification algorithm.
[0197] In step S3, by constructing two parallel branches, namely a global enhancement network and a local attention feature extraction network, image features are effectively captured and amplified. Among them, the global enhancement network is used to extract the global enhancement features of micro-expressions. Specifically as follows:
[0198] To further improve the accuracy and effectiveness of global feature extraction, the present invention innovatively proposes a multi-scale feature extraction network based on channel-spatial global feature enhancement. This module adopts a symmetric multi-scale strategy to gradually expand the receptive field of the channel sub-blocks in two directions respectively, and uses global representation semantics to guide the channel sub-blocks to obtain multi-scale channel enhancement features. Finally, the two multi-scale channel features are concatenated along the channel dimension and dimensionally reduced to obtain enhanced global channel flow semantics to enhance the extraction ability of the global network features.
[0199] This network is composed of a channel feature module and a spatial feature module. In the channel feature module, to fully capture the mutual relationship between different channels and cross-channel interaction information, the feature map F ∈ R output by ResNet18 is used C×H×W as the input of the global feature enhancement branch, and on this basis, a channel shuffle operation is carried out.
[0200] F′ = Shuffle(F)
[0201] Among them, Shuffle(·) represents the channel shuffle operation, and F′ is the channel shuffled feature. To further refine the channel flow information, F′ is grouped into M sub-blocks by channel:
[0202]
[0203] Among them, represents the feature of the i-th channel sub-block. Split C (·) represents the channel partitioning operation. F ∈ R C×H×W means that F is a three-dimensional tensor, and its shape is C×H×W, belonging to the real number set R. C represents the number of channels, that is, the depth of the feature map; H represents the height of the feature map; W represents the width of the feature map. In addition, to measure the channel shuffled feature F′ and the channel sub-block feature Regarding the relationship and expanding the receptive field of the channel sub-blocks, the present invention adopts a symmetric multi-scale strategy to obtain the interaction information between different channel sub-blocks. First, a global representation of F′ is obtained by using a convolution operation:
[0204] g = f 3×3 (F′)
[0205] where f 3×3 (·) represents a convolution operation with a convolution kernel of 3. To obtain multi-scale channel interaction semantics, a symmetric multi-scale strategy is adopted to gradually expand the receptive field of the channel sub-blocks in two directions, and the global representation semantics G is used to guide the channel sub-blocks to obtain multi-scale channel enhanced features. The way to obtain the global channel flow semantic information is as shown in the formula.
[0206]
[0207]
[0208] where, is the multi-scale feature of the i-th channel sub-block calculated from left to right, is the multi-scale feature of the i-th channel sub-block calculated from right to left, G is the global representation semantics, F i ′ is the shuffled feature of the i-th channel, f 3×3 (·) is a convolution operation with a convolution kernel of 3, Concat C (·) is a feature concatenation operation in the channel dimension, f 1×1 (·) is a 1×1 convolution, M is to group the channel shuffled features into M sub-blocks by channel, ω k is the weight for each feed-forward feature and , represents the weighted sum of the cumulative left-side feature information from level 1 to i - 1, represents the weighted sum of the cumulative right-side feature information from i + 1 to M levels.
[0209] The constructed symmetric multi-scale structure gradually expands the receptive field of the channel sub-blocks in two directions, so as to better capture the multi-scale channel interaction semantics and realize the enhancement of the global channel features. Although the symmetric multi-scale channel flow module pays attention to the information interaction between channels, it ignores the capture of spatial information. To obtain the global context information containing spatial semantics, the present invention adopts a maximum operation to obtain the pixel-level space of the global representation semantics G.
[0210]
[0211] where G pixel represents the maximum value feature of G, W max and H maxrespectively represent the maximum functions in the width and height of the feature map. On the other hand, the mean representation of the global representation semantics G is obtained by using the mean operation:
[0212] G c = C Mean (G)
[0213] where G c represents the mean feature of G, and C Mean represents the mean function in the channel dimension. Then, to highlight the significant semantics of the spatial flow, according to the pixel-level spatial representation G pixel and the mean representation G c calculate the spatial position weight:
[0214] λ G = δ(Conv 1×1 (Concat C (G pixel + G c )))
[0215] where δ(·) represents the Sigmoid activation function. Concat C (·) represents the feature concatenation operation in the channel dimension. Finally, the spatial position weight λ G is weighted with the global representation semantics to obtain the enhanced global spatial flow semantics:
[0216]
[0217] where B represents the spatial flow semantics. The spatial flow module fully represents the spatial information of the global features and realizes the enhancement of the global features on the spatial flow. The obtained channel flow semantics A and spatial flow semantics B are fused and used as the global enhanced features of the feature map F:
[0218] F G = A + B
[0219] From the calculation process of the global enhanced feature F G , it can be seen that the constructed channel-spatial global feature enhancement structure, on the one hand, realizes the extraction of channel semantics through the symmetric multi-scale strategy and the guidance of the global representation semantics; on the other hand, it uses the spatial position weight to perform the weighting of spatial semantics. Therefore, it completes the enhancement of the global features from two perspectives of the channel flow and the spatial flow, and improves the discriminability of the global context features.
[0220] Preferably, in the step S3, the local attention feature extraction network focuses on obtaining the local attention features of the micro-expression. Specifically as follows:
[0221] The global enhanced features obtained by the global feature enhancement branch can effectively represent the overall information of the feature map, but they pay insufficient attention to local details. In the task of facial micro-expression recognition, the expression details in local areas are crucial, and there are significant differences in the importance of different areas. For example, areas such as the mouth, eyebrows, and eyes contribute more to expression recognition than areas such as the nose, chin, and forehead. To obtain effective local features, this patent precisely divides the feature map output by ResNet18 into four non-overlapping local feature maps F∈R H×W×C . After performing two 3×3 convolutional operations on each local feature map, this invention patent uses an attention module to extract the attention features of local information. Finally, these attention maps are multiplied element-wise with the input feature map to achieve adaptive refinement of the features. The local attention extraction module of the present invention is as shown in Figure 4 . The output of the local attention feature extraction network is as shown in the formula.
[0222]
[0223] Among them, F is the local feature map after division. f 3×3 (·) represents the convolutional operation with a convolution kernel of 3. f 1×1 (·) represents a 1×1 convolution. represents the element-wise multiplication operation. represents the element-wise addition operation. M c represents the ChannelAttention calculation, and M s represents the Spatial Attention calculation. f down (·) represents the downsampling operation. The local attention module consists of four parallel local attention networks, and each local attention network contains four attention blocks. This module takes four local feature maps with a size of 14×14×128 as input and inputs each local feature map into the corresponding local attention network for processing. After the operation of the local attention module, four local feature maps with a size of 7×7×512 are obtained. Subsequently, these feature maps are concatenated along the spatial axis to form a feature map with a size of 14×14×512, and a global average pooling layer is applied to it, finally generating a feature vector with a dimension of 512.
[0224] Preferably, in step S4, finally, the global enhanced features and the local attention features are fused, and the micro-expression is classified through the fused features. Specifically:
[0225] The present invention concatenates and fuses the obtained global enhancement features and local attention features, and uses the fused feature vector information as the input of a fully connected neural network. The fully connected neural network iteratively updates the model parameters to minimize the objective function, thereby continuously reducing the value of the objective function. To solve the problem of severe imbalance in the number of samples in the task, this patent uses the cross-entropy loss function to adjust the weights of different category data, enabling the model to pay more attention to small samples or samples that are difficult to classify. This method improves the overall classification accuracy by increasing the weights of misclassified samples. The specific form of the cross-entropy loss function is as follows:
[0226]
[0227] Among them, p represents the category probability predicted by the model, y is the true label, and L is the output result of the cross-loss function. The focal loss function introduces a modulation factor γ (γ > 0) on the basis of the cross-entropy loss to enhance the model's attention to difficult-to-classify and misclassified samples. μ is a balance factor used to balance the problem of class imbalance between positive and negative samples. The fully connected neural network incorporating the cross-entropy loss function learns the relationship between the input data and the output label by continuously adjusting the weights and biases, and classifies the newly input micro-expression using the weight relationship between the input data and the label.
[0228] The present invention constructs a new micro-expression recognition algorithm framework, which improves the accuracy of micro-expression recognition. The present invention can be used as a clue for diagnosing certain diseases to help doctors identify some psychological and mental diseases. By monitoring the micro-expressions of patients, the condition can be detected and diagnosed in a timely manner, reducing the costs and time of the hospital and alleviating the pressure on the hospital.
[0229] Next, a micro-expression recognition method based on adaptive amplification of dynamic weak texture information will be described in combination with specific experiments. The following will illustrate the process with specific experiments.
[0230] In this patent, the computer configuration used in the experiment is as follows: The processor is an Intel Core i5-9400F with 8 cores and a main frequency of 3.6 GHz. The system memory is 32 GB. The graphics card is an NVIDIA GeForce RTX 3090 24G. The deep learning training of this research uses the PyTorch 1.10.1 framework, and the optimizer selects the Stochastic Gradient Descent (SGD) algorithm. Among them, the momentum factor is set to 0.9, the weight decay coefficient is 1*10-4, and the initial learning rate is 0.001. To dynamically adjust the learning rate, with the help of the torch.optim.lr_scheduler module in the PyTorch framework, a strategy of gradually decreasing the learning rate as the training epoch increases is implemented. In addition, the SMIC, CASME, and CASMEII micro-expression datasets are also used to evaluate the accuracy of the method of the present invention. These publicly available datasets have annotated the ground truth of the facial expressions of micro-expressions, and have also annotated the start frame, peak frame, and end frame of the expressions.
[0231] To solve the problem of class imbalance, the present invention uses the Unweighted Average Recall (UAR) and Unweighted F1-scores (UF1) to evaluate the method of the present invention.
[0232]
[0233] Among them, TP i , FP i , FN i are the numbers of true positives, false positives, and false negatives in the i-th class respectively, and C is the number of classes. UF1 is calculated by computing the F1 score for each class and taking the average, rather than simply calculating the global accuracy. In this way, UF1 assigns the same weight to each class, so that in the case of class imbalance, the performance of the minority class will not be ignored. Even if the number of samples in some classes is very small, UF1 can ensure that they contribute to the evaluation of the model. UAR first calculates the accuracy for each class and then takes the average of the accuracies, which also solves this problem. It ensures that the prediction results of each class have an equal contribution to the final evaluation. UF1 focuses on the precision and recall of each class, and UAR focuses on the accuracy of each class. Both can provide a more reasonable evaluation of the model performance on class-imbalanced datasets.
[0234] The present invention verifies the effectiveness of the Eulerian video magnification algorithm based on suppressing the static part through the publicly available CASMEII dataset. Figure 1Schematic diagram of amplification based on dynamic weak texture information of the present invention. The three images in the figure are sequentially selected from the sub01 / EP04_02, sub13 / EP09_10, and sub02 / EP08_04 sequences, and the real expressions are disgust, happiness, and happiness respectively. The leftmost column of images is the original image in the dataset, the middle column is the amplified effect diagram of the Eulerian video magnification algorithm, and the rightmost column is the amplified image of the method of the present invention. The part circled by the square in the original image is the main motion area. It can be seen from the figure that after the Eulerian video magnification algorithm is amplified, the background and the noise in the static area of the face are also amplified. However, due to the introduction of the static area amplification suppression module in the method of the present invention, the amplification of the static area is effectively suppressed, so that the amplification effect is only amplified in the motion area, effectively suppressing the generation of noise.
[0235] Figure 6 Schematic diagram of the fixed threshold method of the present invention and the amplification of the method of the present invention. The four images in the figure are sequentially selected from the sub01 / EP04_02, sub13 / EP09_10, sub02 / EP08_04, and sub / EP03_06 sequences, and the real expressions are disgust, happiness, happiness, and surprised expressions respectively. It can be seen from the figure that the sub01 / EP04_02 sequence has the best effect when the amplification factor is 4, while the sub13 / EP09_10 sequence has the best effect when the amplification factor is 6, the sub02 / EP08_04 sequence has the best effect when the amplification factor is 6, and the sub / EP03_06 sequence has the best effect when the amplification factor is 4. It can be seen that for different sequences, different amplification factors need to be set. When the traditional Eulerian video magnification algorithm faces different sequences, it often needs to go through multiple cumbersome tests to determine the appropriate amplification factor, which undoubtedly increases the time cost and operation complexity. In sharp contrast, the method proposed by the present invention innovatively adjusts and determines the appropriate amplification factor according to the growth ratio of high-frequency information. Through the amplified display diagram of the method of the present invention in Figure (g), it can be intuitively seen that the method of the present invention can more excellently amplify the sequence, effectively improving the image quality and detail expressiveness.
[0236] In order to explore that the adaptive optimal amplification factor proposed in the present invention can better adjust the amplification factor, the method proposed in the present invention is compared with the fixed threshold method. The performance comparison table of the comparison results is shown in Table 1. It can be seen from the table that the UF1 and URA of the dataset without amplification processing are the lowest, because the micro-expression texture running information is weak and the movement is local, which makes it difficult for the network to extract the features of the expression and affects the recognition accuracy. As the videos in the dataset are amplified, the accuracy rate will gradually increase. However, when the amplification factor is relatively large, the accuracy rate begins to decrease again. This is because the Euler video magnification algorithm magnifies the high-frequency information. As the magnification factor increases, corresponding high-frequency noise will also be generated, resulting in redundant features being extracted and affecting the recognition accuracy of the algorithm. The proposed method can automatically adjust the amplification factor according to the high-frequency incremental information in the video, reduce the generation of noise, and can select the optimal amplification factor. Therefore, the recognition accuracy of the proposed method is higher than that of the fixed threshold method. It can be seen that the adaptive optimal amplification factor proposed in the present invention can effectively select the amplification factor.
[0237] Table 1 Performance comparison table of the method of the present invention and the fixed threshold algorithm
[0238]
[0239] Table 2 shows the comparison results of various algorithms on the SMIC, CASME, and CASME datasets. It can be seen from the table. As can be seen from the results, the proposed algorithm outperforms other comparison methods in terms of performance on all datasets, demonstrating excellent results. On the SMIC dataset, the method of the present invention achieved 89.07% UF1 and 89.58% UAR, significantly better than MERASTC (UF1 is 84.16% and UAR is 85.74%) and C3Dbed (UF1 is 83.36% and UAR is 83.75%). On the CASME dataset, the method proposed in the present invention obtained the best results again with 87.28% UF1 and 86.87% UAR. It can be seen from the experimental results of the three datasets in the table that the UF1 and UAR of the algorithm proposed in the present invention show significant improvements on the three datasets. This is because the method proposed in the present invention can amplify the weak motion information in micro-expressions and can suppress the noise generated during the amplification process, effectively improving the accuracy of micro-expression recognition.
[0240] Table 2 Comparison result table of various algorithms on the SMIC, CASME, and CASMEII datasets
[0241]
[0242] As can be seen from the above experiments, the proposed method is verified using the publicly available SMIC, CASME, and CASMEII micro-expression datasets. The verification results show that the method proposed in the present invention has an accuracy improvement of about 9.0% compared with the mainstream MERASTC micro-expression recognition algorithm. This indicates that the method proposed in the present invention has good extraction ability and high recognition ability for micro-expression facial features.
[0243] Embodiment 2:
[0244] Corresponding to Embodiment 1 of the present invention, Embodiment 2 of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented:
[0245] Step S1: Collect the facial micro-expression video clip information of the tester according to the vision system and store it in the facial micro-expression data storage memory bank;
[0246] Step S2: Extract the micro-expression emotion video clip from the facial micro-expression data storage memory bank, and amplify the weak motion information of the micro-expression through the designed weak texture information amplification algorithm based on dynamic weak texture information adaptive amplification;
[0247] Step S3: Capture the global enhancement feature and local attention feature of the amplified image through two parallel branches of the global enhancement network and the local attention feature extraction network;
[0248] Step S4: Finally, fuse the global enhancement feature and the local attention feature, and classify the micro-expression through the fused feature.
[0249] The above storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), optical discs, etc. that can store program codes.
[0250] The specific limitations of the steps implemented after the execution of the program in the above computer-readable storage medium can refer to Embodiment 1 and will not be elaborated here.
[0251] Embodiment 3:
[0252] Corresponding to Embodiment 1 of the present invention, Embodiment 3 of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented:
[0253] Step S1: Collect the facial micro-expression video clip information of the tester according to the vision system and store it in the facial micro-expression data storage memory bank;
[0254] Step S2: Extract the micro-expression emotional video clips from the face micro-expression data storage repository, and amplify the weak motion information of the micro-expression through the designed weak texture information amplification algorithm based on adaptive amplification of dynamic weak texture information;
[0255] Step S3: Capture the global enhancement features and local attention features of the amplified image through two parallel branches of the global enhancement network and the local attention feature extraction network;
[0256] Step S4: Finally, fuse the global enhancement features and the local attention features, and classify the micro-expression through the fused features.
[0257] In the technical solution of the present invention, when collecting data, a Point Grey GRAS-03K2C high-speed camera is used to collect the facial video of the tester and store it in the face micro-expression data storage repository. The collected frame rate is 200 frames per second, the resolution is 640×480, and the saving format is uncompressed AVI video format. During the collection, constant illumination and a clean background are ensured, and LED lights are used for supplementary lighting to reduce shadow interference. The collected video stream is transmitted to the image processing and inference module through an interface. This module consists of an Rk3588 development board and an NPU (neural network processing unit), and is responsible for real-time image processing and deep learning inference tasks. Finally, the micro-expression recognition method based on adaptive amplification of dynamic weak texture information proposed by the technical solution of the present invention is used to recognize the micro-expression of the collected data. And the processed results are displayed through a TreePi 4B monitor, and the user can view the results of the expression recognition in real time.
[0258] The technical solution of the present invention aims at the weak motion texture information and local motion of micro-expressions. As a result, traditional algorithms are difficult to accurately identify, and a micro-expression recognition method based on adaptive amplification of dynamic weak texture information is proposed to recognize the micro-expressions of the collected data.
[0259] The present invention has the following technical effects:
[0260] 1. In view of the fact that the traditional Eulerian video magnification algorithm is prone to misidentifying the noise information in the static area as dynamic information and amplifying the noise, which affects the accuracy of micro-expression recognition. For this reason, the present invention proposes an Eulerian video magnification algorithm based on dynamic information amplification. This method identifies the dynamic part in the image through differential technology, constructs a dynamic region mask, and suppresses the amplification of signals outside the mask region, effectively reducing the noise in the amplified image.
[0261] 2. In the traditional Eulerian video magnification algorithm, it is usually necessary to conduct multiple tests on the magnification object to select an appropriate magnification factor. However, this method is not applicable to processing non-fixed data set sequences and online operation. For this reason, the present invention proposes an adaptive optimization method for the magnification factor based on feedback regulation. This method can adaptively select an appropriate magnification coefficient according to different video sequences. This method has a flexible structure and stronger learning ability and higher accuracy compared with the original Eulerian video magnification algorithm.
[0262] 3. In order to further improve the accuracy and effectiveness of global feature extraction, the present invention innovatively proposes a multi-scale feature extraction network based on channel-spatial global feature enhancement. This module adopts a symmetric multi-scale strategy to gradually expand the receptive field of the channel sub-blocks in two directions respectively, and uses global representation semantics to guide the channel sub-blocks to obtain multi-scale channel enhanced features. Finally, the two multi-scale channel features are concatenated along the channel dimension and dimensionally reduced, so as to obtain enhanced global channel flow semantics to enhance the extraction ability of the global network features.
[0263] 4. In order to further improve the extraction of local detail information of micro-expressions, this patent accurately divides micro-expression images into four non-overlapping local feature maps. After performing two convolution operations on each local feature map, an attention module is used to extract the attention features of the local information. Finally, these attention maps are multiplied element by element with the input feature map, so as to realize the adaptive refinement of the features.
[0264] 5. The present invention concatenates and fuses the obtained global enhanced features and local attention features, and uses the fused feature vector information as the input of a fully connected neural network. The fully connected neural network iteratively updates the model parameters to minimize the objective function, so that the value of the objective function continuously decreases. The fully connected neural network learns the relationship between the input data and the output labels by continuously adjusting the weights and biases, and classifies the true facial expressions of the newly input testers by using the weight relationship between the input data and the labels.
[0265] 6. The present invention constructs a new micro-expression recognition algorithm framework. This method not only improves the accuracy of the system, but also improves the real-time performance of the system. The present invention can accurately and quickly recognize the facial micro-expressions of testers. It can be used as a clue for diagnosing certain diseases of testers and helping testers recognize some psychological and mental diseases.
[0266] For the specific limitations on the implementation steps of the computer device above, reference can be made to Embodiment 1, and details will not be described here.
[0267] For the specific implementation solutions of this embodiment, reference can be made to the relevant descriptions in the above embodiments, and details will not be elaborated here.
[0268] It is understandable that the same or similar parts in the above embodiments can be referred to each other, and for the content not described in detail in some embodiments, reference can be made to the same or similar content in other embodiments.
[0269] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "a plurality of" refers to at least two.
[0270] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in sequence, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0271] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0272] Those of ordinary skill in the art of the present technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0273] In addition, each functional unit in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0274] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0275] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0276] The method, device, processor, and computer-readable storage medium for micro-expression recognition based on adaptive amplification of dynamic weak texture information are adopted. The dynamic part in the image is recognized through differential technology, a dynamic region mask is constructed, and the amplification of signals outside the mask region is suppressed, effectively reducing the noise in the amplified image. This method has a flexible structure and stronger learning ability and higher accuracy compared with the original Eulerian video magnification algorithm. The present invention can accurately and quickly recognize the facial micro-expressions of the tester. It can be used as a clue for diagnosing certain diseases of the tester and helping the tester recognize some psychological and mental diseases.
[0277] In this specification, the present invention has been described with reference to its specific embodiments. However, it is obvious that various modifications and transformations can still be made without departing from the spirit and scope of the present invention. Therefore, the specification and the drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method for micro-expression recognition based on adaptive amplification of dynamic weak texture information, characterized in that, The method described above includes the following steps: (1) Acquire the facial micro-expression video segment information of the tester according to the vision system, and store it in the facial micro-expression data storage repository; (2) Extract the micro-expression emotion video segments in the facial micro-expression data storage repository, amplify the micro-expression video information by a weak signal amplification method based on dynamic weak texture information amplification; compare the high-frequency information of the amplified video with the high-frequency information in the original video to obtain the high-frequency increment proportion value; feedback-regulate the amplification factor through the high-frequency increment proportion value; (3) Capture and amplify image features by constructing two parallel branches of a global enhancement network and a local attention feature extraction network; (4) Fuse the global enhancement feature and the local attention feature, and classify the micro-expression through the fused feature.
2. The method for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information according to claim 1, characterized in that, In step (2) described above, extract the micro-expression emotion video segments in the facial micro-expression data storage repository, and amplify the micro-expression video information by a weak signal amplification method based on dynamic weak texture information amplification, specifically: Identify the dynamic part in the image through the differential algorithm, construct a dynamic region mask, and amplify the weak texture information in the dynamic mask region; calculate the brightness value of the pixel point (x, y) in the enhanced image at time t, specifically: Calculate the brightness value of the pixel point (x, y) in the enhanced image at time t according to the following formula: Among them, I(x, y, t) represents the luminance value of the pixel point (x, y) at time t, k represents the level of the pyramid, and I k (x, y, t) represents the pixel point of the k-th layer image at time t, α is the magnification factor, F and F -1 respectively represent the Fourier transform and the inverse transform, and H(f) is a band-pass filter. M(x, y, t) is the mask of the dynamic region. I en (x, y, t) represents the luminance value of the enhanced image pixel point (x, y) at time t.
3. The method for micro-expression recognition based on adaptive amplification using dynamic weak texture information according to claim 1, characterized in that, In step (2) described above, compare the high-frequency information of the amplified video with the high-frequency information in the original video to obtain the high-frequency increment proportion value, specifically: Obtain the high-frequency increment proportion value according to the following formula: where u and v are spatial frequency components respectively, w is frequency information, and H space is a spatial high-pass filter, and H time (w) is a temporal high-frequency filter, F(u, v, w) is the frequency-domain information of the original video after Fourier transform, and F amplified (u, v, w) is the frequency-domain information of the magnified video after Fourier transform, ∈ is a constant with a value of 10 -8 , and α is the magnification factor.
4. The method for micro-expression recognition based on adaptive amplification of dynamic weak texture information according to claim 1, wherein In step (2) described above, feedback-regulate the amplification factor through the high-frequency increment proportion value, specifically including: Iteratively calculate the updated value of the amplification factor to gradually converge the loss function to the optimal solution, specifically: Calculate the updated value of the amplification factor according to the following formula: where α t is the amplification factor at time t, α t+1 is the amplification factor at time t + 1, α t-1 is the amplification factor at time t - 1, E th is the increment ratio threshold constant, β is the learning rate, R(α t ) is the increment ratio at time t, R(α t-1 ) is the increment ratio at time t - 1.
5. The method for micro-expression recognition based on adaptive amplification using dynamic weak texture information according to claim 1, wherein The global enhancement network in step (3) is used to extract the global enhancement features of the micro-expression, specifically including the following steps: Adopt a symmetric multi-scale strategy to gradually expand the receptive field of the channel sub-blocks in two directions respectively, and use the global representation semantics to guide the channel sub-blocks to obtain multi-scale channel enhancement features; splice and downsample the two multi-scale channel features along the channel dimension to obtain the enhanced global channel flow semantics, so as to enhance the extraction ability of the global network features.
6. The method for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information according to claim 5, characterized in that The obtaining of the global channel flow semantic information is specifically: Obtain the global channel flow semantic information according to the following formula: Among them, is the multi-scale feature of the i-th channel sub-block calculated from left to right, is the multi-scale feature of the i-th channel sub-block calculated from right to left, G is the global representation semantics, F i ′ is the shuffled feature of the i-th channel, f 3×3 (·) is a convolution operation with a convolution kernel of 3, Concat C (·) is a feature concatenation operation in the channel dimension, f 1×1 (·) is a 1×1 convolution, M is to group the channel shuffled features into M sub-blocks by channel, ω k is each feed-forward feature and 's weight, represents the weighted sum of the cumulative left-side feature information from level 1 to i-1, represents the weighted sum of the cumulative right-side feature information from level i+1 to M.
7. The method for realizing micro-expression recognition based on dynamic weak texture information adaptive amplification according to claim 1, characterized in that The local attention feature extraction network in step (3) is used to obtain the local attention features of the micro-expression, specifically including the following steps: Divide the micro-expression image into four non-overlapping local feature maps \(F\in\mathbb{R}\) H×W×C , perform two 3×3 convolution operations on each local feature map, use the attention module as the attention feature for extracting local information, perform an element-wise multiplication operation on the attention map and the input feature map, and adaptively refine the features.
8. The method for realizing micro-expression recognition based on adaptive amplification of dynamic weak texture information according to claim 1, characterized in that The local attention feature extraction network in step (3) is specifically: Obtain the output of the local attention feature extraction network according to the following formula: Among them, F is the local feature map after division, and f 3×3 (·) represents a convolution operation with a convolution kernel of 3, and f 1×1 (·) represents a 1×1 convolution, represents an element-wise multiplication operation, represents an element-wise addition operation, and M c represents the ChannelAttention calculation, and M s represents the SpatialAttention calculation, and f down (·) represents a downsampling operation, and F r is the obtained local attention feature.
9. The method for micro-expression recognition based on adaptive magnification using dynamic weak texture information according to claim 1, characterized in that, Step (4) described above specifically includes the following steps: The global enhanced features and local attention features are concatenated and fused, and the fused feature vector information is used as the input of the fully connected neural network. The fully connected neural network minimizes the objective function by iteratively updating the model parameters, causing the value of the objective function to continuously decrease. The fully connected neural network learns the relationship between the input data and the output labels by continuously adjusting the weights and biases, and classifies the newly input facial micro-expressions using the weight relationship between the input data and the labels.
10. An apparatus for micro-expression recognition based on adaptive amplification of dynamic weak texture information, characterized in that, The described device includes: a processor configured to execute computer-executable instructions; a memory storing one or more computer-executable instructions, which, when executed by the processor, implement the steps of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information according to any one of claims 1 to 9.
11. A processor for micro-expression recognition based on adaptive amplification of dynamic weak texture information, characterized in that, The processor is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by the processor to implement the steps of the method for micro-expression recognition based on adaptive amplification of dynamic weak texture information according to any one of claims 1 to 9.