A micro-expression intelligent recognition system based on facial images

By extracting facial motion features and contour jitter frequency, and combining them with forgery tendency representation parameters, forged videos can be accurately identified. This solves the problem of existing technologies relying on explicit defects in forged video identification, and achieves efficient and accurate forged video detection.

CN121074991BActive Publication Date: 2026-01-30GUANGDONG POLICE COLLEGE (GUANGDONG PROVINCIAL PUBLIC SECURITY JUDICIAL MANAGEMENT CADRE COLLEGE)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511605577.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-01-30
Estimated Expiration
2045-11-05

AI Technical Summary

Technical Problem

Existing technologies for identifying fake videos rely on explicit defects, which can easily overlook hidden anomalies and flaws in the video, thus reducing the accuracy of fake video identification.

Method used

The module extracts facial motion features and contour jitter frequency through detail extraction, the module calculates motion state representation values ​​through object analysis, and the module performs frame segmentation and homogenized expression judgment through dynamic evaluation. It constructs a temporal variation curve of contour edge clarity, quantifies the forgery tendency representation parameters of forged videos, and accurately captures the hidden flaws of forged videos.

Benefits of technology

It significantly improves the accuracy and efficiency of fake video identification, reduces the false negative rate, and enhances the reliability and sensitivity of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074991B_ABST
    Figure CN121074991B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent recognition, and more particularly to an intelligent micro-expression recognition system based on facial images. The invention includes a detail extraction module, which retrieves image data corresponding to the observed object in a video image to extract corresponding facial action features; an object analysis module, which combines facial action features and the frequency of contour jitter of the observed object to calculate a representation value of the observed object's action state, thereby classifying the observed object's action abnormality and ambiguity category; a labeling module, which selects to call a conventional labeling module or a dynamic evaluation module based on the action abnormality and ambiguity category; a conventional labeling module, which marks video images as forged videos; and a dynamic evaluation module, which evaluates and labels the observed object. This invention captures facial features and relies on precise quantitative analysis of dynamic details of micro-expressions to improve recognition accuracy and precision. In public security work, it can identify AI-generated facial image videos and combat online and telecommunications fraud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent recognition, and in particular to an intelligent micro-expression recognition system based on facial images. Background Technology

[0002] Facial expression capture technology refers to the use of image processing technology to capture facial expression features in image data for application. For example, in psychological services, it can analyze changes in micro-expressions of visitors in real time, helping to accurately capture real emotional reactions beyond facial expressions. It can also provide a reference for the diagnosis of clinical mental disorders, assisting in early screening and efficacy evaluation.

[0003] For example, Chinese patent application publication number CN111461021A discloses a micro-expression detection method based on optical flow. This method uses the open-source toolkit dlib to perform face recognition on each frame of a video sample, marking the region of interest (ROI) and regions less prone to deformation. It calculates the dense optical flow between two adjacent frames in the video sample, extracts the optical flow within the ROI and the optical flow within the less prone to deformation, and subtracts these two values ​​to remove the influence of head movement. It defines angular regions in polar coordinates, calculates the principal optical flow within each ROI of each frame in the video sample within these defined angular regions, and sequentially represents the principal optical flow trajectory of all frames in the video sample using amplitude and angle. Based on the amplitude and angle trajectories, it captures frames in the video sample where micro-expressions occur and annotates them. This invention can display the movement of micro-expressions in each ROI in real time and does not require data training, offering advantages such as high efficiency, intuitiveness, high accuracy, and stability.

[0004] However, the following problems still exist in the existing technology.

[0005] The identification of forged videos mostly relies on obvious defects such as inter-frame splicing marks and feature blurring, which easily overlooks hidden anomalies and flaws that may exist in the video, thus reducing the accuracy of forged video identification. Summary of the Invention

[0006] To address this issue, the present invention provides a micro-expression intelligent recognition system based on facial images, which overcomes the problem that existing technologies for identifying fake videos mostly rely on obvious defects, easily overlooking potential hidden anomalies and flaws in the video, thus reducing the accuracy of fake video recognition.

[0007] To achieve the above objectives, the present invention provides a micro-expression intelligent recognition system based on facial images, comprising:

[0008] The detail extraction module is used to call the image data corresponding to the observed object in the video image to extract the corresponding facial action features. The facial action features include the uniformity of the change in the action part and the uniformity of the pause duration.

[0009] The object analysis module, which is connected to the detail extraction module, is used to combine the facial action features and the contour shaking frequency of the observed object to calculate the action state representation value of the observed object, so as to classify the action abnormal fuzzy category of the observed object.

[0010] The annotation module is invoked, which is connected to the object analysis module, and is used to select to invoke the regular annotation module or the dynamic evaluation module based on the abnormal fuzzy category of the action.

[0011] A standard annotation module, which is connected to the invoking annotation module, is used to mark the video image as a fake video;

[0012] A dynamic evaluation module, connected to the invocation annotation module, is used to evaluate and annotate the observed object, including:

[0013] The video image is segmented into frames. Based on the homogeneity deviation of facial expressions, it is determined whether the observed object conforms to the facial dynamic benchmark. Based on the video images corresponding to several homogeneous facial expressions in the action change time domain segment, the average synchronization time difference between facial texture and muscle movement and the deviation of action recovery time corresponding to the observed object are extracted. The forgery tendency representation parameters of the video image are evaluated to determine whether the video image should be labeled and the video image is identified and analyzed.

[0014] Furthermore, the object analysis module is used to calculate the action state representation value for the observed object, including:

[0015] The sum of the ratio of the uniform amplitude of the movement part change to the threshold of the uniform amplitude of the movement part change and the ratio of the uniformity of the pause duration to the threshold of the uniformity is used as the first movement state feature.

[0016] The ratio of the frequency of contour jitter of the observed object to the contour jitter frequency threshold is used as the second action state feature;

[0017] The weighted sum of the first action state feature and the second action state feature is used to determine the action state representation value.

[0018] Furthermore, the object analysis module is used to classify the observed object's actions into ambiguous categories, including:

[0019] If the action state representation value of the observed object is greater than or equal to the action state representation threshold, then the action abnormality fuzzy category of the observed object is classified into the low fuzzy category.

[0020] If the action state representation value of the observed object is less than the action state representation threshold, then the action abnormality fuzzy category of the observed object is classified as a high fuzzy category.

[0021] Furthermore, the dynamic evaluation module is used to select between calling the regular annotation module or the dynamic evaluation module, including:

[0022] If the observed object's action is abnormally ambiguous and belongs to the low ambiguity category, then the regular annotation module will be called.

[0023] If the observed object's action is classified as a highly ambiguous category, then the dynamic evaluation module will be invoked.

[0024] Furthermore, the dynamic evaluation module is used to determine whether the observed object conforms to the facial dynamic benchmark, including:

[0025] If the deviation of the homogeneity of the facial expressions corresponding to the observed object is less than the threshold of the deviation of the homogeneity, then the observed object is determined to conform to the facial dynamic benchmark.

[0026] Furthermore, the dynamic evaluation module is used to evaluate the forgery tendency characterization parameters of the video image, including:

[0027] The ratio of the average synchronization time difference between facial texture and muscle movement to the average synchronization time difference threshold is used as the first forgery tendency feature;

[0028] The ratio of the deviation of the action recovery time to the threshold of the action recovery time is used as the second forgery tendency feature;

[0029] The sum of the first forgery tendency feature and the second forgery tendency feature is used as the forgery tendency characterization parameter.

[0030] Furthermore, the dynamic evaluation module is used to determine whether to annotate the video image, including:

[0031] If the forgery tendency characterization parameter of a video image is greater than or equal to the forgery tendency characterization parameter threshold, then the video image is labeled.

[0032] Furthermore, the dynamic evaluation module is used to perform recognition and analysis on the video images, including:

[0033] Construct a time-domain variation curve of the outline edge sharpness corresponding to the observed object, determine the fluctuation value and maximum fluctuation slope of the outline edge sharpness, and determine whether the video image is a fake video.

[0034] Furthermore, the dynamic evaluation module is used to determine the average synchronization time difference between facial texture and muscle movement corresponding to the observed object, including:

[0035] Used to extract the start and end timestamps of the audio time-domain segment in a video image;

[0036] Used to identify the preceding time domain segment corresponding to the start timestamp and the following time domain segment corresponding to the end timestamp;

[0037] Used to call the preceding video image of the preceding time domain segment and the following video image of the following time domain segment;

[0038] Used to determine the preceding and following change timestamps based on the moment corresponding to the maximum morphological change of the mouth in the preceding and following video images;

[0039] The average synchronization time difference is determined based on the average of the differences between the preceding change timestamp and the start timestamp, and the differences between the following change timestamp and the end timestamp.

[0040] Furthermore, the dynamic evaluation module is used to determine whether the video image is a fake video, including:

[0041] If the fluctuation value of the clarity of the contour edge is greater than the fluctuation threshold, and the maximum fluctuation slope is greater than the maximum fluctuation slope threshold, then the video image is determined to be a fake video.

[0042] Compared with existing technologies, this invention extracts corresponding facial action features by calling the image data corresponding to the observed object in the video image; it calculates the action state representation value of the observed object by combining the facial action features and the frequency of contour jitter of the observed object, so as to classify the observed object's action abnormality ambiguity category; based on the action abnormality ambiguity category, it adaptively performs evaluation and analysis, including: marking the video image as a fake video; or, evaluating and labeling the observed object. This invention relies on the precise quantitative analysis of micro-expression dynamic details and significantly improves the recognition efficiency of fake videos through differentiated hierarchical processing, achieving a dual optimization of accuracy and efficiency.

[0043] In particular, this invention includes an object analysis module that captures the dynamic details of micro-expressions of the observed object in video images. Under normal circumstances, the micro-expressions of a real human face have natural amplitude patterns and pause rhythms, and the contour jitters also conform to physiological movement logic. However, the synthesized expressions in deepfake videos often have problems such as inconsistent amplitude of movements, irregular pause times, and chaotic contour jitters. Because the changes in the movement parts of real micro-expressions follow the physiological muscle movement laws, the amplitude changes are uniform and the transitions are smooth, while the synthesized expressions in fake videos are prone to inconsistent amplitude of movements and abrupt changes because the algorithm cannot accurately replicate the muscle movement logic. Based on this, when performing subtle anomaly analysis on the facial area of ​​the observed object, the uniform amplitude of the changes in the movement parts reflects the consistency and regularity of the amplitude of the key facial movement parts during the expression change process. Furthermore, the pauses in genuine facial expressions exhibit a natural temporal regularity, with minimal variation in pause duration among similar expressions. In contrast, due to frame sequence splicing errors and defects in expression template reuse, the pause durations of similar expressions in forged videos tend to fluctuate significantly. The uniformity of pause duration reflects the consistency of pause duration during the transition of facial expressions. In addition, this invention considers not only the subtle changes in key facial movement areas but also the stability of the facial contour in the video image. Essentially, it quantifies the frequency of irregular jittering at the edges of the facial contour. Because of physiological factors such as muscle movement and respiration, the contour jitter of a genuine face exhibits a natural and smooth frequency regularity; however, due to defects in inter-frame synthesis algorithms and feature matching errors, forged videos are prone to abnormal contour jitter frequencies, resulting in excessively frequent jittering and irregular fluctuations. Therefore, this invention calculates a representation value of the action state of the observed object to characterize the degree of dynamic anomaly presented by the observed object in the video image, quantifying the natural deviation between facial movements and contour dynamics, providing data support for subsequent classification of ambiguous action anomaly categories. This invention can accurately capture hidden flaws in fake videos, significantly reduce the false negative rate, and improve the sensitivity of fake video identification.

[0044] In particular, this invention incorporates a dynamic evaluation module for refined analysis of highly blurred videos, avoiding the omission of hidden anomalies in forged videos. Through pre-screening using frame segmentation and homogenized expression judgment, subsequent analysis focuses only on video image segments with questionable facial dynamics of the observed subject, capturing dynamic anomalies through inter-frame temporal comparison. Under normal circumstances, changes in texture, pore contraction, and muscle movement of a real human face are synchronously linked, exhibiting natural synergy under physiological neural control. However, forged videos, due to the algorithm's inability to accurately replicate the physiological correlation between muscle movement and texture changes, are prone to asynchrony. The average synchronization time difference between facial texture and muscle movement characterizes the degree of deviation in the synergy between muscle movement and texture changes. Furthermore, the recovery of expressions on a real human face is influenced by muscle elasticity and physiological inertia, resulting in a fixed time consumption. Forged videos, due to their spliced ​​facial expression templates and abrupt frame sequence transitions, are prone to abnormal recovery times, such as instantaneous restoration to a static state or excessively long recovery times exceeding physiological limits. This invention characterizes the naturalness of the action recovery process by measuring the deviation in action recovery time, reflecting the coherence of dynamic transitions between video frames. Abnormal recovery time is essentially a defect in the inter-frame connection of the synthesized video during facial expression switching, effectively capturing forgery traces such as abrupt facial expressions and distorted transitions. Therefore, this invention uses a forgery tendency characterization parameter to represent the obviousness of forgery traces in video images, quantifying the degree of forgery risk present in the video image, and providing data support for subsequent determination of whether to annotate the video image. This invention can accurately capture subtle forgery traces, significantly reducing the false negative rate of forged videos and improving the reliability of identification.

[0045] In particular, this invention performs in-depth verification of the authenticity of annotated video images. For example, if the initial video image has a high risk of forgery, subsequent verification is conducted by checking for abnormal fluctuations in contour sharpness, thereby improving the reliability of the recognition results. Therefore, this invention addresses the issue that in forged videos, due to algorithmic precision limitations during frame-to-frame stitching and feature fusion, unnatural phenomena such as alternating periods of sharpness and blurriness, and abrupt edge changes in facial contour edges are prone to occur. Specifically, the sharpness of facial contour edges in genuine videos is affected by physiological movement and ambient lighting, exhibiting gradual and regular temporal variations; while the sharpness of contour edges in forged videos fluctuates irregularly due to inter-frame synthesis errors. Therefore, this invention constructs a temporal variation curve of contour edge sharpness, combining the determination of fluctuation value and maximum fluctuation slope to accurately capture inter-frame synthesis defects in forged videos, significantly improving recognition accuracy. Attached Figure Description

[0046] Figure 1 This is a functional block diagram of a micro-expression intelligent recognition system based on facial images, as described in an embodiment of the invention.

[0047] Figure 2A logic decision diagram for classifying the ambiguous categories of abnormal actions of observed objects in an embodiment of the invention;

[0048] Figure 3 A logic decision diagram for selecting to call the conventional annotation module or the dynamic evaluation module in the embodiments of the invention;

[0049] Figure 4 This is a logic diagram for determining whether to annotate video images according to an embodiment of the invention. Detailed Implementation

[0050] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0051] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0052] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0053] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral connection; it can refer to a mechanical connection or an electrical connection. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0054] Please see Figure 1 The diagram shown is a functional block diagram of a micro-expression intelligent recognition system based on facial images according to an embodiment of the present invention. The micro-expression intelligent recognition system based on facial images according to an embodiment of the present invention includes:

[0055] The detail extraction module is used to call the image data corresponding to the observed object in the video image to extract the corresponding facial action features. The facial action features include the uniformity of the change in the action part and the uniformity of the pause duration.

[0056] The object analysis module, which is connected to the detail extraction module, is used to combine the facial action features and the contour shaking frequency of the observed object to calculate the action state representation value of the observed object, so as to classify the action abnormal fuzzy category of the observed object.

[0057] The annotation module is invoked, which is connected to the object analysis module, and is used to select to invoke the regular annotation module or the dynamic evaluation module based on the abnormal fuzzy category of the action.

[0058] A standard annotation module, which is connected to the invoking annotation module, is used to mark the video image as a fake video;

[0059] A dynamic evaluation module, connected to the invocation annotation module, is used to evaluate and annotate the observed object, including:

[0060] The video image is segmented into frames. Based on the homogeneity deviation of facial expressions, it is determined whether the observed object conforms to the facial dynamic benchmark. Based on the video images corresponding to several homogeneous facial expressions in the action change time domain segment, the average synchronization time difference between facial texture and muscle movement and the deviation of action recovery time corresponding to the observed object are extracted. The forgery tendency representation parameters of the video image are evaluated to determine whether the video image should be labeled and the video image is identified and analyzed.

[0061] Specifically, the image data includes facial motion features, the frequency of contour jitter of the observed object, the degree of deviation in the homogeneity of facial expressions, the clarity of contour edges, the average synchronization time difference between the facial texture and muscle movement of the observed object, and the deviation in motion recovery time, etc. The observed object refers to the person in the video image.

[0062] Specifically, there are no restrictions on the specific structure of the detail extraction module, object analysis module, call annotation module, regular annotation module, and dynamic evaluation module. Each module or its units can be composed of logical components or combinations of logical components. Logical components include field-programmable processors, computers, or microprocessors in computers.

[0063] Specifically, the object analysis module is used to calculate the action state representation value for the observed object, including:

[0064] The sum of the ratio of the uniform amplitude of the movement part change to the threshold of the uniform amplitude of the movement part change and the ratio of the uniformity of the pause duration to the threshold of the uniformity is used as the first movement state feature.

[0065] The ratio of the frequency of contour jitter of the observed object to the contour jitter frequency threshold is used as the second action state feature;

[0066] The weighted sum of the first action state feature and the second action state feature is used to determine the action state representation value.

[0067] In this embodiment, the standard deviation of the displacement difference of the action part in consecutive video frames is calculated, and the standard deviation is used as the uniform amplitude of the action change. The smaller the standard deviation, the more uniform the amplitude change and the closer it is to a real expression; the larger the standard deviation, the more disordered the amplitude change.

[0068] To quantify the regularity of pauses observed in the subject and measure the degree of fluctuation in pause duration, the variance of pause duration for multiple similar expressions is calculated to distinguish between natural pauses and abnormal pauses caused by synthetic splicing. Furthermore, the variance is used as the uniformity of the pause duration. The smaller the variance, the better the uniformity of the pause duration, meaning the pauses are more regular; the larger the variance, the worse the uniformity of the pause duration, meaning the pauses are less regular. This will not be elaborated further.

[0069] Specifically, the contour jitter frequency is determined based on the number of times the contour of the observed object jitters in a single video image, which will not be elaborated further here.

[0070] Specifically, the uniformity of the amplitude of changes in movement parts and the uniformity of pause duration directly reflect the dynamic movement logic of the observed object's facial expressions, which are physiological characteristics that deepfake technology struggles to accurately replicate. While the frequency of contour jitter primarily reflects the static edge stability of the facial contour, and can reveal errors in the inter-frame synthesis of forged videos, such as abnormal edge jitter, it is significantly affected by external environmental factors. For example, video compression or sudden changes in lighting can increase the frequency of normal contour jitter, making its distinguishability as a single feature weaker than the first action state feature. Therefore, in implementation, the first action state feature, calculated based on facial action features—namely, the uniformity of the amplitude of changes in movement parts and the uniformity of pause duration—is given a higher weighting coefficient, set to 0.6. Correspondingly, the weighting coefficient for the second action state feature, calculated based on the frequency of contour jitter of the observed object, is set to 0.4.

[0071] In this embodiment, the purpose of setting thresholds for the uniform amplitude of motion changes, uniformity, and contour jitter frequency is to characterize situations where the observed object exhibits a high degree of dynamic abnormality and a high degree of natural deviation between facial movements and contour dynamics. By acquiring a large amount of video sample data of real observed objects in normal facial dynamic scenarios, and by calling the data on the uniform amplitude of motion changes, the uniformity data of pause duration, and the contour jitter frequency data of the observed object, the mean of the uniform amplitude of motion changes, the mean of uniformity, and the mean of contour jitter frequency are calculated and used as the baseline values ​​under normal conditions. The purpose of the threshold is to determine the uniformity threshold of the movement part change as the product of the average uniformity amplitude of the movement part change and the first deviation coefficient, the uniformity threshold as the product of the average uniformity and the second deviation coefficient, and the contour jitter frequency threshold as the product of the average contour jitter frequency and the third deviation coefficient. The first deviation coefficient is selected within the interval [1.2, 1.3], preferably 1.2 in practice; the second deviation coefficient is selected within the interval [1.3, 1.4], preferably 1.3 in practice; and the third deviation coefficient is selected within the interval [1.1, 1.2], preferably 1.2 in practice.

[0072] Specifically, this invention includes an object analysis module that captures the dynamic details of micro-expressions of the observed object in video images. Under normal circumstances, the micro-expressions of a real human face, such as blinking eyes and subtle twitching of the corners of the mouth, have natural amplitude patterns and pause rhythms, and the contour shaking also conforms to physiological movement logic. In contrast, the synthesized expressions in deepfake videos often exhibit problems such as fluctuating amplitude of movements, irregular pause times, and chaotic contour shaking. Because the changes in the movement parts of real micro-expressions, such as the opening and closing of the eyelids when blinking and the upward movement of the corners of the mouth when smiling, follow the laws of physiological muscle movement, the amplitude changes are uniform and the transitions are smooth. However, the synthesized expressions in fake videos, because the algorithm cannot accurately replicate the muscle movement logic, are prone to fluctuating amplitude of movements and abrupt changes, such as a sudden, large upward movement of the corners of the mouth followed by a sharp drop. Based on this, when performing subtle anomaly analysis on the facial area of ​​the observed object, the uniform amplitude of the changes in the movement parts reflects the consistency and regularity of the movement amplitude of key facial movement parts of the observed object, such as the eyes, corners of the mouth, and forehead, during the expression change process. Furthermore, the pauses in genuine facial expressions, such as the duration of a smile or the pause before relaxation after a frown, exhibit natural temporal regularity, with minimal variation in pause duration among similar expressions. In contrast, due to frame sequence splicing errors and defects in expression template reuse, the pause durations of similar expressions in forged videos tend to fluctuate significantly. The uniformity of pause duration reflects the consistency of pause durations during the transition of facial expressions. Moreover, this invention, in addition to considering the subtle changes in key facial movement areas, also considers the stability of the facial contours presented in the video image. Essentially, it quantifies the frequency of irregular jittering at the edges of the facial contours. Because of physiological factors such as muscle movement and respiration, the contour jitter of a genuine face exhibits a natural and smooth frequency regularity; while forged videos, due to defects in inter-frame synthesis algorithms and feature matching errors, are prone to abnormal contour jitter frequencies, resulting in excessively frequent jittering and irregular fluctuations. Therefore, this invention calculates the action state representation value of the observed object to characterize the degree of dynamic anomaly presented by the observed object in the video image, quantifies the natural deviation between facial movements and contour dynamics, and provides data support for subsequent classification of motion anomaly ambiguity categories. This invention can accurately capture hidden flaws in forged videos, significantly reduce the false negative rate, and improve the sensitivity of forged video identification.

[0073] Specifically, please refer to Figure 2 As shown, this is a logic decision diagram for classifying the ambiguous categories of abnormal actions of the observed object according to an embodiment of the present invention. The object analysis module is used to classify the ambiguous categories of abnormal actions of the observed object, including:

[0074] If the action state representation value of the observed object is greater than or equal to the action state representation threshold, then the action abnormality fuzzy category of the observed object is classified into the low fuzzy category.

[0075] If the action state representation value of the observed object is less than the action state representation threshold, then the action abnormality fuzzy category of the observed object is classified as a high fuzzy category.

[0076] The motion state representation threshold is predetermined. The motion state representation value calculated is determined by making the uniform amplitude of the change of the motion part equal to the uniform amplitude threshold of the change of the motion part, the uniformity of the pause duration equal to the uniformity threshold, and the frequency of contour jitter of the observed object equal to the contour jitter frequency threshold.

[0077] Specifically, please refer to Figure 3 As shown, this is a logic decision diagram for selecting to call the conventional annotation module or the dynamic evaluation module in an embodiment of the present invention. The dynamic evaluation module is used to select to call the conventional annotation module or the dynamic evaluation module, including:

[0078] If the observed object's action is abnormally ambiguous and belongs to the low ambiguity category, then the regular annotation module will be called.

[0079] If the observed object's action is classified as a highly ambiguous category, then the dynamic evaluation module will be invoked.

[0080] Specifically, the dynamic evaluation module is used to determine whether the observed object conforms to the facial dynamic benchmark, including:

[0081] If the deviation of the homogeneity of the facial expressions corresponding to the observed object is less than the threshold of the deviation of the homogeneity, then the observed object is determined to conform to the facial dynamic benchmark.

[0082] In this embodiment, the purpose of setting a homogenization quantity deviation threshold is to characterize the situation where the homogenization quantity of the expressions presented by the observed object deviates significantly from the baseline homogenization quantity under normal facial dynamics. By acquiring a large amount of video sample data of real observed objects in normal facial dynamic scenes, calling the homogenization quantity deviation data and homogenization quantity data of facial expressions, the mean of homogenization quantity deviation and the mean of homogenization quantity are solved. Based on the purpose of setting the homogenization quantity deviation threshold, the homogenization quantity deviation threshold is determined as the product of the mean of homogenization quantity deviation and the homogenization deviation coefficient. The homogenization deviation coefficient is selected in the interval [0.9, 0.95], preferably 0.9 in practice. The mean of homogenization quantity is used as the baseline homogenization quantity.

[0083] Specifically, based on several facial expressions of the observed objects in the video images, expressions that meet the criteria of high similarity and the corresponding number of video frames are identified, and the maximum value of the number of video frames is taken as the homogenization quantity.

[0084] Specifically, the high similarity refers to the situation where several expressions of the same observed object in a video image are highly similar. This is determined by quantifying the similarity of key facial features. For example, the coordinate deviation of key facial action parts such as the eyes, corners of the mouth, and forehead in consecutive frames, the cosine similarity of facial textures, and the overlap of contour edges are calculated. If the similarity between any expression and any other expression is greater than the similarity threshold, the two expressions are determined to be homogeneous, and the corresponding number of homogeneous expressions is determined. Based on the purpose of capturing and recognizing highly similar expressions, the similarity threshold is set to 90%, which will not be elaborated further.

[0085] In this embodiment, the homogenization quantity deviation is determined in the following manner:

[0086] Used to calculate the absolute value of the difference between the homogenized quantity and the benchmark homogenized quantity;

[0087] The ratio of the absolute value of the quantity difference to the benchmark homogenized quantity is used as the deviation of the homogenized quantity.

[0088] Specifically, the dynamic evaluation module is used to evaluate the forgery tendency characterization parameters of the video image, including:

[0089] The ratio of the average synchronization time difference between facial texture and muscle movement to the average synchronization time difference threshold is used as the first forgery tendency feature;

[0090] The ratio of the deviation of the action recovery time to the threshold of the action recovery time is used as the second forgery tendency feature;

[0091] The sum of the first forgery tendency feature and the second forgery tendency feature is used as the forgery tendency characterization parameter.

[0092] In this embodiment, the purpose of setting the average synchronization time difference threshold and the motion recovery time deviation threshold is to characterize situations where the forgery traces of video images are relatively obvious and the degree of forgery risk is high. By acquiring a large amount of video sample data of real observed objects in normal facial dynamic scenes, the average synchronization time difference data of facial texture and muscle movement, the motion recovery time deviation data, and the motion recovery time data are called to solve the mean of the average synchronization time difference, the mean of the motion recovery time deviation, and the mean of the motion recovery time, and these are used as the benchmark values ​​under normal conditions. Based on the purpose of setting the above two thresholds, the average synchronization time difference threshold is determined to be the product of the mean of the average synchronization time difference and the synchronization deviation coefficient, and the motion recovery time deviation threshold is determined to be the product of the mean of the motion recovery time deviation and the time deviation coefficient. The synchronization deviation coefficient is selected in the interval [1.1, 1.2], preferably 1.1 in the implementation, and the time deviation coefficient is selected in the interval [1.2, 1.3], preferably 1.2 in the implementation. The mean of the motion recovery time is used as the benchmark motion recovery time.

[0093] In this embodiment, the deviation of the action recovery time is determined in the following manner, including:

[0094] Used to calculate the absolute value of the difference between the motion recovery time and the baseline motion recovery time;

[0095] The ratio of the absolute value of the time difference to the time taken to recover the baseline action is used as the deviation of the action recovery time.

[0096] Specifically, this invention includes a dynamic evaluation module that performs refined analysis on videos with high fuzziness, avoiding the omission of hidden anomalies in forged videos. Through frame segmentation and pre-screening using homogeneous expression judgment, subsequent analysis is performed only on video image segments where facial dynamics of the observed object are questionable. Dynamic anomalies are captured through inter-frame temporal comparison. Under normal circumstances, changes in the texture of a real face, such as skin wrinkles and pore contraction, are synchronized with muscle movements, such as the contraction of the masticatory muscle and the movement of the orbicularis oculi muscle, exhibiting natural synergy under physiological neural control. However, forged videos, because the algorithm cannot accurately replicate the physiological correlation between muscle movement and texture changes, are prone to asynchrony, such as the muscles ceasing movement while the texture continues to change abnormally, or muscle movement preceding texture changes. The average synchronization time difference between facial texture and muscle movement characterizes the degree of deviation in the synergy between muscle movement and texture changes. Furthermore, the recovery of expressions on a real face, such as from laughter to calmness, or from frowning to relaxation, is affected by muscle elasticity and physiological inertia, resulting in a fixed timeframe. Forged videos, due to their abrupt splicing of facial expression templates and jarring frame sequence transitions, are prone to abnormal recovery times, such as instantaneous restoration to a static state or recovery times far exceeding physiological limits. The deviation of action recovery time characterizes the naturalness of the action recovery process, reflecting the coherence of dynamic transitions between video frames. Abnormal recovery time is essentially a defect in the inter-frame connection of the synthesized video during facial expression switching, effectively capturing forgery traces of abrupt facial expressions and distorted transitions. Therefore, this invention uses a forgery tendency characterization parameter to represent the obviousness of forgery traces in video images, quantifying the degree of forgery risk present in the video image, and providing data support for subsequent determination of whether to annotate the video image. This invention can accurately capture subtle forgery traces, significantly reducing the false negative rate of forged videos and improving the reliability of identification.

[0097] Specifically, please refer to Figure 4 As shown, this is a logic diagram for determining whether to annotate a video image according to an embodiment of the present invention. The dynamic evaluation module is used to determine whether to annotate the video image, including:

[0098] If the forgery tendency characterization parameter of a video image is greater than or equal to the forgery tendency characterization parameter threshold, then the video image is labeled.

[0099] If the forgery tendency representation parameter of the video image is less than the forgery tendency representation parameter threshold, then there is no need to annotate the video image.

[0100] The threshold for the forgery tendency characterization parameter is predetermined. The forgery tendency characterization parameter calculated is determined as the threshold when the average synchronization time difference between facial texture and muscle movement is equal to the average synchronization time difference threshold, and the deviation of action recovery time is equal to the deviation of action recovery time threshold.

[0101] Specifically, the dynamic evaluation module is used to identify and analyze the video images, including:

[0102] Construct a time-domain variation curve of the outline edge sharpness corresponding to the observed object, determine the fluctuation value and maximum fluctuation slope of the outline edge sharpness, and determine whether the video image is a fake video.

[0103] It is understandable that a clear contour edge has a narrow and steep gradient distribution, that is, the pixel gradient amplitude is concentrated and has small differences, while a blurry contour edge has a wide and gentle gradient distribution, that is, the pixel gradient amplitude is dispersed and has large differences. Therefore, in this embodiment, the video image is converted to grayscale to obtain the corresponding grayscale image; the Canny edge detection algorithm is used to extract the contour edge of the observed object; the gradient amplitude set of the contour edge pixels is extracted, and the mean and variance of the gradient amplitude set are calculated. The ratio of the mean to the variance is used as the contour edge sharpness.

[0104] The larger the ratio, the more concentrated the gradient amplitude and the sharper the outline, i.e., the higher the clarity; the smaller the ratio, the more dispersed the gradient amplitude and the blurrier the outline, i.e., the lower the clarity.

[0105] Specifically, the dynamic evaluation module is used to construct a time-domain variation curve of the contour edge sharpness corresponding to the observed object, including:

[0106] This is used to construct a Cartesian coordinate system with time as the horizontal axis and the sharpness of the outline edge as the vertical axis.

[0107] In the Cartesian coordinate system, mark the coordinate points of the outline edge sharpness at each moment;

[0108] Connect the coordinate points with a smooth curve to obtain the time-domain variation curve of the edge sharpness of the contour.

[0109] Specifically, there are no restrictions on the method for constructing the time-domain variation curve of contour edge sharpness. For example, the time-domain variation curve of contour edge sharpness can be fitted using Matlab correlation fitting software, which will not be elaborated further.

[0110] Specifically, the dynamic evaluation module is used to determine the average synchronization time difference between facial texture and muscle movement corresponding to the observed object, including:

[0111] Used to extract the start and end timestamps of the audio time-domain segment in a video image;

[0112] Used to identify the preceding time domain segment corresponding to the start timestamp and the following time domain segment corresponding to the end timestamp;

[0113] Used to call the preceding video image of the preceding time domain segment and the following video image of the following time domain segment;

[0114] Used to determine the preceding and following change timestamps based on the moment corresponding to the maximum morphological change of the mouth in the preceding and following video images;

[0115] The average synchronization time difference is determined based on the average of the differences between the preceding change timestamp and the start timestamp, and the differences between the following change timestamp and the end timestamp.

[0116] In this embodiment, the time domain segments formed by extending the start timestamp forward by one unit of time and the end timestamp backward by one unit of time are respectively referred to as the preceding time domain segment and the following time domain segment, which will not be elaborated further.

[0117] Specifically, the maximum morphological change refers to the degree of morphological change of the mouth of the observed object between consecutive video frames in both the preceding and subsequent time domains, quantifying the most obvious range of motion of the mouth during dynamic changes. In practice, the matching degree can be determined by comparing mouth images in adjacent frames, and this matching degree can be used as the morphological change amount to determine the maximum morphological change of the mouth.

[0118] The calculation of the matching degree can be any method that can be implemented in the existing technology, which will not be elaborated further.

[0119] Specifically, the dynamic evaluation module is used to determine whether the video image is a fake video, including:

[0120] If the fluctuation value of the clarity of the contour edge is greater than the fluctuation threshold, and the maximum fluctuation slope is greater than the maximum fluctuation slope threshold, then the video image is determined to be a fake video.

[0121] In this embodiment, the purpose of setting the fluctuation threshold and the maximum fluctuation slope threshold is to characterize the situation where the outline edge of the observed object in the video image is relatively blurry and the possibility of inter-frame synthesis defects is relatively high. By acquiring a large amount of video sample data of real observed objects in normal facial dynamic scenes, calling the fluctuation value data of outline edge clarity and the maximum fluctuation slope data, the mean fluctuation value and the mean maximum fluctuation slope value are calculated and used as the benchmark value under normal conditions. Based on the purpose of setting the above two thresholds, the fluctuation threshold is determined as the product of the mean fluctuation value and the first offset coefficient, and the maximum fluctuation slope threshold is determined as the product of the mean maximum fluctuation slope value and the second offset coefficient. The first offset coefficient is selected in the interval [1.1, 1.15], preferably 1.1 in the implementation, and the second offset coefficient is selected in the interval [1.05, 1.1], preferably 1.05 in the implementation.

[0122] Specifically, this invention performs in-depth verification of the authenticity of labeled video images. For example, if the initial video image has a high risk of forgery, subsequent verification is conducted by checking for abnormal fluctuations in contour sharpness, thereby improving the reliability of the recognition results. Therefore, this invention addresses the issue that in forged videos, due to algorithmic precision limitations during frame-to-frame stitching and feature fusion, unnatural phenomena such as alternating sharpness and blurriness, and abrupt edge changes in facial contour edges are prone to occur. In genuine videos, facial contour edge sharpness is affected by physiological movement and ambient lighting, exhibiting gradual and regular temporal changes; while in forged videos, contour edge sharpness fluctuates irregularly due to inter-frame synthesis errors, such as sharp contours in a few frames followed by sudden blurring in adjacent frames, or abrupt increases or decreases in edge sharpness. Therefore, this invention constructs a temporal variation curve for contour edge sharpness, combining fluctuation values ​​and the maximum fluctuation slope for dual determination, thereby accurately capturing inter-frame synthesis defects in forged videos and significantly improving recognition accuracy.

[0123] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A micro-expression intelligent recognition system based on a human face image, characterized in that, The method comprises the following steps: a detail extraction module is used to call the picture data corresponding to the observation object in the video image to extract the corresponding facial action features, the facial action features including the uniformity of the uniform amplitude of the action part change and the uniformity of the pause duration; an object analysis module is used to calculate the action state representation value of the observation object in combination with the facial action features and the contour jitter frequency of the observation object, to divide the action abnormal fuzzy category of the observation object; an invocation labeling module is used to select to invoke the conventional labeling module or the dynamic evaluation module based on the action abnormal fuzzy category; the conventional labeling module is used to mark the video image as a fake video; the dynamic evaluation module is used to evaluate and label the observation object, including, performing frame number cutting operation on the video image, judging whether the observation object conforms to the facial dynamic benchmark based on the homogenization number deviation of facial expressions, extracting the average synchronization time difference of facial texture and muscle movement and the action recovery time consumption deviation of the observation object corresponding to the video image of a number of homogenized facial expressions in the action change time domain segment, evaluating the fake tendency representation parameter of the video image, and determining whether to label the video image, and performing recognition analysis on the video image.

2. The micro-expression intelligent recognition system based on facial images according to claim 1, characterized in that, The object analysis module is used to calculate the action state representation value of the observation object, including: using the sum of the ratio of the uniform amplitude of the action part change to the uniform amplitude threshold value of the action part change and the ratio of the uniformity of the pause duration to the uniformity threshold value as the first action state feature; using the ratio of the contour jitter frequency of the observation object to the contour jitter frequency threshold value as the second action state feature; using the weighted sum of the first action state feature and the second action state feature as the action state representation value. 3.The micro-expression intelligent recognition system based on facial images according to claim 2, characterized in that, The object analysis module is used to divide the action abnormal fuzzy category of the observation object, including: if the action state representation value of the observation object is greater than or equal to the action state representation threshold value, the action abnormal fuzzy category of the observation object is divided into a low fuzzy category; if the action state representation value of the observation object is less than the action state representation threshold value, the action abnormal fuzzy category of the observation object is divided into a high fuzzy category.

4. The micro-expression intelligent recognition system based on facial images according to claim 3, characterized in that, The dynamic evaluation module is used to select to invoke the conventional labeling module or the dynamic evaluation module, including: if the action abnormal fuzzy category of the observation object is a low fuzzy category, the conventional labeling module is selected to be invoked; if the action abnormal fuzzy category of the observation object is a high fuzzy category, the dynamic evaluation module is selected to be invoked. 5.The micro-expression intelligent recognition system based on facial images according to claim 1, characterized in that, The dynamic evaluation module is used to judge whether the observation object conforms to the facial dynamic benchmark, including: if the homogenization number deviation of the facial expression corresponding to the observation object is less than the homogenization number deviation threshold value, it is judged that the observation object conforms to the facial dynamic benchmark. 6.The micro-expression intelligent recognition system based on facial images according to claim 1, characterized in that, The dynamic evaluation module is used to evaluate the fake tendency representation parameter of the video image, including: using the ratio of the average synchronization time difference of facial texture and muscle movement to the average synchronization time difference threshold value as the first fake tendency feature; a ratio of the action recovery time consumption deviation to an action recovery time consumption deviation threshold value as a second forgery tendency feature; a sum of the first forgery tendency feature and the second forgery tendency feature as the forgery tendency representation parameter.

7. The micro-expression intelligent recognition system based on facial images according to claim 6, characterized in that, The dynamic evaluation module is configured to determine whether to label the video image, including: If the forgery tendency representation parameter of the video image is greater than or equal to a forgery tendency representation parameter threshold value, the video image is labeled. 8.The micro-expression intelligent recognition system based on facial images of claim 1, characterized in that, The dynamic evaluation module is configured to perform recognition analysis on the video image, including: constructing a contour edge sharpness time domain variation curve corresponding to the observation object, determining a fluctuation value of the contour edge sharpness and a maximum fluctuation slope, to determine whether the video image is a fake video. 9.The micro-expression intelligent recognition system based on facial images of claim 8, characterized in that, The dynamic evaluation module is configured to determine an average synchronization time difference of facial texture and muscle movement corresponding to the observation object, including: extracting a start time stamp and an end time stamp of a sound time domain segment in the video image; identifying a pre-sequence time domain segment corresponding to the start time stamp and a post-sequence time domain segment corresponding to the end time stamp; calling pre-sequence video images of the pre-sequence time domain segment and post-sequence video images of the post-sequence time domain segment; based on a pre-sequence change time stamp and a post-sequence change time stamp corresponding to a time instant of a maximum morphological change amount of a mouth in the pre-sequence video images and the post-sequence video images; based on a mean value corresponding to a difference between the pre-sequence change time stamp and the start time stamp and a difference between the post-sequence change time stamp and the end time stamp. 10.The micro-expression intelligent recognition system based on facial images according to claim 8, characterized in that, The dynamic evaluation module is configured to determine whether the video image is a fake video, including: If the fluctuation value of the contour edge sharpness is greater than a fluctuation threshold value, and the maximum fluctuation slope is greater than a maximum fluctuation slope threshold value, the video image is determined to be a fake video.

Citation Information

Patent Citations

  • Micro-expression detection method based on optical flow

    CN111461021A

  • Compression depth counterfeit video detection method based on facial muscle movement

    CN115984740A

  • Video forgery detection method and device, equipment and medium

    CN120032287A