Method, system and device for distinguishing face generated by AI (Artificial Intelligence) and medium
By constructing a facial micro-expression anomaly detection and dynamic face temporal consistency model, and combining convolutional neural networks and recurrent neural networks, the accuracy and robustness problems of AI-generated face detection technology in complex scenarios are solved, and efficient AI-generated face identification is achieved.
Patent Information
- Application Number
- CN202510684609.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-16
AI Technical Summary
Existing AI-generated face detection technology lacks adaptability in complex scenarios, is highly data-dependent, has limited generalization capabilities, and lags behind in technological iteration, resulting in insufficient detection accuracy and robustness, making it difficult to effectively identify AI-generated faces.
A facial micro-expression anomaly detection model and a dynamic face temporal consistency model are constructed. By combining convolutional neural networks and recurrent neural networks, micro-expression details and dynamic temporal consistency are analyzed. The cross-entropy loss function and adaptive optimizer are used to construct a high-quality training set and perform model evaluation and optimization. Finally, detection is performed through dual-model decision fusion.
It significantly improves the detection accuracy and robustness of AI-generated faces, breaks through the limitations of single-dimensional feature detection, enhances the recognition accuracy of generated forged traces and the generalization ability of the model, and reduces the risk of misjudgment.
Smart Images

Figure CN120656212A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition technology, and more specifically relates to a method, system, device and medium for recognizing AI-generated faces. Background Art
[0002] With the development of deep learning technology, AI-generated face recognition has made significant progress. Generative adversarial networks (GANs) can generate highly realistic virtual facial images from random noise through adversarial learning between generators and discriminators. Variational autoencoders (VAEs) generate new samples by learning the characteristic distribution of real facial data. These technologies have broad application value in fields such as artistic creation and game development, but they also pose serious security challenges. Malicious users could use generated fake faces to commit fraud, spread rumors, or infringe on original creative rights, disrupting the industry ecosystem. Therefore, developing high-precision AI-generated face detection technology is key to addressing these risks. Currently, mainstream AI-generated face detection methods are primarily based on physical features and deep learning models. One approach uses image algorithms to analyze multi-dimensional facial information, such as pixel texture and semantic features, to identify traces of AI generation. Another approach builds models such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), relying on large amounts of labeled data to train classifiers and automatically distinguish between real and generated faces. Furthermore, a hybrid framework combining CNNs and visual transformers (ViTs) has demonstrated some generalized recognition capabilities for forged samples generated using unknown techniques, trained on large-scale, multi-source deepfake data. However, this approach still lags behind in technological iteration. Although the above-mentioned mainstream AI-generated face detection methods can realize the recognition of AI-generated faces, such technologies still have many problems in accurately identifying AI-generated faces: First, they lack adaptability to complex scenes. Under low-resolution, light interference or occlusion conditions, traditional feature analysis methods are easily affected by noise, resulting in incomplete feature extraction and a significant decrease in detection accuracy. Second, they are highly data-dependent. Deep learning models require a large amount of high-quality labeled data to support training. Data collection and labeling are costly, and data bias will directly affect model performance. Third, generalization ability is limited. Models trained for specific algorithms or scenarios are difficult to adapt to new generation technologies or face images of diverse scenarios, and are prone to overfitting. Fourth, technology iteration lags behind. Although existing hybrid frameworks have certain out-of-domain recognition capabilities, they are unable to capture new forged features in a timely manner in the face of rapidly updated AI face-changing technologies, and there is a risk of misjudgment or omission. It can be seen that these problems restrict the actual application effect of detection technology. How to improve the accuracy of AI-generated face recognition is an issue that we urgently need to solve. Summary of the Invention
[0003] In response to the above problems, the purpose of the present invention is to provide a method, system, device and medium for identifying AI-generated faces. By integrating micro-expression anomaly detection and dynamic temporal consistency analysis, it effectively combines spatial details and temporal coherence features, and significantly improves the detection accuracy and robustness of AI-generated faces.
[0004] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: In a first aspect, embodiments of the present application provide a method for recognizing AI-generated faces, comprising: Construct facial micro-expression anomaly detection model and dynamic face temporal consistency model; Collecting facial data and AI-generated facial data to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model; Define the loss function, select the optimizer, use the first dataset to train the facial micro-expression anomaly detection model, and use the second dataset to train the dynamic face temporal consistency model; Evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively; Obtain the facial data to be detected, and use the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection respectively, and combine the detection results of the two models to identify AI-generated faces.
[0005] In an optional embodiment, the construction of the facial micro-expression anomaly detection model and the dynamic face temporal consistency model includes: Select a convolutional neural network, recurrent neural network, or long short-term memory-convolutional neural network to build a facial micro-expression anomaly detection model; set model parameters, including the number of convolutional layers, the size and number of filters, and the number of neurons in the fully connected layer; Based on convolutional neural networks, a dynamic facial temporal consistency model is constructed by integrating recurrent neural networks and their variants. An attention mechanism is introduced into the process of recurrent neural networks processing temporal information to guide the model to focus on facial parts that are critical for judging temporal consistency, enhance the model's sensitivity to temporal changes in key areas, and improve the accuracy of temporal modeling.
[0006] In an optional embodiment, the collected facial data and AI-generated facial data are used to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model, respectively, including: Collect real-time facial video and image data containing micro-expressions or from a facial database; collect AI facial video and image data containing micro-expressions generated using AI facial generation technology; Based on the collected real face video and image data, AI face video and image data, face detection algorithms are used to perform face detection, extract the face area, and generate face images. The key feature points of the face are located using the facial key point detection algorithm, and the face image is aligned to the standard position and angle. Based on the aligned facial images, features related to micro-expressions are extracted. The micro-expressions of real faces and those generated by AI are labeled to generate the first dataset, which is then augmented. Collect real face video data from public video databases and AI face video data generated using AI face generation technology; Based on the collected real face video data and AI face video data, a face detection algorithm is used to detect faces in video frames. For the detected face video frames, a face tracking algorithm is used to track the same face in consecutive frames to determine the consecutive video frames with the same face. Normalize the face video frames and cut the consecutive video frames of the same face into time-series segments according to the preset time length; The time sequence segments are labeled to indicate that each time sequence segment is real face video data or AI face video data, and a second data set is generated.
[0007] In an optional embodiment, the method of defining a loss function, selecting an optimizer, using the first data set to train a facial micro-expression anomaly detection model, and using the second data set to train a dynamic face temporal consistency model includes: Define the cross entropy loss function to measure the difference between the predicted results of the facial micro-expression anomaly detection model and the true label; Select the stochastic gradient descent optimizer, Adagrad optimizer, Adadelta optimizer, RMSProp optimizer, or Adam optimizer as the model optimizer to update the parameters of the facial micro-expression anomaly detection model to minimize the cross-entropy loss function; Based on the first dataset, the model is trained using the cross entropy loss function and the model optimizer.
[0008] In an optional embodiment, the method of defining a loss function, selecting an optimizer, using the first data set to train a facial micro-expression anomaly detection model, and using the second data set to train a dynamic face temporal consistency model further includes: The cross entropy loss function is defined as the loss function of the model to measure the difference between the prediction results of the dynamic face temporal consistency model and the true label; the cross entropy loss function is as follows:
[0009] Where N is the number of samples, is the true label, The probability that the prediction result of the dynamic face temporal consistency model is an AI-generated face; The Adam optimizer is selected as the optimizer for the dynamic face temporal consistency model to update the model parameters, and the initial learning rate is set to 0.001; The second dataset is divided into training set, validation set and test set in the ratio of 7:2:1; The time series segment samples of the training set are input into the model, the loss function is calculated, and the parameters of the dynamic face temporal consistency model are updated through the back-propagation algorithm; Use the validation set to evaluate the model's performance after a preset number of training rounds, calculating the accuracy, recall, and F1 value. Adjust the model's hyperparameters based on the calculation results to prevent overfitting. When the performance of the dynamic face temporal consistency model on the validation set no longer improves, training is stopped.
[0010] In an optional embodiment, the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model are evaluated and optimized respectively, including: For the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model, use the corresponding test sets to calculate the evaluation indicators of the models respectively; The evaluation indicators include but are not limited to accuracy, recall rate, F1 value, precision rate, and missed detection rate; Identify model deficiencies based on evaluation metrics and optimize the model by adjusting the model architecture or model hyperparameters.
[0011] In an optional embodiment, the method of acquiring face data to be detected, performing detection using a facial micro-expression anomaly detection model and a dynamic face temporal consistency model, and combining the detection results of the two models to identify AI-generated faces includes: Get the face data to be detected; The face data to be detected is input into the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection to generate prediction results; When the prediction results of both models are that the facial data is not generated by AI, it is determined that the facial data to be detected is not generated by AI; otherwise, it is determined that the facial data to be detected is generated by AI.
[0012] In a second aspect, an embodiment of the present application further provides an AI-generated face recognition system, comprising: Model building module, used to build facial micro-expression anomaly detection model and dynamic face temporal consistency model; A data acquisition module is used to collect facial data and AI-generated facial data, and generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model; A model training module is used to define a loss function, select an optimizer, train a facial micro-expression anomaly detection model using the first dataset, and train a dynamic face temporal consistency model using the second dataset; The model optimization module is used to evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively; The real-time detection module is used to obtain the facial data to be detected, and uses the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection, and combines the detection results of the two models to identify AI-generated faces.
[0013] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the AI-generated face recognition method as described in any one of the above items are implemented.
[0014] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the AI-generated face recognition method as described in any one of the above items are implemented.
[0015] It can be seen from the above technical solutions that the present invention has the following advantages: In the AI-generated face identification method provided in this application, efficient identification of AI-generated faces is achieved by constructing a dual-model architecture of facial micro-expression anomaly detection and dynamic temporal consistency analysis: First, this method breaks through the limitations of traditional single-dimensional feature detection, synchronously captures the differences in micro-expression details and spatiotemporal continuity features, and significantly improves the accuracy of identifying forgery traces; secondly, by fusing the feature extraction capabilities of convolutional neural networks with the temporal modeling advantages of recurrent neural networks, and combining the attention mechanism to enhance the perception of temporal changes in key facial areas, the model's sensitivity to capturing generated artifacts is effectively enhanced; thirdly, a high-quality training set is constructed by collaboratively enhancing real data and generated data, and combined with the cross-entropy loss function and dynamic learning rate optimization mechanism, while ensuring the generalization performance of the model, hyperparameter adaptive adjustment is achieved through validation set monitoring; finally, the dual-model decision fusion mechanism reduces the risk of misjudgment through cross-validation, and significantly enhances the robustness to complex generation algorithms while ensuring high-accuracy detection performance, providing a technical solution that is both innovative and practical for the field of deep fake detection.
[0016] This application breaks through the detection limitations of single static features or temporal features by combining micro-expression detail anomaly analysis with dynamic temporal consistency modeling, achieves dual verification of spatial details and temporal continuity, and significantly improves the ability to identify AI-generated faces.
[0017] This application combines the spatial feature extraction capabilities of convolutional neural networks (CNN) with the temporal modeling capabilities of recurrent neural networks (RNN / LSTM), and introduces an attention mechanism to enhance the perception of temporal changes in key facial areas, thereby improving the model's accuracy in capturing generated artifacts and its ability to model temporal logic.
[0018] This application uses real data and AI-generated data to collaboratively construct a training set, combines data enhancement technology to improve the generalization of the model, and at the same time ensures the model's adaptability to dynamic scenes through time series segment interception and annotation, effectively alleviating the risk of overfitting.
[0019] This application uses the cross-entropy loss function to quantify the prediction difference, combines it with an adaptive optimizer (such as Adam) to dynamically adjust the learning rate, and uses validation set monitoring to achieve hyperparameter tuning, balancing the convergence speed and model stability during training.
[0020] This application integrates the prediction results of micro-expression detection and spatiotemporal consistency analysis to make decisions, reduces the misjudgment rate through complementary verification, and maintains high accuracy and strong robustness in complex generation algorithm scenarios, providing more reliable technical support for deep fake detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 Flowchart of the AI-generated face recognition method provided in this application.
[0023] Figure 2 Schematic diagram of the structure of the AI-generated face recognition system provided in this application.
[0024] Figure 3 This is a schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0025] The specific steps of the AI-generated face recognition method will be described in detail below, and various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather that the present disclosure should be understood to cover all adjustments, equivalents, and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.
[0026] Hereinafter, the terms "include" or "may include" as used in various embodiments of the present disclosure indicate the presence of disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "include," "have," and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing, and should not be understood as excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more features, numbers, steps, operations, elements, components, or combinations of the foregoing.
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] See also Figure 1 FIG. 1 is a flowchart of a method for recognizing an AI-generated face in a specific embodiment, the method comprising: S1: Construct a facial micro-expression anomaly detection model and a dynamic face temporal consistency model.
[0029] In a specific embodiment, the process of constructing a facial micro-expression abnormality detection model includes: Select a deep learning model architecture suitable for micro-expression anomaly detection, such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a combination of these (such as LSTM-CNN) to build a facial micro-expression anomaly detection model. CNNs can effectively extract spatial features from images, while RNNs are well-suited for processing sequential data and can capture the dynamic changes in micro-expressions over time.
[0030] Based on the size and complexity of the dataset used to train the facial micro-expression anomaly detection model, the model parameters should be appropriately set, such as the number of convolutional layers, the size and number of filters, and the number of neurons in the fully connected layers. The optimal parameter settings can be determined through experimentation and parameter tuning to improve model performance.
[0031] In a specific embodiment, the process of constructing a dynamic face temporal consistency model includes: A temporal consistency model for dynamic faces is constructed based on a recurrent neural network (RNN) architecture combined with a convolutional neural network (CNN). Specifically, considering the time series information of dynamic faces, LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) are good choices. They effectively capture dependencies within long sequences and learn the dynamic temporal patterns of faces. Before inputting into the RNN, a CNN is used to extract features from each frame of the facial image. The CNN automatically learns local and global features of facial images, such as facial texture and shape. The feature vectors extracted by the CNN are then fed into the RNN in chronological order to further process the temporal information.
[0032] Furthermore, to allow the dynamic face temporal consistency model to focus more on the parts that are important for determining temporal consistency, an attention mechanism can be incorporated into the model. For example, during RNN processing, the attention mechanism can help the model focus on the temporal changes in key facial features (such as the eyes and mouth).
[0033] S2: Collect facial data and AI-generated facial data to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model.
[0034] In a specific embodiment, the process of generating the first data set includes: First, collect real-life facial videos and images containing micro-expressions in real time or from facial databases. Then, collect AI-generated facial videos and images containing micro-expressions using AI face generation technology. Specifically, first, collect real-life facial videos and images containing various micro-expressions (such as smiles, frowns, surprise, and blinks). This can be obtained from public facial databases (such as CK+ and JAFFE), or augmented by self-filmed videos. Ensure that the dataset includes individuals of diverse ages, genders, ethnicities, and cultural backgrounds to improve the model's generalization capabilities. Second, use a variety of AI face generation technologies (such as GAN and StyleGAN) to generate facial images and videos containing different micro-expressions. During the generation process, various parameters and conditions can be set to increase data diversity. For example, adjust the generator's input noise or change the generation pattern of facial features.
[0035] Then, the collected data is preprocessed to generate a first data set, which specifically includes: Face detection and alignment: Use existing face detection algorithms (such as the Haar cascade detector and MTCNN) to detect faces in collected images and video frames and extract the facial regions. Then, use facial landmark detection algorithms (such as OpenFace and Dlib) to locate key facial features (such as the eyes, nose, and mouth) and align the facial images to a standard position and angle.
[0036] Micro-expression feature extraction: Extract micro-expression-related features based on aligned facial images. Traditional manual feature extraction methods, such as Local Binary Patterns (LBP) and Histogram of Oriented Gradients (HOG), can be used to describe changes in facial texture and shape. Alternatively, deep learning models (such as convolutional neural networks (CNNs)) can be used to automatically extract more advanced feature representations.
[0037] Data Labeling: Label the preprocessed data to identify real-life micro-expressions and AI-generated micro-expressions. Manual labeling can be used to ensure accuracy and consistency. To improve labeling efficiency, you can use labeling tools such as LabelImg and CVAT. Once labeling is complete, the first dataset is generated.
[0038] Data augmentation: To increase the size and diversity of the first dataset and prevent model overfitting, data augmentation is performed. Images can be transformed using random rotation, flipping, scaling, adding noise, and random clipping and interpolation of video frames.
[0039] In a specific embodiment, the process of generating the second data set includes: First, we collected real-world facial video data from public video databases, as well as AI-generated facial video data generated using AI face generation technology. Specifically, we collected a large number of real, dynamic facial videos from various sources, encompassing a wide range of ages, genders, ethnicities, expressions, and movements. For example, these videos could be obtained from public video databases (such as YouTube-FacesDB), or we could film them ourselves, encompassing scenes like daily interactions, changing expressions, and head movements. Furthermore, we utilized current mainstream AI face generation technologies, such as generative adversarial networks (GANs) and variational autoencoders (VAEs), to generate different types of dynamic facial videos. During the generation process, we set a variety of parameters and conditions to simulate various possible false generation scenarios, including varying generation styles and quality.
[0040] Then, the collected data is preprocessed to generate a second data set, which includes: Face detection and tracking: Use reliable face detection algorithms (such as MTCNN) to detect faces in video frames. For detected faces, use face tracking algorithms (such as DeepSORT) to track the same face in consecutive frames to ensure that the face in each frame can be accurately matched.
[0041] Normalization: Normalize the size of the face image and adjust it to a fixed size (such as 128×128 pixels). At the same time, normalize the pixel values, for example, scaling the pixel values to the range of [0, 1] to reduce data variability.
[0042] Time sequence segment capture: Continuous video frames are captured into time sequence segments according to a certain time length (such as 1 second, corresponding to 30 frames, which can be adjusted according to the video frame rate) as input samples for the dynamic face temporal consistency model.
[0043] Labeling data: Each time-series segment is labeled manually or with the help of semi-automatic tools to indicate whether it is from a real face video or an AI-generated face video. Once the labeling is completed, the second dataset is generated.
[0044] S3: Define the loss function, select the optimizer, use the first dataset to train the facial micro-expression anomaly detection model, and use the second dataset to train the dynamic face temporal consistency model.
[0045] In a specific embodiment, the training process of the facial micro-expression abnormality detection model includes: First, define the loss function for the facial micro-expression anomaly detection model: Choose an appropriate loss function to measure the difference between the model's predictions and the true labels. For a binary classification problem (real facial micro-expressions vs. AI-generated facial micro-expressions), you can use the cross-entropy loss function.
[0046] Next, select an optimizer for the facial micro-expression anomaly detection model: Choose an appropriate optimizer to update the model's parameters to minimize the loss function. Common optimizers include stochastic gradient descent (SGD), Adagrad, Adadelta, RMSProp, and Adam. Based on the characteristics of the facial micro-expression anomaly detection model and the availability of training data, select an appropriate optimizer and set appropriate parameters such as the learning rate.
[0047] During training, the first dataset is divided into a training set, a validation set, and a test set. During training, the model is trained on the training set, and the model parameters are updated using the backpropagation algorithm. Simultaneously, the validation set is used to monitor model performance and adjust model hyperparameters to prevent overfitting. Training is terminated when the model's performance on the validation set no longer improves.
[0048] In a specific embodiment, the training process of the dynamic face temporal consistency model includes: First, we define the loss function for the dynamic face temporal consistency model: For a binary classification problem (real or AI-generated), we choose the cross-entropy loss function as the objective function to measure the difference between the prediction results of the dynamic face temporal consistency model and the true label. The formula is:
[0049] Where N is the number of samples, is the true label (0 or 1), is the probability that the prediction result of the dynamic face temporal consistency model is positive (AI generated).
[0050] Next, select an optimizer for the dynamic face temporal consistency model: Use the Adam optimizer to update the model's parameters. This optimizer combines the advantages of the adaptive gradient algorithm (Adagrad) and the root mean square propagation algorithm (RMSProp). It automatically adjusts the learning rate during training, improving training efficiency and convergence speed. Set an appropriate initial learning rate (e.g., 0.001) and adjust it accordingly based on training progress.
[0051] During training, the following process is performed: The second dataset is divided into training set, validation set and test set, and the ratio can be set to 7:2:1 (which can be adjusted according to actual conditions).
[0052] During the training process, the time series segment samples of the training set are input into the model, the loss function is calculated, and the parameters of the model are updated through the back-propagation algorithm.
[0053] After a certain number of training rounds (e.g., 5 rounds), use the validation set to evaluate the model's performance and calculate metrics such as accuracy, recall, and F1 score. Based on the validation set results, adjust the model's hyperparameters (e.g., the number of hidden layer neurons and the learning rate) to prevent overfitting.
[0054] When the model's performance on the validation set no longer improves (for example, the accuracy of the validation set does not improve significantly over 10 consecutive rounds), stop training.
[0055] S4: Evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively.
[0056] In a specific embodiment, the evaluation and optimization process of the facial micro-expression anomaly detection model includes: The trained facial micro-expression anomaly detection model was evaluated using the test set, and evaluation metrics such as precision, recall, and F1 score were calculated. Precision indicates the proportion of samples correctly predicted by the model to the total number of samples, while recall indicates the proportion of positive samples correctly predicted by the model to the actual positive samples. The F1 score is the harmonic mean of precision and recall, comprehensively reflecting the model's performance.
[0057] Based on the evaluation results, analyze the shortcomings of the facial micro-expression anomaly detection model and optimize it accordingly. You can try adjusting the model architecture, increasing or decreasing the amount of training data, and adjusting hyperparameters to improve model performance. Additionally, you can use model optimization techniques such as regularization and dropout to prevent overfitting and improve the model's generalization capabilities.
[0058] In a specific embodiment, the evaluation and optimization process of the dynamic face temporal consistency model includes: Use the test set to comprehensively evaluate the trained dynamic face temporal consistency model. In addition to accuracy, recall, and F1 value, you can also calculate indicators such as precision and missed detection rate to measure the model performance from different perspectives.
[0059] Analyze the incorrect prediction samples of the dynamic face temporal consistency model on the test set to identify the model's weaknesses. For example, if the model is found to be poorly performing dynamic face recognition for certain expressions or movements, you can add relevant data for retraining or adjust the model architecture to enhance its ability to learn these features.
[0060] S5: Obtain the facial data to be detected, and use the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection respectively, and combine the detection results of the two models to identify AI-generated faces.
[0061] In a specific embodiment, when a face to be detected is required to determine whether it is AI-generated, the facial data to be detected is obtained and face recognition is performed using a facial micro-expression anomaly detection model and a dynamic face temporal consistency model. Only when both models detect that the face to be detected is not AI-generated can the face be determined to be not AI-generated. Otherwise, the face is considered AI-generated.
[0062] In this embodiment, a dual-model collaborative detection framework is constructed by integrating micro-expression detail anomaly detection and dynamic temporal consistency analysis, which significantly improves the identification ability of AI-generated faces: on the one hand, it innovatively combines the spatial feature extraction of convolutional neural networks with the temporal modeling capabilities of recurrent neural networks, and introduces an attention mechanism to enhance the perception of temporal changes in key areas, breaking through the limitations of single feature detection; on the other hand, through high-quality data set construction, cross-entropy loss optimization and dynamic learning rate adjustment, the generalization and stability of the model are guaranteed, and finally the risk of misjudgment is reduced through the dual-model decision fusion mechanism, achieving high-precision and robust deep fake detection in complex generation scenarios, providing key technical support for preventing the abuse of AI-generated content.
[0063] like Figure 2 As shown, the following is an embodiment of the AI-generated face recognition system provided by the embodiments of the present disclosure. This system and the AI-generated face recognition methods of the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiments of the AI-generated face recognition system, please refer to the embodiments of the above-mentioned AI-generated face recognition method.
[0064] An AI-generated face recognition system includes: a model construction module, a data acquisition module, a model training module, a model optimization module and a real-time detection module.
[0065] Model building module, used to build facial micro-expression anomaly detection model and dynamic face temporal consistency model.
[0066] The data acquisition module is used to collect facial data and AI-generated facial data, and generate a first data set for training the facial micro-expression anomaly detection model and a second data set for training the dynamic facial temporal consistency model.
[0067] The model training module is used to define the loss function, select the optimizer, use the first data set to train the facial micro-expression anomaly detection model, and use the second data set to train the dynamic face temporal consistency model.
[0068] The model optimization module is used to evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model.
[0069] The real-time detection module is used to obtain the facial data to be detected, and uses the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection, and combines the detection results of the two models to identify AI-generated faces.
[0070] The AI-generated face recognition system provided in this embodiment constructs a two-dimensional detection mechanism by integrating facial micro-expression detail analysis and dynamic temporal consistency modeling. It combines the convolutional-recurrent neural network architecture with the attention mechanism to enhance key feature capture, uses a collaborative enhancement strategy of real and generated data to improve generalization, and adopts adaptive optimization and validation set monitoring to ensure training stability. Ultimately, it reduces the risk of misjudgment through dual-model collaborative decision-making, significantly enhances the accuracy and robustness of AI-generated face recognition, effectively responds to the challenges of complex generation algorithms, and provides a technical solution for deep fake detection that is both innovative and practical.
[0071] Figure 3 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0072] The AI-generated face recognition method provided in the embodiments of the present application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiments of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In the embodiments of the present invention, electronic devices include but are not limited to laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.
[0073] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a button, a camera, a display, and a SIM card interface, etc.
[0074] A processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0075] The processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signal to complete the control of instruction fetching and execution.
[0076] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.
[0077] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of an electronic device. The external memory card communicates with the processor through the external memory interface, enabling data storage. For example, files such as music and videos can be stored on the external memory card.
[0078] Internal memory can be used to store computer-executable program code, which includes instructions. The processor executes the instructions stored in the internal memory to perform various functional applications and data processing of the electronic device. The internal memory can include a program storage area and a data storage area. The internal memory can include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0079] The wireless communication function of an electronic device can be implemented through an antenna, a wireless communication module, a modem processor, and a baseband processor.
[0080] Wireless communication modules can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0081] Electronic devices can implement audio functions through audio modules, speakers, receivers, microphones, headphone jacks, and application processors.
[0082] Electronic devices can achieve shooting functions through ISP, camera, video codec, GPU, display and application processor.
[0083] Electronic devices can achieve display functions through GPU, display screen and application processor.
[0084] A GPU is a microprocessor for image processing that connects the display screen to the application processor. The GPU performs mathematical and geometric calculations for graphics rendering. A processor may include one or more GPUs, which execute program instructions to generate or modify display information.
[0085] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0086] The above-mentioned electronic device implements the AI-generated face recognition method of this application through a two-dimensional detection mechanism that integrates facial micro-expression detail analysis and dynamic temporal consistency modeling, combines the convolutional-recurrent neural network architecture and attention mechanism to enhance the key feature capture capability, adopts a collaborative enhancement strategy of real faces and generated data to improve the generalization of the model, and uses adaptive optimization algorithms and verification set monitoring to ensure training stability. Finally, the risk of misjudgment is reduced through the dual-model collaborative decision-making mechanism, achieving the technical effect of significantly improving the accuracy and robustness of AI-generated face recognition.
[0087] The storage medium provided in this application stores a program product that can implement an AI-generated face recognition method.
[0088] AI-generated face recognition methods include: Construct facial micro-expression anomaly detection model and dynamic face temporal consistency model; Collecting facial data and AI-generated facial data to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model; Define the loss function, select the optimizer, use the first dataset to train the facial micro-expression anomaly detection model, and use the second dataset to train the dynamic face temporal consistency model; Evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively; Obtain the facial data to be detected, and use the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection respectively, and combine the detection results of the two models to identify AI-generated faces. In some possible implementations, the AI-generated face recognition method disclosed herein can be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps described in the above "Exemplary Method" section of this specification according to various exemplary implementations of the present disclosure.
[0089] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0090] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying AI-generated faces, characterized in that: include: Construct facial micro-expression anomaly detection model and dynamic face temporal consistency model; Collecting facial data and AI-generated facial data to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model; Define the loss function, select the optimizer, use the first dataset to train the facial micro-expression anomaly detection model, and use the second dataset to train the dynamic face temporal consistency model; Evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively; Obtain the facial data to be detected, and use the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection respectively, and combine the detection results of the two models to identify AI-generated faces.
2. The AI-generated face recognition method according to claim 1, characterized in that: The construction of the facial micro-expression anomaly detection model and the dynamic face temporal consistency model includes: Select a convolutional neural network, recurrent neural network, or long short-term memory-convolutional neural network to build a facial micro-expression anomaly detection model; set model parameters, including the number of convolutional layers, the size and number of filters, and the number of neurons in the fully connected layer; Based on convolutional neural networks, a dynamic facial temporal consistency model is constructed by integrating recurrent neural networks and their variants. An attention mechanism is introduced into the process of recurrent neural networks processing temporal information to guide the model to focus on facial parts that are critical for judging temporal consistency, enhance the model's sensitivity to temporal changes in key areas, and improve the accuracy of temporal modeling.
3. The AI-generated face recognition method according to claim 2, characterized in that: The face data is collected and the AI generates the face data to generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic face temporal consistency model, respectively, including: Collect real-time facial video and image data containing micro-expressions or from a facial database; collect AI facial video and image data containing micro-expressions generated using AI facial generation technology; Based on the collected real face video and image data, AI face video and image data, face detection algorithms are used to perform face detection, extract the face area, and generate face images. The key feature points of the face are located using the facial key point detection algorithm, and the face image is aligned to the standard position and angle. Based on the aligned facial images, features related to micro-expressions are extracted. The micro-expressions of real faces and those generated by AI are labeled to generate the first dataset, which is then augmented. Collect real face video data from public video databases and AI face video data generated using AI face generation technology; Based on the collected real face video data and AI face video data, a face detection algorithm is used to detect faces in video frames. For the detected face video frames, a face tracking algorithm is used to track the same face in consecutive frames to determine the consecutive video frames with the same face. Normalize the face video frames and cut the consecutive video frames of the same face into time-series segments according to the preset time length; The time sequence segments are labeled to indicate that each time sequence segment is real face video data or AI face video data, and a second data set is generated.
4. The AI-generated face recognition method according to claim 3, characterized in that: The method of defining a loss function, selecting an optimizer, using the first data set to train a facial micro-expression anomaly detection model, and using the second data set to train a dynamic face temporal consistency model includes: Define the cross entropy loss function to measure the difference between the predicted results of the facial micro-expression anomaly detection model and the true label; Select the stochastic gradient descent optimizer, Adagrad optimizer, Adadelta optimizer, RMSProp optimizer, or Adam optimizer as the model optimizer to update the parameters of the facial micro-expression anomaly detection model to minimize the cross-entropy loss function; Based on the first dataset, the model is trained using the cross entropy loss function and the model optimizer.
5. The AI-generated face recognition method according to claim 4, characterized in that: The method further includes defining a loss function, selecting an optimizer, using the first data set to train a facial micro-expression anomaly detection model, and using the second data set to train a dynamic face temporal consistency model. The cross entropy loss function is defined as the loss function of the model to measure the difference between the prediction results of the dynamic face temporal consistency model and the true label; the cross entropy loss function is as follows: Where N is the number of samples, is the true label, The probability that the prediction result of the dynamic face temporal consistency model is an AI-generated face; The Adam optimizer is selected as the optimizer for the dynamic face temporal consistency model to update the model parameters, and the initial learning rate is set to 0.001; The second dataset is divided into training set, validation set and test set in the ratio of 7:2:1; The time series segment samples of the training set are input into the model, the loss function is calculated, and the parameters of the dynamic face temporal consistency model are updated through the back-propagation algorithm; Use the validation set to evaluate the model's performance after a preset number of training rounds, calculating the accuracy, recall, and F1 value. Adjust the model's hyperparameters based on the calculation results to prevent overfitting. When the performance of the dynamic face temporal consistency model on the validation set no longer improves, training is stopped.
6. The AI-generated face recognition method according to claim 5, characterized in that: The trained facial micro-expression anomaly detection model and dynamic face temporal consistency model are evaluated and optimized respectively, including: For the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model, use the corresponding test sets to calculate the evaluation indicators of the models respectively; The evaluation indicators include but are not limited to accuracy, recall rate, F1 value, precision rate, and missed detection rate; Identify model deficiencies based on evaluation metrics and optimize the model by adjusting the model architecture or model hyperparameters.
7. The AI-generated face recognition method according to claim 6, characterized in that: The method of obtaining face data to be detected, performing detection using a facial micro-expression anomaly detection model and a dynamic face temporal consistency model, and combining the detection results of the two models to identify AI-generated faces includes: Get the face data to be detected; The face data to be detected is input into the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection to generate prediction results; When the prediction results of both models are that the facial data is not generated by AI, it is determined that the facial data to be detected is not generated by AI; otherwise, it is determined that the facial data to be detected is generated by AI.
8. An AI-generated face recognition system, characterized in that: The system adopts the AI-generated face recognition method as described in any one of claims 1 to 7; The system comprises: Model building module, used to build facial micro-expression anomaly detection model and dynamic face temporal consistency model; A data acquisition module is used to collect facial data and AI-generated facial data, and generate a first data set for training a facial micro-expression anomaly detection model and a second data set for training a dynamic facial temporal consistency model; A model training module is used to define a loss function, select an optimizer, train a facial micro-expression anomaly detection model using the first dataset, and train a dynamic face temporal consistency model using the second dataset; The model optimization module is used to evaluate and optimize the trained facial micro-expression anomaly detection model and dynamic face temporal consistency model respectively; The real-time detection module is used to obtain the facial data to be detected, and uses the facial micro-expression anomaly detection model and the dynamic face temporal consistency model for detection, and combines the detection results of the two models to identify AI-generated faces.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the AI-generated face recognition method as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the AI-generated face recognition method as claimed in any one of claims 1 to 7 are implemented.