Counterfeit face detection method, device, equipment and medium

By extracting three-dimensional and two-dimensional feature of the target picture and enhancing the features using the multi-head attention mechanism module, the problem of low resolution and low-blocking face detection accuracy in the existing technology is solved, and efficient fake face detection in common scenarios is achieved.

CN120014717AActive Publication Date: 2025-05-16SOUTH CHINA AGRICULTURAL UNIVERSITY

Patent Information

Application Number
CN202510025320.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-16
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing fake face detection methods have low detection accuracy in low-resolution or occluded face images, and rely on specific devices to obtain data, which has poor real-time performance and is difficult to expand to common scenarios.

Method used

By performing three-dimensional feature extraction and two-dimensional feature extraction on the target picture, combined with the multi-head attention mechanism module, the three-dimensional enhanced features and two-dimensional features are spliced ​​and enhanced, and the target output features are generated for detection.

Benefits of technology

It improves the detection accuracy of low-resolution or obscuring face images, reduces dependence on specific devices, and realizes effective fake face detection in common scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014717A_ABST
    Figure CN120014717A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image analysis, and relates to a counterfeited face detection method and device, equipment and a medium. The method comprises the following steps: acquiring a to-be-detected target picture; performing three-dimensional feature extraction on the target picture to obtain a three-dimensional feature of the target picture, and performing feature enhancement on the three-dimensional feature through an attention mechanism to obtain a three-dimensional enhanced feature; performing two-dimensional feature extraction on the target picture to obtain two-dimensional features of the target picture; performing feature splicing and feature enhancement on the three-dimensional enhanced features and the two-dimensional features through a preset multi-head attention mechanism module to obtain target output features; and inputting the target output feature into a preset first activation function to obtain a detection result of the target picture. According to the method and the device, the problems that a reconstruction result is inaccurate, the data acquisition difficulty is relatively high, and some low-resolution or sheltered face images are difficult to accurately detect in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image analysis technology, and in particular, to a forged face detection method, device, equipment and medium. Background Art

[0002] Existing forged face detection methods attempt to enhance the robustness of forged face detection by reconstructing the spatial or other geometric information of face images. Such methods often combine facial key point extraction, deep learning networks, and 3D modeling techniques, and have made some progress. However, they still face several challenges. First, the accuracy and authenticity of reconstruction are highly dependent on the quality of the input image. For low-resolution or occluded face images, the reconstruction results are often not accurate enough. Second, feature reconstruction methods usually only focus on the reconstruction of a single feature, such as facial shape or texture, resulting in weak comprehensive judgment capabilities for deep forgeries. Some methods rely on the use of specific equipment to obtain RGB images or depth images or to build a complete facial 3D model information library. They have limitations such as high difficulty in data acquisition and poor real-time performance, making it difficult to expand to more common forged face detection scenarios. Summary of the invention

[0003] The present application provides a method, device, equipment and medium for detecting fake faces to solve one or more technical problems existing in the prior art and at least provide a beneficial choice or create conditions.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0005] According to one aspect of an embodiment of the present application, a method for detecting a forged face is provided, the method comprising:

[0006] Get the target image to be detected;

[0007] Extracting three-dimensional features of the target image to obtain three-dimensional features of the target image, and enhancing the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features;

[0008] Performing two-dimensional feature extraction on the target image to obtain two-dimensional features of the target image;

[0009] Performing feature concatenation and feature enhancement on the three-dimensional enhanced features and the two-dimensional features through a preset multi-head attention mechanism module to obtain a target output feature;

[0010] The target output feature is input into a preset first activation function to obtain the detection result of the target image.

[0011] In one embodiment of the present application, based on the above solution, extracting three-dimensional features from the target image to obtain the three-dimensional features of the target image includes:

[0012] Extracting the depth features of the target image through a preset first encoder structure;

[0013] Extracting albedo features of the target image through a preset second encoder structure;

[0014] The three-dimensional feature is determined according to the depth feature and the albedo feature.

[0015] In one embodiment of the present application, based on the above solution, the three-dimensional feature is enhanced by the attention mechanism to obtain the three-dimensional enhanced feature, including:

[0016] Performing feature enhancement on the albedo feature by using a preset second activation function and the attention mechanism to obtain an albedo enhancement feature;

[0017] The albedo enhancement feature is fused with the depth feature to obtain the three-dimensional enhancement feature.

[0018] In one embodiment of the present application, based on the above-mentioned solution, the albedo feature is enhanced by using the preset second activation function and the attention mechanism to obtain the albedo enhanced feature, including:

[0019] Performing weight calculation on the albedo feature by using the second activation function to obtain a weight parameter of the albedo feature;

[0020] Performing bias calculation on the albedo feature to obtain a bias parameter of the albedo feature;

[0021] Feature enhancement is performed according to the weight parameter, the bias parameter and the albedo feature of the attention mechanism to obtain the albedo enhanced feature.

[0022] In one embodiment of the present application, based on the aforementioned solution, the multi-head attention mechanism module includes a first multi-head attention mechanism module and a second multi-head attention mechanism module; the three-dimensional enhanced feature and the two-dimensional feature are subjected to feature splicing and feature enhancement by the preset multi-head attention mechanism module to obtain the target output feature, including:

[0023] Performing cross-attention feature extraction on the three-dimensional enhancement feature and the two-dimensional feature respectively through the first multi-head attention mechanism module and the second multi-head attention mechanism module to obtain a three-dimensional attention feature corresponding to the three-dimensional enhancement feature and a two-dimensional attention feature corresponding to the two-dimensional feature;

[0024] The three-dimensional attention feature and the two-dimensional attention feature are concatenated to obtain a target attention feature, and the target attention feature is enhanced by a preset self-attention mechanism module to obtain the target attention enhanced feature;

[0025] The target attention enhancement feature is subjected to layer normalization processing to obtain the target output feature.

[0026] In one embodiment of the present application, based on the aforementioned solution, the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively perform cross-attention feature extraction on the three-dimensional enhanced feature and the two-dimensional feature to obtain a three-dimensional attention feature corresponding to the three-dimensional enhanced feature and a two-dimensional attention feature corresponding to the two-dimensional feature, including:

[0027] Inputting the two-dimensional features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a two-dimensional query vector, a two-dimensional key vector and a two-dimensional value vector corresponding to the two-dimensional features;

[0028] generating the two-dimensional attention feature according to the two-dimensional query vector, the two-dimensional key vector, and the two-dimensional value vector;

[0029] Inputting the three-dimensional enhanced features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a three-dimensional query vector, a three-dimensional key vector and a three-dimensional value vector corresponding to the three-dimensional enhanced features;

[0030] The three-dimensional attention feature is generated according to the three-dimensional query vector, the three-dimensional key vector and the three-dimensional value vector.

[0031] In one embodiment of the present application, based on the above solution, the target image to be detected is obtained by the following steps:

[0032] Acquire video data of a target face to be detected, and convert the video data into a picture frame sequence corresponding to the video data, wherein the picture frame sequence consists of multiple frames of pictures to be detected;

[0033] A frame of the picture to be detected is randomly selected from each frame of the picture to be detected as the target picture.

[0034] According to one aspect of an embodiment of the present application, a forged face detection device is provided, the device comprising:

[0035] An acquisition unit, used for acquiring a target image to be detected;

[0036] A three-dimensional feature extraction unit is used to extract three-dimensional features from the target image to obtain three-dimensional features of the target image, and to enhance the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features;

[0037] A two-dimensional feature extraction unit, used to extract two-dimensional features of the target image to obtain two-dimensional features of the target image;

[0038] A feature enhancement unit, used for performing feature concatenation and feature enhancement on the three-dimensional enhancement feature and the two-dimensional feature through a preset multi-head attention mechanism module to obtain a target output feature;

[0039] The detection unit is used to input the target output feature into a preset first activation function to obtain a detection result of the target image.

[0040] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions. When the executable instructions are executed by a processor, the method described in the above embodiment is implemented.

[0041] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; and a memory for storing executable instructions of the processors, wherein when the executable instructions are executed by the one or more processors, the one or more processors implement the methods described in the above embodiments.

[0042] Beneficial effects of the present application: The present application obtains three-dimensional features by extracting three-dimensional features from a target image, and enhances the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features. The three-dimensional enhanced features obtained in this way can effectively solve the problem of existing methods relying on high-resolution or unobstructed images.

[0043] Furthermore, the three-dimensional enhanced features and the two-dimensional features are feature spliced ​​and enhanced through a multi-head attention mechanism module to obtain a target output feature, and the target output feature obtained in this way fuses the two-dimensional features and the three-dimensional enhanced features, and further enhances the fused features to obtain the final target output feature, and then the target output feature is input into the preset first activation function to obtain the detection result of the target image. The detection result obtained in this way is more accurate, thereby solving the problems of inaccurate reconstruction results, difficulty in data acquisition, and difficulty in accurate detection of some low-resolution or occluded face images in the prior art. At the same time, the present application does not require specific equipment to complete the construction of the three-dimensional face model, and only needs to extract features from the target image to be detected to complete the detection of the face, which can be used for effectiveness analysis of forged faces.

[0044] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0046] Figure 1 is a flow chart of a forged face detection method according to an embodiment of the present application;

[0047] Figure 2 A framework diagram of a learning network structure for a forged face detection method according to an embodiment of the present application;

[0048] Figure 3 It is a logic diagram of a feature front fusion module according to an embodiment of the present application;

[0049] Figure 4 is a logic diagram of a multimodal feature fusion module according to an embodiment of the present application;

[0050] Figure 5 is a block diagram of a forged face detection device according to an embodiment of the present application;

[0051] Figure 6 It is a structural diagram of an electronic device according to the present application. DETAILED DESCRIPTION

[0052] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0053] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the present application.

[0054] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or micro-control node devices.

[0055] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0056] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0057] The following is a detailed introduction to the background technology of the embodiments of the present application:

[0058] In recent years, image generation technology driven by deep learning has developed rapidly, prompting deep face forgery technology to achieve unprecedented realism and create extremely realistic facial images. Face forgery technology has its positive effects. For example, in the film and television production industry, it has become a common and efficient technology. Using face forgery technology, the face of a character in a video can be replaced with the face of another person, adding fun and interactivity. However, this technology also has malicious uses. For example, if lawbreakers use this technology to create false pornographic or speech videos, these videos may be used for false accusations, defamation and revenge porn, causing serious harm to personal portrait rights, reputation rights and privacy rights. Although it has shown excellent visual effects in the field of film and television, the popularization of technology poses a threat to network security, prompting domestic and foreign research and development teams to actively explore ways to combat this technology.

[0059] In related technologies, feature extraction enhancement blocks and mapping blocks are designed, and super-resolution technology and attention mechanisms are combined to restore and enhance the lost spatial texture features in low-quality videos, thereby improving detection performance. In addition, an improved island loss function and regional data enhancement strategy are introduced to further enhance the model's ability to identify low-quality deep fake videos.

[0060] Related technologies also learn the common features of real faces by reconstructing real face images, and explore the essential differences between real faces and forged faces based on classification tasks. Specifically, a reconstruction network is trained using real face images, and the hidden features of the reconstruction network are used to classify real and forged faces. This method not only improves the accuracy of forged face detection, but also improves the robustness of the detection algorithm by designing anti-attack tests.

[0061] With the development of deep learning technology, the traces of forged face images have become increasingly difficult to detect, making it difficult for traditional detection methods to effectively identify these forged images. Therefore, a face forgery detection method based on 3D decomposition is proposed in the related art, which amplifies the forgery traces by decomposing the face image into five components: 3D shape, common texture, identity texture, ambient light, and direct light. First, they regard the face image as the result of the interaction between 3D geometry and lighting environment, and decompose it through a computer graphics view. Then, the researchers found that the forgery traces are mainly hidden in direct light and identity texture, so they proposed to use the facial details of these two parts as clues to detect forgery. Finally, a two-stream structure network (FD2Net) is designed, which uses facial details and original face images as inputs for multimodal tasks, and introduces a supervised attention mechanism to highlight the manipulated areas. Extensive experimental results show that this method has reached the state-of-the-art level in both accuracy and generalization ability of forgery detection.

[0062] Some current forged face detection methods based on feature reconstruction attempt to enhance the robustness of forged detection by reconstructing the spatial or other geometric information of face images. Such methods often combine facial key point extraction, deep learning networks, and 3D modeling technology, and have made some progress. However, they still face several challenges. First, the accuracy and authenticity of the reconstruction are highly dependent on the quality of the input image. For low-resolution or occluded face images, the reconstruction results are often not accurate enough. Second, feature reconstruction methods usually only focus on the reconstruction of a single feature, such as facial shape or texture, and fail to fully integrate multiple feature information, resulting in weak comprehensive judgment capabilities for deep forgeries. Some methods rely on using specific equipment to obtain RGB images and depth images or construct a complete facial 3D model information library, which has limitations such as high difficulty in data acquisition and poor real-time performance, making it difficult to expand to more common forged face detection scenarios.

[0063] Therefore, in view of the problems or defects existing in the above-mentioned related technologies, the present application proposes a forged face detection method to improve the detection capability of forged face images.

[0064] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0065] According to one aspect of the present application, a forged face detection method is provided. Figure 1 This is a flowchart of a forged face detection method according to an embodiment of the present application. The forged face detection method at least includes steps S1 to S5, which are described in detail as follows:

[0066] In step S1, a target image to be detected is obtained.

[0067] Specifically, the target image to be detected is obtained through the following steps:

[0068] Acquire video data of a target face to be detected, and convert the video data into a picture frame sequence corresponding to the video data, wherein the picture frame sequence consists of multiple frames of pictures to be detected;

[0069] A frame of the picture to be detected is randomly selected from each frame of the picture to be detected as the target picture.

[0070] The present application constructs a forged face feature library corresponding to forged face features, and the forged face feature library is constructed by collecting forged face data sets (i.e., face video data forged by various forgery methods such as artificial intelligence). The forged face video data is extracted into a series of picture frames, and the forged face data set is divided into a training set, a test set, and a verification set, and then the forged face data set is trained, tested, and verified, and finally a corresponding forged face feature library is generated. In this way, after inputting the video data of the target face to be detected, the target image to be detected is automatically obtained, so as to perform feature analysis on the target image, and then compare it with the forged face feature library to determine whether the current target face to be detected is a forged face.

[0071] In step S2, three-dimensional features are extracted from the target image to obtain the three-dimensional features of the target image, and the three-dimensional features are enhanced through an attention mechanism to obtain three-dimensional enhanced features.

[0072] In one embodiment of the present application, extracting three-dimensional features from the target image to obtain the three-dimensional features of the target image includes:

[0073] Extracting the depth features of the target image through a preset first encoder structure;

[0074] Extracting albedo features of the target image through a preset second encoder structure;

[0075] The three-dimensional feature is determined according to the depth feature and the albedo feature.

[0076] Specifically, refer to Figure 2 As shown, Figure 2The encoder 1 in the figure is the preset first encoder structure, and the encoder 2 is the preset second encoder structure. Encoder 1 and encoder 2 are trained by the following steps: Use the collected fake face data set to pre-train the face 3D reconstruction network. The face picture (i.e., the target picture) is scaled to a 64*64 format as a training data set, and the input image is decomposed into two factors: depth and albedo. During training, the model only uses single-view images and does not rely on any external supervision, such as 3D models, multi-view images, or any form of annotation information. After the training is completed, a self-encoder-decoder structure (i.e., the first encoder structure and the second encoder structure described in this application) that can extract the depth and albedo features of face pictures is obtained. Figure 2 The feature front fusion module in is used to output three-dimensional features.

[0077] In one embodiment of the present application, the three-dimensional feature is enhanced by an attention mechanism to obtain a three-dimensional enhanced feature, including:

[0078] Performing feature enhancement on the albedo feature by using a preset second activation function and the attention mechanism to obtain an albedo enhancement feature;

[0079] The albedo enhancement feature is fused with the depth feature to obtain the three-dimensional enhancement feature.

[0080] The albedo feature is enhanced by using the preset second activation function and the attention mechanism to obtain the albedo enhancement feature, including:

[0081] Performing weight calculation on the albedo feature by using the second activation function to obtain a weight parameter of the albedo feature;

[0082] Performing bias calculation on the albedo feature to obtain a bias parameter of the albedo feature;

[0083] Feature enhancement is performed according to the weight parameter, the bias parameter and the albedo feature of the attention mechanism to obtain the albedo enhanced feature.

[0084] Specifically, refer to Figure 3 As shown, Figure 3 for Figure 2 The specific schematic diagram of the feature front fusion module shown in, Figure 3 The input 1 is the albedo feature described in this application, Figure 3The input 2 is the depth feature described in this application, where the albedo feature is calculated through two paths, one of which has a convolution layer 1 and a second activation function (this path is used for weight calculation), and the other path has only a convolution layer 2 (this path is used for bias calculation). The weight parameters and bias parameters obtained by calculation are used to adjust the input albedo feature, and finally the albedo feature is enhanced through the SKA attention mechanism (that is, the preset attention mechanism described in this application) to obtain the albedo enhanced feature, and finally the albedo enhanced feature is fused with the depth feature to obtain the three-dimensional enhanced feature. It should be noted that the three-dimensional enhanced feature is used for input Figure 2 The multimodal feature fusion module in

[0085] In step S3, two-dimensional features are extracted from the target image to obtain two-dimensional features of the target image.

[0086] Specifically, the two-dimensional features (i.e. Figure 2 The RGB features shown) are also used for input Figure 2 The multimodal feature fusion module in . That is, Figure 2 The learning network structure shown in the figure can be divided into two paths, one of which is used to process the three-dimensional features, and the obtained three-dimensional enhanced features are input into Figure 2 The other path is the RGB feature (i.e., the two-dimensional feature described in this application), which is input into the multimodal feature fusion module to perform feature splicing and fusion with the three-dimensional enhancement feature for Figure 2 The classifier in outputs the target output feature, and the target output feature is represented by F.

[0087] In step S4, the three-dimensional enhanced features and the two-dimensional features are feature concatenated and enhanced through a preset multi-head attention mechanism module to obtain target output features.

[0088] Specifically, Figure 4 As shown, Figure 4 for Figure 2 The multimodal feature fusion module in the embodiment of the present invention is used to perform feature concatenation and feature enhancement on the three-dimensional enhancement feature and the two-dimensional feature to obtain the target output feature F, and output the target output feature F to the classifier.

[0089] In one embodiment of the present application, the multi-head attention mechanism module includes a first multi-head attention mechanism module and a second multi-head attention mechanism module; the three-dimensional enhanced feature and the two-dimensional feature are subjected to feature splicing and feature enhancement by the preset multi-head attention mechanism module to obtain the target output feature, including:

[0090] Performing cross-attention feature extraction on the three-dimensional enhancement feature and the two-dimensional feature respectively through the first multi-head attention mechanism module and the second multi-head attention mechanism module to obtain a three-dimensional attention feature corresponding to the three-dimensional enhancement feature and a two-dimensional attention feature corresponding to the two-dimensional feature;

[0091] The three-dimensional attention feature and the two-dimensional attention feature are concatenated to obtain a target attention feature, and the target attention feature is enhanced by a preset self-attention mechanism module to obtain the target attention enhanced feature;

[0092] The target attention enhancement feature is subjected to layer normalization processing to obtain the target output feature.

[0093] The method of performing cross-attention feature extraction on the three-dimensional enhanced feature and the two-dimensional feature respectively by the first multi-head attention mechanism module and the second multi-head attention mechanism module to obtain a three-dimensional attention feature corresponding to the three-dimensional enhanced feature and a two-dimensional attention feature corresponding to the two-dimensional feature includes:

[0094] Inputting the two-dimensional features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a two-dimensional query vector, a two-dimensional key vector and a two-dimensional value vector corresponding to the two-dimensional features;

[0095] generating the two-dimensional attention feature according to the two-dimensional query vector, the two-dimensional key vector, and the two-dimensional value vector;

[0096] Inputting the three-dimensional enhanced features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a three-dimensional query vector, a three-dimensional key vector and a three-dimensional value vector corresponding to the three-dimensional enhanced features;

[0097] The three-dimensional attention feature is generated according to the three-dimensional query vector, the three-dimensional key vector and the three-dimensional value vector.

[0098] The specific structure of the multimodal feature fusion module proposed in this application is as follows Figure 4 As shown. It aims to improve the representation of global features and local features through an enhanced semantic attention mechanism. The multimodal feature fusion module includes a dimension adjustment layer, which is used to perform dimension matching so that the three-dimensional enhanced features can achieve a balanced dimension through dimension matching. The dimension adjustment layer includes two linear layers, which are used to match the dimensions of global features and local features respectively, and to adjust the merged feature dimensions back to the dimensions of global features.

[0099] In addition, the multimodal feature fusion module also includes two multi-head self-attention layers (the two multi-head self-attention layers correspond to the first multi-head attention mechanism module and the second multi-head attention mechanism module described in this application, respectively), which are used for global to local and local to global feature attention calculations, respectively, and a self-attention layer for processing merged global and local features. The optional layer normalization layer is used to normalize the enhanced features. During the forward propagation process, the global features and local features are first reshaped into a format suitable for the multi-head self-attention layer, and then the global to local and local to global attention features (i.e., the cross-attention layer) are used to calculate the global to local and local to global attention features. Figure 4 The input in and input 2 are respectively input to the first multi-head attention mechanism module and the second multi-head attention mechanism module) and merged.

[0100] in, Figure 4 Input 1 is a two-dimensional feature, and input 2 is a three-dimensional enhanced feature. The first multi-head attention mechanism module and the second multi-head attention mechanism module both include a query vector layer ( Figure 4 Q in), key vector layer ( Figure 4 K in) and the value vector layer ( Figure 4 V in , therefore, by performing cross-attention calculation on input 1 and input 2 through the cross-attention layer, the corresponding two-dimensional query vector, two-dimensional key vector and two-dimensional value vector can be obtained; the corresponding three-dimensional query vector, three-dimensional key vector and three-dimensional value vector can be obtained, thereby generating two-dimensional attention features and three-dimensional attention features.

[0101] Next, the three-dimensional attention feature and the two-dimensional attention feature are concatenated to obtain the target attention feature, and the target attention feature is feature enhanced by the preset self-attention mechanism module. Finally, layer normalization is optionally applied, and the final enhanced feature (i.e., output target output feature) is returned. The design of the multimodal feature fusion module allows for the effective fusion and enhancement of features of different scales in various deep learning applications to improve the performance of the multimodal feature fusion module.

[0102] The model framework proposed in this application is as follows Figure 2 As shown in the figure, the framework adopts a dual-stream input mechanism to feed the target image to be detected into two independent processing flows. One processing flow is the regular input of the target image, and its two-dimensional features are extracted through a deep neural network; the other flow is aimed at the special codec structure of the target image (i.e., the first encoder structure and the second encoder structure), aiming to extract the three-dimensional features of the face, using Figure 3 The feature front fusion module refers to the fusion of depth features and albedo features in the three-dimensional features. The backbone network uses EfficientNet-b4, and finally uses Figure 4The multimodal feature fusion module refers to the fusion of the features of the two streams to obtain the final enhanced target output features.

[0103] Furthermore, the model learning process is carried out, hyperparameters such as batch and learning rate are set, and cross entropy is used as the loss function. The dataset preprocessing program provided by the DeepfakeBench open source platform is used to extract, save and divide the face images of various original datasets into training sets, test sets and validation sets.

[0104] The following is an exemplary detailed description of the embodiments of the present application:

[0105] First, randomly select face images from the training data set (which can be converted from video data) and input them into Figure 2 The learning network shown in the figure constructs a training data set, scales the size of each image to a 256*256 image format, and then performs image enhancement and normalization on each image.

[0106] The processed face image (i.e., target image) is input into the detection framework, where the first-class image is scaled to 64*64 size and then input into Figure 2 In the encoder 1 and encoder 2 shown, encoder 1 and encoder 2 will output albedo features MSR and depth features Deep respectively.

[0107] The albedo feature and depth feature are pre-fused by the feature pre-fusion module. In the forward propagation, the weights (i.e. Figure 3 The convolutional layer 1 and the second activation function) and the bias ( Figure 3 The albedo feature MSR is adjusted according to the weight parameter and the bias parameter, and then the processed deep feature Deep is added to the adjusted feature. Finally, the feature is enhanced through the attention mechanism and the three-dimensional enhanced feature Feature3d is output.

[0108] The target image is input into two backbone networks to extract features, and Feature_rgb (i.e., two-dimensional features) and Feature3d (three-dimensional enhanced features) are obtained. The multimodal feature fusion module is used to fuse and enhance the two features of Feature_rgb and Feature3d to obtain the final target output feature F.

[0109] Input F into the classification head and the Softmax activation function (i.e., the first activation function described in this application) to obtain the predicted classification result, i.e., the detection result described in this application. The cross entropy loss function is used to calculate the loss value of the output probability and the true label, and the calculation formula is as follows:

[0110]

[0111] Where: n is the total number of classification categories. i ′ is the i-th component of the true label, which takes the value of 0 or 1. i =c|x i) : The probability of the i-th class predicted by the model, usually the output of Softmax.

[0112] The loss value is back-propagated through the network layers, and the gradient of each parameter is calculated using the chain rule. The model parameters are adjusted to reduce the loss, thereby guiding the update of the network weights and optimizing the model performance.

[0113] Repeat the above process until the loss value Loss no longer decreases (reaches convergence);

[0114] During and after model training Figure 2 The performance of the learning network structure shown is evaluated using the indicators AUC and ACC. For AUC calculation, calculate TPR and FPR: for each possible threshold (set by yourself, 0.5 is recommended), calculate the true positive rate (TPR) and false positive rate (FPR). True positive rate (TPR) = number of true positive samples / total number of actual positive samples. False positive rate (FPR) = number of false positive samples / total number of actual negative samples. In the coordinate system, with FPR as the horizontal axis and TPR as the vertical axis, draw a point for each threshold. Then, connect these points to form a ROC curve. Calculate the area under the ROC curve. For ACC calculation, use the formula:

[0115]

[0116] Calculate the classification accuracy of the test image.

[0117] In summary, the design of the multimodal feature fusion module fully utilizes the advantages of the attention mechanism. Through the in-depth interaction between RGB and 3D features, it realizes efficient fusion and semantic enhancement of features, providing strong support for improving the model performance in the forged face detection task. This design can not only capture the subtle correlation between different modalities, but also strengthen the feature representation in an adaptive way, enhancing the model's recognition ability for complex scenes. Specifically, in the feature fusion process, the module first adjusts the dimensions of global and local features through a linear transformation layer to meet the needs of the multi-head attention mechanism. Two cross-attention mechanisms are constructed inside the module: one is the attention from RGB features (2D features) to 3D enhanced features, and the other is the opposite, from 3D enhanced features to RGB features (2D features). These two mechanisms capture the guiding role of RGB features on 3D enhanced features and the refinement of 3D enhanced features of faces on RGB features.

[0118] The framework of this application aims to unsupervisedly reconstruct a high-quality three-dimensional face model from a single two-dimensional image (i.e., target image) without the aid of any additional equipment or data annotation, and use the obtained three-dimensional model features (i.e., target output features) for forged face detection.

[0119] The encoder-decoder structure of the face 3D model reconstruction network obtained by pre-training on a large number of face image datasets can be directly used to extract the 3D model features belonging to the face image. EfficientNet-b4 is introduced as the core feature extraction module, combined with a customized 3D feature pre-fusion and feature weighted aggregation scheme (i.e. Figure 3 and Figure 4 The corresponding feature pre-fusion module and multimodal feature fusion module) can effectively improve the accuracy and generalization performance of the fake face detection network. The present application proposes a fake face detection method based on face picture feature reconstruction to detect fake faces, aiming to solve the problem of low detection accuracy of existing fake face detection methods by introducing three-dimensional face reconstruction, multimodal feature fusion and end-to-end network optimization technology. The present application relates to the fields of computer vision recognition, deep learning, and image processing. Additional features are introduced through a three-dimensional face reconstruction network, and a multimodal feature fusion module is designed to enhance the original features and improve the detection ability of fake face pictures.

[0120] According to one aspect of an embodiment of the present application, a forged face detection device 300 is provided. Figure 5 This is a schematic diagram of a fake face detection device 300 proposed in an embodiment of the present application. The device 300 includes: an acquisition unit 301, a three-dimensional feature extraction unit 302, a two-dimensional feature extraction unit 303, a feature enhancement unit 304 and a detection unit 305.

[0121] An acquisition unit 301 is used to acquire a target image to be detected;

[0122] A three-dimensional feature extraction unit 302 is used to extract three-dimensional features from the target image to obtain three-dimensional features of the target image, and to enhance the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features;

[0123] A two-dimensional feature extraction unit 303 is used to extract two-dimensional features of the target image to obtain two-dimensional features of the target image;

[0124] A feature enhancement unit 304 is used to perform feature concatenation and feature enhancement on the three-dimensional enhancement feature and the two-dimensional feature through a preset multi-head attention mechanism module to obtain a target output feature;

[0125] The detection unit 305 is used to input the target output feature into a preset first activation function to obtain a detection result of the target image.

[0126] As another aspect, the present application further provides a computer-readable storage medium on which a program product capable of implementing the method provided above in this specification is stored. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary implementations of the present application described in the above "Embodiment Method" section of this specification.

[0127] According to the embodiment of the present application, the program product for implementing the above method can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0128] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0129] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0130] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0131] Program code for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0132] Refer to the following Figure 6 The electronic device 400 according to this embodiment of the present application is described. Figure 6 The electronic device 400 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0133] like Figure 6 As shown, the electronic device 400 is in the form of a general computing device. The components of the electronic device 400 may include but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410).

[0134] The storage unit stores program codes, which can be executed by the processing unit 410, so that the processing unit 410 executes the steps described in the above “Example Method” section of this specification according to various exemplary implementations of the present application.

[0135] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 421 and / or a cache memory unit 422 , and may further include a read-only memory unit (ROM) 423 .

[0136] The storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, such program modules 425 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0137] Bus 430 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller node, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0138] The electronic device 400 may also communicate with one or more external devices 1200 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 450. In addition, the electronic device 400 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 460. As shown, the network adapter 460 communicates with other modules of the electronic device 400 via a bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0139] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0140] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0141] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be performed without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for detecting fake faces, characterized in that: The method comprises: Get the target image to be detected; Extracting three-dimensional features of the target image to obtain three-dimensional features of the target image, and enhancing the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features; Performing two-dimensional feature extraction on the target image to obtain two-dimensional features of the target image; Performing feature concatenation and feature enhancement on the three-dimensional enhanced features and the two-dimensional features through a preset multi-head attention mechanism module to obtain a target output feature; The target output feature is input into a preset first activation function to obtain the detection result of the target image.

2. The forged face detection method according to claim 1, characterized in that: The extracting three-dimensional features of the target image to obtain the three-dimensional features of the target image includes: Extracting the depth features of the target image through a preset first encoder structure; Extracting albedo features of the target image through a preset second encoder structure; The three-dimensional feature is determined according to the depth feature and the albedo feature.

3. The forged face detection method according to claim 2, characterized in that: The three-dimensional features are enhanced by the attention mechanism, Get three-dimensional enhanced features, including: Performing feature enhancement on the albedo feature by using a preset second activation function and the attention mechanism to obtain an albedo enhancement feature; The albedo enhancement feature is fused with the depth feature to obtain the three-dimensional enhancement feature.

4. The forged face detection method according to claim 3, characterized in that: The albedo feature is enhanced by using the preset second activation function and the attention mechanism to obtain the albedo enhancement feature, including: Performing weight calculation on the albedo feature by using the second activation function to obtain a weight parameter of the albedo feature; Performing bias calculation on the albedo feature to obtain a bias parameter of the albedo feature; Feature enhancement is performed according to the weight parameter, the bias parameter and the albedo feature of the attention mechanism to obtain the albedo enhanced feature.

5. The forged face detection method according to claim 4, characterized in that: The multi-head attention mechanism module includes a first multi-head attention mechanism module and a second multi-head attention mechanism module; the three-dimensional enhanced feature and the two-dimensional feature are subjected to feature splicing and feature enhancement by the preset multi-head attention mechanism module to obtain the target output feature, including: Performing cross-attention feature extraction on the three-dimensional enhancement feature and the two-dimensional feature respectively through the first multi-head attention mechanism module and the second multi-head attention mechanism module to obtain a three-dimensional attention feature corresponding to the three-dimensional enhancement feature and a two-dimensional attention feature corresponding to the two-dimensional feature; The three-dimensional attention feature and the two-dimensional attention feature are concatenated to obtain a target attention feature, and the target attention feature is enhanced by a preset self-attention mechanism module to obtain the target attention enhanced feature; The target attention enhancement feature is subjected to layer normalization processing to obtain the target output feature.

6. The forged face detection method according to claim 5, characterized in that: The method of performing cross-attention feature extraction on the three-dimensional enhanced feature and the two-dimensional feature respectively by the first multi-head attention mechanism module and the second multi-head attention mechanism module to obtain a three-dimensional attention feature corresponding to the three-dimensional enhanced feature and a two-dimensional attention feature corresponding to the two-dimensional feature includes: Inputting the two-dimensional features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a two-dimensional query vector, a two-dimensional key vector and a two-dimensional value vector corresponding to the two-dimensional features; generating the two-dimensional attention feature according to the two-dimensional query vector, the two-dimensional key vector, and the two-dimensional value vector; Inputting the three-dimensional enhanced features into the first multi-head attention mechanism module and the second multi-head attention mechanism module respectively, to obtain a three-dimensional query vector, a three-dimensional key vector and a three-dimensional value vector corresponding to the three-dimensional enhanced features; The three-dimensional attention feature is generated according to the three-dimensional query vector, the three-dimensional key vector and the three-dimensional value vector.

7. The forged face detection method according to claim 1, characterized in that: The target image to be detected is obtained by the following steps: Acquire video data of a target face to be detected, and convert the video data into a picture frame sequence corresponding to the video data, wherein the picture frame sequence consists of multiple frames of pictures to be detected; A frame of the picture to be detected is randomly selected from each frame of the picture to be detected as the target picture.

8. A forged face detection device, characterized in that: The device comprises: An acquisition unit, used for acquiring a target image to be detected; A three-dimensional feature extraction unit is used to extract three-dimensional features from the target image to obtain three-dimensional features of the target image, and to enhance the three-dimensional features through an attention mechanism to obtain three-dimensional enhanced features; A two-dimensional feature extraction unit, used to extract two-dimensional features of the target image to obtain two-dimensional features of the target image; A feature enhancement unit, used for performing feature concatenation and feature enhancement on the three-dimensional enhancement feature and the two-dimensional feature through a preset multi-head attention mechanism module to obtain a target output feature; The detection unit is used to input the target output feature into a preset first activation function to obtain a detection result of the target image.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the video jitter sequence extraction method based on a fixed scene according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for extracting a video jitter sequence based on a fixed scene according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Face living body detection method and device in motion state scene, equipment and medium

    CN115909467A

  • Illumination guide identification method for deeply-forged face in video and related equipment

    CN119251895A

  • Generative model for 3D face synthesis with HDRI relighting

    US20240020915A1

Cited By

  • Forgery image detection method and device, medium and product

    CN120953779A

  • A method, device, medium and product for detecting a fake image

    CN120953779B