Deep forgery detection method and device based on mixed training
Through the hybrid training method, training sample pairs are constructed and adaptive token mixers are introduced to improve the cross-dataset adaptability and robustness of the deep forged detection model, solving the problems of poor cross-dataset adaptability and high computing cost of existing models, and achieving efficient forged detection.
Patent Information
- Application Number
- CN202510619434.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-25
AI Technical Summary
The existing deep forgery detection models have poor adaptability, insufficient robustness and large model parameters across data sets, making it difficult to effectively identify unknown forgery technologies, and have high computing costs, which limits the deployment on resource-constrained devices.
A deep forgery detection method based on hybrid training is adopted, and mask calculation and front-to-back scene image enhancement are obtained by obtaining the image training set, training sample pairs are constructed, and neural networks are used for iterative training, and adaptive token mixers and regularization losses are introduced to improve the generalization ability and robustness of the model.
It realizes the concise model structure, low computing overhead and high detection accuracy. It is suitable for high-performance image authenticity recognition in multiple types of forgery attack scenarios, especially for resource-constrained devices.
Smart Images

Figure CN120375449A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence security technology, and in particular, to a deepfake detection method, device, storage medium, and electronic device based on hybrid training. Background Art
[0002] Deepfake is a synthetic media technology based on artificial intelligence. Narrowly defined, it refers to the use of deep learning algorithms to manipulate and replace human faces. Compared with other forgery techniques, face deepfakes are more deceptive, capable of highly realistically synthesizing face images or videos, making it difficult for the public to distinguish between authenticity and falsity. In recent years, the abuse of this technology has posed a serious threat to social security, political stability, and personal privacy. Its potential risks are particularly prominent in areas such as financial fraud, public opinion manipulation, and identity theft. Since the human face is a core feature for identity recognition, its forged content is more likely to trigger a trust crisis and exacerbate information pollution. Therefore, there is an urgent need to develop accurate and efficient detection methods to identify and screen forged face content and maintain the health and credibility of the information ecosystem.
[0003] However, face deepfake detection is extremely challenging and different from general image classification tasks. The forgery technology is constantly evolving, and the generation methods are becoming increasingly diverse, making it difficult for the detection model to adapt to unknown forgery types. In addition, forged faces generally remain highly realistic as a whole, with only subtle tampering in local areas, resulting in the model being prone to overfitting to a specific dataset and experiencing a significant decline in performance when faced with cross-datasets and unseen forgery methods. Therefore, improving the generalization ability of the detection model so that it remains robust under different forgery techniques and diverse data distributions is the core problem faced by deep face forgery detection.
[0004] Currently, significant progress has been made in deepfake detection technology. However, existing methods still have various limitations: 1) Poor cross-dataset adaptability. Existing detection models perform well on specific training datasets, but their performance drops significantly when faced with new, unseen datasets. There are significant differences in shooting conditions, image resolutions, compression methods, etc. between different datasets, making it difficult for the model to adapt to the new data distribution, thus reducing the stability and reliability of detection. 2) Insufficient robustness to unknown forgery techniques. Existing detection methods rely on learning known forgery techniques, so when encountering completely new or improved forgery methods, the detection ability drops significantly. For example, deep forgery techniques such as diffusion models are constantly evolving, generating more realistic forged content, making it difficult for traditional detection methods to effectively identify these advanced forgeries, resulting in the detection system being vulnerable to attacks. 3) Large model parameter quantity and high computational cost. Existing detection methods usually improve detection performance by introducing complex neural network modules, but this leads to a significant increase in the quantity of model parameters. Although more complex models can improve the detection ability, their high training costs, computational overheads, and storage requirements limit efficient deployment on resource-constrained devices (such as mobile terminals, embedded systems). Summary of the Invention
[0005] Embodiments of the present application provide a deep forgery detection method, device, storage medium, and electronic device based on hybrid training, which can solve the problems of poor cross-dataset adaptability, insufficient robustness, and large model parameter quantity of existing deep forgery detection models.
[0006] Embodiments of the present application provide a deep forgery detection method based on hybrid training, including: Obtain an image training set, where the image training set includes real images and forged images; Randomly select an original image from the image training set for mask calculation to obtain a grayscale image mask, perform foreground and background image enhancement operations on the original image, and fuse the enhanced original image and the grayscale image mask to obtain a mixed forged image; Construct a training sample pair based on the real image, the forged image, and the mixed forged image; Input the training sample pair into a neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the mixed forged image into an encoder to obtain an original image feature and a mixed forged image feature; classifying based on the original image feature to obtain a classification result; Calculate the cross-entropy loss based on the classification result and the corresponding label; calculate the similarity metric for the original image feature and the mixed forged image feature through a regularization method, and calculate the regularization loss based on the similarity metric result; Iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence; Input the image to be detected into the trained neural network model to obtain the classification result.
[0007] Further, in the above deepfake detection method based on hybrid training, where the mask calculation is performed on the original image to obtain a grayscale image mask, including: Input the original image into the trained face key point detection model to obtain the face key point detection coordinates; Randomly select several coordinate points from all the face key point detection coordinates to obtain a key point set; Generate a convex hull region based on the key point set through the convex hull algorithm; Perform random affine transformation on the convex hull region and then perform Gaussian blur processing to obtain a grayscale image mask.
[0008] Further, in the above deepfake detection method based on hybrid training, where the foreground and background image enhancement operation is performed on the original image to obtain the enhanced original image, and the enhanced original image and the grayscale image mask are fused to obtain a hybrid forged image, including: Divide the original image into a background image and a foreground image, and randomly select one of the background image and the foreground image for data enhancement; Perform mask fusion on the enhanced background image or foreground image and the grayscale image mask to obtain a hybrid forged image.
[0009] Further, in the above deepfake detection method based on hybrid training, where randomly selecting one of the background image and the foreground image for data enhancement is represented by the following formula:
[0010]
[0011] where, is a function for generating random floating-point numbers within the range, is the original image, is the background image, is the foreground image; The process of the mask fusion is represented by the following formula:
[0012] where, is the grayscale image mask, is the hybrid forged image.
[0013] Furthermore, in the above deepfake detection method based on hybrid training, the encoder includes a backbone network and an adaptive token mixer, and the backbone network includes multiple Transformer layers; The processing process of the encoder includes: The original image and the hybrid forged image are input into the backbone network to obtain a first feature and a second feature respectively, and multiple classification tokens are output through the Transformer layer; In the adaptive token mixer, the classification tokens are dynamically weighted by means of one-dimensional convolution; After adding the first feature and the second feature to the corresponding classification tokens respectively and performing normalization, the original image feature and the hybrid forged image feature are obtained.
[0014] Furthermore, in the above deepfake detection method based on hybrid training, the similarity metric between the original image feature and the hybrid forged image feature is calculated by a regularization method, and is calculated by the following formula:
[0015] where represents the dimension of the latent space vector, , is the original image feature under the latent space vector, is the hybrid forged image feature under the latent space vector.
[0016] Furthermore, in the above deepfake detection method based on hybrid training, the cross-entropy loss is:
[0017] where is the cross-entropy loss, represents the size of each MiniBatch, represents the sample set in a MiniBatch, , is the training set, represents the true label corresponding to the original image, when it is 1, it means the original image is a forged image, when it is 0, it means the original image is a real image; The regularization loss is:
[0018] where is the regularization loss, represents the farthest distance parameter.
[0019] The embodiment of the present application further provides a deepfake detection device based on hybrid training, including: An acquisition module, configured to acquire an image training set, where the image training set includes real images and forged images; A hybrid forged image generation module, configured to randomly select an original image from the image training set for mask calculation to obtain a grayscale image mask, perform foreground and background image enhancement operations on the original image, and fuse the enhanced original image and the grayscale image mask to obtain a hybrid forged image; A sample pair generation module, configured to construct training sample pairs based on the real images, the forged images, and the hybrid forged images; A classification module, configured to input the training sample pairs into a neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the hybrid forged image into an encoder to obtain an original image feature and a hybrid forged image feature; performing classification based on the original image feature to obtain a classification result; A training module, configured to calculate a cross-entropy loss based on the classification result and the corresponding label; calculate a similarity metric for the original image feature and the hybrid forged image feature through a regularization method, and calculate a regularization loss based on the similarity metric result; And, configured to iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence; The classification module is further configured to input the image to be detected into the trained neural network model to obtain a classification result.
[0020] The embodiment of the present application further provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above-mentioned deepfake detection methods based on hybrid training.
[0021] The embodiment of the present application further provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above-mentioned deepfake detection methods based on hybrid training.
[0022] The deepfake detection method, device, storage medium, and electronic device provided by the present application have the advantages of simple model structure, low computational overhead, and high detection accuracy. Specifically, the advantages include the following: (1) A lightweight network architecture based on Dinov2-small is proposed, and a transfer learning strategy is applied to verify its effectiveness in the deepfake detection task. While maintaining a low computational cost, this network can achieve high detection accuracy, demonstrating good performance advantages. (2) An image blending strategy based on mask fusion is designed to construct training sample pairs. By fusing the foreground and background masks of an image to generate a blended image, the training sample set is expanded while maintaining semantic consistency, forming "original image - blended image" sample pairs. (3) A hybrid training siamese network architecture is proposed, introducing an encoder with an adaptive token mixing function to balance the expression of latent space features at different levels. At the same time, an image similarity supervision mechanism is introduced in the true / false classification task. By comparing the semantic distances between the original image and the blended image, the model is guided to focus on forgery traces rather than irrelevant regions, thus significantly enhancing the robustness and discriminative ability of the model. (4) The proposed detection method has a simple structure, low implementation cost, and excellent generalization ability, and is particularly suitable for high-performance image authenticity recognition tasks in multi-type forgery attack scenarios. Description of the Drawings
[0023] The following, in conjunction with the drawings, through a detailed description of the specific embodiments of the present application, will make the technical solutions and other beneficial effects of the present application obvious.
[0024] Figure 1 It is a flowchart of the deepfake detection method based on hybrid training provided by an embodiment of the present application.
[0025] Figure 2 It is a flowchart of the generation of training sample pairs provided by an embodiment of the present application.
[0026] Figure 3 It is an example diagram of the training sample pairs provided by an embodiment of the present application.
[0027] Figure 4 It is a schematic structural diagram of the neural network model provided by an embodiment of the present application.
[0028] Figure 5 It is a schematic structural diagram of the deepfake detection device based on hybrid training provided by an embodiment of the present application.
[0029] Figure 6 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0031] An embodiment of the present application provides a deepfake detection method, device, storage medium, and electronic device based on hybrid training. The deepfake detection device provided by the embodiment of the present application can be integrated in an electronic device, and the electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0032] Please refer to Figure 1 , Figure 1 which is a flowchart of the deepfake detection method based on hybrid training provided by the embodiment of the present application. It is applied to an electronic device. The deepfake detection method based on hybrid training includes the following steps: S1. Obtain an image training set, where the image training set includes real images and forged images.
[0033] Specifically, take a data from the training set containing real data and deepfake data, and its true label is . When is real, take 0, and when is forged,
[0034] take 1.
[0035] The original images include real images and forged images. For all the original images and their labels , after applying the hybrid transformation, the hybrid forged images are output, and the label of . Applying the hybrid transformation to the original image is specifically divided into three main processes: mask calculation, foreground and background image enhancement, and mask fusion.
[0036] In one embodiment, the step of performing matte calculation based on the original image in step S2 to obtain a grayscale image matte includes the following steps: S21, input the original image into the trained face key point detection model to obtain face key point detection coordinates; S22, randomly select several coordinate points from all the face key point detection coordinates to obtain a key point set; S23, generate a convex hull region based on the key point set through the convex hull algorithm; S24, perform random affine transformation on the convex hull region and then perform Gaussian blur processing to obtain a grayscale image matte.
[0037] Specifically, first, use the face key point detection model to process the input original image to obtain standard face key point coordinates. To increase the diversity of transformations, randomly select a part of the detected face key points for the construction of subsequent geometric regions. This step can control the shape and coverage range of the matte region. Generate a closed polygon region by calculating the convex hull algorithm for the selected key point set. The convex hull is the smallest convex polygon that contains all the selected points, forming a facial region contour. Apply a random affine transformation to the convex hull region, such as slight translation, rotation, scaling, etc., to create local structural changes. To make the matte edge smooth and natural, perform Gaussian blur processing on the convex hull region after affine transformation. This can avoid obvious artificial edges in the blended image to obtain more difficult training samples. Finally, construct the processed region into a grayscale image matte M of the same size as the original image, and the non-zero region corresponds to the selected face region to be processed.
[0038] In one embodiment, the step of performing foreground and background image enhancement operations on the original image in step S2 to obtain an enhanced original image, and fusing the enhanced original image and the grayscale image matte to obtain a blended forged image includes: S25, divide the original image into a background image and a foreground image, and randomly select one of the background image and the foreground image for data enhancement.
[0039] Specifically, divide into a background image and a foreground image , and randomly select one of or for data enhancement , and this process is expressed as follows:
[0040]
[0041] Among them, is used for generating A function for generating random floating-point numbers within a certain range, is the original image, is the background image, is the foreground image.
[0042] Among them, the specific augmentation strategies of Aug include: randomly shifting the RGB channel values by 20 units with a probability of 30%, randomly adjusting the hue, saturation, and brightness within the range of [-0.3, 0.3], and randomly varying the brightness and contrast within the range of [-0.1, 0.1]. In addition, to simulate image inputs of different qualities, we randomly select between random downsampling (2× or 4×) and sharpening augmentation, thereby enhancing the robustness of the model to different image qualities.
[0043] S26, perform mask fusion on the enhanced background image or foreground image and the grayscale image mask to obtain a blended forged image.
[0044] The process of mask fusion is represented by the following formula:
[0045] Among them, is the grayscale image mask, is the blended forged image.
[0046] S3, construct training sample pairs based on real images, forged images, and blended forged images.
[0047] Figure 2 is the flowchart for generating the training sample pairs provided by the embodiments of the present application, Figure 3 is the example diagram of the training sample pairs provided by the embodiments of the present application, as Figure 2 and Figure 3 shown. Specifically, the training sample pairs include positive sample pairs and negative sample pairs. The positive sample pair is "forged image - blended forged image", and the negative sample pair is "real image - blended forged image". The sample pair labels are consistent with the original image labels i.e., the positive sample pair and the negative sample pair . For the convenience of subsequent description, "original image" is still used to refer to real images and forged images.
[0048] S4, input the original image and the blended forged image into the neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the blended forged image into the encoder to obtain the original image features and the blended forged image features; classifying based on the original image features to obtain the classification result.
[0049] In one embodiment, the encoder includes a backbone network, an adaptive token mixer, and a classification module. The backbone network includes multiple Transformer layers. Figure 4 The structural schematic diagram of the neural network model provided by the embodiment of the present application can be referred to Figure 4 .
[0050] Specifically, use the Dinov2-small backbone network and the adaptive token mixer to construct a feature encoder . Denote the CLS Tokens (classification tokens) outputs of the Transformer layers are respectively
[0051] The processing process of the encoder includes: S41, The original image and the hybrid forged image are input into the backbone network to obtain the first feature and the second feature respectively, and multiple classification tokens are output through the Transformer layers.
[0052] S42, In the adaptive token mixer, the classification tokens are dynamically weighted by means of one-dimensional convolution.
[0053] Use one-dimensional convolution to dynamically weight and calculate the CLS Token, and the calculation formula is as shown in the following formula:
[0054] Where, is the weight of the 𝑖-th layer of the convolutional layer.
[0055] S43, Add the first feature and the second feature to the corresponding classification tokens respectively and then perform normalization to obtain the original image feature and the hybrid forged image feature.
[0056] The final output of the encoder is obtained by adding the output of the last layer of the Transformer to and then performing normalization. Its expression is as shown in the following formula (taking the generation of the original image feature as an example):
[0057] Where, represents the mean value of represents the variance of For generate in the same way.
[0058] In one embodiment, the original image features are classified by a multi-layer perceptron.
[0059] S5. Calculate the cross-entropy loss based on the classification result and the corresponding label; calculate the similarity metric for the original image features and the mixed forged image features by a regularization method, and calculate the regularization loss based on the similarity metric.
[0060] In one embodiment, the similarity metric for the original image features and the mixed forged image features is calculated by a regularization method and calculated by the following formula:
[0061] where represents the dimension of the latent space vector, , is the original image feature under the latent space vector, is the mixed forged image feature under the latent space vector.
[0062] The cross-entropy loss is:
[0063] where is the cross-entropy loss, represents the size of each MiniBatch, represents the sample set in a MiniBatch, , is the training set, represents the true label corresponding to the original image, when is 1, it means the original image is a forged image, when
[0064] where is the regularization loss, represents the farthest distance parameter.
[0065] S6. Iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence.
[0066] Using and the weighted sum to calculate the total loss , using the total loss to train the neural network model, and the total loss calculation formula is shown as follows:
[0067] where represents assigning The weight parameter.
[0068] During the training process of the neural network model, the obtained through the final calculation is used for backpropagation to update the model parameters .
[0069] The following provides an example of the training process: Step 1: Dataset preprocessing and partitioning An overview of each dataset used in the present invention is shown in Table 1. In the data preprocessing stage, 32 frames of images are randomly sampled from each video, and face detection, face cropping, and alignment operations are performed using the DLIB tool. The aligned face images are uniformly adjusted to a size of 256×256 pixels. The image-level labels are consistent with the video labels: if the video is forged, all the image sample labels obtained by sampling are recorded as ; if the video is real, the corresponding image sample labels are recorded as .
[0070] The dataset is partitioned into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set is used for learning the model parameters, the validation set is used for monitoring the model performance during the training process, and the test set is used for evaluating the detection effect of the final model.
[0071] Table 1 Dataset overview
[0072] Step 2: Generation of training sample pairs Before model training, the input data will be transformed into labeled sample pairs through the training sample pair generator module . Sample pairs are generated by combining data augmentation and image mixing strategies. When is a real image, the label is set to 0; when is a forged image, the label is set to 1.
[0073] Figure 2 Shows the flow of the training sample pair generation block. The grayscale image mask is obtained by calculating the convex hull of 81 key points obtained through the face key point predictor and undergoing affine transformation and Gaussian blur processing. The image is divided into a foreground image and a background image , which are used for image mixing operations. In and Randomly select one of them for data augmentation. The augmentation strategies include: randomly offsetting the RGB channels with a 30% probability (amplitude of 20 units), perturbing the hue / saturation / brightness in the interval, and varying the brightness and contrast in the interval. In addition, an image quality perturbation operation is added, that is, randomly selecting one between 2× or 4× downsampling and sharpening enhancement to improve the model's robustness to image quality changes.
[0074] Generate a mixed forged image according to the following formula :
[0075] After the mixed forged image is generated, and will undergo another round of data augmentation, including a random RGB channel offset range , with a 30% probability, hue / saturation / brightness perturbation , and brightness and contrast adjustment . To further improve the model's performance under different image quality conditions, we randomly JPEG compress the image with a 50% probability, and the compression quality range is set to .
[0076] Figure 3 shows the sample generation method. The real image and the mixed image form a negative sample pair, with the label , indicating a significant difference between them and forming a dissimilar supervision signal. The forged image and the mixed image form a positive sample pair, with the label , indicating their similarity and thus forming a similar supervision signal. This label system is compatible with the traditional supervised learning framework of "real is 0, forged is 1", facilitating the reuse and integration of the original data labels.
[0077] Step 3: Neural network model design and loss function definition Figure 3 The neural network structure proposed by the present invention is shown. The model mainly consists of a feature encoder, a latent space contrast regularization module, and a classification head module.
[0078] The backbone network adopts the Dinov2-small architecture, which is based on Vision Transformer (ViT-S / 14). This model consists of 12 Transformer layers, each input image is divided into 256 Patch Tokens, and 1 CLS Token is introduced. The embedding dimension of each Token , the number of heads is 6. The MLP part realizes dimensionality mapping through two linear layers (384 → 1536 → 384). The final output is a 384-dimensional image feature representation.
[0079] During the training process, the input samples go through the Transformer encoder layers to extract the CLS tokens of each layer as hierarchical features. To effectively fuse multi-layer features, an adaptive token mixer is introduced, and the weighted sum of the CLS tokens is calculated through one-dimensional convolution.
[0080] The final image representation and are respectively combined with the output features of the final layer, and obtained through residual connection and normalization. This representation combines the features of the original image and the mixed image, and then is input into the contrast regularization module to calculate the contrast loss , and its formula is as follows:
[0081] where, is the feature distance, is generally set to 1.0 or 1.5.
[0082] In the classification stage, the classification head module only receives for binary classification judgment and calculates the cross-entropy loss :
[0083] The logits are converted into probabilities after passing through Softmax, and the threshold is set to 0.5. If the probability is greater than 0.5, it is judged as forged, and if it is less than 0.5, it is judged as real.
[0084] The total loss function of the model is the weighted sum of the classification loss and the contrast loss:
[0085] where, is the weight coefficient, and usually setting it to 0.5 has the best effect.
[0086] Step 4: Model training and optimization strategy The model is trained using the Adam optimizer, and backpropagation is performed with as the loss function to update the model parameters . During the optimization process, a staged learning rate decay strategy is adopted. After each stage is completed, the learning rate is reduced proportionally . The specific hyperparameter settings are as follows in the table: Table 2 Hyperparameter Setting Table
[0087] S7. Input the image to be detected into the trained neural network model to obtain the classification result.
[0088] According to the method described in the above embodiments, this embodiment will be further described from the perspective of a deepfake detection device based on hybrid training. The deepfake detection device based on hybrid training can be specifically implemented as an independent entity, or integrated in an electronic device, which can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0089] Please refer to Figure 5 , Figure 5 which specifically describes the deepfake detection device based on hybrid training provided in the embodiments of the present application, applied to an electronic device. The deepfake detection device based on hybrid training may include: An acquisition module, configured to acquire an image training set, where the image training set includes real images and forged images; A hybrid forged image generation module, configured to randomly select an original image from the image training set for mask calculation to obtain a grayscale image mask, perform foreground and background image enhancement operations on the original image, and fuse the enhanced original image and the grayscale image mask to obtain a hybrid forged image; A sample pair generation module, configured to construct training sample pairs based on the real images, the forged images, and the hybrid forged images; A classification module, configured to input the training sample pairs into a neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the hybrid forged image into an encoder to obtain an original image feature and a hybrid forged image feature; classifying based on the original image feature to obtain a classification result; A training module, configured to calculate the cross-entropy loss based on the classification result and the corresponding label; calculate the similarity metric for the original image feature and the hybrid forged image feature by a regularization method, and calculate the regularization loss based on the similarity metric result; and, configured to iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence; The classification module is further configured to input the image to be detected into the trained neural network model to obtain a classification result.
[0090] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, please refer to the foregoing method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.
[0091] In addition, an embodiment of the present application further provides an electronic device, which can be a device such as a computer or a tablet computer. The electronic device can implement the steps in any embodiment of the deepfake detection method based on hybrid training provided by the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any of the deepfake detection methods based on hybrid training provided by the embodiments of the present invention. For details, please refer to the foregoing embodiments, which will not be elaborated herein.
[0092] Figure 6 A specific structural block diagram of the electronic device provided by an embodiment of the present invention is shown. The electronic device can be used to implement the deepfake detection method based on hybrid training provided in the above embodiments. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.
[0093] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions. For example, an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, and so on. The RF circuit 510 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.
[0094] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0095] The input unit 530 can be used to receive input digital or character information, and generate keyboards and mice related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode), etc.
[0096] The audio circuit 560, speaker 561, and microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent through the RF circuit 510 to, for example, another terminal, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.
[0097] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential components of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0098] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by calling the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.
[0099] The electronic device 500 further includes a power supply 590 (such as a battery) for powering each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0100] Although not shown, the electronic device 500 further includes a camera (such as a front camera and a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain an image training set, where the image training set includes real images and forged images; Randomly select an original image from the image training set for mask calculation to obtain a grayscale image mask, perform foreground and background image enhancement operations on the original image, and fuse the enhanced original image and the grayscale image mask to obtain a mixed forged image; Construct a training sample pair based on the real image, the forged image, and the mixed forged image; Input the training sample pair into a neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the mixed forged image into an encoder to obtain the original image features and the mixed forged image features; classifying based on the original image features to obtain a classification result; Calculate the cross-entropy loss based on the classification result and the corresponding label; calculate the similarity metric for the original image features and the hybrid forged image features through a regularization method, and calculate the regularization loss based on the similarity metric result; Iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence; Input the image to be detected into the trained neural network model to obtain a classification result.
[0101] In specific implementation, each of the above modules can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of each of the above modules, reference can be made to the foregoing method embodiments, which will not be elaborated herein.
[0102] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium, in which multiple instructions are stored, and these instructions can be loaded by a processor to execute the steps of any one of the embodiments of the deepfake detection method based on hybrid training provided by the embodiments of the present invention.
[0103] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0104] Since the instructions stored in this storage medium can execute the steps in any one of the embodiments of the deepfake detection method based on hybrid training provided by the embodiments of the present invention, the beneficial effects that can be achieved by any of the deepfake detection methods based on hybrid training provided by the embodiments of the present invention can be realized. For details, refer to the foregoing embodiments, which will not be elaborated herein.
[0105] The above has introduced in detail a deepfake detection method, device, storage medium and electronic device based on hybrid training provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A deepfake detection method based on hybrid training, characterized in that, The method includes: Obtaining an image training set, where the image training set includes real images and forged images; Randomly selecting an original image from the image training set for mask calculation to obtain a grayscale image mask, performing foreground and background image enhancement operations on the original image, and fusing the enhanced original image and the grayscale image mask to obtain a mixed forged image; Constructing a training sample pair based on the real image, the forged image, and the mixed forged image; Inputting the training sample pair into a neural network model to obtain a classification result; the processing of the neural network includes: inputting the original image and the mixed forged image into an encoder to obtain an original image feature and a mixed forged image feature; classifying based on the original image feature to obtain a classification result; Calculating the cross-entropy loss based on the classification result and the corresponding label; calculating a similarity metric for the original image feature and the mixed forged image feature by a regularization method, and calculating a regularization loss based on the similarity metric result; Iteratively training the neural network model based on the cross-entropy loss and the regularization loss until convergence; Inputting the image to be detected into the trained neural network model to obtain a classification result.
2. The deepfake detection method based on hybrid training according to claim 1, wherein, The mask calculation based on the original image to obtain a grayscale image mask includes: Inputting the original image into a trained face key point detection model to obtain face key point detection coordinates; Randomly selecting several coordinate points from all the face key point detection coordinates to obtain a key point set; Generating a convex hull region based on the key point set by a convex hull algorithm; Performing a random affine transformation on the convex hull region and then performing Gaussian blur processing to obtain a grayscale image mask.
3. The deepfake detection method based on hybrid training according to claim 1, wherein The foreground and background image enhancement operation on the original image to obtain an enhanced original image, and fusing the enhanced original image and the grayscale image mask to obtain a mixed forged image includes: Dividing the original image into a background image and a foreground image, and randomly selecting one of the background image and the foreground image for data enhancement; Performing mask fusion on the enhanced background image or foreground image and the grayscale image mask to obtain a mixed forged image.
4. The deepfake detection method based on hybrid training according to claim 3, characterized in that, Randomly selecting one of the background image and the foreground image for data enhancement is represented by the following formula: Among them, is a function for generating random floating-point numbers within the range, is the original image, is the background image, is the foreground image; The process of the mask fusion is represented by the following formula: Among them, is a grayscale image mask, is a hybrid forged image.
5. The deepfake detection method based on hybrid training according to claim 1, characterized in that The encoder includes a backbone network and an adaptive token mixer, and the backbone network includes multiple Transformer layers; The processing process of the encoder includes: Inputting the original image and the mixed forged image into the backbone network to respectively obtain a first feature and a second feature, and outputting multiple classification tokens through the Transformer layer; In the adaptive token mixer, dynamically weighting the classification tokens in a one-dimensional convolution manner; Adding the first feature and the second feature to the corresponding classification tokens respectively and then performing normalization to obtain an original image feature and a mixed forged image feature.
6. The deepfake detection method based on hybrid training according to claim 1, characterized in that, Calculating a similarity metric for the original image feature and the mixed forged image feature by a regularization method is calculated by the following formula: Among them, represents the dimension of the latent space vector, , is the original image feature under the latent space vector, is the mixed forged image feature under the latent space vector.
7. The deepfake detection method based on hybrid training according to claim 1, wherein The cross-entropy loss is: Among them, is the cross-entropy loss, represents the size of each MiniBatch, represents the sample set in a MiniBatch, , is the training set, represents the true label corresponding to the original image, when it is 1, it means the original image is a forged image, when it is 0, it means the original image is a real image; The regularization loss is: Among them, is the regularization loss, represents the farthest distance parameter.
8. A deepfake detection device based on hybrid training, characterized in that, Includes: An acquisition module, configured to acquire an image training set, where the image training set includes real images and forged images; A mixed forged image generation module, configured to randomly select an original image from the image training set for mask calculation to obtain a grayscale image mask, perform foreground and background image enhancement operations on the original image, and fuse the enhanced original image and the grayscale image mask to obtain a mixed forged image; A sample pair generation module, configured to construct training sample pairs based on the real images, the forged images, and the mixed forged images; A classification module, configured to input the training sample pairs into a neural network model to obtain classification results; The processing of the neural network includes: inputting the original image and the mixed forged image into an encoder to obtain original image features and mixed forged image features; classifying based on the original image features to obtain classification results; A training module, configured to calculate cross-entropy loss based on the classification results and corresponding labels; calculate a similarity metric for the original image features and the mixed forged image features through a regularization method, and calculate regularization loss based on the similarity metric results; And, configured to iteratively train the neural network model based on the cross-entropy loss and the regularization loss until convergence; The classification module is further configured to input the image to be detected into the trained neural network model to obtain classification results.
9. A computer-readable storage medium, characterized in that, Multiple instructions are stored in the computer-readable storage medium, and the instructions are adapted to be loaded by a processor to execute the deepfake detection method based on hybrid training according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the deepfake detection method based on hybrid training according to any one of claims 1 to 7.
Citation Information
Patent Citations
False image detection method and device based on image self-fusion, equipment and medium
CN119091516A
False face picture detection method based on misleading learning
CN119445342A
Deep forgery detection method based on self-mixing
CN119672816A
Forgery detection of face image
WO2021137946A1