Deep fake detection model training method, deep fake detection method and system

By using a training dataset containing real and fake images to annotate reasons and train a deep fake detection model, the problem of insufficient detection ability for carefully designed fake content is solved, and higher detection accuracy and explainability are achieved.

CN120543952BActive Publication Date: 2025-10-03INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511031987.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-10-03
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

The existing technology has insufficient detection capabilities for carefully designed and high-quality counterfeit content, making it difficult to effectively identify it.

Method used

By obtaining a training dataset containing real images and various types of forged images, annotating the authenticity and reasons for forgery, and training using a deep forgery detection model, including a classifier, image feature extractor, pre-trained language model and feature fuser, a heat map is generated to identify forged areas.

Benefits of technology

Improved detection accuracy for carefully crafted forged content, enabling a more comprehensive understanding of forgery patterns and characteristics, and enhanced accuracy and explainability of forgery detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543952B_ABST
    Figure CN120543952B_ABST
Patent Text Reader

Abstract

This application discloses a training method for a deep fake detection model, a deep fake detection method, and a system, relating to the field of artificial intelligence technology. The method includes obtaining a training data set containing real images and multiple types of forged images, and the images are labeled to indicate authenticity. During the training process, the deep fake detection model is exposed to various types of image samples, thereby learning a wider range of more complex image features and forgery patterns. Forged images are also annotated with the reasons for forgery, which can help the deep fake detection model to deeply understand the essential characteristics of forged content during the training process. Therefore, it can solve the problem of insufficient detection ability for carefully designed and high-quality forged content, and achieve the technical effect of improving the accuracy of forgery detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method for a deep fake detection model, a deep fake detection method, and a system. Background Art

[0002] Deepfake technology leverages deep learning algorithms to create highly realistic manipulation or synthesis of images, videos, or audio, creating indistinguishable fake content. Deepfake technology has broad applications in film, television, entertainment, and virtual characters, but it also poses serious social challenges such as the spread of false information, identity fraud, and privacy violations. To address the threat posed by deepfake technology, researchers are actively developing deepfake detection methods.

[0003] Related technologies include deepfake detection methods based on image quality analysis, which detect forged content by detecting features such as noise and blur. However, this approach can typically only detect relatively obvious forgeries and is insufficient for detecting carefully designed, high-quality forgeries. Summary of the Invention

[0004] The present application provides a deep fake detection method, system, device and electronic device to at least solve the problem in the related art of insufficient detection capability for carefully designed and high-quality fake content.

[0005] This application provides a method for training a deep fake detection model, the method comprising:

[0006] Obtaining a training data set; the training data set includes real images and forged images of different forgeries; the real images and the forged images are each associated with a label indicating the authenticity of the image, and the label corresponding to the forged image also includes a reason for the forgery;

[0007] The deep fake detection model is trained using the training dataset.

[0008] This application also provides a deep fake detection method, which is characterized by comprising:

[0009] Acquire the target image to be detected;

[0010] The target image is input into a deep fake detection model trained using the above-mentioned deep fake detection model training method to obtain an output result indicating the authenticity of the target image.

[0011] This application also provides a deep fake detection system, including:

[0012] A classifier, configured to classify a target image and obtain a classified text description indicating the authenticity of the target image;

[0013] An image feature extractor, configured to extract features from the target image to obtain image features;

[0014] A pre-trained language model is used to generate an output result indicating the authenticity of the target image based on the image features, the classified text description, and the target prompt word;

[0015] A feature fusion device is used to extract features from the output result to obtain text features; and to fuse the text features and the image features to obtain fused features;

[0016] A heat map generator is used to generate a heat map of the target image based on the fusion features; the heat map represents the forged area and the real area in the target image through different pixel values.

[0017] This application also provides a deep fake detection device, the device comprising:

[0018] an acquisition module, configured to acquire a training data set; the training data set includes real images and forged images of different forgeries; the real images and the forged images are each associated with a label indicating the authenticity of the image, and the label corresponding to the forged image also includes a reason for the forgery;

[0019] A training module is used to train the deep fake detection model using the training dataset.

[0020] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above methods when executing the computer program.

[0021] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0022] The present application also provides a computer program product, comprising a computer program, which implements the steps of any of the above methods when executed by a processor.

[0023] This application uses a training dataset containing real images and various types of forged images, with the images labeled to indicate authenticity. During training, the deepfake detection model is exposed to a variety of image samples, thereby learning a wider range of more complex image features and forgery patterns. Furthermore, forged images are annotated with the reasons for forgery, helping the deepfake detection model gain a deeper understanding of the essential characteristics of forged content during training. Therefore, the problem of insufficient detection capabilities for carefully designed, high-quality forgeries can be resolved, achieving the technical effect of improving the accuracy of forgery detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 is a flowchart of a method for training a deep fake detection model provided in an embodiment of the present application;

[0026] Figure 2 Schematic diagram of the architecture of the deep fake detection system provided in an embodiment of the present application;

[0027] Figure 3 is a schematic structural diagram of a training device for a deep fake detection model provided in an embodiment of the present application;

[0028] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0032] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the deep fake detection method depends, the specific application environment architecture or specific hardware architecture is described here.

[0033] The embodiments of the present application provide a method for training a deep fake detection model. The method is described in detail in conjunction with the execution process of the method for training a deep fake detection model.

[0034] Specifically, Figure 1 This is a flowchart of a method for training a deep fake detection model according to an embodiment of the present application.

[0035] like Figure 1 As shown, the deep fake detection method includes: step 110 and step 120.

[0036] Step 110: Obtain a training dataset. The training dataset includes real images and forged images of different forgeries. The real images and forged images are each associated with a label indicating the authenticity of the image. The label corresponding to the forged image also includes the reason for the forgery.

[0037] In the embodiments of the present application, a real image is a naturally captured, unaltered image that retains original real information and can provide a feature sample of the real image for the classifier. The real image may include objects such as people, animals, buildings, etc.

[0038] Fake images are artificially modified or generated through various technical means. They may be visually similar to real images but contain some signs of forgery. Fake images include various types of forgery, such as face replacement, face generation, and facial attribute manipulation. Face replacement involves replacing the face in a target image with another person's face, while retaining other factors from the original image, such as movement, expression, and background. Face generation involves creating a non-existent but realistic-looking "virtual face." Facial attribute manipulation involves modifying specific attributes of a real face image, such as age (making it older / younger), expression (cry to smile), hairstyle, hair color, beard, and makeup. Fake images of various types provide a diverse set of samples for classifier training, enabling the classifier to learn about a variety of possible forgery scenarios.

[0039] In this embodiment, real and fake images can be annotated based on their authenticity, with each image assigned a corresponding label to form a labeled training dataset. The labeling process is relatively straightforward for each real image, allowing it to be clearly labeled as real, providing a standard reference for deepfake detection models to judge real images.

[0040] Forged images should not only be labeled as such, but also have special tags added to the labels. These tags explain why the image was forged. For example, for a forged image with a replaced face, in addition to being labeled as a forged image, a special tag could be added: "There is a difference in skin color between the face and body of the person in the image, and the facial contours are inconsistent with the background lighting. This indicates that it is a face replacement forgery." For forged images generated from faces, special tags could be: "The proportions of the person's facial features do not conform to real human characteristics, and the facial skin texture is too uniform and flawless. This indicates that the face is generated forgery." For images with manipulated facial attributes, special tags could be: "The mouth expression is unnatural, and the raised corners of the mouth are not the normal smile." By adding special tags, the deep fake detection model can learn the specific features and judgment basis of different types of forged images during training.

[0041] Step 130: Use the training dataset to train the deep fake detection model.

[0042] In an embodiment of the present application, a training data set can be input into a deep fake detection model to train the deep fake detection model.

[0043] During training, the deepfake detection model gradually optimizes its parameters through multiple iterations. In each iteration, the deepfake detection model processes images from the training dataset and calculates the error between the deepfake detection model's predictions and the true results based on the labels. This error information is then passed back through the various layers of the deepfake detection model through a backpropagation algorithm, adjusting parameters such as the deepfake detection model's weights and biases to reduce error and improve the deepfake detection model's accuracy, until the deepfake detection model's performance reaches a relatively stable and satisfactory level.

[0044] This application uses a training dataset containing real images and various types of forged images, with the images labeled to indicate authenticity. During training, the deepfake detection model is exposed to a variety of image samples, thereby learning a wider range of more complex image features and forgery patterns. Furthermore, forged images are annotated with the reasons for forgery, helping the deepfake detection model gain a deeper understanding of the essential characteristics of forged content during training. Therefore, the problem of insufficient detection capabilities for carefully designed, high-quality forgeries can be resolved, achieving the technical effect of improving the accuracy of forgery detection.

[0045] In some embodiments, the deep fake detection model includes a classifier, the classifier being configured to classify a target image to be detected and obtain a classified text description characterizing the authenticity of the target image;

[0046] The deepfake detection model is trained using a training dataset, including:

[0047] The difference between the predicted value of the classifier and the true value of the label during training is calculated based on the first loss function, and the parameters of the classifier are updated through the back propagation algorithm based on the calculation result.

[0048] In this embodiment, the target image to be detected can be an image containing a human face, i.e., a facial image. Facial images are widely used in modern society, such as personal photos on social media, real-time images in video conferences, and facial recognition images in identity verification systems. The uniqueness of facial images lies in their high degree of individuality and rich details, making them one of the primary targets of deepfake technology.

[0049] In addition to human faces, target images can also be images with other unique features, such as pet faces, images of iconic buildings, and images of objects with specific textures or patterns. For example, a cat's face image might be used for pet identification, while a partial image of the Eiffel Tower might be used to verify the authenticity of a tourist photo.

[0050] In some embodiments, the target image can be an image uploaded by a user. For example, a user can upload a photo or video frame in a user interface. The user can receive the photo or video frame uploaded in the user interface and use the photo or video frame uploaded in the user interface as the target image. The target image can also be an image uploaded by a camera or an image sent by another client, which is not limited in this embodiment of the present application.

[0051] In the embodiments of the present application, a classifier is a model built based on a machine learning or deep learning algorithm. The classifier is trained on a large amount of labeled data to learn the characteristic difference patterns between real images and forged images, thereby being able to identify the authenticity of the images.

[0052] Taking a deep learning-based convolutional neural network (CNN) classifier as an example, it can include multiple convolutional, pooling, and fully connected layers. This layer automatically extracts low-level texture features, mid-level structural features, and high-level semantic features from an image, and uses these features to determine the image's authenticity. When a target image is input into the classifier, it performs layer-by-layer feature extraction and analysis based on its pre-trained model structure and parameters. First, the image enters the convolutional layer, where the convolution kernel slides across the image to extract basic features such as edges and texture. The pooling layer reduces the dimensionality of these features, preserving key information while reducing computational effort. After multiple layers of convolution and pooling, the target image's features are gradually abstracted into high-level semantic features. These features are then integrated in the fully connected layer and calculated using an activation function. The final output is a probability value, representing the likelihood that the target image is authentic or forged. This probability value can be converted into a corresponding textual description based on a pre-set threshold, completing the image classification process.

[0053] The classification text description is a textual representation of the classifier's authenticity judgment of the target image. It uses concise and clear natural language to indicate whether the target image is real or forged. For example, the classification text description could be "This image is real" or "This image is forged."

[0054] In some embodiments, the classification text description may also include the type of forgery of the forged image, which may include face replacement, face generation, and face attribute manipulation. Face replacement refers to replacing the face in a target image with another person's face, while retaining other factors in the original image, such as the action, expression, and background; face generation refers to generating a non-existent but very real "virtual face" image; and face attribute manipulation refers to modifying specific attributes of a real face image, such as age (making it older / younger), expression (changing from crying to smiling), hairstyle, hair color, beard, makeup, etc.

[0055] In one example, the classification result of a classifier may be output in the form of a probability. For example, the classification result is probability P = [P1, P2, P3, P4], and P1 + P2 + P3 + P4 = 1, where P1 represents the probability that the target image is a real image, P2 represents the probability that the target image is a forged image generated by face replacement, P3 represents the probability that the target image is a forged image generated by a human face, and P4 represents the probability that the target image is a forged image with facial attribute manipulation. If the probability P = [0.1, 0.4, 0.5, 0.1], the classification text description may be that this image is suspected to be forged, and the relevant categories are, in order, forged image generated by a human face, forged image with face replacement, forged image with facial attribute manipulation, and real image. This classification text description provides a richer dimension of information than simple true or false labels.

[0056] In this embodiment, the above-mentioned training data set can be used to train the classifier. During the training process, the classifier extracts key features from the images in the training data set, and based on special markers, guides the classifier to focus on the features of specific areas. For example, when the special markers include "abnormal eye reflection", the classifier will enhance the ability to extract texture and lighting features of the eye area. The classification results are output based on the extracted features, and then compared with the annotated labels, and the error between the predicted results and the true labels is calculated through the first loss function. Based on the error, the parameters of each module in the classifier are adjusted through the back propagation algorithm, and the network structure is optimized so that the results output by the classifier gradually approach the true labels. After repeated training with a large amount of training data, the classifier continuously learns the feature differences between real images and various forged images, and gradually improves the accuracy and reliability of judging the authenticity and forgery types of images.

[0057] In some embodiments, the first loss function is expressed as:

[0058]

[0059] in, represents the first loss function, Indicates the number of image types, Indicates the true value corresponding to the label, represents the predicted value output by the classifier, Indicates the image type The weight of .

[0060] In this embodiment, by collecting real images and multiple forged images, and performing authenticity and forgery reason annotation, the classifier is provided with rich learning samples and semantic information, allowing the classifier to learn the authenticity characteristics of the image and understand the specific types and abnormal manifestations of the forged image, so that the classifier can more accurately identify the forgery traces in the image, which can not only improve the detection accuracy of the classifier, but also enhance the robustness of the classifier, so that the classifier can solve more complex mixed tampering scenarios and improve the detection accuracy of complex forgeries.

[0061] In some embodiments, the image types include real images and forged images of different forgery categories; and the weights of the image types are determined according to the proportions of the respective image types in the training dataset.

[0062] In this embodiment, weights can be assigned to each image type according to the ratio of each image type in the training dataset. For example, assuming that the ratio of the four types of images in the training dataset, namely, real images, forged images generated by human faces, forged images with human face replacement, and forged images with human face attribute manipulation, is 2:4:1:3, then the weight of the real image is , the weight of the fake image generated by the face , the weight of the fake image for face replacement , the weight of the fake image for face replacement .

[0063] In this embodiment, by determining the weights of image types based on the proportions of each image type in the training dataset, the deep fake detection model can perform balanced training on various image types during training, thereby more comprehensively learning the characteristics of different types of fakes and reducing the problem of insufficient detection capabilities for a smaller number of image types due to the imbalance in the number of image samples of each type in the training dataset.

[0064] In some embodiments, the classifier includes a feature extraction network, a multi-classification head, and a label classification mapping module;

[0065] Feature extraction network, used to extract features of the target image;

[0066] A multi-classification head is used to output a classification result representing the authenticity of the target image based on the features extracted by the feature extraction network;

[0067] The label classification mapping module is used to map the classification results to the target template to obtain the classification text description.

[0068] In this embodiment, the feature extraction network is responsible for performing multi-level feature mining on the target image to extract the features of the target image. The feature extraction network can be built based on different deep learning architectures, such as ResNet (Residual Network), VGGNet (Visual Geometry Group Network), EfficientNet (Efficient Convolutional Neural Network), etc. Different deep learning architectures can be selected according to task requirements, and this embodiment of the application is not limited to this.

[0069] The multi-classification head is the decision layer of the classifier. It relies on the feature vector output by the feature extraction network to make in-depth classification decisions and judge the authenticity of the target image.

[0070] In some embodiments, the multi-classification head structure may include a fully connected layer and a Softmax function. The fully connected layer can integrate the high-dimensional feature vectors output by the feature extraction network, and map these features to dimensions corresponding to the number of classification categories through the operation of the weight matrix. For example, in the detection of four types of images, namely, forged images generated by human faces, forged images replaced by human faces, forged images manipulated by human face attributes, and real images, the fully connected layer can map the feature vectors into a four-dimensional vector, and the Softmax function processes the four-dimensional vector and converts it into four probability values, namely P=[P1, P2, P3, P4], and P1+P2+P3+P4=1, P1 represents the probability that the target image is a real image, P2 represents the probability that the target image is a forged image replaced by human faces, P3 represents the probability that the target image is a forged image generated by human faces, and P4 represents the probability that the target image is a forged image manipulated by facial attributes.

[0071] The target template contains textual descriptions for different classification results. After the multi-classification head outputs the probability values ​​for the classification results, the label classification mapping module matches the target template based on the order of these probability values. If the highest probability corresponds to the "forged image generated by a face" category, the label classification mapping module retrieves the corresponding description from the target template and integrates it with the probability order of the other categories. For example, if the multi-classification head outputs probability values ​​P = [0.1, 0.4, 0.5, 0.1], the label classification mapping module generates a detailed classification text description based on these probabilities using the target template: "This image is suspected to be forged. The relevant categories are, in order, forged image generated by a face, forged image with face replacement, forged image with facial attribute manipulation, and real image." In this way, the label classification mapping module converts the probability data output by the multi-classification head into a comprehensible, logically clear, and detection-compliant classification text description. It not only indicates the likelihood of an image being forged but also lists the possible forgery types in descending order of probability.

[0072] The label classification mapping module can also design various target templates based on actual needs to generate more targeted descriptions in different situations. For example, when the probability distribution is relatively dispersed, a description containing multiple possibilities can be generated; when the probability of a certain category is significantly higher than that of other categories, a clear judgment result can be generated.

[0073] In this embodiment, a classifier is constructed by combining a feature extraction network, a multi-classification head and a label classification mapping module. The feature extraction network can extract discriminative features from the target image, the multi-classification head gives the classification results in the form of probabilistic output, and the label classification mapping module converts the output of the model into a human-understandable classification text description. This multi-level, modular classifier design not only improves the detection accuracy, but also makes the detection results more intuitive and easy to understand.

[0074] In some embodiments, the deepfake detection model includes an image feature extractor; the image feature extractor includes a visual encoder and a feature space alignment module;

[0075] Visual encoder, used to encode the target image to be detected;

[0076] The feature space alignment module is used to project the features output by the visual encoder after encoding to a dimension aligned with the text embedding space of the pre-trained language model to obtain image features.

[0077] In this embodiment, the image feature extractor can be constructed using a deep learning algorithm architecture, such as a convolutional neural network or autoencoder. Taking a convolutional neural network as an example, the image feature extractor can include a structure of convolutional layers, pooling layers, and fully connected layers. This structure can simulate the way the human visual system processes images, automatically learning and extracting key image features. The network structure of the image feature extractor has been optimized through training on large amounts of image data, enabling accurate extraction of image features.

[0078] When the target image enters the image feature extractor, the extraction process unfolds according to a pre-trained network structure. Specifically, the target image undergoes convolution operations with different convolution kernels in the convolution layer. By sliding across the image, specific basic features such as edges and textures in the target image are detected and a corresponding feature map is generated. The pooling layer downsamples the feature map, reducing the data dimension while retaining key features and reducing the computational effort. After multiple rounds of convolution and pooling operations, the basic features of the target image are continuously abstracted and integrated, gradually forming higher-level semantic features. The fully connected layer fuses these multi-level features and outputs a set of feature vectors that comprehensively represent the target image, namely, the image features.

[0079] In this embodiment, the visual encoder is a component for image feature extraction, which is responsible for extracting features that can represent the image content from the target image and converting the input image data into high-dimensional feature vectors that contain the visual information of the image.

[0080] Visual encoders can be built on pretrained visual neural network architectures, such as CLIP-ViT (Contrastive Language-Image Pretraining - Vision Transformer) and Swin Transformer (Shifted Window Transformer). Pretrained visual neural networks have been trained on large-scale image datasets and can capture rich semantic information and visual features in images. For example, when a visual encoder uses CLIP-ViT, when a target image is input to the visual encoder, it uses an attention mechanism to capture local and global features of the image and then encodes these features into a high-dimensional vector representation.

[0081] While the feature vector output by the visual encoder contains rich image information, it differs from the text embedding space used by the pre-trained language model in terms of dimensionality and semantic representation. The feature space alignment module can include a fully connected layer that projects the feature vector output by the visual encoder to the same dimensionality as the pre-trained language model text embedding space through a linear transformation. This allows image features to be represented and compared with text information in the same semantic space. In this way, image features can be converted into a format that the pre-trained language model can understand and process.

[0082] During the training process of the image feature extractor, two strategies are adopted to improve the performance and generalization ability of the image feature extractor.

[0083] Strategy 1: Freeze the parameters of the pre-trained visual encoder. Since the visual encoder has been pre-trained on a large-scale dataset and has strong visual feature extraction capabilities, freezing the parameters can preserve this pre-trained visual knowledge, reducing the problem of overfitting on small-scale deepfake detection datasets and allowing the visual encoder to capture common and essential features in the image.

[0084] Strategy 2: Use LoRA (Low-Rank Adaptation) technology to fine-tune the feature space alignment module. LoRA adjusts the output by introducing a low-rank matrix into the fully connected layer, adapting it to new tasks without changing the underlying structure of the fully connected layer. This approach not only reduces the number of training parameters and computational cost, but also enables the feature space alignment module to better map visual features to the text embedding space of the pre-trained language model. This allows the pre-trained language model to directly associate image features with semantic descriptions, improving its ability to understand and analyze image content.

[0085] In this embodiment, through the combination of a pre-trained visual encoder and a feature space alignment module, the image features are aligned with the text embedding space of the pre-trained language model, so that the pre-trained language model can directly use image features for reasoning and judgment. It can not only accurately identify the authenticity of the image, but also provide detailed explanations and basis in combination with text descriptions, thereby improving the accuracy and interpretability of deep fake detection.

[0086] In some embodiments, the deep fake detection model includes a pre-trained language model, which is used to generate an output result indicating the authenticity of the target image based on a classified text description obtained by a classifier classifying the target image to be detected, image features obtained by an image feature extractor extracting features from the target image, and a target prompt word;

[0087] The deepfake detection model is trained using a training dataset, including:

[0088] Obtain training prompt words; the training prompt words are constructed according to the detection task; the training prompt words are used to guide the pre-trained language model to perform the detection task during the training process;

[0089] Freeze the first parameter of the pre-trained language model, input the training prompt word and the training dataset into the pre-trained language model, and update the second parameter of the pre-trained language model through the backpropagation algorithm based on the second loss function; the second loss function represents the difference between the predicted value of the pre-trained language model and the true value of the label.

[0090] In this example, pre-training is a deep learning model training strategy that focuses on using large-scale datasets to initially train the model, enabling it to learn common feature representations. This process is similar to the basic learning stage humans go through before acquiring new knowledge, accumulating experience through extensive reading and observation.

[0091] A pre-trained language model (PLM) is a model designed based on a large-scale corpus (including language training materials such as sentences and paragraphs). The model training task is then trained on a large-scale neural network algorithm structure to achieve learning. The resulting large-scale neural network algorithm structure and parameters are the pre-trained language model. Subsequent tasks can use this model to extract features or fine-tune the model to achieve specific task objectives. The idea behind pre-training is to first train a task to obtain a set of model parameters. This set of model parameters is then used to initialize the network model parameters. The initialized network model is then used to train other tasks to obtain a model adapted for these tasks. By pre-training on a large-scale corpus, neural language representation models can acquire powerful language representation capabilities and extract rich syntactic and semantic information from text. Pre-trained language models can provide token- and sentence-level features containing rich semantic information for use in downstream tasks. Fine-tuning can also be performed directly on the pre-trained model for downstream tasks, quickly and easily obtaining a dedicated downstream model.

[0092] The neural network algorithm structure for training the pre-trained language model can be CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), etc., or it can be a model built on an attention network, such as transformer, BERT (Bidirectional Encoder Representations from Transformer), GPT (Generative Pre-trained Transformer), Clip (Contrastive Language–Image Pretraining), etc., which is not limited in this application. An attention network refers to a network model that uses an attention mechanism for training. The model extracts more important feature information from the input sequence by assigning different weights to each part of the input sequence, so that the model ultimately obtains a more accurate output.

[0093] Fine-tuning refers to performing small-scale training on a pre-trained language model for specific task objectives (downstream tasks) and task data (downstream data), making minor adjustments to the pre-trained language model parameters, and ultimately obtaining a model adapted to the specific task and data. In some embodiments, the pre-trained language model can be fine-tuned for the task objective of "deepfake detection" and the task data of "classified text descriptions" and "image features," making the pre-trained language model more suitable for the task of generating structured query statements.

[0094] The target prompt word is a text instruction used to guide the pre-trained language model to generate a specific output. The target prompt word provides the model with a clear context and task goal, helping the pre-trained language model understand the intent of the input information and reason and generate in the specified direction. In some embodiments, the target prompt word can be a short sentence or question, such as "Determine whether the image is forged based on the provided image features and classification results" or "Can you point out the type of forgery in the image and mark the forged area?" Based on the target prompt word, the pre-trained language model can clearly determine whether the task to be completed is to determine the authenticity of the image, identify forged areas, etc.

[0095] When image features and classified text descriptions are input into the pre-trained language model, the pre-trained language model will associate and integrate these inputs with its own learned knowledge and patterns. Although image features are numerical data, after encoding conversion, the pre-trained language model can understand image features and text information in the same semantic space. Image features provide the visual semantic information of the image, while classified text descriptions provide the pre-trained language model with preliminary judgment clues and analysis basis. The target prompt words drive the pre-trained language model to perform deep reasoning on the input information according to specific logic and task requirements. Based on the language logic, knowledge associations, and reasoning rules learned during the training process, the pre-trained language model conducts a comprehensive analysis of possible forgery traces in image features and suspicious points in classified text descriptions. Through complex neural network calculations and semantic matching, it generates a judgment result on the authenticity of the target image and outputs it in the form of natural language text.

[0096] In this embodiment, the training prompt word functions like an instruction or guidance, which can help the pre-trained language model clarify the goal and focus of the task.

[0097] Detection tasks can include detecting the authenticity of an image or identifying forged areas in an image. Training prompts can be constructed based on the detection task. For example, training prompts could be "Determine whether the following content is a deepfake" or "Analyze features in an image to determine whether a face has been replaced."

[0098] In some embodiments, special annotations from the training dataset, i.e., reasons for forging an image, can also be added to the training prompts. For example, a training prompt could be, "The mouth expression in the image is unnatural; the raised corners of the mouth are not the normal smile. Can you point out the type of forgery in the image and mark the forged areas?" It should be noted that during the inference phase of the pre-trained language model, target prompts do not need to be specially annotated.

[0099] Based on the training prompt words and training data set, the pre-trained language model can learn the association between forgery-related features and forgery reasons in the image.

[0100] Pre-trained language models usually have a huge parameter system, which has been optimized in the pre-training stage to adapt to a variety of natural language processing tasks. However, in the detection task, not all parameters need to be readjusted. For example, the first parameter of the pre-trained language model can be frozen so that the first parameter remains unchanged during the training process. Among them, the first parameter can be a parameter related to the basic language understanding ability of the pre-trained language model, such as the parameters of the word embedding layer, the parameters of some encoder layers, etc. Freezing the first parameter can retain the advantages of the pre-trained model in language understanding, reduce unnecessary adjustments to the first parameter during training, and thus reduce the complexity and computational cost of training.

[0101] Secondary parameters are those that need to be adjusted based on the detection task. These can be parameters in the output layer of a pre-trained language model or in a specific layer related to the task. During training, these secondary parameters are optimized based on the samples and labels in the training dataset.

[0102] Specifically, the training cue words and training dataset are input into the pre-trained language model. The pre-trained language model processes the data according to the cue words and outputs a prediction. A second loss function is then used to calculate the difference between the pre-trained language model's prediction and the true value of the label. This second loss function measures the pre-trained language model's performance on the detection task. For example, the second loss function can measure the difference by calculating the cross-entropy between the predicted probability and the true label.

[0103] Based on the value of the second loss function, the second parameter of the pre-trained language model is updated using a backpropagation algorithm. This algorithm propagates the error from the output layer of the pre-trained language model back to the layer where the second parameter resides, adjusting the second parameter based on the magnitude and direction of the error. This method gradually optimizes the second parameter, making the pre-trained language model more suitable for detection tasks.

[0104] In some embodiments, the second loss function is expressed as:

[0105]

[0106]

[0107] in, represents the second loss function, represents the true value of the label, represents the predicted value of the pre-trained language model, represents the cross entropy loss, Indicates the number of images, Indicates the The true value of the image, Indicates the The predicted value of an image.

[0108] In this embodiment, freezing the first parameter can reduce excessive interference to the core structure of the pre-trained language model during the training process, retain the advantages of the pre-trained language model in language understanding, reduce the complexity and computational cost of training, and by updating the second parameter, enable the pre-trained language model to focus on learning features related to deep fake detection and more accurately identify fake content.

[0109] In some embodiments, the deepfake detection model further includes a feature fuser and a heat map generator;

[0110] Feature fusion device is used to extract features from the output results to obtain text features; and to fuse text features and image features to obtain fused features;

[0111] A heat map generator is used to generate a heat map of the target image based on the fused features; the heat map represents the forged and real areas in the target image through different pixel values.

[0112] In this embodiment, the output of the pre-trained language model is image authenticity judgment and related analysis presented in the form of natural language text, which contains rich semantic information. To extract this information from the output, natural language processing techniques can be used. For example, pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and GPT can be used as feature extractors to encode the output and convert the text into a fixed-dimensional feature vector, i.e., text features.

[0113] Image features are extracted from the target image by the image feature extractor and contain the image's visual information. Text features represent the semantic analysis of the image by the pre-trained language model. Image features and text features describe the target image from different perspectives, but differ in dimensionality and representation. The goal of feature fusion is to combine the image's visual information with the text's semantic information to form a more comprehensive feature representation.

[0114] In some embodiments, the feature fuser can achieve the fusion of text features and image features in different ways. For example, splicing fusion can be used to splice text features and image features in dimensions to form a new vector containing both visual and semantic information. Alternatively, weighted fusion can be used to assign different weights to image features and text features based on their importance, and then perform a weighted sum to obtain a fused feature. Of course, fusion can also be achieved in other ways, which are not limited in the embodiments of the present application.

[0115] In some embodiments, the feature fuser can be a module based on the Transformer architecture, whose core function is to capture the semantic association between text and images through a cross-attention mechanism, and deeply fuse text features and image features to generate fused features containing more information.

[0116] In this embodiment, a heat map is a visualization tool that directly displays the authenticity of different regions in an image, distinguishing between forged and authentic regions in the target image by different pixel values. When generating a heat map, fusion features can be used to evaluate each pixel or region of the target image.

[0117] For example, a heat map generator can be used to generate a heat map. The heat map generator can include an encoder and a decoder. After the fused features and the target image are input into the heat map generator, the encoder can use a convolutional neural network to process the fused features to capture information about the forged and real areas in the image. The decoder is responsible for gradually restoring the encoder's processing results to a heat map with the same resolution as the target image. For example, based on the encoder's processing results, a value representing the likelihood of authenticity can be predicted for each pixel in the image, and these values ​​can be mapped to different colors or grayscale levels to generate a heat map.

[0118] In a heat map, darker colors or higher grayscale values ​​indicate areas that are more likely to be forged, while lighter colors or lower grayscale values ​​indicate areas that are more likely to be real. Of course, a heat map can also be a binary image, where forged areas can be white (pixel value 1) and real areas can be black (pixel value 0), or vice versa.

[0119] In this embodiment, by extracting text features, the semantic information in the output results of the pre-trained language model can be fully mined. Feature fusion combines the visual features of the image with the semantic features to achieve information complementarity and enhancement. The generation of the heat map presents the detection results in an intuitive and visual manner, allowing users to intuitively locate the forged area by observing the heat map, thereby improving the interpretability of the detection results.

[0120] In some embodiments, text features and image features are fused to obtain fused features, including:

[0121] Input text features and image features into the feature fusion device to obtain fused features output by the feature fusion device; the feature fusion device is used to perform cross-attention calculation using the target vector group as a query, the text features as a key, and the image features as a value to obtain fused features;

[0122] Among them, the target vector group is a set of vectors that is initially randomly generated during the feature fuser training process and is obtained through back-propagation iteration during training.

[0123] In traditional cross-attention calculations, image features are usually used as queries (Query, Q), and text features are used as both keys (Key, K) and values ​​(Value, V) to explore the association between images and text.

[0124] In order to further focus attention on forgery-related features during the fusion process and improve the ability to identify forged areas, in this embodiment, a learnable target vector group is designed as the query (Query), and text features are used as keys (Key) and image features as values ​​(Value) for cross-attention calculation to obtain fused features.

[0125] The target vector group is a set of learnable vectors. During the training phase of the feature fusing machine, the training dataset is constructed using a similar approach to that used for training classifiers. The target vector group is initially randomly generated and serves as the model parameter for the feature fusing machine. As training progresses, the loss function is calculated and iteratively updated through backpropagation. During training, the target vector group serves as the query, with text features from the training samples serving as keys and image features serving as values ​​for cross-attention calculations. The parameters of the target vector group are then updated through backpropagation based on the loss function, enabling the target vector group to better guide the feature fusing machine to focus on forgery-related features during the feature fusion process.

[0126] In this embodiment, by introducing a dynamically updated target vector group, the feature fuser focuses on forgery-related features during the feature fusion process, generates more informative fusion features, and thus improves the accuracy of determining forgery areas.

[0127] In some embodiments, training a deep fake detection model using a training dataset includes:

[0128] Initialize the target vector group of the fusion feature detector;

[0129] The training data set is input into the fusion feature detector, and the target vector group is updated by the back propagation algorithm based on the third loss function; the third loss function is expressed as:

[0130]

[0131]

[0132]

[0133]

[0134] in, Represents real text prompts, such as semantic information such as whether the face in the training set image is forged and the forged location; Represents the predicted text hint, that is, the semantic representation corresponding to the text feature predicted by the feature fuser, Indicates the size of the vocabulary, Indicates the height of the mask, Indicates the width of the mask, represents the Sigmod activation function, Represents the true pixel mask value, that is, it indicates which areas in the image belong to the forged area and which areas belong to the real area; represents the predicted pixel mask value, Represents a hyperparameter, which is used to prevent the denominator from being zero.

[0135] in, 、 and Represents the third loss function Different loss terms in and represents a hyperparameter that balances the influence of different loss terms.

[0136] The role of is to constrain the text alignment loss of the feature fusion, so that the feature fusion can trigger the correct segmentation description. During the training process, the model parameters (including the learnable target vector group) are continuously optimized through back propagation to make the predicted text prompts as close to the real text prompts as possible, thus achieving text alignment.

[0137] The role of is to make the output of the feature fusion device have the correct segmentation area. , which can minimize the difference between the predicted pixel mask value and the real pixel mask value during training. For example, in a face replacement image, the real pixel mask value will mark the pixel area of ​​the replaced face as 1 and the others as 0. The feature fusion unit adjusts the model parameters (including the learnable target vector group) through loss back propagation, so that the predicted pixel mask value can accurately fit the distribution of the real pixel mask value and learn the pixel-level positioning rules of the forged area. , further constraining the prediction accuracy of the feature fuser for the forged area, and making the predicted forged area mask and the real mask more closely related at the pixel level from the perspective of overall matching.

[0138] An example 、 and Function: The original background is a photo of a person in a park, and the modified part is a replaced face. The function of is to enable the feature fusion device to fuse image features and text features, accurately capture the semantic association of "face tampering", describe that the face in the image has been tampered with, and achieve text alignment; and The purpose is to make the predicted pixel mask value of the tampered face area accurately cover the area of ​​the replaced face.

[0139] In this embodiment, internal model parameters (including a learnable target vector set) are gradually optimized under the multi-dimensional constraints of a loss function. Each iteration outputs a fused feature based on the input text and image features. The loss function then calculates the difference between the feature and the true label (text hint, pixel mask value), and backpropagates the adjustment parameters. As the iterations proceed, the feature fuser's understanding of text semantics and its ability to detect forged image regions continuously improve. This allows the trained feature fuser to focus on features related to forgery during the fusion process, thereby improving the accuracy of subsequent forged region detection.

[0140] In some embodiments, the output result also includes a basis for determining the authenticity of the target image; if the target image is forged, the output result also includes a forged location of the target image.

[0141] In this embodiment, the judgment basis is explanatory content generated by the pre-trained language model based on the input image features and classified text descriptions, which explains the logic and reasoning behind the pre-trained language model's authenticity judgment. For example, if the model determines that the target image is forged, the judgment basis can point out problems such as unnatural edge transitions, abnormal textures, or lighting effects that do not conform to physical laws. This helps users understand the model's decision-making process and increases the credibility of the detection results. By providing judgment basis, users can more intuitively see signs of forgery and thus better assess the authenticity of the image.

[0142] If the result is forged, the output can also include the forged location of the target image. The forged location is the specific area in the target image that has been tampered with or synthesized. For example, if a face image is forged, the forged location may be the eyes, nose, or other facial features. By indicating the forged location in the output, users can quickly locate the problem area in the image and further verify the possibility of forgery.

[0143] In this embodiment, by guiding the output results of the pre-trained language model to also include judgment basis and forgery location information, richer information is provided to the user, enabling the user to understand the detection results more intuitively, thereby improving the interpretability of the forgery detection results.

[0144] The present application also provides a method for detecting deep fakes, including:

[0145] Acquire the target image to be detected;

[0146] The target image is input into the deep fake detection model trained using the above-mentioned deep fake detection model training method to obtain an output result indicating the authenticity of the target image.

[0147] For the description of the embodiments corresponding to the deep fake detection method, please refer to the relevant description of the embodiments corresponding to the training method of the deep fake detection model, which will not be repeated here.

[0148] This application uses a training dataset containing real images and various types of forged images, with the images labeled to indicate authenticity. During training, the deepfake detection model is exposed to a variety of image samples, thereby learning a wider range of more complex image features and forgery patterns. Furthermore, forged images are annotated with the reasons for forgery, helping the deepfake detection model gain a deeper understanding of the essential characteristics of forged content during training. Therefore, the problem of insufficient detection capabilities for carefully designed, high-quality forgeries can be resolved, achieving the technical effect of improving the accuracy of forgery detection.

[0149] like Figure 2 As shown, the embodiment of the present application also provides a deep fake detection system, including:

[0150] A classifier 210 is used to classify the target image and obtain a classified text description representing the authenticity of the target image;

[0151] An image feature extractor 220 is used to extract features from a target image to obtain image features;

[0152] A pre-trained language model 230 is used to generate an output result based on image features, classification text description, and target prompt words; the output result includes a judgment result on the authenticity of the target image;

[0153] The feature fusion unit 240 is used to extract features from the output results to obtain text features; and to fuse the text features and image features to obtain fused features;

[0154] The heat map generator 250 is used to generate a heat map of the target image based on the fused features; the heat map represents the forged area and the real area in the target image through different pixel values.

[0155] For the description of the embodiments corresponding to the deep fake detection system, please refer to the relevant description of the embodiments corresponding to the deep fake detection method, and will not be repeated here.

[0156] This application uses a classifier to obtain a classified text description of the target image's authenticity, combines it with the image features extracted by the image feature extractor, and then feeds both the classified text description and the image features into a pre-trained language model. With the help of target prompt words, the pre-trained language model's powerful semantic understanding and analysis capabilities are fully utilized. By combining the multimodal feature information of text and image, forgery characteristics can be more comprehensively captured. Therefore, the problem of insufficient detection capabilities for carefully designed, high-quality forgeries can be resolved, achieving the technical effect of improving the accuracy of forgery detection.

[0157] Moreover, by extracting text features, the semantic information in the output results of the pre-trained language model can be fully mined. Feature fusion combines the visual features of the image with the semantic features to achieve information complementarity and enhancement. The generation of heat maps presents the detection results in an intuitive and visual manner, allowing users to intuitively locate the forged area by observing the heat map, thereby improving the interpretability of the detection results.

[0158] The deepfake detection model training method provided in the embodiments of the present application can be executed by a deepfake detection model training device. In the embodiments of the present application, the deepfake detection model training method performed by the deepfake detection model training device is used as an example to illustrate the deepfake detection model training device provided in the embodiments of the present application.

[0159] like Figure 3 As shown, the training device of the deep fake detection model includes:

[0160] Acquisition module 310 is used to acquire a training data set; the training data set includes real images and forged images of different forgeries; the real images and forged images are respectively associated with labels indicating the authenticity of the images, and the labels corresponding to the forged images also include the reason for the forgery;

[0161] The training module 320 is used to train the deep fake detection model using the training dataset.

[0162] Through this application, since the classification text description of the authenticity of the target image is obtained by the classifier, combined with the image features extracted by the image feature extractor, the classification text description and image features are jointly input into the pre-trained language model, and guided by the target prompt word, the powerful semantic understanding and analysis capabilities of the pre-trained language model are fully utilized. By combining the multimodal feature information of text and image, the characteristics of forgery can be captured more comprehensively, and the forged images can be annotated with the reasons for forgery, so that the deep forgery detection model can learn the authenticity characteristics of the image and understand the specific types and abnormal performances of the forged images during the training process. Therefore, the problem of insufficient detection ability for carefully designed and high-quality forged content can be solved, and the technical effect of improving the accuracy of forgery detection can be achieved.

[0163] In some embodiments, the training module 320 may also be used to:

[0164] Calculate the difference between the predicted value of the classifier and the true value of the label during training based on the first loss function, and update the parameters of the classifier through the back propagation algorithm based on the calculation result;

[0165] The first loss function is expressed as:

[0166]

[0167] in, represents the first loss function, Indicates the number of image types, Indicates the true value corresponding to the label, represents the predicted value output by the classifier, Indicates the image type The weight of .

[0168] In some embodiments, the training module 320 may also be used to:

[0169] Obtain training prompt words; the training prompt words are constructed according to the detection task; the training prompt words are used to guide the pre-trained language model to perform the detection task during the training process;

[0170] Freeze the first parameter of the pre-trained language model, input the training prompt word and the training dataset into the pre-trained language model, and update the second parameter of the pre-trained language model through the backpropagation algorithm based on the second loss function; the second loss function represents the difference between the predicted value of the pre-trained language model and the true value of the label.

[0171] In some embodiments, the training module 320 may also be used to:

[0172] Initialize the target vector group of the fusion feature detector;

[0173] The training data set is input into the fusion feature detector, and the target vector group is updated by the back propagation algorithm based on the third loss function; the third loss function is expressed as:

[0174]

[0175]

[0176]

[0177]

[0178] in, 、 and Represents the third loss function Different loss terms in and represents the hyperparameter that balances the influence of different loss terms, Indicates a real text prompt, Represents the predicted text hint, Indicates the size of the vocabulary, Indicates the height of the mask, Indicates the width of the mask, represents the Sigmod activation function, represents the true pixel mask value, represents the predicted pixel mask value, represents a hyperparameter.

[0179] For the description of the embodiment corresponding to the training device for the deep fake detection model, please refer to the relevant description of the embodiment corresponding to the training method for the deep fake detection model, and will not be repeated here.

[0180] The embodiment of the present application also provides an electronic device, such as Figure 4 As shown, it includes a memory 402 and a processor 401, the memory 402 stores a computer program, and the processor 401 is configured to run the computer program to execute any of the above-mentioned training methods for deep fake detection models or the steps in the embodiments of the deep fake detection method.

[0181] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned deep fake detection model training methods or deep fake detection method embodiments when running.

[0182] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0183] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned deep fake detection model training methods or deep fake detection method embodiments.

[0184] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements any of the above-mentioned training methods for deep fake detection models or the steps in the deep fake detection method embodiments.

[0185] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0186] The above is a detailed introduction to the training method of a deep fake detection model, a deep fake detection method and a system provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for training a deep fake detection model, characterized in that: The deepfake detection model includes a pre-trained language model, which is used to generate an output result indicating the authenticity of the target image based on a classified text description obtained by a classifier classifying the target image to be detected, image features obtained by an image feature extractor performing feature extraction on the target image, and a target prompt word; The method comprises: Obtaining a training data set; the training data set includes real images and forged images of different forgeries; the real images and the forged images are each associated with a label indicating the authenticity of the image, and the label corresponding to the forged image also includes a reason for the forgery; The deepfake detection model is trained using the training dataset, including: obtaining training prompt words; constructing the training prompt words according to the detection task; using the training prompt words to guide the pre-trained language model to perform the detection task during training; freezing the first parameters of the pre-trained language model, inputting the training prompt words and the training dataset into the pre-trained language model, and updating the second parameters of the pre-trained language model through a backpropagation algorithm based on a second loss function; the second loss function represents the difference between the predicted value of the pre-trained language model and the true value of the label; The using the training dataset to train the deep fake detection model includes: Calculating the difference between the predicted value of the classifier and the true value of the label during training based on a first loss function, and updating the parameters of the classifier through a back propagation algorithm based on the calculation result; The first loss function is expressed as: in, represents the first loss function, Indicates the number of image types, Represents the true value corresponding to the label, represents the predicted value output by the classifier, Indicates the image type The weight of .

2. The method according to claim 1, characterized in that The image types include real images and forged images of different forgeries; and the weights of the image types are determined according to the proportions of the image types in the training data set.

3. The method according to claim 1, characterized in that The classifier includes a feature extraction network, a multi-classification head and a label classification mapping module; The feature extraction network is used to extract features of the target image; The multi-classification head is used to output a classification result representing the authenticity of the target image based on the features extracted by the feature extraction network; The label classification mapping module is used to map the classification result to a target template to obtain a classification text description.

4. The method according to claim 1, wherein The image feature extractor includes a visual encoder and a feature space alignment module; The visual encoder is used to encode the target image to be detected; The feature space alignment module is used to project the features output by the visual encoder after encoding to a dimension aligned with the text embedding space of the pre-trained language model to obtain image features.

5. The method according to claim 1, wherein The second loss function is expressed as: in, represents the second loss function, represents the true value of the label, represents the predicted value of the pre-trained language model, represents the cross entropy loss, Indicates the number of images, Indicates the The true value of the image, Indicates the The predicted value of an image.

6. The method according to claim 1, characterized in that The deep fake detection model also includes a feature fuser and a heat map generator; A feature fusion device is used to extract features from the output result to obtain text features; and to fuse the text features and the image features to obtain fused features; A heat map generator is used to generate a heat map of the target image based on the fusion features; the heat map represents the forged area and the real area in the target image through different pixel values.

7. The method according to claim 6, characterized in that The fusing the text features and the image features to obtain fused features includes: Inputting the text features and the image features into a feature fusion device to obtain a fusion feature output by the feature fusion device; the feature fusion device is used to perform a cross-attention calculation using a target vector group as a query, the text features as a key, and the image features as a value to obtain the fusion feature; The target vector group is a vector set that is initially randomly generated during the feature fuser training process and is obtained through back-propagation iteration during the training.

8. The method according to claim 7, characterized in that The using the training dataset to train the deep fake detection model includes: Initializing the target vector group of the fusion feature detector; The training data set is input into the fusion feature detector, and the target vector group is updated by the back propagation algorithm based on the third loss function; the third loss function is expressed as: in, 、 and Represents the third loss function Different loss terms in and represents the hyperparameter that balances the influence of different loss terms, Indicates a real text prompt, Represents the predicted text hint, Indicates the size of the vocabulary, Indicates the height of the mask, Indicates the width of the mask, represents the Sigmod activation function, represents the true pixel mask value, represents the predicted pixel mask value, represents a hyperparameter.

9. The method according to claim 1, characterized in that The output result also includes a basis for determining whether the target image is genuine; if the target image is forged, the output result also includes a forged location of the target image.

10. A deep fake detection method, characterized in that: include: Acquire the target image to be detected; The target image is input into a deep fake detection model trained using the method according to any one of claims 1 to 9 to obtain an output result indicating the authenticity of the target image.

11. A deep fake detection system, characterized in that include: A classifier is used to classify a target image and obtain a classified text description characterizing the authenticity of the target image; the classifier is trained according to the following method: obtaining a training data set; The training data set includes real images and forged images of different forgeries; the real images and the forged images are respectively associated with labels indicating the authenticity of the images, and the labels corresponding to the forged images also include the reason for the forgery; Training the classifier using the training data set includes: Calculating the difference between the predicted value of the classifier and the true value of the label during training based on a first loss function, and updating the parameters of the classifier through a back propagation algorithm based on the calculation result; The first loss function is expressed as: in, represents the first loss function, Indicates the number of image types, Represents the true value corresponding to the label, represents the predicted value output by the classifier, Indicates the image type The weight of An image feature extractor, configured to extract features from the target image to obtain image features; A pre-trained language model is used to generate an output result indicating the authenticity of the target image based on the image features, the classified text description, and the target prompt word; A feature fusion device is used to extract features from the output result to obtain text features; and to fuse the text features and the image features to obtain fused features; A heat map generator is used to generate a heat map of the target image based on the fusion features; the heat map represents the forged area and the real area in the target image through different pixel values.

12. A training device for a deep fake detection model, characterized in that: The deepfake detection model includes a pre-trained language model, which is used to generate an output result indicating the authenticity of the target image based on a classified text description obtained by a classifier classifying the target image to be detected, image features obtained by an image feature extractor performing feature extraction on the target image, and a target prompt word; The device comprises: an acquisition module, configured to acquire a training data set; the training data set includes real images and forged images of different forgeries; the real images and the forged images are each associated with a label indicating the authenticity of the image, and the label corresponding to the forged image also includes a reason for the forgery; a training module for training the deepfake detection model using the training dataset, comprising: calculating the difference between the predicted value of the classifier and the true value of the label during training based on a first loss function, and updating the parameters of the classifier using a backpropagation algorithm based on the calculation result; The first loss function is expressed as: in, represents the first loss function, Indicates the number of image types, Represents the true value corresponding to the label, represents the predicted value output by the classifier, Indicates the image type The weight of .

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 10 when executing the computer program.

Citation Information

Patent Citations

  • Multi-image forgery detection method and system based on cross-modal visual large language model

    CN120355985A