Deep pseudo detection model training method and device suitable for multiple attack types

By introducing multi-classification networks into the deep pseudo detection model and adjusting the model parameters, the problem that the existing technology is difficult to identify multiple Deepfake attacks is solved, and a more efficient deep pseudo detection effect is achieved.

CN120356011APending Publication Date: 2025-07-22ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510713538.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing deep pseudo detection tools and technologies are difficult to fully cover the Deepfake technology of multiple attack types, resulting in the inability to effectively identify complex deep pseudo attacks.

Method used

Multi-classification network is introduced in the deep pseudo detection model, and the output of the feature extraction network is used as the input of the binary and multi-classification networks. The model parameters of the feature extraction network are adjusted based on the output results of the binary and multi-classification networks to optimize its performance.

Benefits of technology

The deep pseudo detection model's recognition ability and robustness of various attack types is improved, and the forged features can be extracted more fine-grained, enhancing the detection ability of multiple deep pseudo attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356011A_ABST
    Figure CN120356011A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device suitable for a deep pseudo detection model of various attack types. The method comprises the steps that sample data and a deep pseudo detection model are determined, the deep pseudo detection model comprises a feature extraction network, a dichotomy network and a multi-classification network, the output of the feature extraction network serves as the input of the dichotomy network and the multi-classification network, and the sample data comprises sample images and corresponding dichotomy labels and multi-classification labels; inputting the sample image into a deep pseudo detection model to enable the deep pseudo detection model to extract image features of the sample image through a feature extraction network, performing binary classification processing on the image features through a binary classification network to obtain a binary classification result, and performing multi-classification processing on the image features through a multi-classification network to obtain a multi-classification result; and performing model parameter adjustment on the feature extraction network based on the difference between the dichotomy result and the dichotomy label and the difference between the multi-classification result and the multi-classification label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of information security technology, and in particular, to a training method and device for a deepfake detection model applicable to multiple attack types. Background Art

[0002] Deepfake is a technology for synthesizing or tampering with image, video, and audio content based on AI (Artificial Intelligence) technology. It can replace a person's facial expression, voice, or movement into the image or video of another person, generating highly realistic fake content, and even achieving the effect of being indistinguishable from the real thing, playing an important role in the fields of film and television, entertainment, education, etc.

[0003] Although Deepfake technology has brought significant advantages and innovations, the accompanying risks cannot be underestimated. For example, criminals use Deepfake technology to impersonate corporate executives and instruct employees to transfer funds through video conferencing, resulting in huge financial losses. Currently, although there are some deepfake detection tools and technologies, most of them can only effectively identify specific types of attacks. Due to the wide variety of Deepfake technologies, the current detection methods are difficult to cover comprehensively. Therefore, there is an urgent need for a detection method applicable to multiple attack types. Summary of the Invention

[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:

[0005] According to a first aspect of one or more embodiments of this specification, a training method for a deepfake detection model applicable to multiple attack types is proposed. The method includes:

[0006] Determine sample data and a deepfake detection model. The deepfake detection model includes a feature extraction network, a binary classification network, and a multi-classification network. The output of the feature extraction network serves as the input to the binary classification network and the multi-classification network. The sample data includes sample images and their corresponding binary classification labels and multi-classification labels;

[0007] Input the sample images into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample images through the feature extraction network, performs binary classification processing on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through the multi-classification network to obtain a multi-classification result;

[0008] Based on the differences between the binary classification results and the binary classification labels, as well as the differences between the multi-classification results and the multi-classification labels, adjust the model parameters of the feature extraction network;

[0009] Among them, the binary classification labels and the binary classification results are used to identify whether there is a deepfake attack, and the multi-classification labels and the multi-classification results are used to identify the attack types of deepfake attacks.

[0010] According to the second aspect of one or more embodiments of this specification, a deepfake detection method applicable to multiple attack types is proposed. The method includes:

[0011] Obtain the image to be detected, and input the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected. The deepfake detection model is trained based on the steps described in the first aspect above.

[0012] According to the third aspect of one or more embodiments of this specification, an electronic device is proposed, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor realizes the steps of the method described in the first aspect above by running the executable instructions.

[0013] According to the fourth aspect of one or more embodiments of this specification, a computer-readable storage medium is proposed, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect above are realized.

[0014] According to the fifth aspect of one or more embodiments of this specification, a computer program product is proposed, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in the first aspect above are realized.

[0015] As can be seen from the above embodiments, this specification adds a multi-classification network to the deepfake detection model, uses the output of the feature extraction network as the input of both the binary classification network and the multi-classification network at the same time, and then adjusts the model parameters of the feature extraction network based on the output results of the binary classification network and the multi-classification network to optimize the performance of the feature extraction network. In this way, not only can the feature extraction network be driven to extract more fine-grained forgery features for deepfake attacks of different attack types, but also the recognition ability and robustness of the overall model for deepfake attacks of various attack types can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic diagram of the architecture of a deepfake detection service system provided by an exemplary embodiment.

[0017] Figure 2It is a flowchart of a training method for a deepfake detection model applicable to multiple attack types provided by an exemplary embodiment.

[0018] Figure 3 It is a flowchart of a training method for a deepfake detection model applicable to multiple attack types provided by an exemplary embodiment.

[0019] Figure 4 It is a schematic structural diagram of a deepfake detection model provided by an exemplary embodiment.

[0020] Figure 5 It is a flowchart of a deepfake detection method applicable to multiple attack types provided by an exemplary embodiment.

[0021] Figure 6 It is a schematic structural diagram of a device provided by an exemplary embodiment.

[0022] Figure 7 It is a block diagram of a training device for a deepfake detection model applicable to multiple attack types provided by an exemplary embodiment.

[0023] Figure 8 It is a block diagram of a deepfake detection device applicable to multiple attack types provided by an exemplary embodiment. Detailed implementation manners

[0024] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0025] Deepfake is a technology for synthesizing or tampering with image, video, and audio content based on AI (Artificial Intelligence) technology. It can replace a person's facial expression, voice, or movement into another person's image or video, generating highly realistic false content, and even achieving the effect of being indistinguishable from the real thing, playing an important role in fields such as film and television dramas, entertainment, and education. For example, in a film and television drama, replacing an actor's face with the actor's younger face to improve the viewing effect of the film and television drama; or, by applying a student's face or voice to a video of a historical figure, enabling students to more vividly experience historical events, etc.

[0026] Although Deepfake technology has brought significant advantages and innovations, the accompanying risks cannot be underestimated. For example, criminals use Deepfake technology to impersonate corporate executives and instruct employees to transfer funds through video conferencing, resulting in huge financial losses; another example is using Deepfake technology to create false news, causing damage to the reputation of others, etc.

[0027] Currently, there are already some deepfake detection tools and technologies, but most of them use binary classification models, which can effectively identify specific types of deepfake attacks. However, in the face of a wide variety of Deepfake technologies, this model is difficult to comprehensively cover. Therefore, there is an urgent need for a detection method that can be applied to multiple attack types.

[0028] Based on this, this specification provides a training method for a deepfake detection model applicable to multiple attack types. By introducing a multi-classification network into the deepfake detection model, a deepfake detection model applicable to multiple attack types is trained.

[0029] In implementation, first, determine the sample data and the deepfake detection model. The deepfake detection model includes a feature extraction network, a binary classification network, and a multi-classification network. The output of the feature extraction network serves as the input to the binary classification network and the multi-classification network. The sample data includes sample images and their corresponding binary classification labels and multi-classification labels. Then, input the sample images into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample images through the feature extraction network, performs binary classification processing on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through the multi-classification network to obtain a multi-classification result. Finally, based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, adjust the model parameters of the feature extraction network.

[0030] As can be seen from the above embodiments, this specification adds a multi-classification network to the deepfake detection model, uses the output of the feature extraction network as the input to both the binary classification network and the multi-classification network, and then adjusts the model parameters of the feature extraction network based on the output results of the binary classification network and the multi-classification network to optimize the performance of the feature extraction network. In this way, not only can the feature extraction network extract more fine-grained forgery features for deepfake attacks of different attack types, but also the recognition ability and robustness of the overall model for deepfake attacks of various attack types can be improved.

[0031] Figure 1 It is a schematic diagram of the architecture of a deepfake detection service system provided by an exemplary embodiment. As Figure 1 shown, the system may include a server 11, a network 12, and several electronic devices, such as a PC (Personal Computer) 13, a mobile phone 14, etc.

[0032] Server 11 can be a physical server including an independent host, or the server 11 can be a virtual server hosted by a host cluster. During operation, the server 11 can run the server-side program of a certain application to implement the related functions of the application. For example, when the server 11 runs the program of the deepfake detection service, it can be implemented as the corresponding deepfake detection service platform.

[0033] PC 13 and mobile phone 14 are only some types of electronic devices that users can use. In fact, users can obviously also use electronic devices of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.), etc. One or more embodiments of this specification do not limit this. During operation, the electronic device can run the client-side program of a certain application to implement the related functions of the application. For example, when the electronic device runs the program of the deepfake detection service, it can be implemented as the client of the deepfake detection service. Among them, the application program of the client of the above deepfake detection service can be started and run on the electronic device. The client-side program can be a native application program installed on the electronic device, or the client-side program can be in the form of a small program, a quick application, or other similar forms. Of course, when using web technologies such as HTML5 or similar, the related functions can be implemented through the page displayed by the browser. Here, the browser can be an independent browser application or a browser module embedded in some applications.

[0034] For the network 12 for the interaction between electronic devices such as PC 13 and mobile phone 14 and the server 11, it can be specifically selected to use a wired or wireless network to implement communication based on the communication methods supported by the corresponding electronic devices. This specification does not limit this. For example, PC 13 can support both wired and wireless communication, so wired or wireless network can be used for communication according to needs, while mobile phone 14 usually only supports wireless communication, so wireless network can be used for communication.

[0035] Exemplarily, the training method of the deepfake detection model applicable to multiple attack types provided by the embodiments of this specification can be executed by any one of the above-mentioned client or the above-mentioned server. For example, the server determines sample data and a deepfake detection model, inputs the sample image into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample image through a feature extraction network, performs binary classification processing on the image features through a binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through a multi-classification network to obtain a multi-classification result; then, based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, the model parameters of the feature extraction network are adjusted.

[0036] Exemplarily, the training method of the deepfake detection model applicable to multiple attack types provided by the embodiments of this specification can be completed by the mutual cooperation of the above-mentioned client and the above-mentioned server. For example, the client uses deepfake methods of multiple attack types to generate sample images to construct a sample training set covering multiple attack types, and sends the sample training set to the server; the server uses the sample training set to train the deepfake detection model.

[0037] Similarly, the deepfake detection method applicable to multiple attack types provided by the embodiments of this specification can be executed by any one of the above-mentioned client or the above-mentioned server. For example, the deepfake detection model is deployed to the client, and the client obtains the image to be detected, inputs the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected; or, the deepfake detection model is deployed to the server, and the server obtains the image to be detected, inputs the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected.

[0038] The deepfake detection method applicable to multiple attack types provided by the embodiments of this specification can be realized by the mutual cooperation of the above-mentioned client and the server. Exemplarily, the client collects the image to be detected and sends the image to be detected to the server, and the server inputs the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected.

[0039] Exemplarily, the training method for a deepfake detection model applicable to multiple attack types and the deepfake detection method applicable to multiple attack types provided in the embodiments of this specification can be executed by different parties. For example, the server trains the deepfake detection model and deploys the trained deepfake detection model to the client; then, after the client obtains the image to be detected, it inputs the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected. Another example is that the client trains the deepfake detection model and deploys the trained deepfake detection model to the server; then, after the client obtains the image to be detected, it sends the image to be detected to the server, and the server inputs the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected.

[0040] It should be noted that this specification only exemplarily describes the cooperation method between the client and the server and does not limit it. The specific cooperation method can be set arbitrarily according to actual needs.

[0041] Next, this specification exemplarily describes the training process of the deepfake detection model:

[0042] Figure 2 is a flowchart of a training method for a deepfake detection model applicable to multiple attack types provided by an exemplary embodiment. The execution subject of this method can be Figure 1 the system shown, or Figure 1 any component in the system shown. This specification does not limit the execution subject. This method includes:

[0043] Step S201, determine the sample data and the deepfake detection model. The deepfake detection model includes a feature extraction network, a binary classification network, and a multi-classification network. The output of the feature extraction network is used as the input of the binary classification network and the multi-classification network. The sample data includes sample images and their corresponding binary classification labels and multi-classification labels.

[0044] In this specification, the sample data is the data used to train the deepfake detection model. The sample data includes sample images and their corresponding binary classification labels and multi-classification labels. The sample image can be a real image or a fake image tampered with or generated by deepfake technology; if the sample image is a real image, the corresponding binary classification label and multi-classification label of the sample image both indicate that there is no deepfake attack on the sample image, that is, it indicates that the sample image is a real image; if the sample image is a fake image, the corresponding binary classification label of the sample image indicates that there is a deepfake attack on the sample image, that is, it indicates that the sample image is a fake image, and the corresponding multi-classification label of the sample image indicates the attack type of the deepfake attack in the sample image, that is, it indicates what deepfake technology is used to tamper with or generate the sample image.

[0045] In one embodiment, the binary classification label can be represented by numbers. Exemplarily, the binary classification label can be represented by 0 / 1. For example, when the binary classification label is 0, it indicates that there is no deepfake attack, and the corresponding image is a real image; when the binary classification label is 1, it indicates that there is a deepfake attack, and the corresponding image is a fake image tampered with or generated by deepfake technology. Another example is that when the binary classification label is 0, it indicates that there is a deepfake attack, and the corresponding image is a fake image tampered with or generated by deepfake technology; when the binary classification label is 1, it indicates that there is no deepfake attack, and the corresponding image is a real image.

[0046] In another embodiment, the binary classification label can be represented by text. Exemplarily, the binary classification label can be represented by "real face / deepfake". Exemplarily, the binary classification label can be represented by "real / fake". This specification only gives an exemplary description of the representation form of the binary classification label. The binary classification label can be represented by any distinguishable representation form, and this specification does not limit this.

[0047] The attack types of deepfake attacks are classified according to the technical means of deepfake technology. In one embodiment, based on the differences in the technical means of deepfake technology, the attack types of deepfake attacks can be classified into multiple attack types such as face swap, animation, attribute editing, and stablediffusion generation. Among them, face swap refers to replacing the face of one object with the face of another object; animation refers to generating a dynamic video with natural expression changes or motion changes based on a static picture; attribute editing refers to modifying specific attributes (such as age, gender, expression, hair color, etc.) of an object in an image without changing other parts of the content. For example, changing a person's photo from frowning to smiling, or from young to old, etc. Stablediffusion generation refers to generating an image by a diffusion model.

[0048] It should be noted that under each attack type such as face swap, animation, attribute editing, and stablediffusion generation, there can be many different technical means for deepfake technology. Therefore, when classifying the attack types, the attack types of deepfake attacks can also be classified with a finer or coarser granularity, and this specification does not limit this. Of course, when classifying the attack types of deepfake attacks, it can also be classified according to other criteria or rules, and this specification does not limit this.

[0049] In this specification, multi-class labels are used to identify the attack types of deepfake attacks. Therefore, the number of multi-class labels is determined by the number of attack types of deepfake attacks divided. In one embodiment, the multi-class labels can be represented by numbers. Exemplarily, when the attack types of deepfake attacks are divided into face swapping, animation, attribute editing, and diffusion model generation, the multi-class labels can be represented by numbers from 0 to 4. For example, when the multi-class label is 0, it indicates that there is no deepfake attack, and the corresponding image is a real image; when the multi-class label is 1, it indicates that the attack type of the deepfake attack is face swapping; when the multi-class label is 2, it indicates that the attack type of the deepfake attack is animation; when the multi-class label is 3, it indicates that the attack type of the deepfake attack is attribute editing; when the multi-class label is 4, it indicates that the attack type of the deepfake attack is diffusion model generation.

[0050] In another embodiment, the multi-class labels can be represented by text. Exemplarily, the multi-class labels can be represented by "real / face swapping / animation / attribute editing / diffusion model generation". Exemplarily, the multi-class labels can be represented by "realface / swap / animation / attribute editing / stablediffusion". This specification only gives exemplary descriptions of the representation forms of the multi-class labels. The multi-class labels can be represented by any distinguishable representation form, and this specification does not limit this.

[0051] The deepfake detection model is a model used to detect whether the input image is a real image or a fake image tampered with or generated by deepfake technology. The deepfake detection model usually only includes a binary classification network and only needs to detect whether the input image is a real image or a fake image. However, in this specification, a multi-class classification network is newly added to the deepfake detection model. In this way, the deepfake detection model includes a feature extraction network, a binary classification network, and a multi-class classification network. Among them, the binary classification network and the multi-class classification network share the feature extraction network, that is to say, the output of the feature extraction network is used as the input of the binary classification network and the multi-class classification network.

[0052] In one embodiment, the above-mentioned feature extraction network uses a backbone network, and the backbone network can specifically be a network architecture based on deep learning, which can be used to extract image features from the image input to this backbone network. It should be noted that this specification does not limit the network architecture of the backbone network.

[0053] Exemplarily, the backbone network is composed of multiple layers of convolutional neural networks, and image features are extracted from the input image through the multiple layers of convolutional neural networks. Herein, the specific type of convolutional neural network used to implement the backbone network is not limited in this specification. For example, ResNet (Residual Network) can be used to implement the backbone network of the deepfake detection model. Another example is that VGG (Visual Geometry Group) can be used to implement the backbone network of the deepfake detection model.

[0054] Exemplarily, the backbone network is a neural network based on Transformer (Transformer). Herein, the specific type of Transformer used to implement the backbone network is not limited in this specification. For example, ViT (Vision Transformer) can be used to implement the backbone network of the deepfake detection model. Another example is that Swin Transformer can be used to implement the backbone network of the deepfake detection model.

[0055] Exemplarily, the network architecture of the backbone network can be a hybrid network architecture, that is, it is composed of multiple neural networks with different network architectures. Herein, the network architectures of the multiple neural networks constituting the backbone network are not limited in this specification. For example, a convolutional neural network and ViT are combined as the backbone network for implementing the deepfake detection model.

[0056] In this specification, the backbone network, as the core component for feature extraction, its architecture can be flexibly designed according to specific task requirements. In one embodiment, the backbone network can adopt a single-branch architecture, that is, high-level image features of the input image are gradually extracted through a series of consecutive layers (such as convolutional layers, pooling layers, etc.). In another embodiment, the backbone network can also be designed as an architecture including multiple branches, and the backbone network can extract features from different scales or angles through multiple branches, thereby enhancing the expression ability of the deepfake detection model for the input image.

[0057] The binary classification network is a network used to distinguish whether the input image is a real image or a fake image. The binary classification network can adopt the network architecture of any classification network, and the network architecture of the binary classification network is not limited in this specification. Exemplarily, the binary classification network includes several fully connected layers and activation functions. The binary classification network processes the input image features sequentially through several fully connected layers, and maps the output of the last fully connected layer to the interval [0, 1] through the activation function. When the output value of the activation function approaches 0, it indicates that the input image is a real image, and when the output value of the activation function approaches 1, it indicates that the input image is a fake image.

[0058] A multi-classification network is a network used to identify which category an input image belongs to among multiple predefined categories. In this specification, the multi-classification network can not only distinguish whether the input image is a real image or a fake image, but also further identify which type of deepfake attack the fake image is tampered with or forged by. The multi-classification network can adopt the network architecture of any classification network, and this specification does not limit the network architecture of the multi-classification network. Exemplarily, in the multi-classification network, a classification branch is set for each attack type, and this classification branch is used to determine whether it belongs to this attack type. Among them, each classification branch includes several fully connected layers and activation functions. The input image features are processed sequentially through several fully connected layers, and the output of the last fully connected layer is mapped to the interval [0, 1] through the activation function. When the output value of the activation function approaches 0, it means it does not belong to this attack type, and when the output value of the activation function approaches 1, it means it belongs to this attack type.

[0059] Step S202: Input the sample image into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample image through the feature extraction network, performs binary classification processing on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through the multi-classification network to obtain a multi-classification result.

[0060] Among them, the binary classification result is used to identify whether there is a deepfake attack, and the multi-classification result is used to identify the attack type of the deepfake attack.

[0061] Step S203: Based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, adjust the model parameters of the feature extraction network; among them, the binary classification label and the binary classification result are used to identify whether there is a deepfake attack, and the multi-classification label and the multi-classification result are used to identify the attack type of the deepfake attack.

[0062] That is to say, a multi-task learning mode of simultaneous parallel training of binary classification + multi-classification is adopted to train the feature extraction network. In this way, the supervision signal of multi-classification can drive the feature extraction network to extract more fine-grained forgery features for deepfake attacks of different attack types, and at the same time provide some interpretability for the binary classification result in terms of attack type.

[0063] This specification adds a multi-classification network to the deepfake detection model, takes the output of the feature extraction network as the input of both the binary classification network and the multi-classification network at the same time, and then adjusts the model parameters of the feature extraction network based on the output results of the binary classification network and the multi-classification network to optimize the performance of the feature extraction network. In this way, not only can the feature extraction network be driven to extract more fine-grained forgery features for deepfake attacks of different attack types, but also the recognition ability and robustness of the overall model for deepfake attacks of various attack types can be improved.

[0064] Next, taking the feature extraction network including multiple branches as an example, the training method of the deepfake detection model will be exemplarily described. Figure 3 It is a flowchart of a training method of a deepfake detection model applicable to multiple attack types provided by an exemplary embodiment. The execution subject of this method can be Figure 1 the system shown in Figure 1 any component in the system shown in

[0065] Step S301, determine the sample data and the deepfake detection model. The deepfake detection model includes a feature extraction network, a binary classification network, and a multi-classification network. The feature extraction network includes a global feature extraction network, a local feature extraction network, and a fusion layer. The sample data includes sample images and their corresponding binary classification labels and multi-classification labels.

[0066] In one embodiment, determining the sample data includes: obtaining any sample data from the sample training set, where the sample training set includes multiple sample data. Optionally, the sample training set is any existing public training set. Optionally, considering the numerous technical means of deepfake technology, in order to make the sample training set contain more comprehensive sample data, the sample training set can be generated by deepfake methods of different attack types.

[0067] Exemplarily, the method further includes: generating sample images by deepfake methods of multiple attack types to construct a sample training set covering multiple attack types, and the sample training set is used to train the deepfake detection model. Among them, the sample images generated by any deepfake method of an attack type are all set with a binary classification label for identifying the existence of a deepfake attack, and a multi-classification label for identifying the attack type, so as to obtain sample data including sample images and their corresponding binary classification labels and multi-classification labels.

[0068] Among them, the deepfake methods of multiple attack types can be each existing deepfake method currently, or common deepfake methods, or representative deepfake methods selected according to the forgery characteristics of different deepfake methods. This specification does not limit this.

[0069] In this specification, the global feature extraction network is used to extract the global features of the input image, and this specification does not limit the network architecture of the global feature extraction network. Optionally, the global feature extraction network is a convolutional neural network, for example, ResNet, VGG, etc. Optionally, the global feature extraction network is a neural network based on Transformer. For example, ViT, SwinTransformer, etc.

[0070] In this specification, the local feature extraction network is used to extract the local features of the input image, and the network architecture of the local feature extraction network is not limited in this specification. Optionally, the local feature extraction network is a convolutional neural network, such as ResNet, VGG, etc. Optionally, the local feature extraction network is a neural network based on the local attention mechanism. Optionally, the local feature extraction network is a U-Net (an encoder-decoder structure) network.

[0071] The feature extraction network in this specification includes a global feature extraction network and a local feature extraction network. It should be noted that the number of the global feature extraction network and the local feature extraction network is not limited in this specification. That is to say, the feature extraction network may include one global feature extraction network or multiple global feature extraction networks; the feature extraction network may include one local feature extraction network or multiple local feature extraction networks.

[0072] Step S302: Input the sample image into the deepfake detection model, so that the deepfake detection model extracts the global features of the sample image through the global feature extraction network, extracts the local features of the sample image through the local feature extraction network, fuses the global features and the local features through the fusion layer to obtain the image features of the sample image, performs binary classification processing on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through the multi-classification network to obtain a multi-classification result.

[0073] The global feature extraction network and the local feature extraction network have been described in the above step S301, and will not be elaborated here one by one. Referring to the above step S301, in step S302, only several specific network architectures are used as examples to exemplarily illustrate the data processing process of the global feature extraction network and the local feature extraction network:

[0074] In an illustrated embodiment, the global feature extraction network includes a visual feature extraction network, which is used to extract the global visual features of the input RGB image. Among them, the RGB image refers to an image represented by three color channels of red (R), green (G), and blue (B). In this specification, the sample image is an RGB image. Therefore, when the sample image is input into the visual feature extraction network, the visual feature extraction network can extract the global visual features of the sample image, and the extracted features are the global features of the sample image.

[0075] In one embodiment, the visual feature extraction network is a ViT network. Optionally, the ViT network includes multiple encoding layers, and each encoding layer contains a multi-head self-attention mechanism and a feed-forward neural network. Since each encoding layer contains a self-attention mechanism, when encoding, the relationship between any two image patches in the image can be concerned, so that the global features in the image can be captured. In addition, the ViT network does not need to rely on the network architecture of the traditional convolutional neural network, so it has better performance in some classification tasks.

[0076] It should be noted that this specification only takes "the ViT network includes multiple encoding layers, and each encoding layer contains a multi-head self-attention mechanism and a feed-forward neural network" as an example to exemplarily illustrate the network architecture of the ViT network, and does not limit the network architecture of the ViT network. In actual applications, the network architecture of the ViT network can be adjusted according to needs.

[0077] Next, taking the visual feature extraction network as a ViT network as an example, the process of "extracting the global features of the sample image through the global feature extraction network" will be exemplarily described:

[0078] Exemplarily, extracting the global features of the sample image through the global feature extraction network includes: dividing the input image into multiple image patches (these multiple image patches can be of a fixed size and a fixed shape, such as a square with a fixed side length), and these image patches can overlap or not intersect; then, flattening each image patch into a one-dimensional vector and adding a position encoding to the one-dimensional vector to retain the spatial information of the image patch; then, combining the one-dimensional vectors of all image patches into a sequence according to their positions in the image, and inputting this sequence into the Transformer module of the ViT network. Among them, the Transformer module includes multiple encoding layers, and each encoding layer contains a multi-head self-attention mechanism and a feed-forward neural network; the features output by the last encoding layer can be used as the global features of the sample image; the features output by each encoding layer can also be combined as the global features of the sample image; it is also possible to obtain shallow features, middle features and deep features from multiple encoding layers, and combine the shallow features, middle features and deep features as the global features of the sample image. For example, obtain the features output by the second encoding layer (shallow features), the features output by the twentieth encoding layer (middle features) and the features output by the fiftieth encoding layer (deep features), and combine these features as the global features of the sample image.

[0079] In another embodiment, the visual feature extraction network is a convolutional neural network. Optionally, the convolutional neural network includes multiple convolutional layers and pooling layers, and the global features of the sample image are extracted through the convolutional layers and pooling layers.

[0080] It should be noted that this specification only provides an exemplary description of the network architecture of the visual feature extraction network and does not limit it. Moreover, the global feature extraction network in this specification can be one or multiple. When the global feature extraction network is one, the visual feature extraction network can be any of the above-mentioned network architectures or a network architecture not mentioned in this specification; when the global feature extraction network is multiple, each global feature extraction network can include visual feature extraction networks with different network architectures. Among them, the network architecture of the visual feature extraction network can be any of the above-mentioned ones or a network architecture not mentioned in this specification, and this specification does not limit this.

[0081] In an illustrated embodiment, the local feature extraction network includes at least one of a frequency-domain feature extraction network and a neighboring pixel relationship feature extraction network. Among them, the local feature extraction network can be one or multiple. When the local feature extraction network is one, the local feature extraction network can be a frequency-domain feature extraction network, a neighboring pixel relationship feature extraction network, or a network not mentioned in this specification; when the local feature extraction network is multiple, the local feature extraction network can include at least one of a frequency-domain feature extraction network and a neighboring pixel relationship feature extraction network, or can also include a network not mentioned in this specification, and this specification does not limit this.

[0082] Next, this specification takes the frequency-domain feature extraction network and the neighboring pixel relationship feature extraction network as examples to make an exemplary description of the local feature extraction network:

[0083] Among them, the frequency-domain feature extraction network is used to traverse the input image through a sliding window to extract the frequency-domain features of each local area of the input image. Since the frequency-domain feature extraction network performs feature extraction on the area selected by the sliding window, the features extracted by the frequency-domain feature extraction network are local features rather than global features.

[0084] In an embodiment, traversing the input image through a sliding window to extract the frequency-domain features of each local area of the input image includes: for any local area selected by the sliding window, converting the local area from the pixel domain to the frequency domain to obtain the frequency-domain information of the local area, and performing feature extraction on the frequency-domain information of the local area to obtain the frequency-domain features of the local area.

[0085] Among them, the sliding window can be of any size, such as 8×8 pixels, 16×16 pixels, etc. This specification does not limit this. When converting the local area from the pixel domain to the frequency domain, any transformation method such as DCT (Discrete Cosine Transform), FT (Fourier Transform), DFT (Discrete Fourier Transform), FFT (Fast Fourier Transform) can be used. This specification does not limit this.

[0086] The frequency domain information of the local area can include at least one of DCT coefficients, frequency responses, amplitude spectra, phase spectra, power spectra, energy distributions, frequency band energy ratios, peak frequencies, etc. This specification only gives an exemplary description of the frequency domain information and does not limit the frequency domain information. The frequency domain information can also include other types of information not mentioned in this specification.

[0087] Exemplarily, the frequency domain information includes the frequency response obtained by DCT transformation. For any local area selected by the sliding window, convert the local area from the pixel domain to the frequency domain to obtain the frequency domain information of the local area, and extract features from the frequency domain information of the local area to obtain the frequency domain features of the local area, including: for any local area selected by the sliding window, perform DCT transformation on the pixel values of the local area to obtain the frequency domain response of the local area, and extract features from the frequency domain response of the local area to obtain the frequency domain features of the local area.

[0088] Among them, the frequency response can include at least one of the amplitude response and the phase response. The amplitude response is used to represent the gain or attenuation degree of different frequency signals, usually in decibels (dB); the phase response is used to represent the phase shift introduced to different frequency signals, usually in degrees (°) or radians (rad).

[0089] Optionally, extracting features from the frequency domain response of the local area includes: extracting features from the frequency domain response through a feature extraction layer (such as a convolutional layer, etc.). Optionally, extracting features from the frequency domain response of the local area includes: dividing the frequency band of the frequency response of the local area, obtaining the frequency responses of each frequency band, and determining the frequency domain features of the local area based on the frequency responses of each frequency band. Exemplarily, determining the frequency domain features of the local area based on the frequency responses of each frequency band includes: obtaining the mean value of the frequency responses of each frequency band, and splicing the mean values of the frequency responses of each frequency band into the frequency domain features of the local area. For example, splicing the mean value of the frequency response of each frequency band as a one-dimensional vector to obtain a multi-dimensional vector. This multi-dimensional vector is the frequency domain feature of the local area.

[0090] In one embodiment, when performing frequency band division, the ranges and / or weights of the respective frequency bands may be fixed. In another embodiment, when performing frequency band division, the ranges and / or weights of the respective frequency bands are optimized during the training process of the frequency domain feature extraction network. That is to say, the ranges and / or weights of the respective frequency bands are learnable, and the ranges and / or weights of the respective frequency bands are determined by the first model parameters in the frequency domain feature extraction network, and the first model parameters are model parameters that can be optimized during the training process of the frequency domain feature extraction network. Therefore, as the first model parameters are optimized, the ranges and / or weights of the respective frequency bands are also optimized.

[0091] Among them, the neighboring pixel relationship feature extraction network is used to obtain the neighboring pixel relationship features of the input image based on the differences between each pixel point and its adjacent pixel points in the input image. Since the features extracted by the neighboring pixel relationship feature extraction network are obtained based on the differences between any pixel point and its adjacent pixel points, it can be seen that when the neighboring pixel relationship feature extraction network performs feature extraction, it only focuses on the local part of the input image and does not focus on the global part of the input image. Therefore, the features extracted by the neighboring pixel relationship feature extraction network are local features.

[0092] In one embodiment, the process of obtaining the neighboring pixel relationship features of the input image based on the differences between each pixel point and its adjacent pixel points in the input image by the neighboring pixel relationship feature extraction network may include: traversing the input image through a sliding window, and for any local area selected by the sliding window, calculating the difference between the pixel value of the central pixel point in the local area and the pixel values of other pixel points, and forming a vector with the calculated multiple differences; directly using the vector as the NPR (Neighboring Pixel Relationships) feature, or for each element Vi in the vector, calculating its difference from the specified element Vj, forming a vector with the calculated multiple differences, and using the vector as the NPR feature. By calculating the neighboring pixel relationships, artifacts (such as hair, eyes, beards, etc.) related to image details can be effectively captured for identifying and analyzing specific features or artifacts left during the process of resolution improvement (i.e., upsampling or magnification) in the image.

[0093] Among them, the sliding window can be of any size, such as 2×2 pixels, 3×3 pixels, etc. This specification does not limit this.

[0094] After obtaining the global features and local features, the obtained global features and local features can be fused to obtain image features, and the image features are input into a binary classification network and a multi-classification network for classification. Among them, this specification does not limit the fusion method of the global features and local features. In one embodiment, the fusion layer is a splicing layer; the global features and local features are fused through the fusion layer to obtain the image features of the sample image, including: splicing the global features and local features through the splicing layer to obtain the image features of the sample image. Exemplarily, the visual feature extraction network outputs the global feature [tensor1], the frequency domain feature extraction network outputs the local feature [tensor2], and the neighboring pixel relationship feature extraction network outputs the local feature [tensor3], and each tensor is in vector form. After splicing, [tensor1, tensor2, tensor3] is obtained.

[0095] In another embodiment, the global features and local features are fused through the fusion layer to obtain the image features of the sample image, including: weighted fusion of the global features and local features through the splicing layer to obtain the image features of the sample image. In another embodiment, the fusion layer is a feature extraction layer; the global features and local features are fused through the fusion layer to obtain the image features of the sample image, including: feature extraction of the global features and local features through the feature extraction layer to obtain the image features of the sample image.

[0096] Exemplarily, as Figure 4 shown, the deepfake detection model includes a feature extraction network (Backbone), a binary classification network, and a multi-classification network. The feature extraction network includes three branches and a fusion layer: an RGB branch, an LFS (local frequency statistics) branch, and an NPR (neighboring pixel relationships) branch. Among them, the RGB branch is a visual feature extraction network, the LFS branch is a frequency domain feature extraction network, and the NPR branch is a neighboring pixel relationship feature extraction network. The fusion layer fuses the features output by the three branches, and inputs the fused features into the binary classification network and the multi-classification network for classification through the binary classification network and the multi-classification network.

[0097] It should be noted that in the design stage of this solution, many feature extraction networks were tried (such as using different network architectures, different branch combinations, etc.). After experiments, it was found that the feature extraction network composed of the RGB branch, the LFS branch, and the NPR branch has the best effect, and the RGB branch uses a ViT network trained based on a mask strategy.

[0098] Step S303: Adjust the model parameters of the feature extraction network based on the differences between the binary classification results and the binary classification labels, and between the multi-classification results and the multi-classification labels.

[0099] In this specification, when training the deepfake detection model, both the binary classification results and the multi-classification results are used as the supervision signals of the feature extraction network. In this way, it can not only drive the feature extraction network to extract more fine-grained forgery features for deepfake attacks of different attack types, but also improve the recognition ability and robustness of the overall model for deepfake attacks of various attack types.

[0100] As Figure 4 shown, the first loss value (L_binary) can be determined based on the difference between the binary classification result and the binary classification label, the second loss value (L_muti) can be determined based on the difference between the multi-classification result and the multi-classification label, the third loss value (Loss) can be determined based on the first loss value and the second loss value, and the model parameters of the feature extraction network can be adjusted based on the third loss value.

[0101] Among them, when determining the third loss value based on the first loss value and the second loss value, the first loss value and the second loss value can be weighted averaged to obtain the third loss value; the first loss value and the second loss value can also be averaged to obtain the third loss value; other operations can also be performed on the first loss value and the second loss value to obtain the third loss value, which is not limited in this specification.

[0102] In one embodiment, the global feature extraction network includes a visual feature extraction network. To improve the global feature extraction ability of the visual feature extraction network, this visual feature extraction network can be trained based on a masking strategy. Among them, the masking strategy means: randomly masking some regions in the input image, training the neural network so that the neural network extracts features from the unmasked regions, and the entire input image can be restored according to the extracted features, thereby enhancing the feature extraction ability of the neural network.

[0103] It should be noted that the neural network extracting features from the unmasked regions is the encoding process, and restoring the entire input image according to the extracted features is the decoding process. The visual feature extraction network in this specification only needs to perform encoding without performing decoding. Therefore, in one embodiment, after training the neural network using the masking strategy, the encoding layer in this neural network is used as the visual feature extraction network in this specification; then the visual feature extraction network is further fine-tuned using the method provided in this specification.

[0104] In one embodiment, the global feature extraction network and / or the local feature extraction network in the deepfake detection model are pre-trained with a large number of images (such as hundreds of millions of face images). In this way, when training the deepfake detection model, not only can a trained deepfake detection model be obtained faster, but also the trained deepfake detection model can be ensured to be more accurate.

[0105] Step S304: Based on the difference between the binary classification result and the binary classification label, adjust the model parameters of the binary classification network, and based on the difference between the multi-classification result and the multi-classification label, adjust the model parameters of the multi-classification network.

[0106] It should be noted that this specification only takes "adjusting the model parameters of the binary classification network based on the difference between the binary classification result and the binary classification label, and adjusting the model parameters of the multi-classification network based on the difference between the multi-classification result and the multi-classification label" as an example to exemplarily illustrate the training process of the deepfake detection model. In another embodiment, the binary classification network and the multi-classification network can be pre-trained networks. When training the deepfake detection model, only the feature extraction network needs to be trained, and there is no need to train the binary classification network and the multi-classification network. In another embodiment, the supervision signals of each network in the deepfake detection model are consistent. That is to say, not only based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, the model parameters of the feature extraction network are adjusted, but also based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, the model parameters of the binary classification network and the multi-classification network are adjusted.

[0107] In the above technical solution, considering that the forgery clues and focuses corresponding to deepfake attacks of different attack types are different, in this specification, in the feature extraction stage, the deepfake detection model not only extracts global features but also pays attention to local features at the same time. This method can significantly increase the possibility of capturing forgery clues, so that the classification network can generate more accurate classification results and improve the accuracy of the deepfake detection model.

[0108] Moreover, in the feature extraction stage, the deepfake detection model can also extract features from multiple dimensions such as the RGB domain, the frequency domain, and the relationship between adjacent pixels, so that the deepfake detection model can more comprehensively capture the forgery traces in the image, thereby improving the recognition ability and accuracy of the deepfake detection model.

[0109] Figure 5 is a flowchart of a deepfake detection method applicable to multiple attack types provided by an exemplary embodiment. The execution subject of this method can be Figure 1 the system shown, or, Figure 1For any component in the system shown, this specification does not limit the execution entity. The method includes:

[0110] Step S501, obtain the image to be detected, and input the image to be detected into the deepfake detection model to obtain the deepfake detection result of the image to be detected. The deepfake detection model is trained based on Figure 2 or Figure 3 the steps of the method described in the embodiment shown.

[0111] The deepfake detection method provided in this embodiment can be widely applied to any scenario that requires deepfake detection. For example, in the face recognition scenario or the document verification scenario, this method can effectively identify forged images or videos and prevent identity theft; in another example, in the scenario of news content review, this method can effectively identify whether the news to be published is false content tampered with by deepfake technology, and maintain the authenticity and credibility of the news.

[0112] The above image to be detected can be any image. For example, a face image, a document image, a landscape image, etc. This specification does not limit this. The image to be detected can be an image captured by a local device, an image stored locally, or an image obtained from other devices. This specification does not limit the source of the image to be detected.

[0113] In one embodiment, after obtaining the trained deepfake detection model, the multi-classification network in the deepfake detection model can be removed. That is, in the application stage of the deepfake detection model, only the binary-classification network is needed for classification. Among them, the deepfake detection result of the image to be detected includes the binary-classification result of the image to be detected.

[0114] In another embodiment, after obtaining the trained deepfake detection model, the binary-classification network in the deepfake detection model can be removed. That is, in the application stage of the deepfake detection model, only the multi-classification network is needed for classification. Among them, the deepfake detection result of the image to be detected includes the multi-classification result of the image to be detected.

[0115] In another embodiment, after obtaining the trained deepfake detection model, neither the multi-classification network nor the binary-classification network is removed. That is, the model structure of the deepfake detection model in the application stage is the same as its model structure in the training stage. Among them, the deepfake detection result of the image to be detected includes the binary-classification result and the multi-classification result of the image to be detected. Optionally, the binary-classification result is used for business decision-making. For example, in an identity authentication scenario, if the binary-classification result indicates the existence of a deepfake attack, it is determined that the identity authentication fails. Optionally, the multi-classification result is used for data statistics. For example, in an identity authentication scenario, according to the multi-classification result, the most common types of deepfake attacks are statistically counted. Subsequently, the business can be improved according to this attack type, etc.

[0116] In the above technical solution, since the deepfake detection model adjusts the model parameters of the feature extraction network based on the output results of the binary-classification network and the multi-classification network during the training process to optimize the performance of the feature extraction network, it drives the feature extraction network to extract more fine-grained forgery features for deepfake attacks of different attack types, thereby improving the recognition ability and robustness of the overall model for deepfake attacks of various attack types.

[0117] Figure 6 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 6 , at the hardware level, the device includes a processor 602, an internal bus 604, a network interface 606, a memory 608, and a non-volatile memory 610. Of course, there may also be other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 602 reads the corresponding computer program from the non-volatile memory 610 into the memory 608 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.

[0118] Please refer to Figure 7 , the training device for a deepfake detection model applicable to multiple attack types can be applied to a device such as Figure 6 shown to implement the technical solution of this specification. Among them, the training device for a deepfake detection model applicable to multiple attack types may include:

[0119] A determination unit 701, configured to determine sample data and a deepfake detection model, where the deepfake detection model includes a feature extraction network, a binary-classification network, and a multi-classification network, the output of the feature extraction network is used as the input of the binary-classification network and the multi-classification network, and the sample data includes a sample image and its corresponding binary-classification label and multi-classification label;

[0120] The detection unit 702 is configured to input the sample image into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample image through the feature extraction network, performs binary classification on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification on the image features through the multi-classification network to obtain a multi-classification result;

[0121] The training unit 703 is configured to adjust the model parameters of the feature extraction network based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label;

[0122] Wherein, the binary classification label and the binary classification result are used to identify whether there is a deepfake attack, and the multi-classification label and the multi-classification result are used to identify the attack type of the deepfake attack.

[0123] Optionally, the training unit 703 is further configured to adjust the model parameters of the binary classification network based on the difference between the binary classification result and the binary classification label, and adjust the model parameters of the multi-classification network based on the difference between the multi-classification result and the multi-classification label.

[0124] Optionally, the feature extraction network includes a global feature extraction network, a local feature extraction network, and a fusion layer;

[0125] The detection unit 702 is configured to extract the global features of the sample image through the global feature extraction network; extract the local features of the sample image through the local feature extraction network; and fuse the global features and the local features through the fusion layer to obtain the image features of the sample image.

[0126] Optionally, the global feature extraction network includes a visual feature extraction network, and the visual feature extraction network is configured to extract global visual features of the input RGB image;

[0127] And / or,

[0128] The local feature extraction network includes at least one of the following:

[0129] A frequency domain feature extraction network, which is configured to traverse the input image through a sliding window to extract the frequency domain features of each local area of the input image;

[0130] A neighboring pixel relationship feature extraction network, which is configured to obtain the neighboring pixel relationship features of the input image based on the difference between each pixel point and its adjacent pixel points in the input image.

[0131] Optionally, the visual feature extraction network is trained based on a masking strategy;

[0132] and / or,

[0133] The detection unit 702 is configured to, for any local area selected by the sliding window, obtain the frequency response of the local area based on the pixel values of the local area, perform frequency band division on the frequency response of the local area, obtain the frequency responses of each frequency band, and determine the frequency domain features of the local area based on the frequency responses of each frequency band; wherein, the range and / or weight of each frequency band are optimized during the training process of the frequency domain feature extraction network.

[0134] Optionally, the apparatus further includes:

[0135] A construction unit, configured to generate sample images by using deepfake methods of multiple attack types to construct a sample training set covering multiple attack types, and the sample training set is used to train the deepfake detection model.

[0136] Please refer to Figure 8 , a deepfake detection apparatus applicable to multiple attack types can be applied to a device as shown in Figure 6 to implement the technical solutions of this specification. Among them, the deepfake detection apparatus applicable to multiple attack types may include:

[0137] A detection unit 801, configured to obtain an image to be detected, and input the image to be detected into a deepfake detection model to obtain a deepfake detection result of the image to be detected, and the deepfake detection model is trained based on the steps of the method described in the embodiment shown in Figure 2 or Figure 3 the embodiment shown.

[0138] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method described in any of the above embodiments.

[0139] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0140] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.

Claims

1. A training method for a deepfake detection model applicable to multiple attack types, the method comprising: Determine sample data and a deepfake detection model, the deepfake detection model including a feature extraction network, a binary classification network, and a multi-classification network, the output of the feature extraction network being used as the input of the binary classification network and the multi-classification network, and the sample data including sample images and their corresponding binary classification labels and multi-classification labels; Input the sample images into the deepfake detection model, so that the deepfake detection model extracts the image features of the sample images through the feature extraction network, performs binary classification processing on the image features through the binary classification network to obtain a binary classification result, and performs multi-classification processing on the image features through the multi-classification network to obtain a multi-classification result; Based on the difference between the binary classification result and the binary classification label and the difference between the multi-classification result and the multi-classification label, adjust the model parameters of the feature extraction network; Wherein, the binary classification label and the binary classification result are used to identify whether there is a deepfake attack, and the multi-classification label and the multi-classification result are used to identify the attack type of the deepfake attack.

2. The method according to claim 1, the method further comprising: Based on the difference between the binary classification result and the binary classification label, adjust the model parameters of the binary classification network, and based on the difference between the multi-classification result and the multi-classification label, adjust the model parameters of the multi-classification network.

3. The method according to claim 1, the feature extraction network including a global feature extraction network, a local feature extraction network, and a fusion layer; the extracting the image features of the sample images through the feature extraction network includes: Extract the global features of the sample images through the global feature extraction network; Extract the local features of the sample images through the local feature extraction network; Fuse the global features and the local features through the fusion layer to obtain the image features of the sample images.

4. The method according to claim 3, characterized in that, The global feature extraction network includes a visual feature extraction network, and the visual feature extraction network is used to extract global visual features of the input RGB images; And / or, The local feature extraction network includes at least one of the following: A frequency domain feature extraction network, which is used to traverse the input image through a sliding window to extract the frequency domain features of each local area of the input image; A neighboring pixel relationship feature extraction network, which is used to obtain the neighboring pixel relationship features of the input image based on the difference between each pixel point and its adjacent pixel points in the input image.

5. The method according to claim 4, the visual feature extraction network is trained based on a masking strategy; And / or, Traversing the input image through a sliding window to extract the frequency domain features of each local region of the input image, including: For any local region selected by the sliding window, obtain the frequency response of the local region based on the pixel values of the local region, perform frequency band division on the frequency response of the local region to obtain the frequency responses of each frequency band, and determine the frequency domain features of the local region based on the frequency responses of each frequency band; wherein, the ranges and / or weights of each frequency band are optimized during the training process of the frequency domain feature extraction network.

6. The method according to claim 1, wherein the method further comprises: Generating sample images using deepfake methods of multiple attack types to construct a sample training set covering multiple attack types, and the sample training set is used to train the deepfake detection model.

7. A deepfake detection method applicable to multiple attack types, the method comprising: Obtaining an image to be detected, and inputting the image to be detected into a deepfake detection model to obtain a deepfake detection result of the image to be detected, where the deepfake detection model is trained based on the steps of the method according to any one of claims 1 to 6.

8. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; wherein, the processor realizes the steps of the method according to any one of claims 1-7 by running the executable instructions.

9. A computer-readable storage medium, characterized in that, Stored thereon are computer instructions which, when executed by a processor, realize the steps of the method according to any one of claims 1-7.

10. A computer program product, characterized in that, Comprising a computer program / instructions which, when executed by a processor, realize the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Cross-modal depth forgery detection method based on time-frequency domain visual artifact feature adaptive fusion

    CN114898438A

  • Deep fake face image detection method and system based on attribute guidance

    CN117173761A

  • Generalized deep forgery detection method based on multi-classification guidance

    CN119445679A