Deep pseudo detection method and device suitable for unknown attack type, medium and product
By introducing multiple classified networks with different network architectures into the deep pseudo detection model, the problem that existing tools are difficult to identify unknown attack types Deepfake is solved, and more comprehensive and efficient deep pseudo detection is achieved, improving the robustness of the model and the ability to identify new risks.
Patent Information
- Application Number
- CN202510713511.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-18
AI Technical Summary
Existing deep pseudo detection tools are difficult to effectively resist unknown attack types Deepfake and cannot provide ideal defense effects.
In the deep pseudo detection model, multiple classification networks with different network architectures are introduced, image features are obtained through feature extraction networks, and multiple classification networks are used to classify from different angles, and the deep pseudo detection results are determined by combining the output results of multiple classification networks.
The deep pseudo detection model's detection ability of deep pseudo attacks of unknown attack types is improved, and the model's robustness and identification ability of new risks are enhanced.
Smart Images

Figure CN120339723A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of information security technology, and in particular to a deepfake detection method, device, medium, and product applicable to unknown attack types. Background Art
[0002] Deepfake is a technology for synthesizing or tampering with image, video, and audio content based on AI (Artificial Intelligence) technology. It can replace a person's facial expression, voice, or movement into the image or video of another person, generating highly realistic fake content, and even achieving the effect of being indistinguishable from the real thing, playing an important role in the fields of film and television drama, entertainment, education, etc.
[0003] However, with the development of Deepfake technology, the risks it brings cannot be ignored. For example, criminals use Deepfake technology to impersonate corporate executives and instruct employees to transfer funds through video conferencing, resulting in huge financial losses. At present, there are already some deepfake detection tools and technologies, but in the face of the continuous evolution of Deepfake technology, these tools are still unable to cope with Deepfake of unknown attack types and fail to provide an ideal defense effect. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide the following technical solutions:
[0005] According to the first aspect of one or more embodiments of this specification, a deepfake detection method applicable to unknown attack types is proposed. The method includes:
[0006] In response to receiving an image to be detected, a deepfake detection model is called. The deepfake detection model includes a feature extraction network and multiple classification networks. The output of the feature extraction network is used as the input of the multiple classification networks, and the network architectures of the multiple classification networks are different;
[0007] The image to be detected is input into the deepfake detection model, so that the deepfake detection model extracts features of the image to be detected through the feature extraction network to obtain image features of the image to be detected, and classifies the image features through the multiple classification networks respectively to obtain the forgery probabilities respectively determined by the multiple classification networks. The forgery probability is used to represent the probability that the image to be detected is forged by deepfake technology;
[0008] Statistical processing is performed on the forgery probabilities respectively determined by the multiple classification networks, and the deepfake detection result of the image to be detected is determined based on the statistical processing result.
[0009] According to a second aspect of one or more embodiments of the present specification, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method described in the first aspect above.
[0010] According to a third aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in the first aspect above are implemented.
[0011] According to a fourth aspect of one or more embodiments of the present specification, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in the first aspect above are implemented.
[0012] As can be seen from the above embodiments, multiple classification networks with different network architectures are added to the deepfake detection model in the present specification. The output of the feature extraction network is used as the input of multiple classification networks at the same time. Then, the deepfake detection result is determined by synthesizing the output results of multiple classification networks. Since classification networks with different network architectures can learn and identify image features from different perspectives, adding multiple classification networks with different network architectures to the deepfake detection model can achieve learning and identifying image features from multiple perspectives, enabling the deepfake detection model to more comprehensively identify various forgery clues, thereby effectively improving the detection ability of the deepfake detection model for new risks. That is, it improves the detection ability of the deepfake detection model for deepfake attacks of unknown attack types and improves the robustness of the deepfake detection model. Description of the Drawings
[0013] Figure 1 is a schematic diagram of the architecture of a deepfake detection service system provided by an exemplary embodiment.
[0014] Figure 2 is a flowchart of a deepfake detection method applicable to unknown attack types provided by an exemplary embodiment.
[0015] Figure 3 is a flowchart of a deepfake detection method applicable to unknown attack types based on multi-modal feature extraction provided by an exemplary embodiment.
[0016] Figure 4 is a schematic diagram of the structure of a deepfake detection model proposed by an exemplary embodiment.
[0017] Figure 5 is a schematic diagram of the structure of a device provided by an exemplary embodiment.
[0018] Figure 6It is a block diagram of a deepfake detection device applicable to unknown attack types provided by an exemplary embodiment. Detailed implementation
[0019] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0020] Deepfake is a technology for synthesizing or tampering with image, video, and audio content based on AI (Artificial Intelligence) technology. It can replace a person's facial expression, voice, or movement into the image or video of another person, generating highly realistic false content, even achieving the effect of being indistinguishable from the real thing, and playing an important role in fields such as film and television dramas, entertainment, and education. For example, in a film and television drama, replacing an actor's face with the face of the actor when he was young to improve the viewing effect of the film and television drama; another example is applying a student's face or voice to the video of a historical figure so that students can more vividly experience historical events, etc.
[0021] However, with the development of Deepfake technology, the risks it brings cannot be ignored. For example, criminals use Deepfake technology to impersonate corporate executives and instruct employees to transfer funds through video conferencing, resulting in huge financial losses; another example is using Deepfake technology to create false news, causing damage to the reputation of others, etc.
[0022] Currently, there are already some deepfake detection tools and technologies. However, most of them can only have a good recognition effect on deepfake attacks of known attack types. Facing the continuous evolution of Deepfake technology, these tools are unable to cope with deepfake attacks of unknown attack types and fail to provide an ideal defense effect.
[0023] Based on this, this specification provides a deepfake detection method applicable to unknown attack types. By setting multiple classification networks with different network architectures in the deepfake detection model, the deepfake detection model can learn and identify image features from more perspectives, improving the detection ability of the deepfake detection model for new risks. That is to say, it improves the detection ability of the deepfake detection model for deepfake attacks of unknown attack types.
[0024] In implementation, first, in response to receiving an image to be detected, a deepfake detection model is called. The deepfake detection model includes a feature extraction network and multiple classification networks. The output of the feature extraction network serves as the input to the multiple classification networks, and the network architectures of the multiple classification networks are different. Then, the image to be detected is input into the deepfake detection model, so that the deepfake detection model extracts features of the image to be detected through the feature extraction network to obtain image features of the image to be detected, and classifies the image features through the multiple classification networks respectively to obtain forgery probabilities respectively determined by the multiple classification networks. The forgery probability is used to represent the probability that the image to be detected is forged by deepfake technology. Finally, the forgery probabilities respectively determined by the multiple classification networks are statistically processed, and the deepfake detection result of the image to be detected is determined based on the statistical processing result.
[0025] In the above technical solution, multiple classification networks with different network architectures are added to the deepfake detection model. The output of the feature extraction network is used as the input to the multiple classification networks at the same time. Then, the deepfake detection result is determined by synthesizing the output results of the multiple classification networks. Since the classification networks with different network architectures can learn and identify image features from different perspectives, adding multiple classification networks with different network architectures to the deepfake detection model can achieve learning and identifying image features from multiple perspectives, enabling the deepfake detection model to more comprehensively identify various forgery clues, thereby effectively improving the detection ability of the deepfake detection model for new risks. That is, it improves the detection ability of the deepfake detection model for deepfake attacks of unknown attack types and improves the robustness of the deepfake detection model.
[0026] The deepfake detection method provided in this specification applicable to unknown attack types can be widely applied to any scenario that requires deepfake detection. Next, this specification gives an exemplary illustration of the application scenarios of this deepfake detection method:
[0027] For example: applied to the face recognition scenario:
[0028] In multiple scenarios such as face payment, account registration, account login, or identity verification, subsequent operations can only be performed after face recognition is passed. If the deepfake detection method provided in this specification applicable to unknown attack types is adopted, when a malicious user attempts to deceive using a face image forged by deepfake technology, it can be detected with a higher probability, thereby enhancing the security and reliability of the system.
[0029] Another example: applied to the document verification scenario:
[0030] In multiple scenarios such as account registration and identity verification, subsequent operations can only be carried out after successful document recognition. If the deepfake detection method provided in this specification, which is applicable to unknown attack types, is adopted, when malicious users attempt to deceive using document images forged by deepfake technology, it can be detected with a higher probability, thereby enhancing the security and reliability of the system.
[0031] It should be noted that this specification only provides an exemplary description of the application scenarios of the deepfake detection method and does not limit the application scenarios.
[0032] Figure 1 It is a schematic diagram of the architecture of a deepfake detection service system provided by an exemplary embodiment. As Figure 1 shown, the system may include a server 11, a network 12, and several electronic devices, such as a PC (Personal Computer) 13, a mobile phone 14, etc.
[0033] The server 11 may be a physical server including an independent host, or the server 11 may be a virtual server hosted by a host cluster. During operation, the server 11 may run the server-side program of an application to implement the related functions of the application. For example, when the server 11 runs the program of the deepfake detection service, it can be implemented as a corresponding deepfake detection service platform.
[0034] The PC 13 and the mobile phone 14 are only some types of electronic devices that users can use. In fact, users can obviously also use electronic devices of the following types: tablet devices, laptop computers, personal digital assistants (PDAs), wearable devices (such as smart glasses, smart watches, etc.). One or more embodiments of this specification do not limit this. During operation, the electronic device may run the client-side program of an application to implement the related functions of the application. For example, when the electronic device runs the program of the deepfake detection service, it can be implemented as the client of the deepfake detection service. Among them, the application program of the client of the above deepfake detection service can be started and run on the electronic device. The client-side program may be a native application program installed on the electronic device, or the client-side program may be a small program, a quick application, or other similar forms. Of course, when using web technologies such as HTML5 or similar, the related functions can be implemented through the page displayed by the browser. Here, the browser may be an independent browser application or a browser module embedded in some applications.
[0035] For the network 12 for interaction between electronic devices such as the PC 13 and the mobile phone 14 and the server 11, communication can be specifically implemented using a wired or wireless network based on the communication methods supported by the corresponding electronic devices, and this specification does not limit this. For example, if the PC 13 supports both wired and wireless communication, then wired or wireless network can be used for communication according to needs, while the mobile phone 14 usually only supports wireless communication, so wireless network can be used for communication.
[0036] Exemplarily, the deepfake detection method applicable to unknown attack types provided in the embodiments of this specification can be executed by any one of the above-mentioned client or the above-mentioned server. For example, in response to receiving a to-be-detected image, the terminal invokes the deepfake detection model, inputs the to-be-detected image into the deepfake detection model, so that the deepfake detection model outputs multiple forgery probabilities, performs statistical processing on the multiple forgery probabilities, and determines the deepfake detection result of the to-be-detected image based on the statistical processing result. Another example is that in response to receiving a to-be-detected image, the server invokes the deepfake detection model, inputs the to-be-detected image into the deepfake detection model, so that the deepfake detection model outputs multiple forgery probabilities, performs statistical processing on the multiple forgery probabilities, and determines the deepfake detection result of the to-be-detected image based on the statistical processing result.
[0037] Exemplarily, the deepfake detection method applicable to unknown attack types provided in the embodiments of this specification can be completed by the mutual cooperation of the above-mentioned client and the above-mentioned server. For example, the client sends the to-be-detected image to the server, and in response to receiving the to-be-detected image, the server invokes the deepfake detection model, inputs the to-be-detected image into the deepfake detection model, so that the deepfake detection model outputs multiple forgery probabilities, performs statistical processing on the multiple forgery probabilities, and determines the deepfake detection result of the to-be-detected image based on the statistical processing result.
[0038] It should be noted that this specification only gives an exemplary description of the cooperation method between the client and the server, and does not limit it. The specific cooperation method can be set arbitrarily according to actual needs.
[0039] Next, this specification gives an exemplary description of the implementation process of the deepfake detection method:
[0040] Figure 2 is a flowchart of a deepfake detection method applicable to unknown attack types provided by an exemplary embodiment. The execution subject of this method can be Figure 1 the system shown, or can also be Figure 1 any component unit in the system shown. This specification does not limit the execution subject. This method includes:
[0041] S201. In response to receiving the image to be detected, call the deepfake detection model, which includes a feature extraction network and multiple classification networks. The output of the feature extraction network serves as the input to the multiple classification networks, and the network architectures of the multiple classification networks are different.
[0042] In one embodiment, the image to be detected can be any image, such as a face image, a document image, a landscape image, etc. This specification does not limit this. In another embodiment, the image to be detected is a face image. The deepfake detection method provided in this specification is a face deepfake detection method. This specification does not limit the application object of the deepfake detection method.
[0043] In one embodiment, the image to be detected can be an image captured by a local device, an image stored locally, or an image obtained from other devices. This specification does not limit the source of the image to be detected.
[0044] The deepfake detection model is a model used to detect whether the input image is an image forged by deepfake technology. In this specification, the deepfake detection model includes a feature extraction network and multiple classification networks. The output of the feature extraction network serves as the input to the multiple classification networks (that is, the multiple classification networks share the feature extraction network), and the network architectures of the multiple classification networks are different. Since the network architectures of the multiple classification networks are different, the multiple classification networks can learn and identify image features from different perspectives, enabling the deepfake detection model to more comprehensively identify various forgery clues. Since the multiple classification networks share the feature extraction network, during the training phase of the deepfake detection model, the multiple classification networks can drive the feature extraction network to extract image features from different perspectives, further enabling the deepfake detection model to more comprehensively identify various forgery clues.
[0045] Among them, the feature extraction network can be a feature extraction network with any network architecture. This specification does not limit the network architecture of the feature extraction network. In one embodiment, the feature extraction network can adopt a single-branch architecture, that is, extract the image features of the input image through a series of consecutive layers (such as convolutional layers, pooling layers, etc.). In another embodiment, the feature extraction network can also adopt a multi-branch architecture, extracting features from different scales or angles through multiple branches, thereby enhancing the expression ability of the deepfake detection model for the input image.
[0046] It should be noted that this specification does not limit the number of classification networks in the deepfake detection model. Exemplarily, the number of multiple classification networks can be 2, 3, 4, etc. In one embodiment, the number of multiple classification networks can be set according to actual needs. In scenarios with lower latency requirements, more classification networks can be adopted to improve the accuracy of deepfake detection results; in scenarios with higher latency requirements, fewer classification networks can be adopted to improve the running efficiency of the deepfake detection model.
[0047] Another point to note is that this specification does not limit the network architecture of the classification networks in the deepfake detection model. For example, multiple classification networks can be ConvNeXt (Convolutional Neural Network Next) network, RepLKNet (Representative Large Kernel Network) network, ResNet (Residual Network), SENet (Squeeze-and-Excitation Networks), etc. In one embodiment, multiple classification networks include ConvNeXt network and RepLKNet network.
[0048] In one embodiment, considering that increasing the number of classification networks can improve the accuracy of the deepfake detection model, but at the same time will also slow down the running speed of the model, so a balance needs to be found between the two when designing the deepfake detection model. In this way, both the high accuracy of the deepfake detection results can be ensured and a reasonable processing efficiency can be maintained. Through a large number of experiments, it is found that when the number of multiple classification networks is 2, and one classification network is ConvNeXt network and the other classification network is RepLKNet network, the deepfake detection model can achieve good results in terms of both accuracy and running efficiency.
[0049] S202, Input the image to be detected into the deepfake detection model, so that the deepfake detection model extracts the image features of the image to be detected through the feature extraction network, and classifies the image features through multiple classification networks respectively to obtain the forgery probabilities determined by the multiple classification networks respectively. This forgery probability is used to represent the probability that the image to be detected is forged through deepfake technology.
[0050] S203, Perform statistical processing on the forgery probabilities determined by the multiple classification networks respectively, and determine the deepfake detection result of the image to be detected based on the statistical processing result.
[0051] In one embodiment, the statistical processing result can be directly used as the deepfake detection result. Among them, statistical processing is performed on the forgery probabilities respectively determined by multiple classification networks, and the deepfake detection result of the image to be detected is determined based on the statistical processing result, including: performing statistical processing on the forgery probabilities respectively determined by multiple classification networks, and using the statistical processing result as the deepfake detection result of the image to be detected.
[0052] Optionally, performing statistical processing on the forgery probabilities respectively determined by multiple classification networks, and using the statistical processing result as the deepfake detection result of the image to be detected, including: performing Bayesian fusion on the forgery probabilities respectively determined by multiple classification networks, and using the fusion result as the deepfake detection result of the image to be detected. Exemplarily, the Bayesian formula is used to calculate the forgery probabilities respectively determined by multiple classification networks to obtain the fusion result.
[0053] Optionally, performing statistical processing on the forgery probabilities respectively determined by multiple classification networks, and using the statistical processing result as the deepfake detection result of the image to be detected, including: performing majority voting based on the forgery probabilities respectively determined by multiple classification networks to obtain the deepfake detection result of the image to be detected. Exemplarily, if the forgery probability output by a classification network indicates that the image to be detected is forged by deepfake technology, then vote for the "forged" result; if the forgery probability output by a classification network indicates that the image to be detected is not forged by deepfake technology, then vote for the "real" result; after voting according to the forgery probabilities respectively determined by each classification network, count the number of votes for the "forged" result and the number of votes for the "real" result, and the result with more votes is the deepfake detection result of the deepfake detection model.
[0054] Optionally, performing statistical processing on the forgery probabilities respectively determined by multiple classification networks, and using the statistical processing result as the deepfake detection result of the image to be detected, including: determining the maximum forgery probability determined by multiple classification networks, and using the maximum forgery probability as the deepfake detection result of the image to be detected. In this way, the sensitivity of the deepfake detection model can be improved, and the impact of forged images escaping the deepfake detection model on subsequent processes can be reduced. Moreover, different deepfake technologies will leave different types of traces in the image, and multiple classification networks focus on forgery features from different scales or angles. The selection of the maximum forgery probability can integrate the advantages of these classification networks and adapt to diverse deepfake technology means.
[0055] Optionally, statistical processing is performed on the forgery probabilities respectively determined by multiple classification networks, and the statistical processing result is used as the deepfake detection result of the image to be detected, including: averaging the forgery probabilities respectively determined by multiple classification networks to obtain the deepfake detection result of the image to be detected. A single classification network may have certain biases or limitations. By averaging the forgery probabilities of multiple classification networks, the error influence of individual classification networks can be effectively reduced, thereby improving the accuracy of the final deepfake detection result.
[0056] Optionally, statistical processing is performed on the forgery probabilities respectively determined by multiple classification networks, and the statistical processing result is used as the deepfake detection result of the image to be detected, including: performing weighted averaging on the forgery probabilities respectively determined by multiple classification networks to obtain the deepfake detection result of the image to be detected. In this way, the advantages of each classification network can be utilized more flexibly, improving the accuracy of the deepfake detection result.
[0057] It should be noted that this specification only exemplarily illustrates the statistical processing methods for multiple forgery probabilities and does not limit them. In actual applications, flexible settings can be made according to the requirements of the application scenario, the characteristics of multiple classification networks, etc. Another point to note is that during the training process of the deepfake detection model, the average value of the forgery probabilities respectively determined by multiple classification networks can be used as the deepfake detection result of the deepfake detection model. According to the difference between this deepfake detection result and the sample detection result, the model parameters of the deepfake detection model are adjusted to enable the deepfake detection model to converge more quickly; while during the application process of the deepfake detection model, any statistical processing method provided in this specification can be used to obtain a more accurate deepfake detection result.
[0058] In another embodiment, considering that when the forgery probabilities determined by multiple classification networks vary greatly, it indicates that there are differences in the recognition of the image to be detected by different classification networks. If the deepfake detection result is obtained through statistical processing, the obtained deepfake detection result may not necessarily be correct. To further improve the accuracy rate of the deepfake detection result, in this specification, when the differences among multiple classification networks are small, the statistical processing method is used to obtain the deepfake detection result; when the differences among multiple classification networks are large, the data augmentation mechanism is triggered to obtain more data to be detected, and then a more accurate deepfake detection result is obtained.
[0059] Among them, statistical processing is performed on the forgery probabilities respectively determined by multiple classification networks, and the deepfake detection result of the image to be detected is determined based on the statistical processing result, including: determining the difference between the maximum forgery probability and the minimum forgery probability determined by the multiple classification networks; if the difference is less than the target value, statistical processing is performed on the forgery probabilities respectively determined by the multiple classification networks, and the statistical processing result is used as the deepfake detection result of the image to be detected; if the difference is not less than the target value, a data enhancement mechanism is triggered to obtain additional data to be detected, and the deepfake detection result is determined based on the additional data to be detected. Among them, the data enhancement mechanism is a mechanism for obtaining additional data to be detected and determining the deepfake detection result based on the additional data to be detected.
[0060] Among them, the target value can be any value, for example, 30%, 50%, etc. In one embodiment, the target value can be an empirical value, a value set by a technician, or a value set according to the requirements of the actual application scenario. This specification does not limit the target value.
[0061] Among them, the additional data to be detected can be any type of data. This specification does not limit the additional data to be detected and can be set according to actual needs. Optionally, obtaining additional data to be detected and determining the deepfake detection result based on the additional data to be detected includes: when the image to be detected is any type of physiological feature data, obtaining other types of physiological feature data and determining the deepfake detection result based on the other types of physiological feature data. Exemplarily, when the data to be detected is a face image, the additional data to be detected can be other types of physiological feature data such as iris images and fingerprint images.
[0062] Optionally, obtaining additional data to be detected and determining the deepfake detection result based on the additional data to be detected includes: obtaining the device status data during the process of the device for collecting the image to be detected collecting the image to be detected, and determining the deepfake detection result based on the device status data. Among them, the device status data can be the sensor data collected by the pose sensor, motion sensor, or gravity sensor in the device. When the device collects the image to be detected, the device is usually in a motion state; based on this sensor data, it can be determined whether the device is in an operating state during the process of collecting the image to be detected, and further determine whether the image to be detected is a real image captured by the device or a fake image forged by deepfake technology. Of course, the device status data can also be other types of data, such as the running status of a specified program in the device. This specification does not limit the device status data.
[0063] It should be noted that after obtaining additional data to be detected, the deepfake detection result can be obtained only based on the additional data to be detected, or the deepfake detection result can be obtained by combining the image to be detected and the additional data to be detected. This specification does not limit this. Exemplarily, the image to be detected is any video frame image in the video captured by the device, and the additional data to be detected is the sensor data during the process of the device capturing the video. It can be determined whether the image to be detected is a real image captured by the device or a fake image forged by deepfake technology according to whether the running trajectory indicated by the sensor data matches the movement trajectory of the object captured in the video.
[0064] In the above technical solution, multiple classification networks with different network architectures are added to the deepfake detection model. The output of the feature extraction network is used as the input of the multiple classification networks at the same time. Then, the deepfake detection result is determined by integrating the output results of the multiple classification networks. Since the classification networks with different network architectures can learn and identify image features from different perspectives, adding multiple classification networks with different network architectures to the deepfake detection model can realize learning and identifying image features from multiple perspectives, so that the deepfake detection model can more comprehensively identify various forgery clues, thereby effectively improving the detection ability of the deepfake detection model for new risks. That is, it improves the detection ability of the deepfake detection model for deepfake attacks of unknown attack types and improves the robustness of the deepfake detection model.
[0065] Next, this specification takes the feature extraction network to perform multi-modal feature extraction on the image to be detected as an example to exemplarily illustrate the process of the deepfake detection method. Figure 3 It is a flowchart of a deepfake detection method based on multi-modal feature extraction applicable to unknown attack types provided by an exemplary embodiment. The execution subject of this method can be Figure 1 the system shown, or Figure 1 any component unit in the system shown. This specification does not limit the execution subject. This method includes:
[0066] S301, in response to receiving the image to be detected, call the deepfake detection model, which includes a feature extraction network and multiple classification networks. The output of the feature extraction network is used as the input of the multiple classification networks, and the network architectures of the multiple classification networks are different.
[0067] The above step S301 is the same as the above step S201. Refer to the above step S201 and will not be elaborated here one by one.
[0068] S302. Input the image to be detected into the deepfake detection model, so that the deepfake detection model extracts multi-modal features of the image to be detected through the feature extraction network, obtains the image sub-features of the image to be detected in different modalities, fuses the image sub-features of the image to be detected in different modalities to obtain the image features of the image to be detected, and classifies the image features through multiple classification networks respectively to obtain the forgery probabilities determined by the multiple classification networks respectively. The forgery probability is used to represent the probability that the image to be detected is forged by deepfake technology.
[0069] "Modality" refers to different forms or expressions of image information. Each modality describes the characteristics of the image from a different perspective. These modalities can be based on the original data of the image (such as the pixel domain), data after mathematical transformation (such as the frequency domain, noise domain), or other abstract-level feature representations (such as semantic features, gradient features). By analyzing the performance of the image to be detected in different modalities, more diverse and complementary image sub-features can be extracted, thereby improving the performance of the deepfake detection model.
[0070] In one embodiment, extracting multi-modal features of the image to be detected through the feature extraction network to obtain the image sub-features of the image to be detected in different modalities includes at least one of the following: (1) Extracting features from the pixel domain of the image to be detected through the feature extraction network to obtain the pixel domain image sub-features of the image to be detected; (2) Performing noise transformation on the image to be detected through the feature extraction network to obtain the noise domain representation of the image to be detected, and extracting features from the noise domain representation to obtain the noise domain image sub-features of the image to be detected; (3) Performing frequency domain transformation on the image to be detected through the feature extraction network to obtain the frequency domain representation of the image to be detected, and extracting features from the frequency domain representation to obtain the frequency domain image sub-features of the image to be detected.
[0071] Among them, when performing noise transformation on the image to be detected, any noise transformation method can be used. For example, noise separation methods based on statistical models (such as Gaussian noise modeling, Poisson noise modeling, mixed noise modeling, etc.), filtering-based methods (such as Gaussian filtering, median filtering, bilateral filtering, etc.), deep learning-based methods (such as autoencoders, noise prediction networks, etc.), etc. This specification does not limit this. Exemplarily, the filtering layer of SRM (steganalysis rich model) is used as the noise transformer, and the image to be detected is subjected to noise transformation through this noise transformer to obtain the noise domain representation of the image to be detected.
[0072] When performing frequency domain transformation on the image to be detected, any frequency domain transformation method can be used. For example, Fourier transform, discrete cosine transform, Laplace transform, etc. This specification does not limit this.
[0073] In one embodiment, the feature extraction network includes multiple branches, and different branches are used to extract image sub-features of the input image in different modalities.
[0074] Exemplarily, the feature extraction network includes a pixel domain branch and a noise domain branch. The pixel domain branch extracts features from the original image to be detected to obtain the pixel domain image sub-features of the image to be detected; the noise domain branch performs noise transformation on the image to be detected to obtain the noise domain representation of the image to be detected, and then extracts features from the noise domain representation to obtain the noise domain image sub-features of the image to be detected. Then, the pixel domain image sub-features and the noise domain image sub-features are fused (such as concatenation, averaging, weighted averaging, extraction by a feature extraction layer, etc.) to obtain the image features of the image to be detected.
[0075] Exemplarily, the feature extraction network includes a pixel domain branch, a noise domain branch, and a frequency domain branch. The pixel domain branch extracts features from the original image to be detected to obtain the pixel domain image sub-features of the image to be detected; the noise domain branch performs noise transformation on the image to be detected to obtain the noise domain representation of the image to be detected, and then extracts features from the noise domain representation to obtain the noise domain image sub-features of the image to be detected; the frequency domain branch performs frequency domain transformation on the image to be detected to obtain the frequency domain representation of the image to be detected, and then extracts features from the frequency domain representation to obtain the frequency domain image sub-features of the image to be detected. Then, the pixel domain image sub-features, the noise domain image sub-features, and the frequency domain image sub-features are fused (such as concatenation, averaging, weighted averaging, extraction by a feature extraction layer, etc.) to obtain the image features of the image to be detected.
[0076] In another embodiment, the feature extraction network includes a single branch, which obtains the image representations of the image in different modalities and extracts the image sub-features of the image in different modalities through this single branch.
[0077] Exemplarily, perform noise transformation on the image to be detected to obtain the noise domain representation of the image to be detected, concatenate the pixel domain representation of the image to be detected (i.e., the original image to be detected) with the noise domain representation, and then extract features from the concatenated image representation through the feature extraction network to obtain the image features of the image to be detected. These image features are the image features of the image to be detected in two modalities, namely the pixel domain and the noise domain.
[0078] Exemplarily, perform noise transformation on the image to be detected to obtain the noise domain representation of the image to be detected; perform frequency domain transformation on the image to be detected to obtain the frequency domain representation of the image to be detected; concatenate the pixel domain representation, the noise domain representation, and the frequency domain representation of the image to be detected, and then extract features from the concatenated image representation through the feature extraction network to obtain the image features of the image to be detected. These image features are the image features of the image to be detected in three modalities, namely the pixel domain, the noise domain, and the frequency domain.
[0079] In S303, statistical processing is performed on the forgery probabilities respectively determined by multiple classification networks, and a deepfake detection result of the image to be detected is determined based on the result of the statistical processing.
[0080] The above step S303 is the same as the above step S203, and reference can be made to the above step S203, which will not be elaborated here one by one.
[0081] It should be noted that Figure 4 A deepfake detection model is shown. The deepfake detection model obtains a pixel-domain representation and a noise-domain representation of the image to be detected, splices the pixel-domain representation and the noise-domain representation and inputs them into a Multi-domain Adapter for information fusion, outputs a three-channel tensor, and then inputs the tensor into two classifiers based on the ConvNeXt network and the RepLKNet network to obtain the forgery probabilities predicted by the two classifiers respectively, and determines the deepfake detection result based on these two forgery probabilities. Through experiments, it is found that the deepfake detection model can achieve a detection rate of 98% in the test set of known attack types and also 90% in the test set of unknown attack types. It can be seen that the deepfake detection model has a high detection ability for deepfake attacks of unknown attack types. Moreover, the deepfake detection model only contains two classification networks, enabling the deepfake detection model to maintain a high operating efficiency while having high robustness.
[0082] In the above technical solution, multiple classification networks with different network architectures are added to the deepfake detection model, and the output of the feature extraction network is used as the input of multiple classification networks at the same time. Then, the deepfake detection result is determined by comprehensively considering the output results of multiple classification networks. Since classification networks with different network architectures can learn and identify image features from different perspectives, adding multiple classification networks with different network architectures to the deepfake detection model can realize learning and identifying image features from multiple perspectives, enabling the deepfake detection model to more comprehensively identify various forgery clues, thereby effectively improving the detection ability of the deepfake detection model for new risks, that is, improving the detection ability of the deepfake detection model for deepfake attacks of unknown attack types and improving the robustness of the deepfake detection model.
[0083] Moreover, the deepfake detection model can also extract pixel-domain image sub-features and noise-domain image sub-features of the input image and fuse them, enabling the deepfake detection model to learn different features from different modalities of the image, enabling the deepfake detection model to more comprehensively understand the authenticity of the image, improving the recognition ability of the deepfake detection model for various deepfake technologies, and improving the generalization of the deepfake detection model.
[0084] Figure 5It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 5 , at the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in a software manner. For example, the processor 502 reads the corresponding computer program from the non-volatile memory 510 into the memory 508 and then runs it. Of course, in addition to the software implementation method, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or a logical device.
[0085] Please refer to Figure 6 , the deepfake detection device applicable to unknown attack types can be applied to a device as shown in Figure 5 to implement the technical solutions of this specification. Among them, the deepfake detection device applicable to unknown attack types may include:
[0086] A calling unit 601, configured to call a deepfake detection model in response to receiving an image to be detected. The deepfake detection model includes a feature extraction network and multiple classification networks. The output of the feature extraction network is used as the input of the multiple classification networks, and the network architectures of the multiple classification networks are different;
[0087] A detection unit 602, configured to input the image to be detected into the deepfake detection model, so that the deepfake detection model extracts features of the image to be detected through the feature extraction network to obtain the image features of the image to be detected, and classifies the image features through the multiple classification networks respectively to obtain the forgery probabilities respectively determined by the multiple classification networks. The forgery probability is used to represent the probability that the image to be detected is forged by deepfake technology;
[0088] A statistical unit 603, configured to perform statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and determine the deepfake detection result of the image to be detected based on the statistical processing result.
[0089] Optionally, the statistical unit 603 is configured to perform statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and use the statistical processing result as the deepfake detection result of the image to be detected; or,
[0090] A statistical unit 603 is configured to determine the difference between the maximum forgery probability and the minimum forgery probability determined by the multiple classification networks. If the difference is less than a target value, statistical processing is performed on the forgery probabilities respectively determined by the multiple classification networks, and the result of the statistical processing is used as the deepfake detection result of the image to be detected. If the difference is not less than the target value, a data augmentation mechanism is triggered to obtain additional data to be detected, and the deepfake detection result is determined based on the additional data to be detected.
[0091] Optionally, the statistical unit 603 is configured to, when the image to be detected is any type of physiological feature data, obtain physiological feature data of other types, and determine the deepfake detection result based on the physiological feature data of other types; or,
[0092] The statistical unit 603 is configured to obtain device status data of the device used to collect the image to be detected during the process of collecting the image to be detected, and determine the deepfake detection result based on the device status data.
[0093] Optionally, the statistical unit 603 is configured to perform Bayesian fusion on the forgery probabilities respectively determined by the multiple classification networks, and use the fusion result as the deepfake detection result of the image to be detected; or,
[0094] The statistical unit 603 is configured to perform a majority vote based on the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected; or,
[0095] The statistical unit 603 is configured to determine the maximum forgery probability determined by the multiple classification networks, and use the maximum forgery probability as the deepfake detection result of the image to be detected; or,
[0096] The statistical unit 603 is configured to average the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected; or,
[0097] The statistical unit 603 is configured to perform a weighted average on the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected.
[0098] Optionally, the detection unit 602 is configured to perform multi-modal feature extraction on the image to be detected through the feature extraction network to obtain image sub-features of the image to be detected in different modalities; and fuse the image sub-features of the image to be detected in different modalities to obtain the image feature of the image to be detected.
[0099] Optionally, the detection unit 602 is configured to perform at least one of the following:
[0100] Feature extraction is performed on the pixel domain of the image to be detected through the feature extraction network, and pixel domain image sub-features of the image to be detected are obtained;
[0101] The image to be detected is subjected to noise transformation through the feature extraction network to obtain a noise domain representation of the image to be detected, and feature extraction is performed on the noise domain representation to obtain noise domain image sub-features of the image to be detected;
[0102] The image to be detected is subjected to frequency domain transformation through the feature extraction network to obtain a frequency domain representation of the image to be detected, and feature extraction is performed on the frequency domain representation to obtain frequency domain image sub-features of the image to be detected.
[0103] Optionally, the multiple classification networks include a ConvNeXt network and a RepLKNet network.
[0104] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor runs the executable instructions to implement the steps of the method as described in any one of the above embodiments.
[0105] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.
[0106] Based on the same concept as the above method, this specification also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method as described in any one of the above embodiments are implemented.
Claims
1. A deepfake detection method applicable to unknown attack types, the method comprising: In response to receiving an image to be detected, calling a deepfake detection model, the deepfake detection model including a feature extraction network and multiple classification networks, the output of the feature extraction network serving as the input of the multiple classification networks, and the network architectures of the multiple classification networks being different; Inputting the image to be detected into the deepfake detection model, so that the deepfake detection model extracts features of the image to be detected through the feature extraction network to obtain image features of the image to be detected, and classifying the image features through the multiple classification networks respectively to obtain forgery probabilities respectively determined by the multiple classification networks, the forgery probability being used to represent the probability that the image to be detected is forged by deepfake technology; Performing statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and determining the deepfake detection result of the image to be detected based on the statistical processing result.
2. The method according to claim 1, wherein the performing statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and determining the deepfake detection result of the image to be detected based on the statistical processing result includes: Performing statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and taking the statistical processing result as the deepfake detection result of the image to be detected; Or, Determining the difference between the maximum forgery probability and the minimum forgery probability determined by the multiple classification networks, if the difference is less than a target value, performing statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and taking the statistical processing result as the deepfake detection result of the image to be detected; If the difference is not less than the target value, triggering a data augmentation mechanism to obtain additional data to be detected, and determining the deepfake detection result based on the additional data to be detected.
3. The method according to claim 2, wherein the obtaining additional data to be detected and determining the deepfake detection result based on the additional data to be detected includes: When the image to be detected is physiological feature data of any type, obtaining physiological feature data of other types, and determining the deepfake detection result based on the physiological feature data of other types; or, Obtaining device status data of the device during the process of collecting the image to be detected when collecting the image to be detected, and determining the deepfake detection result based on the device status data.
4. The method according to claim 2, wherein the performing statistical processing on the forgery probabilities respectively determined by the multiple classification networks, and taking the statistical processing result as the deepfake detection result of the image to be detected includes: Performing Bayesian fusion on the forgery probabilities respectively determined by the multiple classification networks, and taking the fusion result as the deepfake detection result of the image to be detected; Or, Performing majority voting based on the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected; or, Determining the maximum forgery probability determined by the multiple classification networks, and taking the maximum forgery probability as the deepfake detection result of the image to be detected; or, Average the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected; or, Perform a weighted average on the forgery probabilities respectively determined by the multiple classification networks to obtain the deepfake detection result of the image to be detected.
5. The method according to claim 1, wherein the step of extracting the image features of the image to be detected by the feature extraction network to obtain the image features of the image to be detected includes: Performing multi-modal feature extraction on the image to be detected by the feature extraction network to obtain image sub-features of the image to be detected in different modalities; Fusing the image sub-features of the image to be detected in different modalities to obtain the image features of the image to be detected.
6. The method according to claim 5, wherein the step of performing multi-modal feature extraction on the image to be detected by the feature extraction network to obtain image sub-features of the image to be detected in different modalities includes at least one of the following: Performing feature extraction on the pixel domain of the image to be detected by the feature extraction network to obtain pixel domain image sub-features of the image to be detected; Performing noise transformation on the image to be detected by the feature extraction network to obtain a noise domain representation of the image to be detected, and performing feature extraction on the noise domain representation to obtain noise domain image sub-features of the image to be detected; Performing frequency domain transformation on the image to be detected by the feature extraction network to obtain a frequency domain representation of the image to be detected, and performing feature extraction on the frequency domain representation to obtain frequency domain image sub-features of the image to be detected.
7. The method according to claim 1, wherein the multiple classification networks include a ConvNeXt network and a RepLKNet network.
8. An electronic device, comprising: A processor; A memory for storing processor-executable instructions; wherein, the processor realizes the steps of the method according to any one of claims 1-7 by running the executable instructions.
9. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are realized.
10. A computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-7 are realized.