Deep forgery detection method and device, computer equipment and storage medium

By combining the CLIP ViT-L/14 visual encoder, global feature adapter, and local feature adapter, an interactive fusion classifier is used for deep fake detection, which solves the problem of insufficient generalization ability of traditional detection technology and achieves higher detection accuracy and generalization ability.

CN120726342APending Publication Date: 2025-09-30SHENZHEN TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510807921.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Traditional deepfake detection techniques lack generalization capabilities when faced with diverse, unknown, or new forgery methods (especially advanced GANs and diffusion models).

Method used

The CLIP ViT-L/14 visual encoder is used for multi-layer feature extraction, combined with the global feature adapter and the local feature adapter for feature extraction, and an interactive fusion classifier is used for deep fake detection. A DFA framework is constructed to enhance the detection ability.

Benefits of technology

Significantly enhances the generalization and accuracy of deep fake detection, improving detection effectiveness when facing diverse, unknown, or new fake technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726342A_ABST
    Figure CN120726342A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of image processing, and relates to a deep forgery detection method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a deep forgery detection request which is sent by a user terminal and carries to-be-detected image data; inputting the to-be-detected image data into a CLIP ViT-L / 14 visual encoder to carry out multi-layer feature extraction operation so as to obtain multi-level feature data; inputting the multi-level feature data into a global feature adapter to perform global feature extraction operation to obtain global feature data; inputting the multi-level feature data into a local feature adapter for facial feature extraction operation to obtain facial feature data; inputting the global feature data and the facial feature data into an interactive fusion classifier to carry out deep forgery detection operation to obtain a deep forgery detection result; and sending the deep forgery detection result to the user terminal. According to the invention, the generalization ability and accuracy of deep counterfeiting detection are significantly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a deep fake detection method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of artificial intelligence (AI), deepfake technology can generate highly realistic facial images and videos. Its misuse poses an increasingly serious challenge to social security, political stability, and economic order. Therefore, it is crucial to develop technologies that can accurately and efficiently identify these fake contents.

[0003] Existing deepfake detection techniques primarily rely on deep learning methods, which train binary classifiers to distinguish between real and forged content. These methods typically operate at the image or video level. Image-level methods focus on assessing the probability of forgery in a single frame, often employing convolutional neural network (CNN) architectures (such as ResNeXt, ConvNeXt, and Xception) for feature extraction and classification. Xception is widely used due to its deep separable convolutional structure's excellent performance in efficiently capturing subtle forgery traces. Video-level methods aim to leverage temporal information, typically first using CNNs to extract spatial features from each frame. These are then combined with models capable of processing sequential data (such as the Vision Transformer (ViT)) to model global context and identify potential cross-frame inconsistencies in forged videos.

[0004] However, despite the progress made by these deep learning methods, they often exhibit limited generalization capabilities when faced with synthetic facial images and videos generated by unknown or novel techniques, particularly more advanced generative adversarial networks (GANs) or emerging diffusion models. Traditional media forensics techniques, such as those that rely on analyzing signal-level clues (such as double JPEG compression artifacts), physical artifacts, or semantic-level inconsistencies (such as metadata analysis), have proven insufficiently reliable and effective in identifying increasingly complex and versatile deepfake content. Even deep learning-based binary classifiers often exhibit poor generalization capabilities when faced with samples outside the training distribution, particularly those generated by unknown or novel forgery techniques (especially diffusion models).

[0005] This shows that traditional Deepfake detection technology has insufficient generalization capabilities when faced with diverse, unknown or new forgery methods (especially advanced GAN and diffusion models). Summary of the Invention

[0006] The purpose of the embodiments of the present application is to propose a deepfake detection method, apparatus, computer device, and storage medium to address the problem of insufficient generalization capability of traditional deepfake detection technology when faced with diverse, unknown, or new forgery methods (especially advanced GANs and diffusion models).

[0007] To solve the above technical problems, the present invention provides a method for detecting deep fakes, which adopts the following technical solutions: Receiving a deep fake detection request including image data to be detected from a user terminal; Inputting the image data to be detected into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data; Inputting the multi-level feature data into a global feature adapter to perform a global feature extraction operation to obtain global feature data; Inputting the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data; Inputting the global feature data and the facial feature data into an interactive fusion classifier to perform a deep fake detection operation to obtain a deep fake detection result; Sending the deep fake detection result to the user terminal.

[0008] Furthermore, after the step of receiving a deep fake detection request carrying image data to be detected from a user terminal and before the step of inputting the image data to be detected into a CLIP ViT-L / 14 visual encoder for performing a multi-layer feature extraction operation to obtain multi-level feature data, the following step is further included: A preprocessing operation is performed on the image data to be detected.

[0009] Furthermore, after the step of inputting the multi-level feature data into the global feature adapter to perform a global feature extraction operation to obtain the global feature data, the following steps are also included: Get the preset query embedding vector; Inputting the global feature data and the query embedding vector into a multilayer perceptron respectively to obtain an attention bias signal; The attention bias signal is sent to the CLIP ViT-L / 14 visual encoder to dynamically guide the CLIP ViT-L / 14 visual encoder to perform an attention allocation operation according to the attention bias signal.

[0010] Furthermore, the step of inputting the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data specifically includes the following steps: Acquire facial key point coordinate data corresponding to the image data to be detected; Inputting the multi-level feature data and the facial key point coordinate data into a key point mask generator to perform a spatial attention mask generation operation to obtain spatial attention mask data; The spatial attention mask data and the image data to be detected are input into a convolutional neural network to perform a facial area feature extraction operation to obtain the facial feature data.

[0011] Furthermore, the step of inputting the global feature data and the facial feature data into an interactive fusion classifier to perform a deepfake detection operation to obtain a deepfake detection result specifically includes the following steps: Performing a splicing operation on the global feature data and the facial feature data along a sequence dimension to obtain spliced ​​feature data; Inputting the spliced ​​feature data into a Transformer encoder for fusion operation to obtain fused feature data; Performing a pooling operation on the fused feature data to obtain a pooled feature vector; The pooled feature vector is input into the fully connected classification layer for binary classification to obtain the deep fake detection result.

[0012] Furthermore, the global feature adapter, the local feature adapter, and the interactive fusion classifier are trained using a multi-task learning paradigm, wherein the multi-task learning paradigm is expressed as:

[0013] Among them, Wglobal represents the weight coefficient of the loss function loss1 of the global feature adapter, Wlocal represents the weight coefficient of the loss function loss2 of the local feature adapter, and Wfusion represents the weight coefficient of the loss function loss3 of the interactive fusion classifier.

[0014] To solve the above technical problems, the present application also provides a deep fake detection device, which adopts the following technical solutions: a request receiving module, configured to receive a deep fake detection request sent by a user terminal and carrying image data to be detected; A multi-layer feature extraction module is used to input the image data to be detected into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data; A global feature extraction module, configured to input the multi-level feature data into a global feature adapter to perform a global feature extraction operation to obtain global feature data; A facial feature extraction module, configured to input the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data; a deepfake detection module, configured to input the global feature data and the facial feature data into an interactive fusion classifier to perform a deepfake detection operation and obtain a deepfake detection result; A result output module is used to send the deep fake detection result to the user terminal.

[0015] Furthermore, the device further comprises: The preprocessing module is used to perform preprocessing operations on the image data to be detected.

[0016] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the deep fake detection method as described above are implemented.

[0017] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the deep fake detection method described above.

[0018] This application provides a deepfake detection method, comprising: receiving a deepfake detection request from a user terminal containing image data to be detected; inputting the image data to be detected into a CLIP ViT-L / 14 visual encoder for multi-layer feature extraction to obtain multi-level feature data; inputting the multi-level feature data into a global feature adapter for global feature extraction to obtain global feature data; inputting the multi-level feature data into a local feature adapter for facial feature extraction to obtain facial feature data; inputting the global feature data and the facial feature data into an interactive fusion classifier for deepfake detection to obtain a deepfake detection result; and transmitting the deepfake detection result to the user terminal. Compared to the prior art, this application constructs two collaborative adapter streams around a parameter-frozen CLIP visual encoder: a global feature adapter (Global) and a local feature adapter (Local). Finally, an interactive fusion classifier (IFC) is used to integrate the information from these two streams to make a final authenticity judgment, thereby significantly enhancing the generalization and accuracy of deepfake detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] Figure 1 is an exemplary system architecture diagram to which the present application may be applied; Figure 2 This is a flowchart of the implementation of the deep fake detection method provided in the embodiment of the present application; Figure 3 is a schematic structural diagram of a deep fake detection device provided in an embodiment of the present application; Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0022] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0024] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0025] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0026] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0027] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0028] It should be noted that the deep fake detection method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the deep fake detection device is generally set in the server / terminal device.

[0029] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0030] Continue to refer Figure 2 , which shows a flow chart of an embodiment of a deep fake detection method according to the present application. The deep fake detection method includes: step S201, step S202, step S203, step S204, step S205, and step S206.

[0031] In step S201, a deep fake detection request carrying image data to be detected is received from a user terminal.

[0032] In the embodiments of the present application, the user terminal refers to a terminal device used to execute the image processing method for preventing document abuse provided by the present application. The user terminal can be a mobile terminal such as a mobile phone, a smart phone, a laptop computer, a digital broadcast receiver, a PDA (personal digital assistant), a PAD (tablet computer), a PMP (portable multimedia player), a navigation device, etc., as well as a fixed terminal such as a digital TV, a desktop computer, etc. It should be understood that the examples of user terminals here are only for convenience of understanding and are not used to limit the present application.

[0033] In step S202, the image data to be detected is input into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data.

[0034] In the embodiment of the present application, the image frame of the image data to be detected is fed into the frozen CLIP ViT-L / 14 visual encoder to extract a multi-level feature map clip-feature rich in semantic information. These features are simultaneously fed into the global feature adapter and the local feature adapter.

[0035] In some optional implementations of the embodiments of the present application, the present application may explore the use of other large-scale pre-trained vision or vision-language basic models other than CLIP ViT-L / 14 as the backbone network for parameter freezing.

[0036] In step S203, the multi-level feature data is input into a global feature adapter to perform a global feature extraction operation to obtain global feature data.

[0037] In this embodiment, this module aims to adapt CLIP's powerful general-purpose visual representation capabilities to the task of Deepfake detection, focusing on identifying global inconsistencies in overall image content that may indicate forgery. It optimizes forgery-related global contextual cues through a specific interaction mechanism with the CLIP layer (including feature fusion and attention bias feedback), generating a global feature map Gfmp.

[0038] In this embodiment of the present application, the global feature adapter contains a lightweight Transformer structure and initializes a set of learnable query embedding vectors. It receives feature maps of different depths from the CLIP model and combines these external features with its internal representation using a dedicated fusion layer at a specific level.

[0039] In step S204, the multi-level feature data is input into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data.

[0040] In this embodiment, facial forgeries often introduce local anomalies or inconsistencies in specific semantic regions (such as the eyes, mouth, nose, and facial contours). This module aims to enhance the model's ability to perceive local forgery cues within these key regions by explicitly leveraging prior knowledge of facial structure. It leverages anatomical prior knowledge derived from facial landmarks to isolate and amplify inconsistencies in key regions, thereby addressing the limitations of traditional methods in local perception.

[0041] In step S205 , the global feature data and the facial feature data are input into the interactive fusion classifier to perform a deep fake detection operation to obtain a deep fake detection result.

[0042] In some optional implementations of the embodiments of the present application, the global feature adapter, local feature adapter, and interactive fusion classifier are trained using a multi-task learning paradigm, where the multi-task learning paradigm is expressed as:

[0043] Among them, Wglobal represents the weight coefficient of the loss function loss1 of the global feature adapter, Wlocal represents the weight coefficient of the loss function loss2 of the local feature adapter, and Wfusion represents the weight coefficient of the loss function loss3 of the interactive fusion classifier.

[0044] In some optional implementations of the embodiments of the present application, the present application may adopt not only learnable loss weights but also manually set fixed weight values.

[0045] In step S206, the deep fake detection result is sent to the user terminal.

[0046] In an embodiment of the present application, a deepfake detection method is provided, comprising: receiving a deepfake detection request from a user terminal carrying image data to be detected; inputting the image data to be detected into a CLIP ViT-L / 14 visual encoder for multi-layer feature extraction to obtain multi-level feature data; inputting the multi-level feature data into a global feature adapter for global feature extraction to obtain global feature data; inputting the multi-level feature data into a local feature adapter for facial feature extraction to obtain facial feature data; inputting the global feature data and the facial feature data into an interactive fusion classifier for deepfake detection to obtain a deepfake detection result; and sending the deepfake detection result to the user terminal. Compared with the prior art, the present application constructs two collaborative adapter streams around a parameter-frozen CLIP visual encoder: a global feature adapter (Global) and a local feature adapter (Local). Finally, an interactive fusion classifier (IFC) is used to integrate the information of these two streams to make a final authenticity judgment, thereby significantly enhancing the generalization ability and accuracy of deepfake detection.

[0047] In some optional implementations of the embodiments of the present application, after the step of receiving a deep fake detection request carrying image data to be detected from a user terminal, and before the step of inputting the image data to be detected into the CLIP ViT-L / 14 visual encoder for performing a multi-layer feature extraction operation to obtain multi-level feature data, the following steps are further included: Perform preprocessing operations on the image data to be detected.

[0048] In this example, the input video first undergoes preprocessing steps such as frame sampling, face detection and alignment, image normalization, and facial key point extraction. A default frame count of 32 is selected, and the Dlib library is used for face detection and 81 facial key point extraction. The final facial image is resized to a uniform 256x256 pixel resolution, and the key point coordinates are normalized accordingly.

[0049] In some optional implementations of the embodiments of the present application, after the step of inputting the multi-level feature data into the global feature adapter to perform a global feature extraction operation to obtain the global feature data, the following steps are also included: Get the preset query embedding vector; The global feature data and query embedding vector are input into the multi-layer perceptron respectively to obtain the attention bias signal; The attention bias signal is sent to the CLIP ViT-L / 14 visual encoder to dynamically guide the CLIP ViT-L / 14 visual encoder to perform attention allocation operations according to the attention bias signal.

[0050] In an embodiment of the present application, after being processed by the Transformer module inside the adapter, features related to the query embedding are processed by a multi-layer perceptron (MLP) network to generate an attention bias signal, which is integrated into the self-attention calculation process of the original CLIP model, thereby dynamically guiding the attention allocation of CLIP without modifying the pre-trained weights.

[0051] In some optional implementations of the embodiments of the present application, the step of inputting the multi-level feature data into the local feature adapter to perform a facial feature extraction operation to obtain facial feature data specifically includes the following steps: Obtaining facial key point coordinate data corresponding to the image data to be detected; Input the multi-level feature data and facial key point coordinate data into the key point mask generator to perform a spatial attention mask generation operation to obtain spatial attention mask data; The spatial attention mask data and the image data to be detected are input into the convolutional neural network to perform facial area feature extraction operation to obtain facial feature data.

[0052] In an embodiment of the present application, the Local module receives an image frame, a clip-feature, and the corresponding facial landmark coordinates as input. A landmark mask generator uses the input 81 key points and clip-feature to generate a corresponding spatial attention mask for each region based on predefined facial groupings (such as eyebrows, eyes, nose, and lips). The Local module also uses a convolutional neural network (CNN) independent of CLIP—specifically, a ResNeXt-50 with the final classification layer removed—as its visual feature extraction backbone to process the input image to learn and extract a feature map Lfmp that is highly correlated with the facial region.

[0053] In some optional implementations of the embodiments of the present application, the independent CNN backbone network used in the Local module, in addition to ResNeXt-50, can also use other CNN architectures with comparable computational efficiency and performance, such as the EfficientNet series, MobileNet series, etc.

[0054] In some optional implementations of the embodiments of the present application, the step of inputting the global feature data and the facial feature data into the interactive fusion classifier to perform a deepfake detection operation and obtain a deepfake detection result specifically includes the following steps: Performing a splicing operation on the global feature data and the facial feature data along the sequence dimension to obtain spliced ​​feature data; Input the spliced ​​feature data into the Transformer encoder for fusion operation to obtain fused feature data; Perform pooling operation on the fused feature data to obtain a pooled feature vector; The pooled feature vector is input into the fully connected classification layer for binary classification to obtain the deep fake detection result.

[0055] In this embodiment of the present application, the interactive fusion classifier (IFC) module receives the feature map Gfmp from the global feature adapter and the feature map Lfmp from the local module. These two feature maps are concatenated along the sequence dimension. A Transformer encoder performs the key fusion step, deeply interacting and fusing the complex dependencies between global and local forgery cues. After deep Transformer fusion, the resulting enhanced feature sequence is aggregated into a single feature vector through a pooling operation (such as global average pooling). This vector is fed into a fully connected classification layer to generate the final binary classification prediction result fusion.preds.

[0056] In some optional implementations of the embodiments of the present application, the IFC module uses Transformer Encoder for feature fusion, but other fusion methods may also be used, such as simpler feature splicing followed by direct MLP processing, or using attention-based weighted summation.

[0057] In summary, the present application provides a deepfake detection method that, by adopting a DFA framework, significantly enhances the generalization capability of deepfake detection models in the face of diverse, unknown, or novel forgery techniques. By efficiently adapting parameters to a powerful base model and aided by local feature adapters, the model demonstrates strong robustness against diverse forgery techniques. On mixed datasets and independent test sets, the present invention demonstrates superiority across multiple evaluation metrics, including accuracy, precision, AUC, and error rate, surpassing several existing baseline methods. Through a carefully designed adapter interaction strategy, the method effectively leverages the powerful prior knowledge of a large pre-trained model (CLIP) while maintaining the frozen state of the CLIP model parameters. This method requires only training a relatively lightweight adapter module, significantly reducing the computational resources, time, and storage space required for training and lowering the cost of model training and deployment. By capturing overall anomalies through a global feature adapter and focusing on details in key facial regions through local feature adapters, the interactive fusion classifier integrates these features, enabling a more comprehensive detection of forgery clues and improving detection reliability.

[0058] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0059] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0060] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware using computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0061] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0062] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a deep fake detection device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0063] like Figure 3 As shown, the deep fake detection device 200 of the embodiment of the present application includes: a request receiving module 210 for receiving a deep fake detection request including image data to be detected, sent by a user terminal; The multi-layer feature extraction module 220 is used to input the image data to be detected into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data; The global feature extraction module 230 is used to input the multi-level feature data into the global feature adapter to perform a global feature extraction operation to obtain global feature data; The facial feature extraction module 240 is used to input the multi-level feature data into the local feature adapter to perform facial feature extraction operations to obtain facial feature data; A deepfake detection module 250 is configured to input the global feature data and the facial feature data into an interactive fusion classifier to perform a deepfake detection operation and obtain a deepfake detection result; The result output module 260 is used to send the deep fake detection results to the user terminal.

[0064] In an embodiment of the present application, a deep fake detection device 200 is provided, including: a request receiving module 210, for receiving a deep fake detection request carrying image data to be detected sent by a user terminal; a multi-layer feature extraction module 220, for inputting the image data to be detected into the CLIP ViT-L / 14 visual encoder for performing a multi-layer feature extraction operation to obtain multi-level feature data; a global feature extraction module 230, for inputting the multi-level feature data into a global feature adapter for performing a global feature extraction operation to obtain global feature data; a facial feature extraction module 240, for inputting the multi-level feature data into a local feature adapter for performing a facial feature extraction operation to obtain facial feature data; a deep fake detection module 250, for inputting the global feature data and the facial feature data into an interactive fusion classifier for performing a deep fake detection operation to obtain a deep fake detection result; and a result output module 260, for sending the deep fake detection result to the user terminal. Compared with the existing technology, this application builds two collaborative adapter streams around a parameter-frozen CLIP visual encoder: the Global Feature Adapter (Global) and the Local Feature Adapter (Local). Finally, an Interactive Fusion Classifier (IFC) is used to integrate the information of these two streams to make the final authenticity judgment, thereby significantly enhancing the generalization ability and accuracy of deep fake detection.

[0065] In some optional implementations of the embodiments of the present application, the deep fake detection apparatus 200 further includes: The preprocessing module is used to perform preprocessing operations on the image data to be detected.

[0066] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device according to an embodiment of the present application.

[0067] The computer device 300 includes a memory 310, a processor 320, and a network interface 330 that are interconnected through a system bus. It should be noted that the figure only shows the computer device 300 having components 310-330, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0068] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0069] The memory 310 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 310 may be an internal storage unit of the computer device 300, such as a hard disk or memory of the computer device 300. In other embodiments, the memory 310 may also be an external storage device of the computer device 300, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 310 may also include both the internal storage unit of the computer device 300 and its external storage device. In the embodiment of the present application, the memory 310 is generally used to store the operating system and various application software installed on the computer device 300, such as computer-readable instructions for the deep fake detection method. In addition, the memory 310 can also be used to temporarily store various data that has been output or is about to be output.

[0070] In some embodiments, the processor 320 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 320 is generally used to control the overall operation of the computer device 300. In the embodiment of the present application, the processor 320 is used to execute computer-readable instructions or process data stored in the memory 310, such as computer-readable instructions for executing the deepfake detection method.

[0071] The network interface 330 may include a wireless network interface or a wired network interface. The network interface 330 is generally used to establish a communication connection between the computer device 300 and other electronic devices.

[0072] The computer device provided in this application builds two collaborative adapter streams around a parameter-frozen CLIP visual encoder: a global feature adapter (Global) and a local feature adapter (Local). Finally, an interactive fusion classifier (IFC) is used to integrate the information of these two streams to make the final authenticity judgment, thereby significantly enhancing the generalization ability and accuracy of deep fake detection.

[0073] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the deep fake detection method as described above.

[0074] The computer-readable storage medium provided in this application constructs two collaborative adapter streams around a parameter-frozen CLIP visual encoder: a global feature adapter (Global) and a local feature adapter (Local). Finally, an interactive fusion classifier (IFC) is used to integrate the information of these two streams to make the final authenticity judgment, thereby significantly enhancing the generalization ability and accuracy of deep fake detection.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.

[0076] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A deep fake detection method, characterized in that: The steps include: Receiving a deep fake detection request including image data to be detected from a user terminal; Inputting the image data to be detected into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data; Inputting the multi-level feature data into a global feature adapter to perform a global feature extraction operation to obtain global feature data; Inputting the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data; Inputting the global feature data and the facial feature data into an interactive fusion classifier to perform a deep fake detection operation to obtain a deep fake detection result; Sending the deep fake detection result to the user terminal.

2. The deep fake detection method according to claim 1, characterized in that After the step of receiving a deep fake detection request carrying image data to be detected from a user terminal and before the step of inputting the image data to be detected into a CLIP ViT-L / 14 visual encoder for performing a multi-layer feature extraction operation to obtain multi-level feature data, the following step is also included: A preprocessing operation is performed on the image data to be detected.

3. The deep fake detection method according to claim 1, characterized in that After the step of inputting the multi-level feature data into the global feature adapter to perform a global feature extraction operation to obtain the global feature data, the following steps are also included: Get the preset query embedding vector; Inputting the global feature data and the query embedding vector into a multilayer perceptron respectively to obtain an attention bias signal; The attention bias signal is sent to the CLIP ViT-L / 14 visual encoder to dynamically guide the CLIP ViT-L / 14 visual encoder to perform an attention allocation operation according to the attention bias signal.

4. The deep fake detection method according to claim 1, characterized in that The step of inputting the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data specifically includes the following steps: Acquire facial key point coordinate data corresponding to the image data to be detected; Inputting the multi-level feature data and the facial key point coordinate data into a key point mask generator to perform a spatial attention mask generation operation to obtain spatial attention mask data; The spatial attention mask data and the image data to be detected are input into a convolutional neural network to perform a facial area feature extraction operation to obtain the facial feature data.

5. The deep fake detection method according to claim 1, characterized in that The step of inputting the global feature data and the facial feature data into an interactive fusion classifier to perform a deep fake detection operation to obtain a deep fake detection result specifically includes the following steps: Performing a splicing operation on the global feature data and the facial feature data along a sequence dimension to obtain spliced ​​feature data; Inputting the spliced ​​feature data into a Transformer encoder for fusion operation to obtain fused feature data; Performing a pooling operation on the fused feature data to obtain a pooled feature vector; The pooled feature vector is input into the fully connected classification layer for binary classification to obtain the deep fake detection result.

6. The deep fake detection method according to claim 1, characterized in that The global feature adapter, the local feature adapter, and the interactive fusion classifier are trained using a multi-task learning paradigm, wherein the multi-task learning paradigm is expressed as: Among them, Wglobal represents the weight coefficient of the loss function loss1 of the global feature adapter, Wlocal represents the weight coefficient of the loss function loss2 of the local feature adapter, and Wfusion represents the weight coefficient of the loss function loss3 of the interactive fusion classifier.

7. A deep fake detection device, characterized in that: include: a request receiving module, configured to receive a deep fake detection request sent by a user terminal and carrying image data to be detected; A multi-layer feature extraction module is used to input the image data to be detected into the CLIP ViT-L / 14 visual encoder to perform a multi-layer feature extraction operation to obtain multi-level feature data; A global feature extraction module, configured to input the multi-level feature data into a global feature adapter to perform a global feature extraction operation to obtain global feature data; A facial feature extraction module, configured to input the multi-level feature data into a local feature adapter to perform a facial feature extraction operation to obtain facial feature data; a deepfake detection module, configured to input the global feature data and the facial feature data into an interactive fusion classifier to perform a deepfake detection operation and obtain a deepfake detection result; A result output module is used to send the deep fake detection result to the user terminal.

8. The deep fake detection device according to claim 7, characterized in that: The device further comprises: The preprocessing module is used to perform preprocessing operations on the image data to be detected.

9. A computer device comprising a memory and a processor, characterized in that: The memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the deep fake detection method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the deep fake detection method according to any one of claims 1 to 6.