Pluggable forged face detection method based on test time domain adaptation

By introducing a pluggable adaptation module based on test time domain adaptation on the deep forgery detector, the problems of weak generalization and low domain adaptation efficiency in the prior art are solved, and efficient adaptation and generalization performance improvement of new identity information and forgery algorithms are achieved.

CN120048008AActive Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510115842.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing deep falsification detection methods have problems such as weak generalization, overfitting and difficulty in adapting to unknown falsification data, and the domain adaptation method is inefficient and has a small application range.

Method used

Using a pluggable forged face detection method based on test time domain adaptation, the pluggable adaptation module is introduced on the detector trained in the source domain to improve the test accuracy of the detector in the target domain. The method includes a feature conversion layer and a nearest neighbor corrector, dynamically updates the features and corrects the prediction results, and improves generalization and testing efficiency.

Benefits of technology

It improves the generalization performance of deep forgery detection, enhances the detector's adaptability to new identity information and forgery algorithms, and does not require retraining the detector. It is suitable for online testing scenarios and improves testing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048008A_ABST
    Figure CN120048008A_ABST
Patent Text Reader

Abstract

The invention provides a pluggable fake face detection method based on test time domain adaptation, and the method comprises the steps: employing a detector which is trained in a source domain as a basic detector, and improving the test accuracy of the basic detector in a target domain through a pluggable adaptation module; according to the pluggable depth forgery detection module based on test time domain adaptation, the generalization performance of an existing detector is improved in the test stage, the structure of the detector does not need to be known, meanwhile, the detector does not need to be retrained, in the test time domain adaptation process, parameters of a basic detector are frozen and are not changed, only parameters of a feature conversion layer are updated, and the detection efficiency is improved. Old features can better adapt to new forged data, prediction variance can be reduced by weighting results of nearest neighbor samples of the old features, and prediction consistency constraints are constructed to relieve the influence of noise samples on a feature conversion layer during updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field, and in particular, to a method for detecting pluggable forged faces based on test-time adaptation. Background Art

[0002] With the continuous development of deep generative models, face deep fake images or videos have become increasingly popular on information networks and social media platforms; these generative technologies use neural networks to learn to generate realistic face images and videos, and can even forge the facial features and expressions of specific individuals; although deep fake technologies have broad application prospects in entertainment fields such as film production, virtual reality, and games, and can reduce production costs and enhance visual effects, their abuse has also brought many social problems; specifically, deep fakes may violate the privacy of public figures and accelerate the spread of false information; therefore, the widespread application of deep fakes poses new challenges to information security.

[0003] For the detection of deep fake content, existing methods mainly regard deep fake detection as a binary classification task and achieve it by training neural networks on a dataset containing paired real and forged images; currently, a considerable number of trained detectors are available for the identification of forged faces; however, these detection methods have the problem of weak generalization, specifically manifested as that the detector can achieve extremely high detection accuracy on data with the same training set distribution, but the detection accuracy on data with different distributions will drop significantly.

[0004] On the one hand, these detection methods are prone to overfitting to features irrelevant to forgery traces, including image backgrounds and face identity information, etc. When facing face images of different identities, the detection accuracy will be interfered by the new identity information.

[0005] On the other hand, deep fake algorithms are constantly iterating and updating, and different deep fake algorithms will leave different forgery features on the forged content, and detectors using specific forgery algorithms will be difficult to detect unknown forged images.

[0006] Therefore, a pluggable test-time adaptation method is needed to enable the detector to achieve good detection performance on content with new identity information and forgery features without retraining the existing detector, and improve its generalization. To solve the above technical problems, the prior art has designed the following technical solutions, for example:

[0007] ①The patent document with the publication number of CN118968269A, "A Deepfake Detection Method Combining Spatial Domain Texture Differences and Frequency Domain Information", combines spatial domain information and frequency domain information to improve the generalization of the detector; regarding the texture difference information in the spatial domain, that is, the textures of the background image and the foreground image in the face-swapped image may be different. Based on this, the common information shared by different forgery methods can be mined, and the frequency domain information is further fused to improve the detection accuracy.

[0008] ②The patent document with the publication number of CN118470585A, "A Deepfake Detection Method with Multi-Domain Fusion", proposes a feature fusion method for spatial domain information and frequency domain information; in the model training stage, a suitable convolutional neural network is selected for feature extraction, and after feature fusion, a traditional machine learning model such as SVM is used for classification, which improves the generalization of the detector.

[0009] ③The patent document with the publication number of CN118379608B, "A Highly Robust Deepfake Detection Method Based on Adaptive Learning", uses adaptive technology to solve the robustness problem of deepfake detection, which aims at the problem that it is difficult for the detector to identify the forged images after quality degradation; this method first adaptively generates degraded face images of different qualities based on the degradation generation algorithm to supplement the quality diversity of the existing forged face image dataset, and at the same time combines the adaptive sampling network to coordinate the learning signals of face images of different qualities, dynamically capturing the forgery features of face images with unknown quality, thereby improving the detection performance of the deepfake detection model for low-quality forged images.

[0010] ④The patent document with the publication number of CN118397440A, "Deepfake Detection Method and System Based on Unsupervised Domain Adaptation", combines domain adaptation technology to improve the generalization of deepfake detection; this method uses unsupervised domain adaptation technology to identify known and unknown deepfake categories in the target domain without the need for target domain labels, thereby improving the generalization ability of the model.

[0011] However, the above technologies still have design deficiencies and solution defects. For example:

[0012] (1) Over-reliance on prior knowledge of existing forgery methods;

[0013] Existing deepfake detection methods dedicated to improving generalization require known forgery methods to construct the training dataset and still have difficulty adapting to unknown forged data.

[0014] (2) Low efficiency of domain adaptation and small application scope;

[0015] The deepfake detection method based on domain adaptation can only process a large number of data to be detected offline and can only be applied to specific detector structures. Summary of the Invention

[0016] To solve the above technical problems, the present invention provides a pluggable forged face detection method based on test time domain adaptation, which improves the generalization of deepfakes. Aiming at the problem of over-reliance on the prior knowledge of existing forgery methods, a domain adaptation scheme is used to alleviate the dependence of the detector on training knowledge; aiming at the problems of low efficiency and small application scope of existing domain adaptation methods, a pluggable test time domain adaptation method is proposed, which can be used to improve the generalization of different trained detectors without retraining the detector, and can also be used in online test scenarios to improve test efficiency.

[0017] In a first aspect, a pluggable forged face detection method based on test time domain adaptation uses a detector trained in the source domain D s as a base detector, and improves the test accuracy of the base detector in the target domain D t through a pluggable adaptation module;

[0018] As an example, the pluggable adaptation module only processes the output results of the base detector without modifying the structure of the base detector or retraining.

[0019] The specific design of the detection method includes:

[0020] Step 1: Prepare the base detector and test samples;

[0021] The base detector includes: a feature extractor and a classifier Set the test sample x i ;

[0022] The feature extractor: used to extract the deep features of the image

[0023] The classifier: calculates the output score according to the deep features

[0024] As an example, the base detector uses: a trained ViT model as the base detector.

[0025] As a preferred example, the ViT model is trained on a dataset of paired real and forged face images;

[0026] The methods used for forging face images include: Deepfakes and FaceSwap algorithms; the dataset to be tested includes faces of other identities, including forged face images generated using the Face2Face forgery algorithm.

[0027] As an example, the classifier is: a learnable prototype classifier.

[0028] Step 2: Input the test sample x i into the basic detector;

[0029] Use the historical record M = {(F, L)} to store the output results of the basic detector and perform filtering;

[0030] The specific solution includes the following two sub-steps:

[0031] Step 2.1: Initialize the historical record M = {(F, L)}, where the deep features are initialized as the weight values of the classifier, and the output scores are initialized as {0, 1};

[0032] Step 2.2: Update the historical record according to the output results of the basic detector. The updated historical record M' = M ∪ (F i , L i );

[0033] Calculate the entropy of each test sample according to the output scores:

[0034] H(L i ) = -∑σ(L i )log(σ(L i ));

[0035] where: σ is the softmax function. Subsequently, filter the K samples with the largest entropy and maintain the size of the historical record at N m ;

[0036] Step 3: The classifier calculates the class prototypes based on the historical record M = {(F, L)} and calculates the prediction results based on this. Subsequently, use the pseudo-labels to calculate the standard cross-entropy loss to update the feature transformation layer;

[0037] The specific solution includes the following sub-steps:

[0038] Step 3.1: The formula for calculating the class prototypes is designed as follows:

[0039]

[0040] where: I(·) is the indicator function, and j ∈ {1,..., N m} is the index of the element in the historical record; a feature transformation layer T is innovatively introduced r , r ∈ {1,..., N t}, including N t fully connected networks; the feature transformation layer can be dynamically updated during the testing process, enabling the features of the trained basic detector to adapt to new unknown forged data;

[0041] Subsequently, calculate the adapted features and class prototypes according to the feature transformation layer:

[0042]

[0043] And design through the following formula to calculate the similarity as the prediction result:

[0044]

[0045] Where: S(·) is the cosine similarity function, and P i k (r) is the predicted value of class r at the k-th transformation layer; P i k is the predicted result averaged over all feature transformation layers;

[0046] Step 3.2: Dynamically update the feature transformation layer according to the predicted result;

[0047] Use the softmax value σ(L i ) of the output score of the basic detector to guide the predicted result P i k , and at the same time set a threshold Conf to ignore samples with low confidence to alleviate the impact of noisy labels:

[0048] L i = CE(σ(L i ), P i );

[0049]

[0050] Where: CE is the standard cross-entropy loss function.

[0051] Step 4: Select the sample in the historical record that is the nearest neighbor of the test sample in the feature space, and use the nearest neighbor sample to further correct the predicted result of the test sample;

[0052] The specific solution includes the following sub-steps:

[0053] Step 4.1: Select N f nearest neighbor samples according to the feature distance:

[0054]

[0055] where: dis(·) is the distance function, and β(N f ) is the distance between the deep feature F of the test sample i and the deep feature of the N f -th nearest neighbor sample; subsequently, N f nearest neighbor prediction results P i (n), n ∈ {1,..., N f} can be obtained;

[0056] Step 4.2: Further correct the prediction result of the test sample according to the nearest neighbor prediction result to obtain the final prediction result:

[0057]

[0058] As an example, to further reduce the impact of noisy labels on the update of the feature transformation layer, a nearest neighbor consistency constraint is introduced to force the test sample and its corresponding nearest neighbor samples to output similar results:

[0059]

[0060] Step 5: After calculating the final prediction result, calculate the loss according to the following formula and update the feature transformation layer K using the backpropagation algorithm s Step:

[0061] L = L lcpc + αL nfc ;

[0062] where: α is a hyperparameter to balance the importance of the consistency constraint.

[0063] In a second aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.

[0064] In a third aspect, the present application discloses a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.

[0065] In a fourth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by the processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.

[0066] Advantages of the present invention:

[0067] 1. Plug-and-Play Deepfake Detection Module Based on Test-Time Domain Adaptation: The proposed solution in this application constructs a plug-and-play deepfake detection module based on test-time domain adaptation, which enhances the generalization performance of existing detectors during the test phase without the need to know the detector's structure and without retraining the detector. This module operates in an online test scenario, improving the test efficiency compared to offline domain adaptation.

[0068] 2. Learnable Prototype Classifier: The proposed solution in this application introduces a feature transformation layer to adaptively transform the deep features in the base detector during testing. During the test-time domain adaptation process, the parameters of the base detector are frozen and only the parameters of the feature transformation layer are updated. Through feature transformation, the old features can better adapt to new forged data, including new identity data and new forgery algorithms.

[0069] 3. Nearest Neighbor Corrector: The proposed solution in this application constructs a nearest neighbor-based prediction corrector to correct the prediction results of a certain test sample. By weighting the results of its nearest neighbor samples, the prediction variance can be reduced. At the same time, a prediction consistency constraint is constructed to mitigate the impact of noisy samples on the update of the feature transformation layer. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is the overall structural design diagram of a method for detecting pluggable forged faces based on test-time domain adaptation of the present invention.

[0071] Figure 2 A method for detecting pluggable forged faces based on test-time domain adaptation of the present invention

[0072] Figure 3 A method for detecting pluggable forged faces based on test-time domain adaptation of the present invention DETAILED DESCRIPTION OF THE INVENTION

[0073] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. Refer to Figures 1 to 3 as shown in

[0074] In the first aspect, a method for detecting pluggable forged faces based on test-time domain adaptation uses a detector trained in the source domain D s as the base detector, and improves the test accuracy of the base detector in the target domain D t through a pluggable adaptation module.

[0075] As an example, the pluggable adaptation module only processes the output results of the basic detector without modifying the structure of the basic detector or retraining it.

[0076] The specific design of the detection method includes:

[0077] Step 1: Prepare a basic detector and test samples;

[0078] The basic detector includes: a feature extractor and a classifier Set the test sample x i ;

[0079] The feature extractor: used to extract the deep features of the image

[0080] The classifier: calculates the output score according to the deep features

[0081] As an example, the basic detector uses: a trained ViT model as the basic detector.

[0082] As a preferred example, the ViT model is trained on a dataset of paired real and fake face images;

[0083] The methods used for the fake face images include: Deepfakes and FaceSwap algorithms; the test dataset to be included faces of other identities, including fake face images generated using the Face2Face forgery algorithm.

[0084] As an example, the classifier is: a learnable prototype classifier.

[0085] Step 2: Input the test sample x i into the basic detector;

[0086] Use the history record M = {(F, L)} to store the output results of the basic detector and perform filtering;

[0087] The specific scheme includes the following two sub-steps:

[0088] Step 2.1: Initialize the history record M = {(F, L)}, where the deep features are initialized to the weight values of the classifier, and the output scores are initialized to {0, 1};

[0089] Step 2.2: Update the history record according to the output results of the basic detector, and the updated history record M' = M ∪ (F i , L i );

[0090] Calculate the entropy of each of the test samples based on the output scores:

[0091] H(L i ) = -Σσ(L i ) log(σ(L i ));

[0092] where: σ is the softmax function. Subsequently, filter the K samples with the largest entropy and maintain the size of the historical record at N m ;

[0093] Step 3: The classifier calculates the class prototypes based on the historical record M = {(F, L)}, and calculates the prediction result based on this. Subsequently, use the pseudo labels to calculate the standard cross-entropy loss to update the feature transformation layer;

[0094] The specific solution includes the following sub-steps:

[0095] Step 3.1: The formula for calculating the class prototypes is designed as follows:

[0096]

[0097] where: I(·) is the indicator function, and j ∈ {1,..., N m} is the element index in the historical record; The feature transformation layer T r , r ∈ {1,..., N t} is innovatively introduced, including N t fully connected networks; The feature transformation layer can be dynamically updated during the test, enabling the features of the trained base detector to adapt to new unknown forged data;

[0098] Subsequently, calculate the adapted features and class prototypes based on the feature transformation layer:

[0099]

[0100] And calculate the similarity as the prediction result through the following formula design:

[0101]

[0102] where: S(·) is the cosine similarity function, and P i k (r) is the predicted value for class r at the k-th transformation layer; P i k is the average of the prediction results of all feature transformation layers;

[0103] Step 3.2: Dynamically update the feature transformation layer according to the prediction result;

[0104] Use the softmax value σ(L i ) of the output score of the base detector to guide the prediction result P i k , and set a threshold Conf to ignore samples with low confidence to mitigate the impact of noisy labels:

[0105] L i = CE(σ(L i ), P i );

[0106]

[0107] where: CE is the standard cross-entropy loss function.

[0108] Step 4: Select the sample in the history that is the nearest neighbor of the test sample in the feature space, and use the nearest neighbor sample to further correct the prediction result of the test sample;

[0109] The specific solution includes the following sub-steps:

[0110] Step 4.1: Select N f nearest neighbor samples according to the feature distance:

[0111]

[0112] where: dis(·) is the distance function, and β(N f ) is the distance between the deep feature F i of the test sample and the deep feature of the N f -th nearest neighbor sample; subsequently, N f nearest neighbor prediction results P i (n), n ∈ {1,..., N f};

[0113] Step 4.2: Further correct the prediction result of the test sample according to the nearest neighbor prediction result to obtain the final prediction result:

[0114]

[0115] As an example, to further reduce the impact of noisy labels on the update of the feature transformation layer, introduce the nearest neighbor consistency constraint to force the test sample and its corresponding nearest neighbor sample to output similar results:

[0116]

[0117] Step 5: After calculating the final prediction result, calculate the loss according to the following formula and use the backpropagation algorithm to update the feature transformation layer Ks Step:

[0118] L = L lcpc + αL nfc ;

[0119] Where: α is a hyperparameter to balance the importance of the consistency constraint.

[0120] In a second aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.

[0121] In a third aspect, the present application discloses a non - transitory computer - readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the method described in any of the above aspects.

[0122] In a fourth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by the processor of the electronic device, enabling the electronic device to execute the method described in any of the above aspects.

[0123] To better illustrate the design principle of the present invention, the detection method of the present invention is exemplified through specific embodiments as follows:

[0124] Embodiment 1:

[0125] Application of the detection method

[0126] Step 1: Use the trained ViT model as the basic detector, which is trained on a dataset of paired real - face and fake - face images. The methods used for generating fake - face images include Deepfakes and FaceSwap algorithms. The test dataset includes faces of other identities, including fake faces generated using the Face2Face forgery algorithm.

[0127] Step 2: Set the input batch size during testing to 32, and sequentially input all the data to be tested into the basic detector. For the ViT model, use the Class Token as the deep - layer feature F i , and use the output of the fully - connected layer for classification as the output score L i . Store the output results in the historical record, and set the maximum value N m of the saved samples to 1000.

[0128] Step 3: Calculate the class prototype according to the sample output results in the historical record, and set the number N tIt is 20. Subsequently, the deep features of the class prototype and the test samples are input into the feature transformation layer to obtain the adapted prototype and features, and their similarity is calculated to obtain the prediction result. In order to adapt to the new identity information and forgery algorithm in the test data, the feature transformation layer needs to be dynamically updated, and the filtering threshold Conf is set to 0.7 when calculating the cross-entropy loss.

[0129] Step 4: First, use the Euclidean distance function to screen the nearest neighbor samples from the historical records, and the number of nearest neighbors can be set to N f It is 16. Subsequently, the corrected prediction result can be calculated according to the nearest neighbor samples This result is used as the detection result for identification. Finally, the consistency constraint loss value is calculated.

[0130] Step 5: Set the hyperparameter α to 10.0, and calculate the overall loss value L. Assuming that the weight parameter of the feature transformation layer is θ, the parameter can be updated according to the following backpropagation formula:

[0131]

[0132] where: l r is the learning rate. In the specific implementation, the Adam optimizer can be selected, and the learning rate is set to 0.0001. In the update step K of this batch s is 1. Subsequently, return to Step 2 until all samples are tested.

[0133] The present invention greatly improves the generalization performance of deepfake detection; to verify the improvement of the generalization performance of the proposed method for existing deepfake detectors, the ViT model trained on the FF++ c23 dataset is used as the basic detector, and tests are carried out on other datasets including CDF-v2, DFD, and DFDC. Using AUC as the evaluation index, the method in the proposed method can achieve an average performance improvement of 9%, improving the efficiency of domain adaptation;

[0134] Existing domain adaptation methods for deepfake detection mostly adopt the offline test method, that is, multiple rounds of iteration are carried out on the target test dataset. The method in the proposed method can work in the online test scenario; slightly increasing the time consumption of the detector during each round of inference. For Specific Example 1, the increased time consumption is only 0.035 s, and the efficiency of domain adaptation is greatly increased compared with offline testing.

[0135] Note that, for method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily essential for this application.

[0136] Optionally, an embodiment of this application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0137] An embodiment of this application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0138] The figure is a block diagram of an electronic device 800 shown in this application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0139] Refer to Figure 2 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0140] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0141] The memory 804 is configured to store various types of data to support the operation of the device 800; examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0142] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0143] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0144] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0145] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0146] The sensor assembly 814 includes one or more sensors for providing status assessment of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0147] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0148] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0149] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0150] Figure 3FIG. 0 is a block diagram of another electronic device 1900 shown in the present application. For example, the electronic device 1900 may be provided as a server.

[0151] Referring to Figure 3 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above-described method.

[0152] The electronic device 1900 may also include a power component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.

[0153] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0154] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units may refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0155] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0156] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0157] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0158] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0159] The above are only the preferred embodiments of the present invention. It should be understood that the description of the above embodiments is only used to help understand the method and its core idea of the present invention, and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A pluggable forged face detection method based on test time domain adaptation, using a s The trained detector is used as the basic detector, and the pluggable adaptation module is used to improve the basic detector in the target domain D t The test accuracy of Specific detection methods include: Step 1: Prepare basic detector and test samples; The basic detector includes: a feature extractor and a classifier Set the test sample x i ; The feature extractor: used to extract deep features of the image The classifier: calculates the output score based on the deep features Step 2: The test sample x i input to the base detector; Use the historical record M = {(F, L)} to store the output results of the basic detector and filter them; Step 3: The classifier calculates the class prototype according to the historical record M={(F,L)}, and calculates the prediction result based on it, and then updates the feature conversion layer by calculating the standard cross entropy loss using the pseudo label; Step 4: Selecting a sample that is the nearest neighbor of the test sample in the feature space from the historical records, and using the nearest neighbor sample to further correct the prediction result of the test sample; Step 5: After calculating the final prediction result, calculate the loss according to the following formula and use the back propagation algorithm to update the feature conversion layer K s step: L=L lcpc +αL nfc ; Where: α is a hyperparameter to balance the importance of the consistency constraint.

2. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The pluggable adaptation module only processes the output result of the basic detector without modifying the structure of the basic detector or retraining it.

3. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The basic detector uses the trained ViT model as the basic detector; The ViT model is trained on a dataset of paired real and fake face images; The methods used to forge facial images include: Deepfakes and FaceSwap algorithms; the data set to be tested includes faces of other identities, including forged facial images generated using the Face2Face forging algorithm.

4. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The classifier is a learnable prototype classifier.

5. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The specific scheme of step 2 includes the following two sub-steps: Step 2.1: Initialize the history record M = {(F, L)}, wherein the deep features are initialized as the weight values ​​of the classifier, and the output score is initialized as {0, 1}; Step 2.2: Update the historical record according to the output result of the basic detector, and the updated historical record M'=M∪(F i ,L i ); Calculate the entropy of each of the test samples based on the output scores: H(L i )=-∑σ(L i )log(σ(L i )); Where: σ is the softmax function, and then the K samples with the largest entropy are filtered to maintain the size of the historical record at N m .

6. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The specific scheme of step 3 includes the following sub-steps: Step 3.1: The calculation prototype formula is designed as follows: Where: I(·) is the indicator function, and j∈{1,...,N m } is the element index in the historical record; the feature conversion layer T is innovatively introduced r ,r∈{1,...,N t }, including N t A fully connected network; The feature conversion layer can be dynamically updated during the test process, so that the trained basic detector features can adapt to new unknown forged data; The adapted features and class prototypes are then calculated according to the feature conversion layer: And the similarity is calculated as the prediction result through the following formula design: Where: S(·) is the cosine similarity function, and P i k (r) is the predicted value for category r at the kth transformation layer; P i k is the average prediction result of all feature conversion layers; Step 3.2: Dynamically update the feature conversion layer according to the prediction results; Using the softmax value σ(L i ) to guide the prediction result P i k , and set a threshold Conf to ignore samples with low confidence to alleviate the impact of noise labels: IT i =CE(σ(L i ),P i ); Where: CE is the standard cross entropy loss function.

7. The pluggable forged face detection method based on test time domain adaptation according to claim 1 is characterized in that: The specific scheme of step 4 includes the following sub-steps: Step 4.1: Choose N based on feature distance f A sample of nearest neighbors: Where: dis(·) is the distance function, β(N f ) is the deep feature F of the test sample i and Nth f The distance of the deep features of the nearest neighbor samples; then N f The nearest neighbor prediction result P i (n),n∈{1,...,N f }; Step 4.2: According to the nearest neighbor prediction result, further correct the prediction result of the test sample to obtain the final prediction result:

8. The pluggable forged face detection method based on test time domain adaptation according to claim 7 is characterized in that: In order to further reduce the impact of noise labels on the update of the feature conversion layer, the nearest neighbor consistency constraint is introduced to force the test sample and its corresponding nearest neighbor sample to output similar results:

9. An electronic device, characterized in that: The electronic device comprises: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the method as claimed in any one of claims 1 to 8 above.

10. A non-transitory computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A highly robust deepfake detection method based on adaptive learning

    CN118379608B

  • Depth forgery detection method and system based on unsupervised domain self-adaption

    CN118397440A

  • Deep forgery detection method based on multi-domain fusion

    CN118470585A

  • Depth forgery detection method fusing spatial domain texture difference and frequency domain information

    CN118968269A

  • Method, device and equipment for identifying face image and storage medium

    CN109670491A

Cited By

  • Deep forgery detection model training method, deep forgery detection method and deep forgery detection system

    CN120543952A