Face Forgery Detection Method and System Based on Efficient Fine-Tuning of Visual Pre-Trained Models
By inserting the adapter of the central differential convolution operator in the ConvNext-V2 model for fine-tuning, the demand for a large number of training data and computing resources in the prior art is solved, and efficient face forgery detection under a small number of data sets is achieved.
Patent Information
- Application Number
- CN202410057960.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-01-15
AI Technical Summary
In the case of a small number of training data sets, the prior art needs to retrain the tens of millions of parameters of the entire detection network, resulting in strict requirements in computing resources and data set scales, making it difficult to effectively identify fake face images.
The adapter based on the central differential convolution operator is used to insert the inverse bottleneck module of the initial visual pre-trained model ConvNext-V2, and only fine-tune the parameters of the adapter module, reduce the parameter learning scale, and combine the FaceForensics++ and Celeb-DF datasets for training and verification.
The performance of face forgery detection is significantly improved under a small number of training data sets, the requirements for the number of training data sets are reduced, the requirements for computing resource requirements are reduced, and complex framework structures are designed without relying on human prior experience.
Smart Images

Figure CN117912078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of face forgery identification, and specifically, to a face forgery detection method, system, terminal and medium based on an efficiently fine-tuned visual pre-training model. Background Art
[0002] With the development of deep learning, especially generative network technology, the quality of face forgery images generated based on deep forgery technology has been significantly improved, making it difficult to effectively identify them relying on human vision or traditional recognition techniques. The spread and use of forged face images on social media will have a serious impact on personal information security and even threaten the stable development of the country and society. Therefore, in order to minimize the impact caused by forged images as much as possible, it has become increasingly important to develop the ability to identify forged images.
[0003] In recent years, most forgery identification techniques are designed based on human prior knowledge for the network structure and are supervised and trained with a large number of labeled datasets, enabling the network to have the ability to identify prior forgery features.
[0004] According to different human prior experiences, the identification methods can be roughly divided into three types.
[0005] First, based on the prior knowledge that there are easily differences between the noise statistical features (local noise statistics) of genuine and forged images, Zhou et al. (Zhou, P., Han, X., Morariu, V.I., Davis, L.S.: Two-stream neuralnetworks for tampered face detection. In: CVPR Workshop (2017)) designed a network structure focusing on extracting local noise features to improve forgery identification performance;
[0006] Second, based on the prior knowledge that there are easily differences between the frequency domain features of genuine and forged images, Qian et al. (Qian, Y., Yin, G., Sheng, L., Chen, Z., Shao, J.: Thinking in frequency: Face forgery detectionby mining frequency-aware clues. In: ECCV (2020)) designed a network structure to extract frequency domain features using tools such as the discrete cosine transform (DCT);
[0007] Thirdly, based on the prior knowledge that forged images have spatial domain artifacts, Zhao et al. (Zhao, H., Wei, T., Zhou, W., Zhang, W., Chen, D., Yu, N.: Multi-attentional deepfake detection. In: CVPR (2021)) redesigned the spatial attention module to enhance the model's ability to extract local subtle features.
[0008] Although the above detection network designed based on human prior experience can achieve strong forgery discrimination ability, since it is necessary to redesign the detection network structure according to prior knowledge and train all tens of millions of parameters of the entire detection network from scratch, there are more stringent requirements for the quantity scale of the training dataset and the GPU computing power. In the actual scenario where only a small amount of training dataset is provided, the effect achieved by the above method will be greatly reduced. Summary of the Invention
[0009] Aiming at the defect of retraining all parameters in the network depending on a large number of labeled training datasets in the prior art, the purpose of the present invention is to provide a face forgery detection method, system, terminal and medium based on efficient fine-tuning of a visual pre-trained model.
[0010] According to one aspect of the present invention, there is provided a face forgery detection method based on efficient fine-tuning of a visual pre-trained model, including:
[0011] Obtain the initial visual pre-trained model ConvNext-V2;
[0012] Insert an adapter based on the central difference convolution operator into the inverse bottleneck module of the initial visual pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model;
[0013] Obtain task training data and divide it into a training set and a validation set;
[0014] Use the training set to train the fine-tuned ConvNext-V2 model, and use the validation set to verify whether the training is completed;
[0015] Use the trained fine-tuned ConvNext-V2 model for face forgery detection.
[0016] Preferably, the adapter based on the central difference convolution operator includes parameters Θ, W up , W down , Θ ∈ R 9 represents the parameters of the convolution operator in the adapter, W up ∈ R c’×C represents the parameter matrix, W down ∈ RC×c’ Represents a parameter matrix.
[0017] Preferably, inserting the adapter based on the central difference convolution operator into the inverse bottleneck module of the initial visual pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model includes:
[0018] Obtain the hidden layer feature H in the initial visual pre-trained model ConvNext-V2, H ∈ R H’×W’×C , where H′ and W′ are the height and width, and C is the number of channels;
[0019] Perform a linear transformation on the hidden layer feature H using the parameter matrix W down to reduce the dimension and activate it using a non-linear activation function to obtain H down ∈ R H’×W’×c’ ;
[0020] Perform central difference convolution operation on H down using the central difference convolution operator and the parameter Θ;
[0021] Perform a linear transformation on H up using the parameter matrix W down to increase the dimension and activate it using a non-linear activation function to obtain ΔH ∈ R H’×W’×C ;
[0022] Update the hidden layer feature H according to the formula H ← H + sΔH, where s is a hyperparameter coefficient.
[0023] Preferably, the obtained task training data is the face forgery detection benchmark dataset of FaceForensics++ and Celeb-DF.
[0024] Preferably, training the fine-tuned ConvNext-V2 model using the training set includes:
[0025] Fix all the parameters in the detection network except Θ, W up , W down , and perform iterative learning on the parameters Θ, W up , W down on the training set.
[0026] Preferably, the verification of whether the training is completed using the validation set includes:
[0027] When the value of the negative log-likelihood loss function in the current round on the validation set increases by more than a set value compared to the minimum negative log-likelihood loss function value during the training process, stop the training.
[0028] Preferably, using the trained fine-tuned ConvNext-V2 model for face forgery detection includes:
[0029] Input the face image to be detected into the trained fine-tuned ConvNext-V2 model, and output the true or false attribute of the face image.
[0030] According to the second aspect of the present invention, a face forgery detection system based on an efficiently fine-tuned visual pre-trained model is provided, including:
[0031] A basic module that obtains the initial visual pre-trained model ConvNext-V2;
[0032] A fine-tuning module that inserts an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial visual pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model;
[0033] A data module that obtains task training data and divides it into a training set and a validation set;
[0034] A training module that trains the fine-tuned ConvNext-V2 model using the training set and validates whether the training is completed using the validation set;
[0035] An application module that uses the trained fine-tuned ConvNext-V2 model for face forgery detection.
[0036] According to the third aspect of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can be used to execute any of the above methods or run the above system.
[0037] According to the fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute any of the above methods or run the above system.
[0038] Compared with the prior art, the embodiments of the present invention at least have the following beneficial effects:
[0039] In the face forgery detection method and system based on an efficiently fine-tuned visual pre-trained model in the embodiments of the present invention, an adapter (CDC-Adapter) based on the central difference convolution operator is creatively designed and introduced into the inverted bottleneck module of the ConvNext-V2 model. When fine-tuning the parameters of the ConvNext-V2 model in the face forgery detection scenario, only the parameters of the adapter module (CDC-Adapter) need to be fine-tuned, greatly reducing the scale of parameter learning.
[0040] The face forgery detection method and system based on an efficient fine-tuning visual pre-training model in the embodiments of the present invention have been experimentally verified on benchmark datasets such as FaceForensics++ and Celeb-DF. The results show that it can reduce the requirement for the quantity scale of the training dataset. In the actual scenario where only a small amount of training dataset is provided, this method and system are significantly superior to the existing face forgery detection methods based on prior experience.
[0041] The face forgery detection method and system based on an efficient fine-tuning visual pre-training model in the embodiments of the present invention do not need to rely on human prior experience to design a complex framework structure. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0043] Figure 1 is a flowchart of the face forgery detection method based on an efficient fine-tuning visual pre-training model in an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of the overall structure of the adapter (CDC-Adapter) based on the central difference convolution operator in a preferred embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of the data flow of the adapter (CDC-Adapter) based on the central difference convolution operator in a preferred embodiment of the present invention;
[0046] Figure 4 is a structural framework diagram of the face forgery detection system based on an efficient fine-tuning visual pre-training model in an embodiment of the present invention;
[0047] Figure 5 is the experimental result of the face forgery detection method based on an efficient fine-tuning visual pre-training model in a specific embodiment of the present invention under different training dataset quantity scales. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0049] The present invention provides an embodiment, a face forgery detection method based on an efficient fine-tuning visual pre-training model, asFigure 1 As shown, the main steps are as follows:
[0050] S100, obtain the initial vision pre-trained model ConvNext-V2;
[0051] S200, insert an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial vision pre-trained model ConvNext-V2 obtained in S100 to obtain a fine-tuned ConvNext-V2 model;
[0052] S300, obtain task training data and divide it into a training set and a validation set;
[0053] S400, use the training set in S300 to train the fine-tuned ConvNext-V2 model in S200, and use the validation set in S300 to verify whether the training is completed;
[0054] S500, use the trained fine-tuned ConvNext-V2 model for face forgery detection.
[0055] In the above embodiment, a method based on efficiently fine-tuning a vision pre-trained model is used to solve face forgery detection. Only the ConvNext-V2 model needs to be fine-tuned, and there is no need to rely on human prior experience to design a complex framework structure.
[0056] On the basis of considering the model performance, the embodiment of the present invention selects the pre-trained ConvNext-V2 as the basic model. This basic model integrates the fully convolutional MAE framework and the global response normalization layer GRN, further strengthening the channel-wise feature economy in the ConvNext architecture. This design significantly improves the performance of pure ConvNext on various recognition benchmarks. Although its convolutional architecture is relatively simple, its performance is not inferior to that of Transformer. This basic model not only ensures high recognition performance, but also helps to reduce the scale of parameter learning, thus achieving more efficient model training and lower computational costs.
[0057] The basic model contains an inverted bottleneck module, which can partially reduce the parameter scale of the model while slightly improving the accuracy, and improve the overall performance of the model.
[0058] In order to further reduce the scale of parameter learning, in a preferred embodiment of the present invention, a preferred structure of the adapter based on the central difference convolution operator is provided, as Figure 2 and Figure 3 shown, the CDC-Adapter module contains parameters Θ, W up , W down , Θ ∈ R 9Represents the parameters of the 3×3 convolutional operator in the CDC-Adapter, W up ∈R c’×C Represents the parameter matrix, W down ∈R C ×c’ Represents the parameter matrix. When the initial vision pre-trained model ConvNext-V2 passes through the adapter, it first undergoes downsampling, activation function, differential center convolutional operator, activation function, and upsampling.
[0059] Based on the adapter of the central difference convolutional operator in the above embodiments, fine-tune the initial vision pre-trained model ConvNext-V2. In a preferred embodiment, S200 is implemented, and the adapter based on the central difference convolutional operator is inserted into the inverted bottleneck module of the initial vision pre-trained model. The design feature of the inverted bottleneck module is small in the middle and large at both ends, which can effectively avoid data loss. Further, as Figure 2 and Figure 3 shown, the main steps of S200 are:
[0060] S201, Obtain the hidden layer feature H in ConvNext-V2, H ∈ R H’×W’×C , where H′, W′ are the height and width, and C is the number of channels;
[0061] S202, Use the parameter matrix W down to perform linear transformation and dimensionality reduction on the hidden layer feature H, and use a non-linear activation function for activation to obtain H down ∈R H’×W’×c’ ;
[0062] S203, Use the central difference convolutional operator and parameter Θ to perform central difference convolution operation on H down ;
[0063] S204, Use the parameter matrix W up to perform linear transformation and dimensionality increase on H down and use a non-linear activation function for activation to obtain ΔH ∈ R H’×W’×C ;
[0064] S205, According to the formula H ← H + sΔH, update the hidden layer feature H, where s is the hyperparameter coefficient.
[0065] Further, in S203, it includes two convolutional layers. One convolutional layer is for H obtained in S202 downPerform a convolution operation with a convolution operator of size 3×3 and parameters 1 - Θ; another convolutional layer performs a convolution operation on the features that have completed the central difference, with a convolution operator of size 3×3 and parameters Θ, and finally add the results of the two convolutional layers. In this embodiment, a fine-tuned ConvNext-V2 model is obtained. When the model is fine-tuned for downstream tasks, the model carries the knowledge of the downstream tasks. Therefore, in the subsequent training process, only the parameters of the adapter module (CDC-Adapter) need to be fine-tuned, greatly reducing the scale of parameter learning.
[0066] In a preferred embodiment of the present invention, S300 is implemented to obtain task training data and divide it into a training set and a validation set. The task training data is a face forgery detection benchmark dataset such as FaceForensics++ and Celeb-DF.
[0067] In a preferred embodiment of the present invention, S400 is implemented to train the fine-tuned ConvNext-V2 model using the training set and verify whether the training is completed using the validation set. Specifically,
[0068] Fix all the parameters in the detection network except Θ, W of the adapter (CDC-Adapter) based on the central difference convolution operator up 、W down Outside, perform iterative learning on the parameters Θ, W up 、W down on the training set, and when the value of the negative log-likelihood loss function in the current round on the validation set increases by more than a set value compared to the minimum negative log-likelihood loss function value during the training process, stop the training to obtain the determined parameters of the face forgery detection network. In some specific embodiments, the set value can be set to 3%.
[0069] In this embodiment, when fine-tuning the parameters of the ConvNext-V2 model in the face forgery detection scenario, only the parameters of the adapter module (CDC-Adapter) need to be fine-tuned, greatly reducing the scale of parameter learning.
[0070] In a preferred embodiment of the present invention, 500 is implemented to use the trained fine-tuned ConvNext-V2 model for face forgery detection. Specifically, the input is the face image to be detected, and the output is the true or false attribute of the input image determined by the detection model
[0071] Based on the same inventive concept, in other embodiments of the present invention, a face forgery detection system based on an efficiently fine-tuned visual pre-trained model is provided, such as Figure 4As shown in the figure, it includes a basic module, a fine-tuning module, a data module, a training module, and an application module. The basic module obtains the initial visual pre-training model ConvNext-V2; the fine-tuning module inserts an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial visual pre-training model to obtain the fine-tuned ConvNext-V2 model; the data module obtains task training data and divides it into a training set and a validation set; the training module uses the training set to train the fine-tuned ConvNext-V2 model and uses the validation set to verify whether the training has been completed; the application module uses the trained fine-tuned ConvNext-V2 model for face forgery detection.
[0072] In the above examples of the present invention, the specific implementation techniques of each module / unit can refer to the corresponding steps of the face forgery detection method based on the efficient fine-tuning of the visual pre-training model in the above embodiments, which will not be elaborated here.
[0073] Based on the same inventive concept, in other embodiments of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can be used to execute any one of the face forgery detection methods based on the efficient fine-tuning of the visual pre-training model, or to run the face forgery detection system based on the efficient fine-tuning of the visual pre-training model.
[0074] Based on the same inventive concept, in other embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used to execute any one of the face forgery detection methods based on the efficient fine-tuning of the visual pre-training model, or to run the face forgery detection system based on the efficient fine-tuning of the visual pre-training model.
[0075] To verify the performance and effect of the face forgery detection method based on the efficient fine-tuning of the visual pre-training model in the embodiments of the present invention, in a specific embodiment, the following three groups of experiments are carried out:
[0076] Experimental Example 1
[0077] Experimental verification is carried out on face forgery detection benchmark datasets such as FaceForensics++ and Celeb-DF in the scenario of few training data to verify the performance of the method in the embodiments of the present invention in the scenario of few training data.
[0078] The experimental results are as Figure 5 shown, where Figure 5 (The first one on the left) is the experimental result under FaceForensics++ (C23, low compression scenario), Figure 5 (The second one on the left) is the experimental result under FaceForensics++ (C40, high compression scenario), Figure 5(The first one on the right) shows the experimental results under Celeb-DF. With only 1 / 256 to 1 / 32 of the entire training set, the method of the embodiment of the present invention has improved the performance by at least 3%-10% compared with the existing methods. The method of the embodiment of the present invention can achieve performance close to that under the entire training set by using only 1 / 32 of the training set. This result verifies that the method in the embodiment of the present invention, by fine-tuning the visual pre-trained model, utilizes the pre-trained knowledge of the fine-tuned visual pre-trained model, and greatly reduces the dependence on the scale of the training data set.
[0079] Experimental Example 2
[0080] The method in the above embodiment was experimentally verified with the existing methods on face forgery detection benchmark data sets such as FaceForensics++ and Celeb-DF under the entire training set.
[0081] 1. Comparison with existing face forgery detection methods:
[0082] Table 1 shows the experimental results of the comparison with existing face forgery detection methods in the scenario of the entire training set
[0083]
[0084] According to the experimental results in Table 1, the number of parameters to be trained by the method of the embodiment of the present invention is only about 380,000. Compared with the training parameter number of more than 20 million of other existing methods, the method in the embodiment of the present invention greatly reduces the training parameter number, thereby reducing the requirement for the GPU computing power.
[0085] 2. Comparison with other efficient fine-tuning methods:
[0086] Table 2 shows the experimental results of the comparison with other efficient fine-tuning methods in the scenario of the entire training set
[0087]
[0088] According to the experimental results in Table 2, the existing methods available for efficiently fine-tuning the visual pre-trained model are not applicable to the special application scenario of face forgery detection, while the method of the above embodiment ensures the face forgery detection performance while meeting the requirement of fine-tuning with extremely few parameters. The method in the above embodiment leads the other existing efficient fine-tuning methods by at least 10% in performance, and even leads the full-parameter fine-tuning (FT) under multiple benchmark data sets.
[0089] Experimental Example 3
[0090] The method of the embodiment of the present invention was applied to other types of visual pre-trained models to verify the applicability of the method.
[0091] Table 3 shows the experimental results of applying the method of the embodiments of the present invention to various types of visual pre-training models
[0092]
[0093] According to the experimental results in Table 3, the method in the embodiments of the present invention is applicable to visual pre-training models with different structures and different pre-training strategies, and can achieve performance equivalent to that of the full-parameter fine-tuning method. The adapter based on the central difference convolution operator (CDC-Adapter) proposed in the embodiments of the present invention can be applied to both the visual attention mechanism network (Transformer) structure and the convolutional network (ConvNets), and is also compatible with different pre-training strategies such as supervised learning and self-supervised learning.
[0094] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which do not affect the essence of the present invention. The above preferred features can be combined arbitrarily without conflict.
Claims
1. A face forgery detection method based on an efficiently fine-tuned visual pre-trained model, characterized in that, Including: Obtain the initial vision pre-trained model ConvNext-V2; Insert an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial vision pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model; Obtain task training data and divide it into a training set and a validation set; Use the training set to train the fine-tuned ConvNext-V2 model, and use the validation set to verify whether the training is completed; Use the trained fine-tuned ConvNext-V2 model for face forgery detection; The adapter based on the central difference convolution operator includes parameters Θ, W up , W down , Θ ∈ R 9 represents the parameters of the convolution operator in the adapter, W up ∈ R c'×C represents the parameter matrix, W down ∈ R C×c' represents the parameter matrix; The step of inserting an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial vision pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model includes: Obtain the hidden layer feature H in the initial vision pre-training model ConvNext-V2, where H ∈ R H'×W'×C , H' and W' are the height and width, and C is the number of channels; Using the parameter matrix W down Perform a linear transformation to reduce the dimension of the hidden layer feature H and activate it using a non-linear activation function to obtain H down ∈R H'×W'×c' ; Perform central difference convolution operation on H using the central difference convolution operator and parameter Θ down ; Using the parameter matrix W up for H down perform a linear transformation to increase the dimension and activate it using a non-linear activation function to obtain ΔH ∈ R H'×W'×C ; Update the hidden layer feature H according to the formula H←H + s△H, where s is a hyperparameter coefficient.
2. The face forgery detection method based on an efficiently fine-tuned visual pre-trained model according to claim 1, characterized in that, The obtained task training data is the face forgery detection benchmark dataset of FaceForensics++ and Celeb-DF.
3. A face forgery detection method based on an efficiently fine-tuned visual pre-trained model according to claim 1, characterized in that Using the training set to train the fine-tuned ConvNext-V2 model includes: Fix all the parameters in the fine-tuned ConvNext-V2 model except Θ, W up , W down except Θ, W up , W down and perform iterative learning on the training set for the parameters Θ, W 4. The face forgery detection method based on an efficiently fine-tuned visual pre-trained model according to claim 3, characterized in that, The step of using the validation set to verify whether the training is completed includes: When the value of the negative log-likelihood loss function in the current round on the validation set increases by more than a set value compared to the minimum negative log-likelihood loss function value during the training process, stop the training.
5. The face forgery detection method based on an efficiently fine-tuned visual pre-trained model according to claim 1, characterized in that, The step of using the trained fine-tuned ConvNext-V2 model for face forgery detection includes: Input the face image to be detected into the trained fine-tuned ConvNext-V2 model, and output the true or false attribute of the face image.
6. A face forgery detection system based on an efficiently fine-tuned visual pre-trained model, characterized in that, Including: A basic module that obtains the initial vision pre-trained model ConvNext-V2; A fine-tuning module that inserts an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial vision pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model; A data module that obtains task training data and divides it into a training set and a validation set; A training module that uses the training set to train the fine-tuned ConvNext-V2 model and uses the validation set to verify whether the training is completed; An application module that uses the trained fine-tuned ConvNext-V2 model for face forgery detection; Among them, the adapter based on the central difference convolution operator includes parameters Θ and W up and W down , Θ ∈ R 9 represents the parameters of the convolution operator in the adapter, and W up ∈ R c'×C represents the parameter matrix, and W down ∈ R C×c' represents the parameter matrix; The step of inserting an adapter based on the central difference convolution operator into the inverted bottleneck module of the initial vision pre-trained model ConvNext-V2 to obtain a fine-tuned ConvNext-V2 model includes: Obtain the hidden layer feature H in the initial vision pre-training model ConvNext-V2, where H ∈ R H'×W'×C , H' and W' are the height and width, and C is the number of channels; Using the parameter matrix W down Perform a linear transformation on the hidden layer feature H to reduce its dimension, and use a non-linear activation function to activate it to obtain H down ∈R H'×W'×c' ; Perform central difference convolution operation on H using the central difference convolution operator and parameter Θ down ; Using the parameter matrix W up for H down perform a linear transformation to increase the dimension and use a non-linear activation function for activation to obtain △H ∈ R H'×W'×C ; Update the hidden layer feature H according to the formula H←H + s△H, where s is a hyperparameter coefficient.
7. A terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it can be used to execute the method described in any one of claims 1-5, or, run the system described in claim 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it can be used to execute the method described in any one of claims 1-5, or, run the system described in claim 6.
Citation Information
Patent Citations
Remote sensing image change detection method
CN116310851A
General deep forgery detection method based on generative model
CN117238015A