A Method for Tracing Deepfake Video Technology Based on Image Frequency Domain Information
By introducing image frequency domain information into the traceability of deep forgery video technology, and integrating image features and frequency domain features, the problem of low traceability accuracy in the existing technology is solved, and higher traceability accuracy and classification capabilities of deep forgery technology are achieved.
Patent Information
- Application Number
- CN202210586229.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-05-27
AI Technical Summary
The prior art has low accuracy in the traceability of deep fake video technology. Different forgery methods in a single original image are similar, making it difficult to effectively distinguish and trace the information.
The deep fake video technology traceability method based on image frequency domain information is adopted, and the original image information is supplemented through frequency domain information. The image features and frequency domain features are fused through the fusion method to obtain the fusion features, which are used for the deep fake technology traceability model to classify different forgery methods.
The accuracy of traceability of deep forgery technology has been improved, the shortcomings of manual features and deep learning methods are overcome, and the model's ability to classify different forgery technologies is enhanced.
Smart Images

Figure CN115188039B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for tracing the source of deepfake videos based on image frequency-domain information, belonging to the fields of deep learning and computer vision. Background Art
[0002] In recent years, computer vision technology and deep neural network technology have developed rapidly. Especially the development of generative adversarial networks (GANs) and variational autoencoders (VAEs) in neural network models has achieved amazing results in image and video generation. In 2017, a foreign forum user used a generative adversarial network (GAN) to forge a realistic video and posted it on the Internet, and thus this technology was called deepfake technology.
[0003] Specifically, deepfake technology mainly forges or edits the face part. Existing deepfake technologies can be mainly divided into four categories: reproduction, replacement, editing, and generation. Reproduction is to use the behavior of the original face to drive the target face so that the behavior of the target face is the same as that of the original face. Replacement means replacing the target face with the original face. Editing is to change the attributes of the target face. For example, changing the age, gender, skin color, etc. of the target face. Generation is to create a complete face that does not exist in reality through a generative adversarial network (GAN).
[0004] In the early stage when deepfake technology was proposed, making a deepfake video required the producer to have relevant professional knowledge and a large amount of computing resources. However, with the development of deepfake technology, some easy-to-use mobile or computer software has emerged on the Internet, enabling ordinary people without relevant professional knowledge and computing resources to easily make high-quality deepfake videos using a computer or mobile phone. Moreover, due to the lack of effective screening and review mechanisms, there are currently a large number of deepfake videos on the Internet. Some well-produced forged videos cannot be accurately identified by professionals, and it is even more difficult for ordinary people to distinguish the authenticity of the videos, and they are more likely to be misled and harmed by the forged videos. In major events or sensitive issues, deepfake videos may cause serious adverse effects. Therefore, tracing the source of deepfake videos and accurately identifying their production technology or software can help staff block the spread of forged videos from the source and avoid adverse effects on society.
[0005] There is little research on the traceability of deepfake technology. The current methods mainly use manual features (such as co-occurrence matrices) or deep learning models to extract features for technology traceability. Only using manual feature extraction for technology traceability, the extracted features are fixed and often cannot fully utilize the forgery information in deepfake images. Deep learning models tend to learn the high-level semantic information in images. The high-level semantic information (such as face shape, face size, etc.) of the forged faces generated by different deepfake methods is extremely similar. Therefore, only using deep learning models for technology traceability of deepfakes has unsatisfactory results. In the upsampling process of deep convolutional networks, checkerboard artifacts will inevitably be left in the images, and these checkerboard artifacts will cause changes in the high-frequency information of the images. Different forgery methods use different model structures and training parameters, and the generated checkerboard artifacts are also different, leaving more obvious differences in the forgery traces in the frequency domain.
[0006] Therefore, in the current existing technology, the forgery information of different forgery methods in a single original image is similar, resulting in a low traceability accuracy. Summary of the Invention
[0007] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art and provide a method for tracing deepfake video technology based on image frequency domain information, which uses frequency domain information to supplement the original image information, fuses image features and frequency domain features through a fusion method to obtain fusion features, and is used for the deepfake technology traceability model to classify different forgery methods. Compared with the manual feature method and the deep learning method alone, the accuracy of deepfake technology traceability is greatly improved.
[0008] The technical solution adopted by the present invention: A method for tracing deepfake video technology based on image frequency domain information, comprising the following steps:
[0009] Step 1: Decompose the input deepfake video into video frames and extract frames to obtain the extracted video frames;
[0010] Step 2: Apply the RetinaFace model to the video frames extracted in Step 1 for face detection. If there is a face in the frame image of the video frame, obtain the face key point coordinates in the frame image, perform affine transformation on the face key point coordinates in the frame image to align and scale them with the standard face key point coordinates, and then crop the aligned and scaled face area to obtain an RGB face image;
[0011] Step 3: Convert the RGB face image obtained by cropping in Step 2 into a grayscale image, and then use the discrete cosine Fourier transform DCT to obtain the frequency-domain amplitude image corresponding to the cropped RGB face image; use the frequency-domain cropping algorithm to crop the low-frequency part in the frequency-domain amplitude image, only retain the high-frequency part in the frequency-domain amplitude image, and finally perform the inverse discrete cosine Fourier transform on the cropped frequency-domain image to obtain the high-frequency frequency-domain feature of the RGB face image;
[0012] Step 4: Concatenate the RGB face image obtained in Step 2 and the high-frequency frequency-domain feature obtained in Step 3 along the channel direction to obtain a 4-channel concatenated feature, and then perform information exchange and fusion on the 4-channel concatenated feature through a convolutional layer with a kernel size of 1×1 in the channel direction to obtain a 4-channel frequency-domain fusion feature;
[0013] Step 5: Use the Xception deep convolutional network as the backbone network, take the frequency-domain fusion feature obtained in Step 4 as the input, and finally output a one-dimensional forgery trace feature, which is used for the final feature classification;
[0014] Step 6: Pass the one-dimensional forgery trace feature obtained in Step 5 through a multi-classification system, that is, composed of multi-class fully connected layers, and the output of each class corresponds to a deep forgery technique, to obtain the probability that the RGB face image belongs to each deep forgery technique. Finally, average and fuse the output results of the RGB face images from the same video to obtain the traceability result of the deep forgery technique of the input deep forgery video.
[0015] In the said Step 1, decomposing the input deep forgery video into video frames and extracting frames to obtain the extracted video frames is specifically as follows: Decompose the input deep forgery video into single-frame images. For video frames with a frame number not less than 60, uniformly extract 60 frame images, and for video frames with a frame number less than 60, extract all the video frames.
[0016] In the said Step 3, obtaining the high-frequency frequency-domain feature of the RGB face image is specifically as follows:
[0017] Use the frequency-domain cropping algorithm to crop the low-frequency part in the frequency-domain amplitude image, and the cropped frequency-domain image P C , and the calculation formula is as follows:
[0018] P C =F(P B )
[0019] F is the cropping algorithm, which sets the values in the upper left corner area of the frequency-domain amplitude image P B to 0, where the range of the upper left corner area is based on P BAn isosceles right triangle with a right-angled side length equal to one-third of the side length. The area within this triangle is the low frequency of the frequency domain amplitude image.
[0020] The specific cropping algorithm F is as follows:
[0021] First, construct the cropping occlusion, and the calculation formula is as follows:
[0022]
[0023] where H is the cropping occlusion, and H i,j is the feature point value corresponding to the coordinate (i, j) in the cropping occlusion, and is the side length of the frequency domain amplitude image P B ;
[0024] Then multiply the cropping occlusion H and the frequency domain amplitude image P B point by point to obtain the high-frequency frequency domain amplitude image P C , that is, P C = F(P B );
[0025] Finally, perform the inverse discrete cosine Fourier transform on the obtained high-frequency frequency domain amplitude image P C to obtain the high-frequency frequency domain features P D of the RGB face image.
[0026] In step 4, the 4-channel frequency domain fusion feature is P E , and the formula is as follows:
[0027] P E = R(B(Conv 1×1 (Cat(P A , P D ))))
[0028] where B is the batch normalization layer Batch Normal, and R is the ReLU activation function; P A is the RGB face image.
[0029] In step 5, the Xception deep convolutional network is used as the backbone network to extract one-dimensional forgery trace features, specifically as follows:
[0030] Change the input of the original Xception deep convolutional network to 299×299×4 to adapt to the size of the frequency domain fusion feature in step 4; use the frequency domain fusion feature obtained in step 4 as the input of the modified Xception deep convolutional network; output one-dimensional forgery trace features with 2048 channels.
[0031] The advantages and effects of the present invention compared with the prior art are as follows:
[0032] (1) While extracting the features of the original RGB image, the present invention introduces frequency domain features as supplementary features, which can not only extract the forgery traces in the RGB image, but also obtain the forgery features in the frequency domain. By using the above two features, a classification model with excellent performance can be obtained for the technical traceability of deepfake videos. By combining the image information and its frequency domain information for the technical traceability of deepfake technology, the flexibility and accuracy of traceability are improved.
[0033] (2) Compared with the method using manual features, the present invention uses a convolutional neural network to extract features, which improves the flexibility of feature extraction. Compared with the method using only a deep learning model, the introduction of frequency domain information improves the classification ability of the model for different forgery techniques.
[0034] (3) The present invention overcomes the problem in the existing research technology of lacking the discrimination and traceability of forgery methods. A multi-classification system is used to classify the forgery techniques of videos, which helps relevant personnel to locate the video source faster, block its dissemination process, and reduce the impact of malicious face forgery videos on society. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the implementation flowchart of the method of the present invention;
[0036] Figure 2 is the schematic diagram of the frequency domain cropping algorithm in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0038] As Figure 1 shown, the method of the present invention is divided into three parts: image preprocessing, image feature extraction, and feature classification, and the specific implementation steps are as follows:
[0039] Image preprocessing:
[0040] Step 1: Extract frames from the original video
[0041] Videos on the Internet often reach more than a thousand frames. If each frame of the video is detected, the time and computational resource overhead are unbearable. Therefore, in this invention, first, the OpenCV computer vision software library is used to decompose the video into video frames. Then, for each video with more than 60 frames, 60 frames of images are extracted, and for videos with less than 60 frames, all video frames are retained for the detection of deepfake technology traceability, that is, as the input images of the traceability model.
[0042] Step 2: Face detection and cropping
[0043] Most deepfake videos modify or forge human faces, and the forgery traces are mainly concentrated in the face area. Moreover, there may be cases where there is no face or the face area accounts for a small proportion in some video frame images. These useless background information will affect the model's extraction of forgery trace features and thus affect the technical traceability performance of the model. Therefore, in order to avoid the interference of background information on traceability, it is necessary to perform face detection and cropping on video frames. Moreover, the faces in video frames may have different angles and poses. In order to make the model focus on the forgery traces on the face rather than the pose and angle of the face, it is necessary to align the detected faces to ensure that the faces are in the same position and size in the image. Therefore, in the present invention, first, the RetinaFace face detection algorithm is used to detect the key points I of the face in the video frame image A =[x 1 ,y 1 ,x 2 ,y 2 ,x 3 ,y 3 ,x 4 ,y 4 ,x 5 ,y 5 , and the face is aligned to the standard face key point I B by using affine transformation to obtain the aligned face image P A .
[0044] Image feature extraction:
[0045] Step three: Calculate the frequency-domain image of the face image
[0046] The frequency information of an image represents the change rate of the gray value of the image at spatial points and is the gradient of the gray level in the plane space. First, the gray image of the original image is obtained, and then the gray image is used for calculation to obtain its frequency-domain information. The formula is as follows:
[0047] P B =D(G(P A ))
[0048] where G is the gray-scale transformation that transforms the original image P A into a gray image. D is the discrete cosine transform (DCT) that transforms the gray image into a frequency-domain amplitude image. Its center represents the low-frequency information of the image, and the periphery represents the high-frequency information of the image.
[0049] When generating forged images, deepfake technologies all need to go through an upsampling stage, and the upsampling processes of different technologies are different. Therefore, different forgery technologies will leave different checkerboard artifacts on the images. These checkerboard artifacts change violently and the patterns repeat in the image space, so they will leave forgery traces in the high-frequency region of the frequency-domain image. To make the model focus on the forgery traces in its high-frequency information, the method of the present invention crops the low-frequency information, and the formula is as follows:
[0050] P C =F(P B )
[0051] F is a cropping algorithm that sets the values in the upper-left corner area of the frequency-domain image P B to 0. Among them, the range of the upper-left corner area is an isosceles right triangle with a right-angled side length of 1 / 3 of the side length of P B . The area within this triangle is the low-frequency and mid-frequency parts of the frequency-domain image.
[0052] As Figure 2 shown, the specific cropping algorithm is as follows:
[0053] First, construct a cropping mask, and the calculation formula is as follows:
[0054]
[0055] where H is the cropping mask, and H i,j is the value of the feature point corresponding to the coordinates (i, j) in the cropping mask, and is the side length of the frequency-domain amplitude image P B ;
[0056] Then multiply the cropping mask H point by point with the frequency-domain amplitude image P B to obtain the high-frequency frequency-domain amplitude image P C .
[0057] Since the convolutional neural network cannot directly process the frequency-domain image, finally, perform the inverse discrete cosine transform on P C to obtain the face frequency-domain feature P D . The overall formula process of this step is as follows:
[0058] P D =D -1 (P C )
[0059] Step Four: Combine the RGB original image information and the frequency-domain information
[0060] To utilize both the forgery information in the original image and the forgery information in the frequency-domain image, the original image and the frequency-domain image are concatenated along the channel direction to obtain a 4-channel concatenated feature, and then the two types of information are further fused through a convolutional layer with a kernel size of 1*1 to obtain a 4-channel fused feature P E , and the formula is as follows:
[0061] P E = R(B(Conv 1×1 (Cat(P A , P D ))))
[0062] where B is the batch normalization layer (Batch Normal), and R is the ReLU activation function.
[0063] Step Five: Extract forgery trace features
[0064] Use the deep convolutional network Xception as the backbone network to extract forgery trace features. The input size of the original Xception network is 299×299×3. Since the frequency-domain features are fused in the present invention and it has 4 channels, the input of the original network is changed to 299×299×4. The finally output forgery trace feature is a one-dimensional feature vector with 2048 channels.
[0065] Feature classification:
[0066] Step Six: Classify using the extracted features
[0067] Then, the present invention adopts a multi-classification system to classify the features output in Step Five, where the output of each category corresponds to a deep forgery technique. This classification system includes a multi-classification fully connected layer. Its input feature dimension is 2048, and the output feature dimension is the number of technology types n to be traced. Finally, the output features of the multi-classification fully connected layer pass through the Softmax layer, and its output is that the sum of n probabilities is 1, indicating the probability that the video frame is forged using each technology.
[0068] To obtain the overall technology tracing result of the video, the present invention finally averages the detection results belonging to the same video to obtain the probability that the video is forged using each technology.
[0069] The present invention can be applied to the tracing of deep forgery technology in Internet videos in real scenarios, and the tracing classification effect is accurate, which can help relevant personnel accurately locate the video technology method.
[0070] In summary, the present invention uses a method for tracing deep forgery video technology based on the fusion of the frequency domain and the original image, overcomes the problem of poor tracing effect using only the original image, and improves the accuracy of tracing deep forgery videos.
[0071] The parts not described in detail in the present invention belong to the well-known technology in the art.
[0072] Although the specific implementation methods of the present invention are described above, those skilled in the art should understand that these are only examples. Without departing from the principles and implementation of the present invention, various changes or modifications can be made to these implementation schemes. Therefore, the protection scope of the present invention is defined by the appended claims.
Claims
1. A method for tracing deepfake video technology based on image frequency domain information, characterized in that, it includes the following steps: Step 1: Decompose the input deepfake video into video frames and extract frames to obtain the extracted video frames; Step 2: Apply the RetinaFace model to the video frames extracted in Step 1 for face detection. If there is a face in the frame image of the video frame, obtain the face key point coordinates in the frame image, perform affine transformation on the face key point coordinates in the frame image to align and scale them with the standard face key point coordinates, and then crop the face area after alignment and scaling to obtain an RGB face image; Step 3: Convert the RGB face image cropped in Step 2 into a grayscale image, and then use the discrete cosine Fourier transform DCT to obtain the frequency domain amplitude image corresponding to the cropped RGB face image; Use the frequency domain cropping algorithm to crop the low-frequency part in the frequency domain amplitude image, only retain the high-frequency part in the frequency domain amplitude image, and finally perform the inverse discrete cosine Fourier transform on the cropped frequency domain image to obtain the high-frequency frequency domain feature of the RGB face image; Step 4: Concatenate the RGB face image obtained in Step 2 and the high-frequency frequency domain feature obtained in Step 3 along the channel direction to obtain a 4-channel concatenated feature, and then perform information exchange and fusion on the 4-channel concatenated feature through a convolutional layer with a convolution kernel size of 1×1 in the channel direction to obtain a 4-channel frequency domain fusion feature; Step 5: Use the Xception deep convolutional network as the backbone network, take the frequency domain fusion feature obtained in Step 4 as the input, and finally output a one-dimensional forgery trace feature, which is used for the final feature classification; Step 6: Pass the one-dimensional forgery trace feature obtained in Step 5 through a multi-classification system, that is, composed of multi-classification fully connected layers, and the output of each category corresponds to a deepfake technology, obtain the probability that the RGB face image belongs to each deepfake technology, and finally average and fuse the output results of the RGB face images from the same video to obtain the tracing result of the deepfake technology of the input deepfake video; In the said Step 3, to obtain the high-frequency frequency domain feature of the RGB face image, specifically as follows: The low-frequency part in the frequency-domain amplitude image is cropped using the frequency-domain cropping algorithm, and the cropped frequency-domain image , and the calculation formula is as follows: For the cropping algorithm, set the values of the upper left corner region of the frequency domain amplitude image to 0. Among them, the range of the upper left corner region is an isosceles right triangle with a right side length of 1 / 3 of the side length. The region within this triangle is the low frequency of the frequency domain amplitude image; The described cropping algorithm Specifically as follows: First, construct a cropping occlusion, and the calculation formula is as follows: Among them, is for cropping and occlusion, is the numerical value of the feature point corresponding to the coordinates in the cropping and occlusion, w is the side length of the frequency domain amplitude image ; of the side length; Then perform cropping and occlusion with the frequency-domain amplitude image point by point to obtain the high-frequency frequency-domain amplitude image , that is ; Finally, the obtained high-frequency frequency-domain amplitude image , is subjected to inverse discrete cosine Fourier transform, and thus the high-frequency frequency-domain features of the RGB face image are obtained ; In the said step 4, the frequency-domain fusion feature of the 4 channels is , and the formula is as follows: Among them, is the batch normalization layer Batch Normal, is the ReLU activation function; is the RGB face image.
2. The method for tracing deepfake video technology based on image frequency domain information according to claim 1, characterized in that: In the said Step 1, decomposing the input deepfake video into video frames and extracting frames to obtain the extracted video frames, specifically as follows: Decompose the input deepfake video into single-frame images. For video frames with no less than 60 frames, uniformly extract 60 frame images, and for video frames with less than 60 frames, extract all the video frames.
3. The method for tracing deepfake video technology based on image frequency domain information according to claim 1, characterized in that: In the said Step 5, using the Xception deep convolutional network as the backbone network to extract a one-dimensional forgery trace feature, specifically as follows: Change the input of the original Xception deep convolutional network to 299×299×4 to adapt to the frequency-domain fusion feature size in Step 4; use the frequency-domain fusion feature obtained in Step 4 as the input of the modified Xception deep convolutional network; output a one-dimensional forged trace feature with 2048 channels.
Citation Information
Patent Citations
Deep fake face reverse traceability method and system
CN114093013A
Program encoding and counterfeit tracking system and method
US20060015464A1