Lightweight deep forgery detection system
By designing a lightweight deep forgery detection system, using the pyramid expansion convolution module and feature shuffling feature fusion device, combined with the KAN network classifier, the shortcomings in the existing lightweight detection network in terms of accuracy and generalization capabilities are solved, and efficient forgery detection is achieved.
Patent Information
- Application Number
- CN202510189319.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing lightweight deep forgery detection networks have poor performance in detection accuracy and generalization capabilities, and it is difficult to balance low computing requirements, high precision and strong generalization capabilities.
A lightweight deep forgery detection system is designed, including backbone network, pyramid expansion convolution module, feature shuffle feature fusion device and classifier. The pyramid expansion convolution module extracts local features through global average pooling and multi-scale expansion convolution. The feature shuffle feature fusion device can separate convolution and feature shuffle efficient fusion features through depth, and the classifier uses KAN network for nonlinear fitting.
It realizes that the detection accuracy and generalization ability are improved when the computing requirements are low, solves the redundancy problem between global and local feature extraction, and improves the fusion efficiency of feature information.
Smart Images

Figure CN120123728A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of forgery detection, specifically a lightweight deepfake detection system. Background Art
[0002] Deepfake detection aims to identify forged videos generated by deep learning techniques. With the rapid development of deep learning techniques, deepfake techniques for generating forged face videos have also made great progress. These techniques have brought many benefits in aspects such as cultural exchange, entertainment, and film and television production, but their abuse has also caused serious problems, such as information tampering, privacy leakage, and public opinion manipulation, posing a threat to social stability. Therefore, developing deepfake detection techniques is crucial for protecting information authenticity and ensuring social security.
[0003] In contrast, some researchers adopt frame-based methods to detect forgeries by analyzing individual frames extracted from videos. Such methods usually utilize image forensics, tampering trace analysis, or face detail extraction to capture local tampering features, which can effectively reduce the demand for computing resources but rely on large-scale training data. In addition, a large amount of redundant information is often generated when extracting local and global features. Since these methods usually only focus on isolated global or local features and ignore the deep relationships between feature information, their generalization ability in cross-dataset detection tasks is limited.
[0004] Currently, the research on lightweight deepfake detection networks is relatively limited. The detection accuracy and generalization ability of the existing lightweight deepfake detection networks are not outstanding. Therefore, how to balance lightweight design, high accuracy, and robust generalization remains a major challenge in current research. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to overcome the above technical defects.
[0006] To solve the above problems, the technical solution of the present invention is: a lightweight deepfake detection system, the detection system includes a backbone network, a pyramid dilated convolution module, a feature shuffle feature fuser, and a classifier;
[0007] The backbone network is connected to the pyramid dilated convolution module, and both are spliced with the feature shuffle feature fuser. The feature shuffle feature fuser is connected to the classifier, and an input end is connected to the backbone network;
[0008] The pyramid dilated convolution module includes global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7). A 1x1 convolution is connected to the global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7).
[0009] The feature shuffle feature fusion module includes feature shuffle. A 1x1 depthwise separable convolution is connected to the feature shuffle. A 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2) are connected to the 1x1 depthwise separable convolution. A 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2) are connected to the 1x1 depthwise separable convolution. The two 3x3 depthwise separable dilated convolutions (dilation rate is 2) are concatenated with the 1x1 depthwise separable convolution, and a 1x1 convolution is connected to them. A 3x3 convolution, normalization, and an activation function are connected to the 1x1 convolution. The 3x3 convolution is sequentially connected to the normalization and the activation function.
[0010] Furthermore, the backbone network includes backbone module 1, backbone module 2, backbone module 3, backbone module 4, and backbone module 5. Backbone module 1 is connected to the global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7). Backbone module 5 is concatenated with the 1x1 convolution.
[0011] Furthermore, the classifier includes two KAN networks. The two KAN networks are connected, and a fully connected layer is connected to them. A Real channel and a Fake channel are connected to the fully connected layer.
[0012] The advantages of the present invention compared with the existing technologies are as follows:
[0013] The lightweight deepfake detection network proposed by the present invention achieves better detection accuracy and generalization while the computational requirements are significantly lower than the classical methods in the field. The designed pyramid dilated convolution module enhances local features without significantly increasing computational resources, improves detection accuracy, and effectively solves the redundancy between global and local feature extraction.
[0014] The shuffle feature fusion module can efficiently fuse features, solve the problem that information between channel groups of lightweight modules cannot be effectively fused, and handle complex dependencies in feature extraction with almost no increase in computing resources. For the first time, the KAN network is introduced into the field of deepfake detection, performing fine-grained non-linear fitting on the input data, accurately approximating complex distributions, improving the non-linear ability of the features already extracted by the network, and thus improving the detection accuracy and generalization ability. Brief Description of the Drawings
[0015] Figure 1 is a flowchart of the lightweight deepfake detection system of the present invention. Detailed Embodiments
[0016] The following further describes the detailed embodiments of the present invention with reference to the accompanying drawings.
[0017] In order to make the content of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0018] As Figure 1 shown, a lightweight deepfake detection system, the detection system includes a backbone network, a pyramid dilated convolution module, a shuffle feature fusion module, and a classifier;
[0019] The backbone network is connected to the pyramid dilated convolution module, and both are concatenated with the shuffle feature fusion module. The shuffle feature fusion module is connected to the classifier, and an input end is connected to the backbone network;
[0020] The pyramid dilated convolution module includes global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7). A 1x1 convolution is connected to the global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7);
[0021] The shuffle feature fusion module includes feature shuffle. A 1x1 depthwise separable convolution is connected to the feature shuffle. A 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2) are connected to the 1x1 depthwise separable convolution. A 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2) are connected to the 1x1 depthwise separable convolution. The two 3x3 depthwise separable dilated convolutions (dilation rate is 2) are concatenated with the 1x1 depthwise separable convolution, and a 1x1 convolution is connected thereto. A 3x3 convolution, normalization, and an activation function are connected to the 1x1 convolution. The 3x3 convolution is sequentially connected to normalization and the activation function.
[0022] The backbone network includes backbone module 1, backbone module 2, backbone module 3, backbone module 4, and backbone module 5. Backbone module 1 is connected to global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5), and 3x3 dilated convolution (dilation rate is 7). Backbone module 5 is concatenated with 1x1 convolution.
[0023] The classifier includes two KAN networks. The two KAN networks are connected, and a fully connected layer is connected thereto. A Real channel and a Fake channel are connected to the fully connected layer.
[0024] In specific use, there are the backbone network, the pyramid dilated convolution module, the feature shuffle feature fusion device, and the classifier part. The backbone network can extract features, but it focuses on global features. The feature data output from the first layer of the backbone network is transmitted to the pyramid dilated convolution module. First, global average pooling operation is performed to prevent feature loss. Then, local features are extracted through 1×1 convolution operation. Then, three 3×3 convolutions with different dilation rates are used to expand the receptive field and extract more local features at different scales. Then, the fused feature output after 1×1 convolution and the final feature output of the backbone network are concatenated. The concatenated features are input into the shuffle feature fusion device. First, feature shuffling is performed to solve the problem that information cannot be exchanged between different feature groups caused by grouped convolution in the backbone network. After sufficient feature fusion, it passes through a feature extraction module composed of depthwise separable convolutions to further extract the features after the fusion of global features and local features. Here, using the feature extraction combination composed of depthwise separable convolutions can achieve the purpose of effective feature extraction with only a small increase in computational cost. The features extracted after mixing are input into the classifier. Ordinary detection classifiers usually only use the fully connected layer to achieve the classification function, but the classifier of TinyDF introduces the KAN network, which can process the already extracted features, enhance the non-linear ability of the already extracted features, improve the perception ability of non-linear features, and can dynamically adjust the sensitivity to different features through the adaptive network, assign a greater weight to more important features, and improve the detection performance and generalization ability.
[0025] The above describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. All in all, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative work without departing from the gist of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A lightweight deep fake detection system, characterized by: The detection system includes a backbone network, a pyramid dilation convolution module, a feature shuffling feature fusion device and a classifier; The backbone network is connected to the pyramid expansion convolution module, and both are spliced with the feature shuffle feature fusion device, the feature shuffle feature fusion device is connected to the classifier, and the backbone network is connected to an input terminal; The pyramid dilated convolution module includes global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5) and 3x3 dilated convolution (dilation rate is 7), and the global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5) and 3x3 dilated convolution (dilation rate is 7) are connected with 1x1 convolution; The feature shuffling feature fuser includes feature shuffling, the feature shuffling is connected with a 1x1 depthwise separable convolution, the 1x1 depthwise separable convolution is connected with a 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2), the 1x1 depthwise separable convolution is connected with a 1x1 depthwise separable convolution and a 3x3 depthwise separable dilated convolution (dilation rate is 2), two of the 3x3 depthwise separable dilated convolutions (dilation rate is 2) are spliced with the 1x1 depthwise separable convolution, and a 1x1 convolution is connected thereto, the 1x1 convolution is connected with a 3x3 convolution, normalization and activation function, and the 3x3 convolution is sequentially connected with the normalization and activation function.
2. The lightweight deep fake detection system according to claim 1, characterized in that: The backbone network includes backbone module 1, backbone module 2, backbone module 3, backbone module 4 and backbone module 5. The backbone module 1 is connected with global average pooling, 1x1 convolution, 3x3 dilated convolution (dilation rate is 3), 3x3 dilated convolution (dilation rate is 5) and 3x3 dilated convolution (dilation rate is 7), and the backbone module 5 is spliced with 1x1 convolution.
3. The lightweight deep fake detection system according to claim 1, characterized in that: The classifier includes two KAN networks, the two KAN networks are connected, and a fully connected layer is connected thereon, and a Real channel and a Fake channel are connected thereon.