Face counterfeiting recognition detection method and device based on image high-frequency noise and GhostNet network
By combining image high-frequency noise with the face forgery recognition and detection method of GhostNet network, the problem of identifying artificial intelligence to generate face forgery images in the prior art is solved, and efficient and accurate forgery image recognition and classification are achieved.
Patent Information
- Application Number
- CN202411853210.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively identify and detect fake face images generated by artificial intelligence, resulting in personal privacy and security threats, dissemination of false information and legal ethics issues.
The face forgery recognition and detection method based on image high-frequency noise and GhostNet network is adopted, and the high-frequency noise in the face image is extracted through the high-pass filter SRM, and the image is extracted in depth feature with the GhostNet module, and a cross-modal attention fusion device and perceptron adapter are designed to realize the classification of real and false images.
It improves the accuracy and generalization of face forgery recognition detection, enhances the model's adaptability in deep forgery detection tasks, and effectively recognizes and classifies face forgery images generated by artificial intelligence.
Smart Images

Figure CN120164263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence generated content, and specifically, to a method and device for face forgery recognition and detection based on image high-frequency noise and GhostNet network. Background Art
[0002] With the rapid development and popularization of artificial intelligence (AI) related technologies, artificial intelligence generated content (AIGC) has become a highly regarded field in today's society. AIGC refers to content generated using artificial intelligence technologies such as deep learning, including but not limited to text, audio, pictures, and videos. In recent years, on the one hand, big data technology has been continuously mature, and the basic computing power has been greatly improved. AIGC technologies based on generative artificial intelligence such as generative adversarial networks and diffusion models have been rapidly iterated, completely breaking the limitations of templated and formulaic generation methods, and can flexibly generate rich multi-modal content. For example, for text generation, in the early days, most methods were based on templates, such as early dialogue systems, which filled in relevant information based on symbolic rules and templates. Nowadays, large language models (LLMs) can perform controllable text generation, and the content length continues to break through. On the other hand, with the continuous deepening of the integration of digital and real (the deep integration of digital technology and the real economy), the demand for the quality and richness of digital content by humans has reached an unprecedented high. The emergence of a large number of AI-generated works has lowered the professional threshold for the production of high-quality digital content.
[0003] While AIGC technology brings convenience, it also has some harmful effects. In terms of image data, face images generated by AI are called forged faces, mainly including face synthesis, face editing, face replacement (Deepfake), and face reproduction. Face synthesis generates realistic face images through deep learning models such as generative adversarial networks (GANs). Face editing modifies existing images to change expressions, ages, etc. Face replacement replaces the face in a video with another person, and face reproduction allows a person to perform different actions in a video. However, this technology also brings a series of hazards and risks: it poses a threat to personal privacy and security, and can deceive face recognition through synthetic videos for illegal profit or fraud; it may lead to the widespread dissemination of false information, endangering social stability and trust; the abuse of this technology has raised legal and moral issues, such as infringement of portrait rights, copyright, and social trust crises; the reduction of technical thresholds increases the risk of abuse.
[0004] To solve the above existing problems, people have been seeking an ideal technical solution. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a face forgery recognition and detection method and device based on image high-frequency noise and GhostNet network.
[0006] To achieve the above object, the first aspect of the present invention provides a face forgery recognition and detection method based on image high-frequency noise and GhostNet network, including the following steps: Step 1, select a vision transformer as the backbone network of the two-stream network to initially extract image features, introduce GhostNet modules into the two branches of the two-stream network respectively to perform deep feature extraction on the image, and introduce a cross-modal attention fusion device behind the backbone network to integrate the two-stream features; and add a perceptron adapter behind the cross-modal attention fusion device to realize the classification of genuine and fake images; Step 2, train the two-stream network using the masked image modeling self-supervised pre-training method; Step 3, extract the high-frequency noise in the face image through a high-pass filter SRM, and use the high-frequency noise and the low-frequency texture of the image as the inputs of the two-stream network respectively to send into the two-stream network to obtain the face forgery recognition result.
[0007] Furthermore, in order to transfer complementary features from one modality to another modality, the present invention provides a cross-domain bidirectional adapter between the two branches of the backbone network to solve the problem of mutual complementarity and transfer between the RGB stream and the noise stream.
[0008] Specifically, the cross-domain bidirectional adapter is respectively connected to the GhostNet modules in the two branches; wherein, the cross-domain bidirectional adapter includes a downsampling layer, an activation function layer and an upsampling layer connected in sequence.
[0009] Furthermore, in order to achieve efficient and effective data fusion of high-frequency noise and visible light images, the cross-modal attention fusion device includes two Transformer modules, a feature splicing module and a three-layer self-attention mechanism module; wherein, one Transformer module is used to extract features of the visible light image by using the self-attention mechanism, and the other Transformer module is used to extract features of the high-frequency noise image by using the self-attention mechanism; The feature splicing module is used to splice the features extracted by the two Transformer modules and send them into the three-layer self-attention mechanism module for deep extraction.
[0010] To achieve the above object, the second aspect of the present invention provides a face forgery recognition and detection device based on image high-frequency noise and GhostNet network, including: The dual-stream network construction module is used to select a vision transformer as the backbone network of the dual-stream network to initially extract image features, introduce GhostNet modules in the two branches of the dual-stream network to perform deep feature extraction on the images, and introduce a cross-modal attention fusion module after the backbone network to integrate the dual-stream features; and a perceptron adapter is added after the cross-modal attention fusion module to achieve the classification of genuine and fake images; The dual-stream network training module is used to train the dual-stream network using the masked image modeling self-supervised pre-training method; The high-frequency noise acquisition module is used to extract high-frequency noise in the face image through the high-pass filter SRM; The face forgery recognition module is used to use the high-frequency noise and the low-frequency texture of the image as the inputs of the dual-stream network and send them into the dual-stream network to obtain the face forgery recognition result.
[0011] To achieve the above object, a third aspect of the present invention provides a computer device, which is characterized in that it includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, it implements the steps of the method described in the first aspect.
[0012] To achieve the above object, a fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0013] To achieve the above object, a fifth aspect of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0014] The beneficial effects of the present invention are as follows: The present invention uses a vision transformer (ViT) trained by the MIM self-supervised pre-training method as the backbone network of the dual-stream network; extracts high-frequency noise in the face image through the high-pass filter SRM, and uses the high-frequency noise and the low-frequency texture of the image as the inputs of the dual-stream network respectively; in the two branches of the dual-stream network, GhostNet is introduced to perform deep feature extraction on the images; a cross-modal attention fusion module is designed and placed after the backbone network to integrate the dual-stream features, further model the interaction between the noise stream and the spatial stream features, and improve the adaptability of the model in the deep fake detection task; and a perceptron adapter (MA) is added after the cross-modal fusion module to achieve the classification of genuine and fake images. Description of the Drawings
[0015] Figure 1It is a schematic flowchart of the face forgery recognition and detection method described in Embodiment 1 of the present invention.
[0016] Figure 2 It is a schematic structural diagram of the two-stream network described in Embodiment 1 of the present invention.
[0017] Figure 3 It is a schematic structural diagram of the cross-domain bidirectional adapter described in Embodiment 1 of the present invention.
[0018] Figure 4 It is a schematic structural diagram of the cross-modal attention fusion module described in Embodiment 1 of the present invention.
[0019] Figure 5 It is a schematic structural diagram of the perceptron adapter described in Embodiment 1 of the present invention.
[0020] Figure 6 It is a schematic flowchart of the model self-learning in Embodiment 1 of the present invention.
[0021] Figure 7 It is a result diagram of the verification embodiment of the present invention.
[0022] Figure 8 It is a schematic diagram of the principle of the face forgery recognition and detection system described in Embodiment 2 of the present invention.
[0023] Figure 9 It is a structural diagram of the computer device described in the present invention. Detailed implementation manners
[0024] The technical solutions of the present invention will be further described in detail below through specific implementation manners.
[0025] In the technology of identifying artificially intelligent forged faces, there are several significant deficiencies. First, these technologies often overly focus on low-level visual features of images, such as edges and corners, which may lead to the model overfitting to the noise and unrepresentative details of the images during the training process. Second, these methods perform poorly in terms of generalization, are difficult to adapt to image changes in different scenarios and conditions, and lack the necessary robustness. In addition, although the methods based on frequency-domain features have good robustness to transformations such as image compression, they often ignore the association between spatial-domain features and frequency-domain features, which may lead to performance degradation when dealing with detection tasks that rely on color and texture information. Finally, the current pre-training methods are difficult to learn universal frequency-domain features, which limits the performance of these methods in cross-domain detection tasks.
[0026] To solve the above problems, the present invention proposes a face forgery recognition and detection method based on image high-frequency noise and GhostNet network. This method uses a Vision Transformer (ViT) trained by the Masked Image Modeling (MIM) self-supervised pre-training method as the backbone network of the two-stream network. The high-frequency noise in the face image is extracted by the high-pass filter SRM, and the high-frequency noise and the low-frequency texture of the image are respectively used as the inputs of the two-stream network to maximize the extraction of the edge features of the image content. In the two branches, this model introduces GhostNet to perform deep feature extraction on the image, and can obtain better feature extraction effects at a lower time cost and hardware cost. To further model the interaction between the noise stream and the spatial stream features and improve the adaptability of the model in the deep forgery detection task, this model designs a cross-modal attention fusion module which is placed after the backbone network to integrate the two-stream features. Finally, a Multilayer Perceptron Adapter (MA) is added after the cross-modal fusion module to realize the classification of genuine and fake images.
[0027] The technical solution of the present invention will be further described in detail below through specific embodiments.
[0028] Embodiment 1 This embodiment provides a face forgery recognition and detection method based on image high-frequency noise and GhostNet network, which is characterized in that, as Figure 1 shown, it includes the following steps: Step 1, select a Vision Transformer as the backbone network of the two-stream network to initially extract image features, introduce GhostNet modules in the two branches of the two-stream network respectively to perform deep feature extraction on the image, and introduce a cross-modal attention fusion module after the backbone network to integrate the two-stream features; and add a Multilayer Perceptron Adapter after the cross-modal attention fusion module to realize the classification of genuine and fake images; the network structure of the two-stream network is as Figure 2 shown; Step 2, train the two-stream network using the Masked Image Modeling self-supervised pre-training method; Step 3, extract the high-frequency noise in the face image by the high-pass filter SRM, and use the high-frequency noise and the low-frequency texture of the image as the inputs of the two-stream network respectively and send them into the two-stream network to obtain the face forgery recognition result.
[0029] Current deep forgery detection methods often overfit to the color textures of specific methods, while using high-frequency noise can remove the color texture features of the image itself and expose more imperceptible characteristics. In addition, there are often mixed forgery traces at the edges of forged face images, which are difficult to identify in the RGB domain but are obvious in the noise space.
[0030] Therefore, in this embodiment, in order to improve the accuracy and generalization of the detection method, the SRM high-pass filter is used to extract high-frequency noise from the input image. As Figure 2 shown, in this embodiment, the high-frequency noise features after passing the spatial domain image through the SRM high-pass filter are used as the input of the noise stream.
[0031] Specifically, RGB domain image x is input into a two-dimensional convolution Conv2d , and then the output value is mapped to the interval [-3, 3] through the Hardtanh () activation function. The formula is as follows:
[0032] is the output value.
[0033] Among them, in this embodiment, the SRM high-pass filter adopts a classic design, and three convolution kernels with the same parameters are set in Conv2d . Among them, each kernel is divided by a fixed value and the sum of the matrix parameters is ensured to be 0, so as to keep the brightness of the filtered image unchanged and highlight the high-frequency part in the image, that is, the edge information.
[0034] Specifically, the parameters of the convolution kernel are as follows:
[0035] It can be understood that the SRM high-pass filter captures the regions with large frequency changes in the image through convolution operations, thereby highlighting the high-frequency content of the image, including edge and detail information, while suppressing low-frequency regions such as color textures and smooth backgrounds.
[0036] Deep convolutional neural networks usually contain a large number of convolution operations, which leads to huge computational costs. Although existing research works, such as MobileNet and ShuffleNet, have introduced depthwise separable convolutions or shuffling operations to build efficient CNNs using smaller convolution filters (floating-point operations), the remaining convolutional layers still occupy a considerable amount of memory and floating-point operation counts.
[0037] In view of the extensive redundancy existing in the intermediate feature maps of mainstream CNN computations, the GhostNet module is introduced in this embodiment. First, compared with the units in research works that widely use 1×1 pointwise convolutions, the main convolution in the GhostNet module can have a custom kernel size. Existing methods mostly use pointwise convolutions to process cross-channel features and then depthwise separable convolutions to process spatial information. In contrast, the GhostNet module first generates several intrinsic feature maps through ordinary convolutions, and then uses low-cost linear operations to enhance the features and increase the number of channels. In previous efficiency architectures such as MobileNet, the operations for processing each feature map are limited to depthwise separable convolutions or shift operations, while the linear operations in the GhostNet module have greater diversity. Finally, in the GhostNet module, the identity mapping is implemented in parallel with the linear transformation to retain the intrinsic feature maps.
[0038] Low-frequency textures usually contain the main shape and structural information in an image, while high-frequency noise may contain edges, details, or possible artifacts. To transfer complementary features from one modality to another, the cross-domain bi-directional adapter BCA (Bi-directional Cross-modal Adapter) is designed between the two data streams in this embodiment. BCA solves the problem of mutual complementation and transfer between the RGB stream and the noise stream. Figure 3 The structural schematic diagram of the BCA module is shown.
[0039] It can be seen that the cross-domain bi-directional adapter is respectively connected to the GhostNet modules in the two branches.
[0040] In one embodiment, the cross-domain bi-directional adapter includes a downsampling layer, an activation function layer, and an upsampling layer connected in sequence.
[0041] The activation function selected for the activation function layer can be set according to requirements, such as the Sigmoid function, the Relu function, etc., and no more restrictions will be imposed here.
[0042] Furthermore, the present method designs a cross-modal attention fusion unit that fuses the features of visible light images and high-frequency noise features, as Figure 4 shown. The core of this module is a feature processing unit composed of two Transformer modules. First, the module receives two sets of inputs: one is a visible light image, and the other is the corresponding high-frequency noise image. These two sets of inputs respectively enter their respective Transformer modules for feature extraction.
[0043] In the visible light image feature processing part, the Transformer module uses the self-attention mechanism to deeply analyze the image and extract key features. Similarly, the high-frequency noise image is also processed by another Transformer module with the same self-attention mechanism to extract the deep features of the noise. The outputs of these two modules will jointly form the basis for data fusion.
[0044] Next, the output features of the two Transformer modules are concatenated to form a new fused feature vector. This vector combines the detailed information of the visible light image and the characteristics of high-frequency noise, providing a rich data foundation for subsequent deep processing.
[0045] In order to further deepen the expression of features, the concatenated fusion features are sent to a network containing a three-layer self-attention mechanism for deep extraction. The first layer of the self-attention mechanism aims to mine the correlation information between features, the second layer further strengthens key features and suppresses irrelevant features, and the third layer performs fine-grained extraction of features to improve the distinguishability of features.
[0046] Finally, after deep processing of the three layers of self-attention mechanism, the module outputs a highly fused and highly discriminative feature vector. This vector not only contains rich information of the original image, but also incorporates noise features, providing powerful feature support for subsequent visual tasks such as target detection or image classification. In general, the data fusion module designed in this embodiment achieves efficient and effective data fusion through a sophisticated structure and deep extraction technology.
[0047] Finally, the perceptron adapter is connected in series after the feature fusion module. Figure 5 As shown in Figure 1, the module consists of a global average pooling, a layer normalization, and a multi-layer perceptron. The specific calculation process can be expressed as follows:
[0048] in, Indicates input, Pool () represents global average pooling, and Norm() represents layer normalization; W 1 , W 2 Represents the connection weight matrix between neurons from the input layer to the hidden layer and from the hidden layer to the output layer of the multilayer perceptron, Represents the bias of each layer; represent GELU Activation function; is the model output, representing the probability that the image is generated by AI.
[0049] Average pooling can effectively compress the dimension of the feature map while retaining the global spatial information. This means that regardless of where the object in the image is located, average pooling can capture the context information of the entire image, which is very important for image classification tasks. Normalization can accelerate the model training process and help stabilize the learning process of the neural network. It reduces the problem of internal covariate shift and improves the generalization ability of the model by normalizing each feature of each sample.
[0050] It should be noted that with the rapid development of artificial intelligence generation technology, the iterative update of algorithms has become a norm. To ensure the accuracy and reliability of the output content, the research team of this study continuously optimizes the algorithm to meet the needs of users. In this process, we fully recognize that user feedback is an important way to improve the performance of the algorithm. Therefore, the present invention provides a correction function so that users can make corrections in a timely manner when they find that there are errors in the algorithm output content.
[0051] When the user uses the correction function, the corrected data and its corresponding labels will be recorded in real time and stored in our dataset. The continuous accumulation of this dataset is of great significance for us to optimize the algorithm. When the dataset accumulates to a certain scale, the system will automatically start the program of retraining the model when it detects that the system resources are idle. In this way, the algorithm can continuously learn and evolve to improve the accuracy and adaptability of its output.
[0052] In addition, the present invention also implements a model version management function. Each time the model is retrained and successfully loaded, the system will generate a new version number. In this way, when problems occur in the application of the new model, it can be quickly rolled back to the previous stable version to ensure the continuity and stability of the service. The record of model versions also helps us track the optimization path of the algorithm and provides valuable data support for subsequent research. The model self-learning process is as Figure 6 shown.
[0053] Specifically, manually judge whether the face forgery recognition result is correct. If it is incorrect, correct the face forgery recognition result and add the face forgery recognition result to the updated dataset; When the updated dataset reaches the preset scale and the system resources are detected to be idle, train the two-stream network using the masked image modeling self-supervised pre-training method based on the updated dataset to update the weights of the two-stream network; After the two-stream network is retrained and successfully loaded, generate a new version number for the two-stream network so that when problems occur in the application of the new version of the two-stream network, it can be rolled back to the previous version.
[0054] To verify the method described in this embodiment, the verification experiment was conducted on the public dataset DeepFaceGen. There were a total of 813,847 images, including 463,583 real face images and 350,264 forged images. Among them, 560,000 images were randomly selected for training, and the remaining images were used for evaluation and testing. The change in accuracy during the model training process is as shown in Figure 7 shown.
[0055] Through a large number of experiments, it can be proved that the accuracy of the face forgery recognition model proposed by the present invention on the test set is 99%, which has practical value.
[0056] Embodiment 2 Based on the same inventive concept, this embodiment of the present application also provides a face forgery recognition detection device based on image high-frequency noise and GhostNet network for implementing the above-mentioned face forgery recognition detection method involving image high-frequency noise and GhostNet network. The implementation solutions provided by this device to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more device embodiments provided below can refer to the limitations on the face forgery recognition detection method based on image high-frequency noise and GhostNet network in Embodiment 1, and will not be repeated here.
[0057] Specifically, as shown in Figure 8 shown, the face forgery recognition detection device based on image high-frequency noise and GhostNet network includes: A dual-stream network construction module, which is used to select a vision transformer as the backbone network of the dual-stream network to initially extract image features, introduce GhostNet modules in two branches of the dual-stream network respectively to perform deep feature extraction on the image, and introduce a cross-modal attention fusion device after the backbone network to integrate the dual-stream features; and a perceptron adapter is added behind the cross-modal attention fusion device to realize the classification of genuine and forged images; A dual-stream network training module, which is used to train the dual-stream network using the masked image modeling self-supervised pre-training method; A high-frequency noise acquisition module, which is used to extract high-frequency noise in the face image through a high-pass filter SRM; A face forgery recognition module, which is used to use the high-frequency noise and the low-frequency texture of the image as the inputs of the dual-stream network respectively and send them into the dual-stream network to obtain the face forgery recognition result.
[0058] Embodiment 3 This embodiment provides a computer device, which can be a terminal, and its internal structure diagram can be as shown in Figure 9As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it realizes a face forgery recognition and detection method based on image high-frequency noise and the GhostNet network. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0059] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0060] Embodiment 4 This embodiment provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, realizes the steps of the face forgery recognition and detection method based on image high-frequency noise and the GhostNet network described in Embodiment 1.
[0061] Embodiment 5 Based on the above embodiments, this embodiment provides a computer program product, including a computer program, which when executed by a processor implements the following steps: Step 1, select a vision transformer as the backbone network of the two-stream network to initially extract image features, introduce GhostNet modules in the two branches of the two-stream network respectively to perform deep feature extraction on the image, and introduce a cross-modal attention fusion module after the backbone network to integrate the two-stream features; and add a perceptron adapter after the cross-modal attention fusion module to achieve classification of genuine and fake images; Step 2, train the two-stream network using the masked image modeling self-supervised pre-training method; Step 3, extract high-frequency noise in the face image through a high-pass filter SRM, and use the high-frequency noise and the low-frequency texture of the image as the inputs of the two-stream network respectively and send them into the two-stream network to obtain the face forgery recognition result.
[0062] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0063] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them; although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or perform equivalent replacements for some technical features; without departing from the spirit of the technical solutions of the present invention, they should all be covered by the scope of the technical solutions claimed in the present invention.
Claims
1. A face forgery recognition and detection method based on image high-frequency noise and GhostNet network, characterized in that: The following steps are involved: Step 1: Select the visual transformer as the backbone network of the two-stream network to preliminarily extract image features, introduce the GhostNet module in the two branches of the two-stream network to extract deep features of the image, and introduce the cross-modal attention fuser after the backbone network to integrate the two-stream features; And add a perceptron adapter after the cross-modal attention fuser to achieve classification of real and fake images; Step 2: Use the mask image modeling self-supervised pre-training method to train the two-stream network; Step 3: extract the high-frequency noise in the face image through the high-pass filter SRM, and send the high-frequency noise and the low-frequency texture of the image into the two-stream network as the input of the two-stream network respectively to obtain the face forgery recognition result.
2. According to claim 1, a method for face forgery recognition and detection based on image high-frequency noise and GhostNet network is characterized in that: A cross-domain bidirectional adapter is provided between two branches of the backbone network, and the cross-domain bidirectional adapter is respectively connected to the GhostNet modules in the two branches; The cross-domain bidirectional adapter includes a downsampling layer, an activation function layer, and an upsampling layer connected in sequence.
3. The face forgery recognition and detection method based on image high-frequency noise and GhostNet network according to claim 1 or 2, characterized in that: The cross-modal attention fuser includes two Transformer modules, a feature concatenation module, and a three-layer self-attention mechanism module; Among them, one Transformer module is used to extract features from visible light images using the self-attention mechanism, and the other Transformer module is used to extract features from high-frequency noise images using the self-attention mechanism; The feature splicing module is used to splice the features extracted by the two Transformer modules and send them to the three-layer self-attention mechanism module for deep extraction.
4. The method for detecting face forgery based on image high-frequency noise and GhostNet network according to claim 1 or 2, characterized in that: The perceptron adapter includes a global average pooling, a layer normalization and a multi-layer perceptron, and the specific expression is: , in, Indicates input, Pool () represents global average pooling, and Norm() represents layer normalization; W 1 , W 2 Represents the connection weight matrix between neurons from the input layer to the hidden layer and from the hidden layer to the output layer of the multilayer perceptron, b 1 , b 2 Represents the bias of each layer; represent GELU Activation function; is the model output, representing the probability that the image is generated by AI.
5. The method for face forgery recognition and detection based on image high-frequency noise and GhostNet network according to claim 1 or 2, characterized in that: Extracting high-frequency noise from face images through high-pass filter SRM includes: Will RGB Domain Image x Input 2D convolution conv2d Then through Hardtanh ()The activation function maps the output value to the range [-3,3].
6. The method for face forgery recognition and detection based on image high-frequency noise and GhostNet network according to claim 1, characterized in that: Manually determine whether the forged face recognition result is correct. If it is wrong, correct the forged face recognition result and add the forged face recognition result to the update data set; When the update data set reaches a preset size and the system resources are detected to be idle, the two-stream network is trained based on the update data set using a mask image modeling self-supervised pre-training method to update the weights of the two-stream network; After the two-stream network is retrained and successfully loaded, a new version number is generated for the two-stream network so that when problems occur during the application of the new version of the two-stream network, it can be rolled back to the previous version.
7. A face forgery recognition and detection device based on image high-frequency noise and GhostNet network, characterized in that: include: The two-stream network construction module is used to select the visual transformer as the backbone network of the two-stream network to preliminarily extract image features, introduce the GhostNet module in the two branches of the two-stream network to extract deep features of the image, and introduce the cross-modal attention fuser after the backbone network to integrate the two-stream features; And add a perceptron adapter after the cross-modal attention fuser to achieve classification of real and fake images; Two-stream network training module, used to train two-stream networks using a mask image modeling self-supervised pre-training method; A high-frequency noise acquisition module is used to extract high-frequency noise in the face image through a high-pass filter SRM; The face forgery recognition module is used to send the high-frequency noise and the low-frequency texture of the image as the input of the two-stream network respectively, to obtain the face forgery recognition result.
8. A computer device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory, used to store computer programs; A processor, configured to implement the steps of the method according to any one of claims 1 to 6 when executing a program stored in a memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Deep fake face image detection method based on double-flow CNN and ViT hybrid model
CN120833637A