Active defense method based on face identity watermark and mixed attention module
Through the active defense method based on face identity watermark and hybrid attention module, the problem of degraded local tamper detection performance and insufficient robustness in Deepfake technology is solved, and accurate detection and traceability of face tampering is achieved, and the stability and adaptability of the model are improved.
Patent Information
- Application Number
- CN202510819520.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
When the prior art faces local facial tampering by Deepfake technology, the detection efficiency decreases, and the active defense method is not robust enough in image processing operations, making it difficult to maintain stability in complex environments.
Active defense method based on face identity watermark and hybrid attention module is adopted, stable features are extracted through retinal face detection, combined with average hash and hybrid attention encoder module, invisible watermark is embedded, and weighted cross entropy, multi-scale image pixel loss and adaptive message loss are trained to realize the detection and traceability of face tampering.
It improves the practicality and functionality of the model, can maintain stability in common image processing operations, and realizes accurate detection and traceability of face tampering.
Smart Images

Figure CN120337309A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to an active defense method based on face identity watermark and hybrid attention module. Background Art
[0002] With the continuous upgrading of Deepfake technology, the degree of falsification of forged images has continuously broken through the original upper limit, which poses an unprecedented challenge to traditional image detection methods. DeepFake is composed of "Deep Learning" and "Fake". This cutting-edge technology is based on deep learning models such as generative adversarial networks (GANs) to build a sophisticated face forgery system. Through deep learning of a large amount of image data, the model can accurately capture biological features such as the facial bone structure, muscle movement patterns, and skin texture, and can even restore the facial light and shadow changes under different lighting conditions. In the forgery stage, it deeply embeds the feature parameters of the target face into the source image and generates false image content that is difficult to distinguish from the real one through pixel-level intelligent adjustment. The face forgery technology uses artificial intelligence algorithms to accurately lock the facial area in the image for tampering, and the generated false content is highly similar to the real image at the visual level, posing a major threat to social order, information security, and privacy protection.
[0003] To break through the technical bottleneck of passive detection, the academic and industrial circles have turned their attention to active defense methods. The core of active defense lies in pre-embedding invisible digital watermarks, feature identifiers and other signals in the original image to enhance the anti-tampering ability of the image from the source. This defense strategy is particularly prominent in traceability tracking, and can effectively record the original information of the image, providing key clues for identifying the source of forgery. Compared with the "post-verification" of passive detection, active defense realizes "prevention before the event", significantly improving the anti-attack ability and robustness of the system.
[0004] Deficiencies of the prior art: Limited detection scope: Traditional image tampering detection methods focus on extracting the overall features of the image and have a certain detection ability for global tampering. However, when faced with common local facial tampering in Deepfake technology, the detection efficiency drops significantly, and it is difficult to accurately identify forgery traces that only modify the facial area while other areas such as the background remain unchanged, resulting in a decrease in detection accuracy and robustness.
[0005] Poor adaptability to complex environments: The invisible watermarks embedded in existing active defense methods are too sensitive to common image processing operations such as compression, blurring, and noise. Once a Deepfake image undergoes such conventional processing, the detection performance is difficult to be stable, severely limiting its application effect in actual scenarios. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an active defense method based on face identity watermark and hybrid attention module, which realizes the functions of detecting and tracing face tampering, and greatly improves the practicability and functionality of the model.
[0007] The present invention adopts the following technical solutions to achieve the invention purpose: An active defense method based on face identity watermark and hybrid attention module, characterized by comprising the following steps: S1: Extraction of face identity code; S2: Generation of face identity code; S3: Hybrid attention encoder module; S4: Design of loss function; S5: Analysis of face feature correlation.
[0008] As a further limitation of the technical solution, the specific steps of S1 are as follows: The primary link in the generation of face identity watermark is to accurately extract stable face features from the face region of the input image. The retina face detection algorithm in the face detection and alignment algorithm based on deep learning is used to detect and locate the face in the input image, and the main loss for locating the face region is the face classification loss: (1); Where: is the true label; is the probability that the prior box predicted by the model contains a face; is the weight of the positive sample; is the focusing parameter for adjusting the weights of positive and negative samples; Construct a two-dimensional Mel face feature extraction structure for image texture feature extraction. This two-dimensional Mel face feature extraction structure mainly consists of three parts: multi-resolution stacked maps, multi-scale two-dimensional Mel filters, and logarithmic energy spectrum feature extraction; The multi-resolution stacked maps downsample the image at different resolutions to form a series of images at different scales. This process mainly includes Gaussian smoothing and downsampling operations; The core of Gaussian smoothing is to perform a convolution operation between the Gaussian kernel and the image. The expression of the Gaussian kernel is: (2); Where: is the coordinate of the pixel in the Gaussian kernel; is the standard deviation of the Gaussian function; The construction formula of the multi-scale two-dimensional Mel filter is as follows: (3); Where: ; ; is the wavelength of the sine function; is the direction of the filter; represents the phase shift; is a variable parameter that controls the spatial localization degree of the filter; represents the spatial aspect ratio and controls the shape of the filter in different directions; For images with different resolutions after multi-resolution stacking processing, two-dimensional Mel filters of different scales are used to perform convolution operations with the images at this layer to obtain response images at different scales. Finally, the frequency distribution and energy information of the images are extracted using the logarithmic energy spectrum extraction method. The extraction formula is as follows: Energy spectrum calculation: (4); Where: represents the multi-scale response image; represents the frequency domain coordinates of the corresponding image, represents taking the absolute value; Logarithmic extraction: (5); Where: Energy represents the face information extracted at the image frequency domain coordinates ; is a small constant.
[0009] As a further limitation of this technical solution, the specific steps of S2 are: The average hashing algorithm is used to convert the extracted face feature vector into a binary hash value with a specified length. The generation formula is as follows: The face feature vector array is: (6); The calculation formula for its average value E is: (7); For each element in the face feature array , if , then the corresponding binary value is 1, otherwise it is 0. The mathematical expression is as follows: (8).
[0010] As a further limitation of this technical solution, the specific steps of S3 are: S31: Multi-scale identity watermark preprocessing module; In the identity code watermark transformation module, after processing, the binary watermark information with length L is converted into a format consistent with the dimensions of the image tensor. First, the watermark is reshaped into a two-dimensional array with a size of (H / 2, W / 2). , and then multiple scale factors are selected to perform upsampling. The upsampling formula is: (9); where: represents the upsampling operation; represents different scale factors; After multi-scale upsampling, the feature maps after upsampling at different scales are weighted and fused. The formula for weighted fusion is: (10); where: represents the weight corresponding to the th scale factor; S32: Hybrid attention embedding module; This module receives the input original image and the watermark message after multi-scale preprocessing , constructs a visual Mamba-like linear attention U-shaped network that combines a kind of Mamba-like linear attention mechanism with a U-shaped network to specifically handle the watermark embedding work. This model mainly consists of three parts: a channel attention-based feature initial processing module (SE-Stem) for initial feature extraction, a linear attention module, and a multi-scale dilated downsampling convolutional block; S321: Channel attention-based feature initial processing module SE-Stem; Construct a two-branch structure to separately process the input image and the watermark. Each branch processes the image and the watermark through convolutional operations and the channel attention mechanism respectively, gradually reducing their spatial dimensions and increasing the channel dimensions while retaining important information. Finally, the two feature information is concatenated to prepare for subsequent operations. The formula is as follows: (11); where: represents the channel attention operation; and represent the consecutive convolutional operations; represents feature concatenation; represents the input source image; represents the preprocessed identity watermark; S322: Linear attention module; Input the fused features of the source image and the identity watermark after preliminary processing by the channel attention-based feature initial processing module into the Mamba-like linear attention block for further processing; Replace the two linear blocks in the original Mamba-like linear attention module with a row-wise feature collaboration module and a column-wise feature collaboration module to form a new MLLA structure linear attention module; The row-wise feature collaboration module consists of a feature extraction convolutional layer, a horizontal position encoding layer, and a linear layer. The input feature vector first passes through a convolutional layer to extract row-local features: (12); Where: is the extracted row-local feature, is the feature map preliminarily processed by SE-Stem, Conv () represents the convolution operation; Then it enters the horizontal position encoding layer for matrix operations: (13); Where: is The position encoding in the horizontal direction, and its calculation process is as follows: Input feature map , then the horizontal position encoding parameter , is the preset maximum horizontal position number, indicating the maximum position range in the horizontal direction that the model can handle. Select the part corresponding to the current feature map width from the horizontal position encoding parameter W ; ; For the th row feature vector , the calculation of the linear layer is: (14); Where: , and are projection matrices, is the processed row feature; , , respectively represent the weights obtained by multiplying the th row feature vector by three different projection matrices, represents linear activation of the key matrix, softmax represents a linear activation function; The column-wise feature collaboration module consists of a feature extraction convolutional layer, a vertical position encoding layer, and a linear layer. The input feature vector first passes through a convolutional layer to extract column-local features: (15); Wherein: is the extracted column local feature, and is the feature preliminarily processed by SE-Stem; Then it enters the vertical position encoding layer for matrix operation: (16); Wherein: is the position encoding in the vertical direction, and its calculation process is as follows: Input feature map , then the vertical position encoding parameter , is the preset maximum vertical position number, indicating the maximum position range in the vertical direction that the model can handle. Select the part corresponding to the height H of the current feature map from the vertical position encoding parameter ; ; For the column feature vector , the calculation of the linear layer is: (17); Wherein: is the processed column feature; respectively represent the weights obtained by multiplying the column feature vector by three different projection matrices, represents the linear activation of the key matrix; S323: Multi-scale dilated downsampling convolutional block; By combining dilated convolution and depthwise separable convolution, it can expand the receptive field without significantly reducing the resolution. Dilated convolution captures features at three different scales of local, medium, and global by using different dilation rates in the convolution kernel, while retaining the details of the image. Adopting the idea of dynamically adjusting the dilation rate, introduce learnable parameters to dynamically determine the optimal dilation rate of each convolutional layer. The parameter is dynamically adjusted according to the features of the input image, so as to capture more context information at different scales. The feature analysis layer analyzes the statistical information of the input features; First, calculate the gradient of the input feature, and then calculate the dynamic dilation rate according to the gradient and the parameter : (18); Finally, perform operations using dilated convolutions with different dilation rates dynamically selected according to d; int() represents integerization; The entire downsampling process can be expressed as: (19); Where: represents the finally output feature map; represents the flattening layer; represents the convolutional layer 2; represents the multi-scale dilated convolution; represents the depthwise separable convolution; represents the convolutional layer 1; represents the preliminary reshaping of the input features; represents the input features.
[0011] As a further limitation of this technical solution, the specific steps of S4 are: The loss function consists of three parts: weighted cross-entropy loss, multi-scale image pixel loss, and adaptive message loss; S41: Weighted cross-entropy loss; The weighted cross-entropy loss is used to solve the class imbalance problem by assigning different weights to different classes. The formula is as follows: (20); Where: is the true class label of the face region detected in the i th face sample; is the positive class probability predicted by the model; is the weighted weight; S42: Multi-scale image pixel loss; By introducing the idea of multi-scale in the mean square error and calculating the mean square error at different scales, the details and structural information of the image can be better captured. The formula is as follows: (21); Where: is the number of scales; is the th number of pixels at a scale; and are the representations of the watermarked image and the original image at the th scale, respectively; S43: Adaptive texture feature message loss; Construct an innovative adaptive loss function; First, the local texture features of the image need to be extracted. The local texture features help to understand the complexity of the image, thereby determining in which regions to embed stronger watermarks. The formula is as follows: (22); Where: and are the gradients of the input image respectively; is the local texture feature of the image position ; Then, according to the local texture feature, an adaptive weight is calculated. In the regions with rich texture, the watermark embedding strength is increased; in the smooth regions, the watermark embedding strength is decreased; (23); Wherein: is the adaptive weight of the image position ; and are the minimum and maximum values of the texture feature respectively; Finally, the definition of the adaptive message loss function: (24); Wherein: is the original watermark message; is the extracted watermark message; is the classification label, is the decoding threshold, is the decoding confidence; is the adaptive weight at position i, is the total number of pixels of the watermark; S44: The overall loss function is the weighted sum of the weighted cross-entropy loss, the multi-scale image loss and the adaptive message loss: (25); Wherein: , and are the weights of each loss term.
[0012] As a further limitation of this technical solution, the specific steps of S5 are: Because two decoding methods are used in this model: decoding the watermark embedded in the face image using a decoder and regenerating the watermark in the face area of the face image using the same face recognition algorithm and the two-dimensional Mel face extraction structure. The traditional correlation comparison includes comparing the Hamming distance. However, if there are a small number of bit differences between two vectors, the Hamming distance may increase significantly. Even if the two watermarks are very similar visually, due to the differences in individual bits, the Hamming distance may lead to misjudgment as dissimilar. In order to accurately evaluate the similarity between the watermarks decoded by the two decoding methods, the regularized Pearson correlation coefficient is used to quantify the linear association between them, and the formula is as follows: (26); Wherein: is a regularization term; and respectively represent the and th dimensional element of the feature vector; is the dimension of the feature vector; and are the average values of the feature vectors and ; When it is 1, it means that the two vectors are completely positively correlated, and when it is 0, it means that there is no linear relationship between the two feature vectors.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are: 1. Compared with the traditional method, the present invention can simultaneously realize the functions of detecting and tracing face tampering through different decoding structures, greatly improving the practicability and functionality of the model.
[0014] 2. The present invention proposes a new method for extracting face identity watermarks, which has been experimentally verified to have good robustness against conventional noise operations.
[0015] 3. The present invention proposes VM-UNet (a linear attention UNet similar to the state space model), which is a new type of architecture. By innovatively combining linear attention, channel attention, the Mamba model and the Unet architecture, it is specifically designed to handle the watermark embedding task, and the new architecture has achieved a high improvement in the natural vision and robustness of the watermarked images. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the overall structure diagram of the present invention.
[0017] Figure 2 is the VM-Unet structure diagram of the present invention.
[0018] Figure 3 is the schematic diagram of the two-dimensional Mel face feature extraction structure of the present invention.
[0019] Figure 4 is the schematic diagram of the SE-Stem module structure of the present invention.
[0020] Figure 5 is the schematic diagram of the multi-scale dilated downsampling convolutional block (MSDC) structure of the present invention.
[0021] Figure 6 is the schematic diagram of the HLMLLA module structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0022] The following will describe in detail a specific embodiment of the present invention in conjunction with the accompanying drawings. It should be understood that the protection scope of the present invention is not limited by the specific embodiment.
[0023] In the present invention, a method for actively defending against face forgery based on face identity watermark and hybrid attention module is proposed. By embedding an exclusive face identity watermark in the face image, this method can simultaneously detect and trace the forgery behavior. Specifically, this method first inputs a source target face image, uses face recognition technology to divide the face area of the face image, and generates a face identity code using the rich texture features of the face area. Then, the extracted identity code is processed through shape adjustment and average hashing and used as a robust watermark I A and embedded into the entire image. During the embedding process, a new framework VM-Unet that combines a linear attention mechanism (MLLA) similar to the state space model and the Unet network is proposed to specifically handle the watermark embedding work. This architecture enhances the feature processing and watermark embedding accuracy by combining the advantages of linear attention, channel attention, and state space model (SSM), supplemented by the efficient symmetric sampling structure of Unet. After that, the image undergoes a series of processes to enhance the robustness of the model, such as passing through a noise layer (compression, blurring, cropping), and an image containing the face identity code watermark is generated. In the construction of the watermark decoder, this method uses two different decoding methods for tampering detection and traceability analysis respectively: on the one hand, the identity code watermark I B is extracted from the face area of the tampered image through the same face recognition algorithm as used to generate the face identity code before, and on the other hand, the decoder is used to extract the embedded identity code watermark I A’, from the image. By comparing the correlation between I B and I A’ , it is determined whether the image has been tampered with. On the other hand, by comparing I A and I A’ , the image source can be further traced to achieve traceability analysis.
[0024] The present invention includes the following steps: S1: Extraction of face identity code.
[0025] The specific steps of S1 are as follows: The primary link in generating the face identity watermark is to accurately extract stable face features from the face area of the input image. The RetinaFace algorithm, which is a face detection and alignment algorithm based on deep learning, is used to detect and locate the face in the input image. The main loss for locating the face area is the face classification loss: (1); Where: is the true label; is the probability that the prior box predicted by the model contains a face; is the weight of the positive sample; is the focusing parameter that adjusts the weights of positive and negative samples; Inspired by the Mel frequency spectrum coefficients in audio anti-counterfeiting, after the input image is subjected to face region localization by the RetinaFace detection algorithm, a two-dimensional Mel face feature extraction structure for extracting image texture features is constructed. This structure is as Figure 3 shown. This two-dimensional Mel face feature extraction structure uses the method of multi-scale feature extraction. By extracting face features at different scales, it captures the multi-scale texture information in the image, making the extracted face features more stable and still having strong robustness under various distortions such as JPEG compression, noise, and scaling. This two-dimensional Mel face feature extraction structure mainly consists of three parts: a multi-resolution stacked atlas, a multi-scale two-dimensional Mel filter, and logarithmic energy spectrum feature extraction; The multi-resolution stacked atlas downsamples the image at different resolutions to form a series of images at different scales. This process mainly includes Gaussian smoothing and downsampling operations. Gaussian smoothing refers to performing Gaussian filtering on the face region image to remove high-frequency noise in the image and make the image smooth. Downsampling is to downsample the image after Gaussian smoothing, selecting pixel points every other row and column, thereby obtaining images at different resolution scales.
[0026] The core of Gaussian smoothing is to perform a convolution operation between the Gaussian kernel and the image. The expression of the Gaussian kernel is: (2); where: are the coordinates of the pixels in the Gaussian kernel; is the standard deviation of the Gaussian function; The construction formula of the multi-scale two-dimensional Mel filter is as follows: (3); where: ; ; is the wavelength of the sine function; is the direction of the filter; represents the phase shift; is a variable parameter that controls the spatial localization degree of the filter; represents the spatial aspect ratio, controlling the shape of the filter in different directions; For images of different resolutions after multi-resolution stacking processing, two-dimensional Mel filters of different scales are used to perform convolution operations with the images at this layer to obtain response images at different scales. These response images reflect the texture features and energy information of the images at different scales and directions. Finally, the frequency distribution and energy information of the images are extracted using the method of logarithmic energy spectrum extraction. The extraction formula is as follows: Energy spectrum calculation: (4); Where: represents the multi-scale response image; represents the frequency domain coordinates of the corresponding image, represents taking the absolute value; Logarithmic extraction: (5); Where: Energy represents the face information extracted at the image frequency domain coordinates ; is a small constant used to prevent numerical problems in logarithmic operations.
[0027] S2: Generation of face identity codes.
[0028] The specific steps of S2 are as follows: The face feature vectors generated and extracted through the above steps are too long, and directly embedding them into the image will affect the watermark embedding quality. Therefore, the average hashing algorithm is used to convert the generated face feature vectors into a binary hash value of a specified length. The generation formula is as follows: The face feature vector array is: (6); The calculation formula for its average value E is: (7); For each element in the face feature array , if , then the corresponding binary value is 1, otherwise it is 0. The mathematical expression is as follows:
[0029] S3: Hybrid attention encoder module.
[0030] The specific steps of S3 are as follows: In the process of embedding the image identity code watermark, in order to ensure a high degree of integrity of the image in terms of vision and quality, maintain the original details and clarity of the image as much as possible, and at the same time ensure the strong robustness of the embedded watermark so that it can still exist stably and completely when facing common image processing attacks such as compression, cropping, filtering, or being interfered by noise, the encoder module consists of two parts: the multi-scale identity watermark preprocessing module and the hybrid attention embedding module.
[0031] S31: Multi-scale identity watermark preprocessing module; In order to enable the binary hash watermark generated by the face feature code to better fuse with the image features, in the identity code watermark transformation module, the binary watermark information of length L is processed and then converted into a format consistent with the image tensor dimension. This conversion process ensures that the watermark information can be effectively fused with the image features in the same dimension. Specifically, the watermark is first reshaped into a two-dimensional array with a size of (H / 2, W / 2). , and then multiple scale factors are selected to perform upsampling. The upsampling formula is: (9); where: represents the upsampling operation; represents different scale factors; After multi-scale upsampling, the feature maps after upsampling at different scales are weighted and fused. The weighted fusion formula is: (10); where: represents the weight corresponding to the th scale factor; By preprocessing the watermark through multiple scales, the watermark has a certain representation at different scales, which can effectively enhance the redundancy and robustness of the watermark, and at the same time ensure the dimensional consistency with the image features.
[0032] S32: Hybrid attention embedding module; This module receives the input original image (R represents the image dataset, 3 indicates that the image is input in the RGB three-channel format, and H and W respectively represent the height and width of the image) and the watermark message after multi-scale preprocessing . In order to more comprehensively represent the content of the image, a visual Mamba-like linear attention U-Net (Vision-Mamba-U-Net, VM-UNet) combining a Mamba-like linear attention mechanism (MLLA) and a U-Net (Unet) is constructed to specifically handle the watermark embedding work, as Figure 2As shown in the figure, the model mainly consists of three parts: a feature initial processing module based on channel attention for initial feature extraction (Squeeze-and-Excitation Network Stem, SE-Stem) (SE-Stem), a linear attention module (HLMLLA), and a multi-scale dilated downsampling convolutional block (MSDC); S321: Feature initial processing module SE-Stem based on channel attention; This module is designed as an initial processing module. To enable the proposed architecture to be specifically used for watermark embedding, we improved the Stem module. Specifically, a dual-branch structure is constructed to separately process the input image and the watermark. As Figure 4 shown, each branch processes the image and the watermark respectively through convolutional operations and the channel attention mechanism (SEnet), gradually reducing their spatial dimensions and increasing the channel dimensions while retaining important information. Finally, the two feature information are concatenated to prepare for subsequent operations. The formula is as follows: (11); Where: represents the channel attention operation; and represent the successive convolutional operations; represents feature concatenation; represents the input source image; represents the preprocessed identity watermark; S322: Linear attention module (HLMLLA); The fused feature obtained by preliminarily processing the source image and the identity watermark through the feature initial processing module based on channel attention is input into the Mamba-Like Linear Attention (MLLA) block for further processing; MLLA (Mamba-Like Linear Attention) is an improved method that combines the Mamba model and the linear attention mechanism, aiming to improve the performance of the model in visual tasks. MLLA mainly integrates two key factors in Mamba: the "forget gate" and the block design. However, the forget gate must use recursive calculations, which may not be very suitable for modeling non-causal data, such as images. Images are two-dimensional spatial data, and their information is mainly reflected in the spatial relationships and local features between pixels, unlike sequential data that has an obvious time order or causal relationship and does not inherently require recursion. To effectively model the spatial features in face images and avoid the unnecessary time consumption and complexity that may occur when recursive calculations process non-causal data.
[0033] Replace the two linear blocks in the original Mamba-like linear attention module with a row feature collaboration module and a column feature collaboration module to form a new linear attention module of the MLLA structure; as Figure 6 shown, these two components enhance the model's understanding of pixel position relationships through the linear attention mechanism, combined with horizontal position encoding and vertical position encoding, without relying on a recurrent structure to capture image sequence information.
[0034] The row feature collaboration module consists of a feature extraction convolutional layer, a horizontal position encoding layer, and a linear layer. The input feature vector first extracts row local features through a convolutional layer: (12); where: is the extracted row local feature, is the feature map preliminarily processed by the SE-Stem, Conv () represents the convolution operation; Then it enters the horizontal position encoding layer for matrix operations: (13); where: is the position encoding in the horizontal direction, and its calculation process is as follows: Input feature map , , , represent the channels, height, and width of the feature map respectively. Then the horizontal position encoding parameter , is the preset maximum horizontal position number, indicating the maximum position range in the horizontal direction that the model can handle. Select the part corresponding to the current feature map width from the horizontal position encoding parameter W ; ; For the th row feature vector (14); where: , and are projection matrices, is the processed row feature; , , respectively represent the th The weights obtained by multiplying with three different projection matrices represent linear activation of the key matrix, softmax indicating a linear activation function; The columnar feature collaboration module consists of a feature extraction convolutional layer, a vertical position encoding layer, and a linear layer. The input feature vector first extracts column local features through a convolutional layer: (15); Where: is the extracted column local feature, is the feature preliminarily processed by SE-Stem; Then it enters the vertical position encoding layer for matrix operations: (16); Where: is the position encoding in the vertical direction, and its calculation process is as follows: Input feature map , then the vertical position encoding parameter , is the preset maximum vertical position number, indicating the maximum position range in the vertical direction that the model can handle. Select the part corresponding to the height H of the current feature map from the vertical position encoding parameter ; ; For the column feature vector , the calculation of the linear layer is: (17); Where: is the processed column feature; respectively represent the weights obtained by multiplying the column feature vector with three different projection matrices, representing linear activation of the key matrix; S323: Multi-scale dilated downsampling convolutional block; Common downsampling methods use max pooling and average pooling to achieve downsampling. Although these methods can reduce the resolution, they often lose some feature information. By combining dilated convolution and depthwise separable convolution, the receptive field can be expanded without significantly reducing the resolution. The specific structure is as Figure 5As shown, dilated convolution captures features at three different scales: local, medium-scale, and global, by using different dilation rates in the convolution kernel, while preserving the details of the image. To improve the adaptability and performance of dilated convolution, the idea of dynamically adjusting the dilation rate is adopted, and learnable parameters are introduced to dynamically determine the optimal dilation rate for each convolutional layer. The parameter is dynamically adjusted according to the features of the input image, so as to capture more context information at different scales. The feature analysis layer analyzes the statistical information of the input features; First, calculate the gradient of the input features , and then according to the gradient and the parameter calculate the dynamic dilation rate : (18); Finally, dilated convolutions with different dilation rates are dynamically selected for operation according to d; int() represents integerization; Its entire downsampling process can be expressed as: (19); Where: represents the finally output feature map; represents the flattening layer; represents convolutional layer 2; represents multi-scale dilated convolution; represents depthwise separable convolution; represents convolutional layer 1; represents the initial reshaping of the input features; represents the input features.
[0035] S4: Loss function design.
[0036] The specific steps of the said S4 are as follows: The loss function consists of three parts: weighted cross-entropy loss, multi-scale image pixel loss, and adaptive message loss; ensuring that while maintaining the high fidelity of the image, the watermark message can be effectively embedded and extracted; S41: Weighted cross-entropy loss; In the face detection stage, in order to reduce the class imbalance problem caused by different face categories and positions, weighted cross-entropy loss is used to solve the class imbalance problem by assigning different weights to different classes. The formula is as follows: (20); Where: is the true class label of the face region detected by the i th face sample; is the positive class probability predicted by the model; is the weighted weight such that the loss contribution of each category can be adjusted according to its importance; S42: Multi-scale image pixel loss; The image loss is used to measure the difference between the original image and the watermarked image, ensuring that the watermark embedding process does not significantly reduce the image quality. The image loss is usually calculated using the mean square error (MSE). However, MSE usually only focuses on the pixel value differences and ignores the structural and texture information of the image. By introducing the idea of multi-scale in the mean square error, calculating the mean square error at different scales can better capture the detailed and structural information of the image. The formula is as follows: (21); Where: is the number of scales, usually set to 1, 2, 4; is the th number of pixels at the scale; and are the representations of the watermarked image and the original image at the th scale respectively; S43: Adaptive texture feature message loss; To ensure that the watermark message can be accurately extracted, the message loss function plays a crucial role in the watermark embedding process. However, traditional message loss functions may cause over-strong watermark embedding in some cases, thus having a negative impact on the image quality. To solve this problem, an innovative adaptive loss function is constructed, which can flexibly adjust the watermark embedding strength according to the local texture features of the image. This adaptive mechanism not only helps to maintain the overall quality of the image but also significantly improves the robustness of the watermark, ensuring that the watermark information can be effectively extracted under various conditions; First, the local texture features of the image need to be extracted. The local texture features can help understand the complexity of the image, thus determining in which regions a stronger watermark can be embedded. The formula is as follows: (22); Where: and are the gradients of the input image respectively; is the local texture feature at the image position ; Then, according to the local texture features, the adaptive weight is calculated. In regions with rich texture, the watermark embedding strength is increased; in smooth regions, the watermark embedding strength is decreased; (23); Where: is the adaptive weight at the image position ; and are the minimum and maximum values of the texture feature respectively; Finally, the definition of the adaptive message loss function: (24); where: is the original watermark message; is the extracted watermark message; is the classification label, which is used to convert the watermark message bit into a binary classification problem (1 or -1), is the decoding threshold, is the decoding confidence, indicating the difference between the extracted watermark message bit and the threshold ; is the adaptive weight at position i, is the total number of pixels of the watermark; S44: The overall loss function is the weighted sum of the weighted cross-entropy loss, the multi-scale image loss, and the adaptive message loss: (25); where: , and are the weights of each loss term.
[0037] S5: Analysis of the correlation of face features.
[0038] The specific steps of the above S5 are as follows: Because two decoding methods are used in this model: decoding the watermark embedded in the face image using a decoder and regenerating the watermark in the face area of the face image using the same face recognition algorithm and the two-dimensional Mel face extraction structure. Traditional correlation comparison includes comparing the Hamming distance. However, if there are a small number of bit differences between two vectors, the Hamming distance may increase significantly. Even if two watermarks are very similar visually, due to individual bit differences, the Hamming distance may lead to misjudgment as dissimilar. To accurately evaluate the similarity between the watermarks decoded by the two decoding methods, the regularized Pearson correlation coefficient is used to quantify their linear association. The formula is as follows: (26); where: is a regularization term, which can avoid the case of a zero denominator and improve numerical stability; and respectively represent the and th dimensional elements of the feature vectors; is the dimension of the feature vector; and is the eigenvector and the average value of; When it is 1, it means that the two vectors are completely positively correlated, and when it is 0, it means that there is no linear relationship between the two eigenvectors.
[0039] Experimental settings Training and testing were carried out on the CelebA face dataset, and all images were cropped to a size of 128×128. The division ratio of the training set to the test set is 0.8:0.2.
[0040] In terms of the configuration of training parameters, it was divided into 100 Epochs. In the first 50 Epochs of training, a relatively high initial learning rate was adopted. This strategy helps the model quickly explore the parameter space in the initial stage of training, thus accelerating the convergence process. Subsequently, in the remaining 50 Epochs, we lowered the learning rate so that the model can be more finely adjusted in the later stage of training.
[0041] The batch size set is 64. It can maintain sufficient sample diversity during each gradient update, thus effectively improving the generalization ability of the model. The Adam optimizer was selected as the optimizer.
[0042] In the construction of the loss function, we set weight coefficients in the loss function, and the initial weights are respectively , , .
[0043] Experimental results: Table 1 Comparison of visual quality of watermarked images
[0044] Table 2 Robustness test of conventional noise
[0045] Table 3 Generalization test: trained on the CelebA dataset and tested on the FFHQ dataset
[0046] The specific embodiments of the present invention disclosed above are only for illustration. However, the present invention is not limited thereto, and any changes that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. An active defense method based on face identity watermark and hybrid attention module, characterized in that It includes the following steps: S1: Extraction of face identity code; S2: Generation of face identity code; S3: Hybrid attention encoder module; S4: Design of loss function; S5: Analysis of face feature correlation.
2. The active defense method based on face identity watermark and hybrid attention module according to claim 1, characterized in that: The specific steps of S1 are as follows: The primary link in generating face identity watermark is to accurately extract stable face features from the face region of the input image. The RetinaFace detection algorithm in the deep learning-based face detection and alignment algorithm is used to detect and locate the face in the input image. The main loss for locating the face region is the face classification loss: (1); Wherein: is the true label; is the probability that the prior box predicted by the model contains a face; is the weight of the positive sample; is the focusing parameter for adjusting the weights of positive and negative samples; Construct a two-dimensional Mel face feature extraction structure for image texture feature extraction. This two-dimensional Mel face feature extraction structure mainly consists of three parts: multi-resolution stacked atlas, multi-scale two-dimensional Mel filter, and logarithmic energy spectrum feature extraction; The multi-resolution stacked atlas downsamples the image at different resolutions to form a series of images at different scales. This process mainly includes Gaussian smoothing and downsampling operations; The core of Gaussian smoothing is to perform a convolution operation on the image using a Gaussian kernel. The expression of the Gaussian kernel is: (2); Wherein: are the coordinates of the pixels in the Gaussian kernel; is the standard deviation of the Gaussian function; The construction formula of the multi-scale two-dimensional Mel filter is as follows: (3); Wherein: ; ; is the wavelength of the sine function; is the direction of the filter; represents the phase shift; is a variable parameter that controls the spatial localization degree of the filter; represents the spatial aspect ratio and controls the shape of the filter in different directions; For the images at different resolutions after multi-resolution stacking processing, use two-dimensional Mel filters at different scales to perform convolution operations on this layer of images to obtain response images at different scales. Finally, use the logarithmic energy spectrum extraction method to extract the frequency distribution and energy information of the image. The extraction formula is as follows: Energy spectrum calculation: (4); Wherein: represents the multi-scale response image; represents the frequency domain coordinates of the corresponding image, represents taking the absolute value; Logarithmic extraction: (5); where: Energy represents the face information extracted at the image frequency domain coordinates ; is a small constant.
3. The active defense method based on face identity watermark and hybrid attention module according to claim 2, characterized in that: The specific steps of S2 are as follows: Use the average hashing algorithm to convert the extracted face feature vector into a binary hash value of a specified length. The generation formula is as follows: The face feature vector array is: (6); The calculation formula for its average value E is: (7); For each element in the face feature array , if , then the corresponding binary value is 1, otherwise it is 0. The mathematical expression is as follows: (8)。 4. The active defense method based on face identity watermark and hybrid attention module according to claim 3, characterized in that: The specific steps of S3 are as follows: S31: Multi-scale identity watermark preprocessing module; In the identity code watermark transformation module, after being processed, the binary watermark information with length L is converted into a format consistent with the dimensions of the image tensor. The watermark is first reshaped into a two-dimensional array with a size of (H / 2, W / 2). , and then multiple scale factors are selected for upsampling. The upsampling formula is as follows: (9); Wherein: represents an upsampling operation; represents different scale factors; After multi-scale upsampling, the feature maps after upsampling at different scales are weighted and fused. The weighted fusion formula is: (10); Wherein: represents the weight corresponding to the th scale factor; S32: Hybrid attention embedding module; This module receives the input original image and the watermark message after multi-scale preprocessing , constructs a visual Mamba-like linear attention U-shaped network that combines a Mamba-like linear attention mechanism with a U-shaped network to specifically handle the watermark embedding work. This model mainly consists of three parts: a channel attention-based feature initial processing module (SE-Stem) for initial feature extraction, a linear attention module, and a multi-scale dilated downsampling convolutional block; S321: Feature initial processing module SE-Stem based on channel attention; Construct a two-branch structure to separately process the input image and the watermark. Each branch processes the image and the watermark through convolution operations and channel attention mechanisms respectively, gradually reducing their spatial dimensions and increasing the channel dimensions while retaining important information. Finally, the two feature information is concatenated for subsequent operations. The formula is as follows: (11); Wherein: represents channel attention operation; and represents consecutive convolution operations; represents feature concatenation; represents the input source image; represents the preprocessed identity watermark; S322: Linear attention module; Input the fused features of the source image and the identity watermark after preliminary processing by the feature initial processing module based on channel attention into the class Mamba linear attention block for further processing; Replace the two linear blocks in the original class Mamba linear attention module with a row-wise feature collaboration module and a column-wise feature collaboration module to form a new MLLA structure linear attention module; The row-wise feature collaboration module consists of a feature extraction convolutional layer, a horizontal position encoding layer, and a linear layer. The input feature vector first extracts row local features through a convolutional layer: (12); Wherein: is the extracted local line feature, is the feature map preliminarily processed by SE-Stem, Conv () represents a convolution operation; Then enter the horizontal position encoding layer for matrix operations: (13); Wherein: is the position code in the horizontal direction, and its calculation process is as follows: Input feature map , the horizontal position encoding parameter , is the preset maximum horizontal position number, indicating the maximum position range in the horizontal direction that the model can handle. Select the part corresponding to the current feature map width from the horizontal position encoding parameter W ; ; For the row feature vector , the calculation of the linear layer is as follows: (14); Wherein: , and are projection matrices, is the processed row feature; , , respectively represent the weights obtained by multiplying the row feature vector by three different projection matrices, represents linear activation of the key matrix, softmax represents a linear activation function; The columnar feature collaboration module consists of a feature extraction convolutional layer, a vertical position encoding layer, and a linear layer. The input feature vector first extracts columnar local features through a convolutional layer: (15); Wherein: is the extracted column local feature, and is the feature preliminarily processed by SE-Stem; Then it enters the vertical position encoding layer for matrix operations: (16); Wherein: is the position encoding in the vertical direction, and its calculation process is as follows: Input feature map , the vertical position encoding parameter , is the preset maximum number of vertical positions, indicating the maximum position range in the vertical direction that the model can handle. Select the part corresponding to the height H of the current feature map from the vertical position encoding parameter ; ; For the column eigenvector , the calculation of the linear layer is as follows: (17); Wherein: is the processed column feature; respectively represent the column feature vectors and the weights obtained by multiplying with three different projection matrices, representing the linear activation of the key matrix; S323: Multi-scale dilated downsampling convolutional block; By combining dilated convolution and depthwise separable convolution, it is possible to expand the receptive field without significantly reducing the resolution. Dilated convolution captures features at three different scales of local, medium, and global by using different dilation rates in the convolutional kernel, while retaining the details of the image. Adopting the idea of dynamically adjusting the dilation rate, learnable parameters are introduced to dynamically determine the optimal dilation rate for each convolutional layer. The parameters are dynamically adjusted according to the features of the input image, so as to capture more context information at different scales. The feature analysis layer analyzes the statistical information of the input features; First, calculate the gradient of the input feature , and then calculate the dynamic dilation rate based on the gradient and the parameter : (18); Finally, dilated convolutions with different dilation rates are dynamically selected for operations according to d; int() represents integerization; Its entire downsampling process can be expressed as: (19); Wherein: Represents the final output feature map; Represents the flattening layer; Represents convolutional layer 2; Represents the multi-scale dilated convolution; Represents the depthwise separable convolution; Represents convolutional layer 1; Represents the preliminary reshaping of the input features; Represents the input features.
5. The active defense method based on face identity watermark and hybrid attention module according to claim 4, characterized in that: The specific steps of S4 are as follows: The loss function consists of three parts: weighted cross-entropy loss, multi-scale image pixel loss, and adaptive message loss; S41: Weighted cross-entropy loss; The weighted cross-entropy loss is used to solve the class imbalance problem by assigning different weights to different classes. The formula is as follows: (20); Wherein: is the true class label of the face region detected by the i th face sample; is the positive class probability predicted by the model; is the weighted weight; S42: Multi-scale image pixel loss; By introducing the idea of multi-scale in the mean square error and calculating the mean square error at different scales, it can better capture the details and structural information of the image. The formula is as follows: (21); Wherein: is the number of scales; is the number of pixels on the th scale; and are the representations of the watermarked image and the original image on the th scale, respectively. S43: Adaptive texture feature message loss; Construct an innovative adaptive loss function; First, the local texture features of the image need to be extracted. The local texture features help understand the complexity of the image, so as to determine in which regions to embed stronger watermarks. The formula is as follows: (22); Wherein: and are respectively the gradients of the input image ; is the local texture feature of the image position ; Then, according to the local texture features, calculate the adaptive weights. In regions with rich textures, increase the watermark embedding strength; in smooth regions, reduce the watermark embedding strength; (23); Wherein: is the image position of the adaptive weight; and are respectively the minimum value and the maximum value of the texture feature; Finally, the definition of the adaptive message loss function: (24); Wherein: is the original watermark message; is the extracted watermark message; is the classification label, is the decoding threshold, is the decoding confidence; is the adaptive weight at position i, is the total number of pixels of the watermark; S44: The overall loss function is the weighted sum of the weighted cross-entropy loss, multi-scale image loss, and adaptive message loss: (25); Wherein: , and the weights of each loss item.
6. The active defense method based on face identity watermark and hybrid attention module according to claim 5, characterized in that: The specific steps of S5 are as follows: Because two decoding methods are used in this model: decoding the watermark embedded in the face image using a decoder and regenerating the watermark in the face area of the face image using the same face recognition algorithm and the two-dimensional Mel face extraction structure. Traditional correlation comparisons include comparing the Hamming distance. However, if there are a small number of bit differences between two vectors, the Hamming distance may increase significantly. Even if the two watermarks are very similar visually, due to individual bit differences, the Hamming distance may lead to misjudgment as dissimilar. To accurately evaluate the similarity between the watermarks decoded by the two decoding methods, the regularized Pearson correlation coefficient is used to quantify the linear association between them. The formula is as follows: (26); Wherein: is a regularization term; and respectively represent the and th dimensional element of the feature vectors; is the dimension of the feature vector; and are the average values of the feature vectors and ; When it is 1, it means that the two vectors are completely positively correlated, and when it is 0, it means that there is no linear relationship between the two feature vectors.
Citation Information
Patent Citations
Interaction method, system and device for information
CN107515717A
Screen sharing method and device, storage medium and electronic equipment
CN111767554A
Depth image watermarking method based on mixed frequency domain channel attention
CN115272044A
Multi-scale image tampering detection method based on mixed attention mechanism
CN115578626A
Active defense detection method based on face key point watermark
CN117474741A