Diffusion model customized privacy protection method and system based on mask attention mechanism elimination
Through the method of mask attention mechanism elimination, the face mask is segmented and the attention of the diffusion model is erased. Combined with the correction training of the tri-sampling function, the problem of malicious utilization of the diffusion model customization method is solved, and effective privacy protection and image disturbance effects are achieved.
Patent Information
- Application Number
- CN202510218238.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
The existing method of customizing diffusion models is easily exploited by malicious users and violates personal privacy. The existing forged detection technology has problems such as data set restrictions, detection later than release, high computational complexity and iterative optimization. The existing adversarial attack technology cannot completely prevent the diffusion model from learning face features.
The mask attention mechanism is used to eliminate the mask, and the face mask mask is segmented by segmentation model, and the attention of the diffusion model is erased by using the mask cross attention and self-attention erase loss function. Combined with the triad sampling function correction training process, the protective image is generated.
Effectively disrupt the output of the diffusion model, protect personal privacy, and the generated images cannot recognize facial features, which improves the disruptive effect of generating adversarial samples, and reduces the computational complexity and computing resource requirements.
Smart Images

Figure CN120337271A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and privacy protection, and particularly relates to a diffusion model customization privacy protection method and system based on mask attention mechanism elimination. Background Art
[0002] With the wide application of artificial intelligence-based image generation technology in production and life, the research on large image generation models has also shown a continuous growth trend. Traditional generation models based on generative adversarial networks and variational autoencoders have made significant progress in image generation, but still cannot meet the industrial demand for high-resolution and high-quality image generation. As an emerging generation model, the diffusion model can generate high-quality, high-resolution and artificially controllable images, so it has attracted much attention and research in the industrial and academic fields.
[0003] To meet the market demand for customized generation of private images, researchers have actively explored customized generation technologies based on diffusion models. Among them, methods such as DreamBooth and text inversion are the most representative and have been widely used. DreamBooth associates a specific object with a specific identity prompt by fine-tuning the parameters of the diffusion model, so as to achieve customized generation of specific objects or portraits; while the text inversion technology maps the object image into a specific text embedding vector and uses this vector to generate customized pictures. However, with the wide application of customization technologies, some malicious users have started to use these technologies to create illegal images and infringe on the portrait rights of others, as shown in Figure 1 (1). They may use the face pictures of public figures to train the model and then generate indecent related photos to spread on the Internet, which seriously violates the personal portrait rights.
[0004] Faced with this problem, researchers have taken a variety of technical means to deal with it. On the one hand, to prevent the spread of illegal images, researchers have developed forgery detection technologies to judge whether an image is a synthetic image and take measures to stop its spread in time. On the other hand, to prevent the customized model from learning face information, researchers have adopted adversarial attack technologies. By adding adversarial perturbations to face images, the learning process of the model is disrupted, and the model is avoided from accurately learning face features, further protecting user privacy. This method not only effectively prevents the generation and spread of illegal images, but also provides a safer and more reliable guarantee for users when using customization technologies. The present invention belongs to the second technical means.
[0005] Forgery detection technologies have played a certain role in identifying and preventing the spread of synthetic images, but they also have some defects and limitations, mainly including the following points:
[0006] (1) Dataset Limitations: Forgery detection techniques typically require a large number of real images and synthetic fake images for training and validation. However, due to the diverse types and forms of synthetic images, and the continuous changes in the generation methods of malicious users, it is very challenging to construct a comprehensive and accurate dataset. Current datasets may not cover all possible types of synthetic images, resulting in limitations in the applicability of detection techniques in practical applications.
[0007] (2) Detection Delayed after Publication: The process of forgery detection is delayed after the publication of synthetic images. By the time fake images are detected, these images have already been spread on the Internet for some time, and may even have attracted extensive attention and dissemination, while the forgery detection technology only identifies and confirms these images at a later time. Due to the delay in forgery detection technology after publication, malicious users have more time to improve their synthetic techniques to avoid detection. This makes the technology more complex and difficult to counter, and also poses greater challenges to researchers and technicians.
[0008] (3) Computational Complexity: Some forgery detection techniques may require a large amount of computational resources and time for image analysis and feature extraction, especially when detecting on large-scale datasets. This computational complexity may limit the practical application scope of the technology and increase the cost and time consumption of the detection process.
[0009] (4) Iterative Optimization: Malicious users may take advantage of the limitations of forgery detection techniques to avoid detection by continuously optimizing the generation methods of synthetic images. They can gradually narrow the loopholes and limitations of detection techniques through repeated trials and adjustments, making the detection techniques less robust and reliable in practice.
[0010] On the other hand, the face privacy protection method based on adversarial attack technology can add adversarial perturbations to images before they are published, making it impossible for customized models to learn face features and generating distorted images, thus blocking the spread of this information at the source. Adversarial attack technology usually requires experiments and adjustments to find appropriate adversarial perturbation parameters so that the generated images can protect privacy while ensuring visual quality. However, diffusion models have high robustness and strong learning ability, especially for complex face structures and features. Existing adversarial attack technologies may not be able to completely prevent diffusion models from learning key face attributes, such as facial contours, expression features, etc. These technologies lack research on the internal architecture and generation process of diffusion models, and these contents affect the generation quality and attack effect of the models. Although adding adversarial perturbations can cause a certain degree of distortion in face images, there is still a risk that the model can learn face information. Summary of the Invention
[0011] The purpose of the present invention is to solve the problem that the current customization method based on the diffusion model can be used by malicious users to violate privacy. By disturbing the diffusion model customization method and disturbing the output of the model for the customized person, personal privacy is protected.
[0012] The technical solution adopted by the present invention is as follows:
[0013] A diffusion model customization privacy protection method based on the elimination of masked attention mechanism, comprising the following steps:
[0014] Given the original image to be protected, use a segmentation model to segment out the mask of the face part;
[0015] Use the masked cross-attention erasing loss function to erase the cross-attention corresponding to the identity text prompt in the diffusion model;
[0016] Use the masked self-attention erasing loss function to erase the self-attention in the diffusion model to interfere with the relationship between the face pixel points of the generated image;
[0017] Use the masked cross-attention erasing loss function and the masked self-attention erasing loss function to train the adversarial perturbation, and add the adversarial perturbation to the original image to obtain the protected image.
[0018] Furthermore, the masked cross-attention erasing loss function aims to minimize the attention energy in the masked area while maximizing the energy outside the masked area; the masked cross-attention erasing loss function is expressed as:
[0019]
[0020] Where M is a binary mask, where M(x,y)=1 represents the area containing identity information, and M(x,y)=0 represents the background area; is the attention map of the identity text prompt, Where A c [:,∶,∶,w s represents taking the w c th feature in the text dimension of the cross-attention map A s , w s represents the position of the identity prompt in the entire text.
[0021] Furthermore, the masked self-attention erasing loss function is expressed as:
[0022]
[0023] Where A s represents the self-attention map, A sThe dimension of is B×H×W×D, where D = H×W represents the number of self-attention image pixels; to erase the corresponding values in each attention map, the mask M is resized from the H×W shape to the D×1 shape and denoted as M r .
[0024] Furthermore, a hybrid quality score is used to measure the impact of the time steps of the diffusion model on adversarial attacks. The hybrid quality score is calculated using the following steps:
[0025] Calculate the gradient G on the image during the reverse propagation of the diffusion model t , and then calculate the L1 norm N of the gradient G t ; t
[0026] Convert G t to a probability map p through the softmax function t , and then calculate the entropy H of p t ; t
[0027] Calculate the hybrid quality score: HQS t = norm(N t ) - norm(H t ), where norm represents the normalization process for all elements in the matrix.
[0028] Furthermore, a cubic sampling function is used to correct the sampling time steps during the training of the diffusion model, so that the diffusion model can learn at smaller time steps as much as possible, thereby enhancing the effect of adversarial attacks.
[0029] Furthermore, the cubic sampling function is expressed as:
[0030] t′ = T * (1 - (t / T) 3 )
[0031] where t represents the current sampling time step, and T represents the maximum sampling time step of the diffusion model.
[0032] Furthermore, the adversarial perturbation is trained using the masked cross-attention erasure loss function and the masked self-attention erasure loss function. Its training objective is to maximize the triple loss, including the DreamBooth loss function L DB , the masked cross-attention erasure loss function L MCAE , and the masked self-attention erasure loss function L MSAE , ensuring comprehensive protection of face information. That is, the total loss function is: L total = L DB + λ1L MSAM + λ2L MCAE , where λ1 and λ2 are weight coefficients; the process of adversarial training includes iteratively updating the perturbation to ensure that both cross - attention and self - attention are effectively minimized in the masked region; in each training cycle, first train the surrogate model using the clean image set, then use the projected gradient descent method to iteratively train the adversarial samples multiple times, and finally retrain the model using the adversarial samples.
[0033] A customized privacy - protection system for a diffusion model based on masked attention mechanism elimination, which includes:
[0034] A segmentation module, which is used to segment the mask of the face part from the given original image to be protected by using a segmentation model;
[0035] A cross - attention erasure module, which is used to erase the cross - attention corresponding to the identity text prompt in the diffusion model by using the masked cross - attention erasure loss function;
[0036] A self - attention erasure module, which is used to erase the self - attention in the diffusion model by using the masked self - attention erasure loss function to interfere with the relationship between the face pixel points of the generated image;
[0037] A protected image generation module, which is used to train the adversarial perturbation by using the masked cross - attention erasure loss function and the masked self - attention erasure loss function, and add the adversarial perturbation to the original image to obtain the protected image.
[0038] The beneficial effects of the present invention are as follows:
[0039] Compared with the existing methods, the present invention proposes a novel customized privacy - protection method for a diffusion model based on masked attention elimination, which can effectively improve the disturbing effect of the generated adversarial samples on the output of the diffusion model. Given an original image to be protected, the present invention first segments the mask of the face part by using a segmentation model, and then designs a masked - based loss function to erase the cross - attention and self - attention of the model to generate a protected image. The protected image can, under the condition of ensuring sufficient similarity to the original image, not allow the diffusion model to learn the face features it contains, thereby disturbing the output of the diffusion model.
[0040] The present invention includes two innovative points. First, the present invention analyzes the cross-attention of each word in the text prompt during the generation process of the diffusion model, that is, the degree of attention of the model to each word. The present invention finds that the model particularly focuses on the special words introduced by the customization technology (i.e., identity prompt words). Based on this, the present invention uses a semantic segmentation model to first segment the face information in the image, and then introduces a loss function that erases the cross-attention of the identity prompt words on the segmented foreground (face) to train the adversarial perturbation. In addition, the self-attention mechanism of the diffusion model plays an important role in maintaining the relationship between pixel points when generating images. The present invention constructs a segmentation-based self-attention loss function to eliminate the self-attention of the model, thereby reducing the quality of the generated images. In addition, the present invention analyzes the relationship between the specific sampling time steps of the diffusion model and the attention mechanism. At smaller sampling time steps, the model pays more attention to the detailed information of the face and is more suitable for performing adversarial attacks on face information. Therefore, the present invention designs a cubic sampling function to correct the sampling time steps in model training, so that the model learns at smaller time steps as much as possible, thereby enhancing the effect of adversarial attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 In (1), the results of protected data and unprotected data passing through the malicious customization system are shown, and in (2), the present invention enables the customization system to generate pictures in which the faces are unrecognizable.
[0042] Figure 2 is the flowchart of the adversarial perturbation training of the present invention. Among them, K self represents the self-attention of the diffusion model, and K cross represents the cross-attention of the diffusion model, X adv represents the adversarial sample (the picture protected by adding the adversarial perturbation), P represents the trainable adversarial perturbation, and X represents the original picture.
[0043] Figure 3 is the comparison of the protection results for different noise budgets.
[0044] Figure 4 is the comparison of the protection results for different customization methods. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and the accompanying drawings.
[0046] The object of the present invention is to solve the problem that the current customization method based on the diffusion model can be used by malicious users to violate privacy, and it can be applied to two scenarios, such as Figure 1As shown in (2). The first scenario is portrait protection. In this scenario, the present invention enables the protected image to be safely published on the Internet. Since the adversarial perturbation threshold is very small, the protected image is almost visually indistinguishable from the original image. However, when a malicious user fine-tunes the diffusion model using the downloaded image, the model cannot normally output a human face. The second scenario is identity hiding. Since the present invention interferes with the generation of the attention mechanism of the diffusion model, the model cannot determine the identity information of the target, thereby hiding the identity information of the protected human face. The ultimate goal of the present invention is to protect personal privacy by disturbing the diffusion model customization method and disturbing the model's output for customized characters.
[0047] As Figure 2 shown, the diffusion model customization privacy protection method based on cross-attention elimination proposed by the present invention mainly consists of three modules. The first module is masked cross-attention mechanism erasure. In the generation process of the diffusion model, the present invention calculates the cross-attention of the model and designs a masked cross-attention erasure (MCAE) loss function to minimize the masked cross-attention corresponding to the identity text prompt. Similarly, in the second module, the present invention designs a masked self-attention erasure (MSAE) loss function to erase the self-attention of the model, thereby disturbing the relationship between the face pixels of the generated image and making the diffusion model unable to generate a complete-structured image. The third module is the cubic sampler. The present invention selects a cubic function as the sampler instead of sampling uniformly. This sampler samples the diffusion model time steps that are relatively early in the adversarial perturbation learning process, which is beneficial for the learning of adversarial perturbations. Finally, the present invention continuously iteratively updates the training adversarial samples through projected gradient descent. The adversarial samples trained by the present invention have a strong attack ability against the diffusion model customization technology, making the output of the diffusion model distorted and generating indistinguishable human faces.
[0048] (I) Background knowledge:
[0049] (1) Diffusion models are a class of probability models that gradually sample (denoise) from Gaussian noise to generate images. The Stable Diffusion model (SD) has achieved excellent performance in text-to-image generation. It uses a variational autoencoder to map images to a smaller latent space and designs a diffusion and denoising process in the latent space, significantly reducing the computational cost. Formally, given an image x and a text prompt y, the Stable Diffusion model is trained using the mean squared error loss:
[0050]
[0051] where z tis the vector in the latent space obtained by encoding the image through the autoencoder E at time step t. ∈~N(0, I) represents Gaussian noise, ∈ θ is a UNet model, c θ is the text encoder that maps the text y to a text embedding vector.
[0052] (2) DreamBooth technology achieves user customization by fine-tuning the parameters of the diffusion model. The training dataset of this technology includes a set X of specific subjects s and a set X of images belonging to the category of this subject c , and is accompanied by corresponding prompt statements y s and y c , such as "a photo of sks[class noun]" and "a photo of[class noun]". Among them, X s contains various private images for customization, while X c contains images of the same category as X s to alleviate overfitting of the model during training. Therefore, DreamBooth uses two parts of loss to train the diffusion model:
[0053]
[0054] Among them, represents the latent space vector mapped from the images in X s , represents the latent space vector mapped from the images in X c .
[0055] (3) In the diffusion model, the attention mechanism is crucial for capturing the dependencies between various spatial elements in the input data. The calculation of attention can be expressed as:
[0056]
[0057] Among them, Q and K represent the query and key respectively, and d k represents the dimension of the key.
[0058] In the Stable Diffusion model, there are two mechanisms: Cross-Attention, A c ), and self-Attention, A s ). For cross-attention, the text prompt y is first encoded into a text embedding vector and then extracted as the key K. The vector z t in the latent space is also partitioned and projected as the query Q, and at this time, A c can be calculated through the above formula. Similarly, the vector z tIt is also used to generate the query Q and is also used to generate the key K, and the same calculation method is adopted to generate the self-attention A s A c only represents the interaction relationship between the text prompt and the generated image, and A s reflects the internal structural relationship of the image, ensuring that the generated image is visually coherent and natural.
[0059] (4) The deep learning-based adversarial attack technology manipulates the model by introducing imperceptible perturbations to the input data. In image generation, its goal is to distort or invalidate the output of the generation model through the optimal perturbation δ. The perturbation is restricted within an η-ball according to the distance metric l p , where η represents the maximum allowable perturbation amplitude. This technology ensures that the output is different from the true value y by maximizing the loss function. The optimization of δ true is carried out in the following way: adv The optimization of δ
[0060] δ adv = arg max L(f(x + δ), y true ), s.t. ||δ p || ≤ η (4)
[0061] where δ adv represents the optimized adversarial perturbation, L represents the loss function of the training model, f() represents the deep learning model, x represents the input clean sample image, and δ p represents the p-norm of the adversarial perturbation.
[0062] Projected gradient descent is a widely used adversarial attack method. The perturbation is calculated by iteratively perturbing the input along the direction of formula (4). Each iteration process of projected gradient descent updates the adversarial sample x′ k as follows:
[0063] x′0 = x, (5)
[0064]
[0065] where x′ k represents the adversarial sample after k optimizations, represents the gradient on the sample x calculated according to the backpropagation of the gradient.
[0066] Here, ∏ x,η (z) restricts the pixel values of z within the η-ball, γ represents the step size, and k represents the number of iterations.
[0067] (2) Masked attention erasure:
[0068] To effectively conceal identity-related information in the generated images, the present invention proposes a Masked Cross-Attention Erasure (MCAE) loss. This loss function aims to eliminate the cross-attention in the diffusion model. By weakening the cross-attention related to identity prompt words in the model, this loss function can prevent the model from learning specific identity features.
[0069] Cross-attention map A c has the dimension of B×H×W×N, where B is the batch size, H and W are the height and width of the cross-attention map respectively, and N is the number of prompt words in the text sentence. The attention map of the identity text prompt word can be expressed as:
[0070]
[0071] where, A c [:,∶,∶,w s represents taking the w c -th feature of the A s matrix in the text dimension, and w s represents the position of the identity prompt word in the entire text.
[0072] Assume that the binary mask is represented as M, where M(x,y) = 1 represents the region containing identity information (such as the face), and M(x,y) = 0 represents the background region. The MCAE loss aims to minimize the attention energy in the masked region (i.e., the region where M(x,y) = 1), while maximizing the energy outside these regions. This can be mathematically expressed by the following formula:
[0073]
[0074] By erasing the cross-attention in the masked region, it prevents the model from associating the visual features corresponding to the identity prompt words with these regions, making it more difficult for the model to reproduce identity information related to the human face.
[0075] In addition to the cross-attention mechanism, the present invention also performs erasure on the self-attention mechanism of the diffusion model. Therefore, the present invention designs a Masked Self-Attention Erasure (MSAE) loss, aiming to suppress the correlations within the masked regions in the self-attention map, further reducing the model's ability to infer identity-related features. Assume A s represents the self-attention map, and the dimension of A s is B×H×W×D, where D = H×W represents the number of pixels in the self-attention map. To erase the corresponding values in each attention map, the mask M is resized from the H×W shape to the D×1 shape and denoted as M r .
[0076] The formula of the MSAE loss function is as follows:
[0077]
[0078] By combining the dual methods of eliminating masked cross-attention and self-attention, it helps to more comprehensively remove the identity features in the image.
[0079] (3) Cubic time step sampling:
[0080] The present invention uses a hybrid mass fraction to measure the impact of the diffusion model time step on adversarial attacks.
[0081] To calculate the hybrid mass fraction, the present invention first calculates the gradient G on the image during the backpropagation of the diffusion model t , and then calculates the L1 norm of the gradient through the following formula:
[0082]
[0083] Among them, represents the absolute value of the i-th element of the gradient matrix G t .
[0084] Then, G t is converted into a probability map p t through a softmax function. Then, calculate the entropy of p t :
[0085]
[0086] Among them, represents the value of the i-th element of the probability matrix p t .
[0087] The hybrid mass fraction (Hybrid Quality Score, HQS) can be calculated by the following formula:
[0088] HQS t = norm(N t ) - norm(H t ) (12)
[0089] Among them, norm represents the normalization process for all elements in the matrix.
[0090] The hybrid mass fraction reflects the importance of the diffusion model sampling time step for adversarial attacks. The present invention observes that in the early sampling time steps, the HQS is relatively high. As the time step increases, the HQS decreases, and these time steps are not conducive to the training of adversarial samples. Therefore, the present invention uses the following function for sampling:
[0091] t′ = T * (1 - (t / T) 3 ) (13)
[0092] Where t represents the current sampling time step, and T represents the maximum sampling time step of the diffusion model.
[0093] Finally, the training objective of the present invention is to maximize the triple loss, including the DreamBooth loss, the Masked Self - Attention Erasure (MSAE) loss, and the Masked Self - Attention Erasure (MSAE) loss, to ensure comprehensive protection of face information. That is, the total loss function is:
[0094] L total = L DB + λ1L MSAE + λ2L MCAE (14)
[0095] Where λ1 and λ2 are weight coefficients, and the values of λ1 and λ2 are preferably 0.1.
[0096] The process of adversarial training includes iteratively updating the perturbation to ensure that both cross - attention and self - attention are effectively minimized in the masked region. In each training cycle, first, the surrogate model is trained using a clean image set. Then, the Projected Gradient Descent (PGD) method is used for multiple - iteration training of adversarial samples, and finally, the model is retrained using these adversarial samples.
[0097] Finally, the adversarial perturbation trained using the above - mentioned loss function is added to the original image to obtain the protected image.
[0098] Key points of the present invention:
[0099] 1. The present invention proposes a customized privacy - protection method based on diffusion. Different from previous methods, the present invention studies the structure and training characteristics of the diffusion model itself, and trains stronger adversarial samples to protect privacy.
[0100] 2. The present invention deeply studies the relationship between images and texts, and proves the importance of texts in generation guidance. On this basis, a novel masked - attention - elimination module is proposed to "erase" the attention map related to the topic - recognizer tokens.
[0101] 3. Based on the relationship between the diffusion - model sampling process and the Projected Gradient Descent method, the present invention introduces a cubic function as a sampler corrector, enabling the model to sample more small time steps during training.
[0102] 4. Experimental results show that, in terms of two facial benchmarks and common prompts, the present invention has significantly improved the face recognition rate and image quality compared to the state-of-the-art methods on average.
[0103] Effects of the present invention:
[0104] 1. Datasets: The method of the present invention was quantitatively evaluated on two benchmark datasets: CelebA-HQ and VGGFace2. CelebA-HQ contains 30,000 images with a resolution of 1024×1024. VGGFace2 contains approximately 3.31 million images of about 9,131 individuals. For each dataset, 50 faces were selected as protected subjects. Each subject contains 8 different images and was evenly divided into two subsets: a clean image set and a target protection set. All images were center-cropped and resized to a unified resolution of 512×512.
[0105] 2. Evaluation metrics: The images protected by the present invention were fine-tuned for DreamBooth, and 16 images were generated for each subject and test prompt. Then, these images were evaluated using a set of metrics. The FID and BRISQUE metrics were used to test the image quality to test the perturbation results. Since the images generated by the perturbed DreamBooth model should lack detectable faces, the RetinaFace detector was used to quantify this, and this evaluation metric is called the Face Detection Failure Rate (FDFR). The ArcFace recognizer was also used to extract face features and calculate the cosine distance from the face features of the entire original dataset, and this evaluation metric is the Identity Score Matching (ISM).
[0106] 3. Experimental Results: To verify the effectiveness of the present invention, the present invention was compared with two currently best-performing diffusion model customization attack methods, Mist and Anti-DreamBooth. The source codes of them were reproduced, and adversarial samples were trained with equivalent noise thresholds. Then, these protected adversarial samples were used to fine-tune the diffusion model. Finally, the prompts "a photo of sks person" and "a dslr portrait of sks person" were input into the trained diffusion model, and the generated images were tested. As can be seen from Table 1, the method proposed by the present invention shows significant performance in four metrics and two prompts. A higher FDFR rate means fewer faces are detected, while a lower ISM rate indicates a lower similarity between the output image and the original image in terms of faces. In terms of image quality, higher FID and BRISQUE indicate lower image quality, which also shows that the output of the model customized by DreamBooth is significantly distorted.
[0107] 4. Robustness Experimental Analysis: To verify the noise robustness of the adversarial perturbations trained by the present invention, a series of ablation analysis experiments were conducted by the present invention.
[0108] (1) Noise budget. The noise budget η represents the allowed perturbation amplitude, and a larger noise budget is easily perceptible to the human eye. The present invention uses Learned Perceptual Image Patch Similarity (LPIPS) to measure the image quality before and after adding perturbations. Figure 3Among them, the first row is the original image protected by the present invention, the second row shows the magnified details within the red frame of the first row, and the third row shows the corresponding output generated by the customized diffusion model. As the noise budget η increases from left to right, the output quality of the diffusion model gradually deteriorates, and the identity information in the image also gradually disappears. This is because larger perturbations enhance the effects of the MSAE and MCAE loss functions. When the noise budget increases, the fluctuations and distortions in the image become more obvious, thus affecting the visual quality of the image. To quantify this impact, the LPIPS (Learned Perceptual Image Patch Similarity) metric is adopted to measure the perceptual similarity between the original image and the perturbed image, thereby providing an objective measure of quality degradation. In Table 1, η = 0.05 is set as the default value, and at this time the LPIPS is still very low, which means that the added adversarial perturbation is almost invisible. As shown in Table 2, when the budget is as small as 0.03, both the FID and BRISQUE scores of the output of the customized model generated by the present invention are relatively high, indicating that the privacy protection effect of the present invention is still effective under a very low noise budget. As the noise budget increases, the protection performance of the present invention becomes stronger.
[0109] (2) Privacy protection for different customization methods. To verify the generality of the present invention in various customized generation methods, the privacy protection effect of the present invention was tested on customized methods other than DreamBooth. These customized methods include: Low-Rank Adaptation (LoRA), Textual Inversion, and IP-Adapter. Adversarial samples were trained for LoRA, TI, and IP-Adapter respectively, and then the generation effects of these adversarial samples were tested on these customized methods. For IP-Adapter without an identity text prompt, image prompt features were used as a substitute during the training process. As Figure 4 shown, the present invention successfully disrupted the image outputs of all methods and effectively erased the identity information in the images generated by LoRA and IP-Adapter. However, for the images generated by Textual Inversion, the identity information is still partially retained. This may be due to its unique embedding mechanism: features related to the human face are precisely optimized and encoded into a high-dimensional latent feature, and it is not very effective to erase the identity from it.
[0110] (3) Robustness against various image preprocessings. Before being uploaded to the Internet or shared, digital images usually undergo various preprocessing techniques, such as compression and noise reduction, to optimize storage and improve visual quality. Among them, JPEG compression is a widely used technique that reduces the file size by removing redundant information. Similarly, Gaussian blur is a commonly used filtering method aimed at reducing image noise and details. Table 3 shows the robustness of the present invention against these preprocessing techniques. JPEG compression and Gaussian blur reduce the privacy protection effect of the present invention. Nevertheless, the BRISQUE value of the model output attacked by the present invention is still high, and the quality of the generated images is still poor, indicating that the adversarial samples trained by the present invention still maintain a certain degree of robustness after image preprocessing.
[0111] Table 1 Comparison of Privacy Protection in the Face Scenario
[0112]
[0113] Table 2 Influence of Noise Budget Amplitude on the Privacy Protection of the Present Invention
[0114] η LPIPS FDFR ISM FID BRISQUE 0 - 0.07 0.61 226 17.97 0.01 0 0.08 0.53 272 33.93 0.03 0.01 0.59 0.37 427 40.20 0.05 0.03 0.78 0.27 464 44.63 0.07 0.07 0.82 0.26 477 45.65 0.09 0.11 0.85 0.16 496 47.39
[0115] Table 3 Robustness Comparison of Image Post-Processing Methods
[0116]
[0117] Another embodiment of the present invention provides a customized privacy protection system for a diffusion model based on mask attention mechanism elimination, which includes:
[0118] A segmentation module for segmenting the mask of the face part from the given original image to be protected by using a segmentation model;
[0119] A cross-attention erasure module for erasing the cross-attention corresponding to the identity text prompt in the diffusion model by using a mask cross-attention erasure loss function;
[0120] A self-attention erasure module for erasing the self-attention in the diffusion model by using a mask self-attention erasure loss function to interfere with the relationship between the face pixels of the generated image;
[0121] A protected image generation module for training adversarial perturbations by using a mask cross-attention erasure loss function and a mask self-attention erasure loss function, and adding the adversarial perturbations to the original image to obtain a protected image.
[0122] The division of the above modules is only for illustrative purposes. In actual applications, the above functions can be assigned to different functional modules according to needs to complete all or part of the functions described in the foregoing method. For the specific working processes of the above modules, reference can be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0123] Another embodiment of the present invention provides a computer device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the method of the present invention.
[0124] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disc). When the computer program stored in the computer-readable storage medium is executed by a computer, the various steps of the method of the present invention are implemented.
[0125] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A customized privacy protection method for diffusion models based on the elimination of masked attention mechanism, characterized in that Including the following steps: Given the original image to be protected, use a segmentation model to segment the mask of the face part; Use the masked cross-attention erasure loss function to erase the cross-attention corresponding to the identity text prompt in the diffusion model; Use the masked self-attention erasure loss function to erase the self-attention in the diffusion model to interfere with the relationship between the face pixels of the generated image; Use the masked cross-attention erasure loss function and the masked self-attention erasure loss function to train the adversarial perturbation, and add the adversarial perturbation to the original image to obtain the protected image.
2. The method according to claim 1, wherein The masked cross-attention erasure loss function aims to minimize the attention energy in the masked area while maximizing the energy outside the masked area; the masked cross-attention erasure loss function is expressed as: Where M is a binary mask, where M(x,y)=1 represents the area containing identity information, and M(x,y)=0 represents the background area; is the attention map of the identity text prompt word, Among them A c [:,∶,∶,w s ] represents the cross attention map A c In the text dimension w s Features, w s Indicates the position of the identity cue word in the entire text.
3. The method according to claim 2, characterized in that, The masked self-attention erasure loss function is expressed as: Among them, A s represents the self-attention map, and the dimension of A s is B×H×W×D, where D = H×W represents the number of self-attention map pixels; in order to erase the corresponding values in each attention map, the mask M is resized from the H×W shape to the D×1 shape and denoted as M r .
4. The method according to claim 1, wherein Use the hybrid quality score to measure the impact of the time step of the diffusion model on the adversarial attack. The hybrid quality score is calculated by the following steps: Calculate the gradient G on the image during the backpropagation of the diffusion model t , and then calculate the L1 norm N of the gradient G t ; t ; Convert G to probability map p through the softmax function t and then calculate the entropy H of p t ; t t Calculate the mixed mass fraction: HQS t = norm(N t ) - norm(H t ), where norm represents the normalization process for all elements in the matrix.
5. The method according to claim 1, characterized in that, Use the cubic sampling function to correct the sampling time step in the diffusion model training, so that the diffusion model can learn at a smaller time step as much as possible, thereby enhancing the effect of the adversarial attack.
6. The method according to claim 5, wherein The cubic sampling function is expressed as: t′ = T * (1 - (t / T) 3 ) Where t represents the current sampling time step, and T represents the maximum sampling time step of the diffusion model.
7. The method according to claim 1, wherein Training adversarial perturbations using the masked cross-attention erasure loss function and the masked self-attention erasure loss function, the training objective is to maximize the triple loss, including the DreamBooth loss function L DB , the masked cross-attention erasure loss function L MCAE , the masked self-attention erasure loss function L MSAE , to ensure comprehensive protection of face information, that is, the total loss function is: L total = L DB + λ1L MSAE + λ2L MCAE , where λ1 and λ2 are weight coefficients; the process of adversarial training includes iteratively updating the perturbation to ensure that both cross-attention and self-attention are effectively minimized in the masked region; in each training cycle, first train the surrogate model using a clean image set, then use the projected gradient descent method to perform multiple iterations of training on adversarial samples, and finally retrain the model using the adversarial samples.
8. A customized privacy protection system for a diffusion model based on the elimination of the masked attention mechanism, characterized in that, Including: A segmentation module for, given the original image to be protected, using a segmentation model to segment the mask of the face part; A cross-attention erasure module for using the masked cross-attention erasure loss function to erase the cross-attention corresponding to the identity text prompt in the diffusion model; A self-attention erasure module for using the masked self-attention erasure loss function to erase the self-attention in the diffusion model to interfere with the relationship between the face pixels of the generated image; A protected image generation module for using the masked cross-attention erasure loss function and the masked self-attention erasure loss function to train the adversarial perturbation, and adding the adversarial perturbation to the original image to obtain the protected image.
9. A computer device, characterized in that, Including a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Image text recognition method and system based on mask diffusion model, storage medium and equipment
CN121095960A
Vision-language unified processing method based on mask puzzle
CN121211502A
An adversarial attack method based on diffusion model and structure perception partition reconstruction
CN122551413A
An adversarial attack method based on diffusion model and structure perception partition reconstruction
CN122551413B