Multi-target countermeasure attack method and device based on text description
Through the multi-objective adversarial attack method based on text description, the preset model is used to extract features and generate adversarial perturbation images, which solves the problem of low attack efficiency of multi-objective class in the prior art, and achieves efficient and flexible adversarial sample generation.
Patent Information
- Application Number
- CN202511012504.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The existing deep neural network adversarial sample generation methods are inefficient and inflexible in multi-objective class attack scenarios, and cannot efficiently generate adversarial perturbations against multiple target classes.
Using a multi-objective adversarial attack method based on text description, features are extracted through the text encoder and image encoder in the preset model, and an opposition perturbation image is generated using the semantic fusion module and generator. Combining the cross entropy loss function, the KL divergence loss function and the regularization term, a multi-objective adversarial sample is generated.
Improves the efficiency and flexibility of adversarial sample generation, and eliminates the need for retraining the generator. It can dynamically generate adversarial perturbation images for multiple target categories, improving attack success rate and robustness.
Smart Images

Figure CN120510487A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a multi-target counterattack method and device based on text description. Background Art
[0002] With the widespread application of deep neural networks (DNNs) in fields such as image classification and speech recognition, their security has received increasing attention. Research has found that DNNs are significantly vulnerable to adversarial examples, which can mislead the model's decision-making by adding perturbations that are imperceptible to humans.
[0003] In related technologies, the use of generative adversarial networks (GANs) to generate adversarial examples can significantly improve the efficiency of adversarial attacks. However, this approach typically only targets a single target category and requires retraining the generative model for each target category, increasing computational costs. This approach is particularly inefficient and lacks flexibility when attacking multiple target categories. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide a multi-target adversarial attack method based on text description, which improves the attack efficiency and flexibility by targeting adversarial disturbances of multiple target categories.
[0005] To achieve the above technical objectives, on the one hand, the present invention provides a multi-target adversarial attack method based on text description, including: extracting text features in the target text based on a text encoder in a preset model, and the target text has a corresponding target category; extracting image features in the original image based on an image encoder in the preset model, and the category of the original image is inconsistent with the target category; based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features; based on the generator in the preset model, the semantic features are generated into an adversarial perturbation image; the adversarial perturbation image is added to the original image to generate a corresponding adversarial image sample.
[0006] Specifically, the semantic fusion module based on the preset model fuses the text features and the image features to obtain fused semantic features, including: utilizing the linear transformation in the semantic fusion model, embedding the text features as a query, and using the latent encoding of the text features as keys and values to determine an attention map in the spatial and channel dimensions; based on the attention map, determining the parameters of the linear transformation, and utilizing the parameters to map the text features to the latent subspace of the corresponding preset model, wherein the latent subspace is a subspace of the image features, so that the text features are fused with the image features.
[0007] Specifically, the generator based on the preset model generates an adversarial perturbation image from the semantic features, including: based on the text feature distribution of the target category, the generator generates an adversarial perturbation image from the semantic features, so as to offset the semantic features of the adversarial image sample to the feature center of the target category through the adversarial perturbation image.
[0008] Specifically, the generating of the adversarial perturbation image from the semantic features based on the generator in the preset model includes: decoding the semantic features based on the decoder in the generator to obtain the adversarial perturbation image.
[0009] In addition, the method also includes: generating a composite loss function based on the cross entropy loss function, the KL divergence loss function and the regularization term; and generating the generator based on the composite loss function.
[0010] Wherein, the preset model includes: a comparative language-image pre-training CLIP framework.
[0011] In addition, the method also includes: using the encoder in the generator to extract image features of the original image; after obtaining the fused semantic features, splicing the semantic features with the extracted image features; decoding the spliced image features through the decoder in the generator to obtain the adversarial perturbation image.
[0012] In addition, the method also includes: receiving an image of a first preset size for carrying the text feature and an image of a first preset size for carrying the image feature through a multi-scale fusion module in the preset model, and generating three feature images of different scales corresponding to the text feature image and the image feature image respectively through an adaptive pooling operation; respectively for the text feature image and the image feature image, copying the corresponding three feature images of different scales along the spatial dimension to generate a feature image of the first preset size; receiving the original image, and using a lightweight network to extract features of the original image to determine the fusion weight of the multi-scale features, which acts on feature maps of different scales; respectively for the text feature image and the image feature image, performing weighted fusion on the corresponding three feature images of different scales and the corresponding feature image of the first preset size based on the fusion weight to obtain a feature image of the second preset size to fuse the text feature and the image feature.
[0013] In addition, the method further includes: after generating the adversarial image sample, obtaining a high-frequency component of the image sample, and constraining the high-frequency component.
[0014] On the other hand, the present invention provides a multi-target adversarial attack device based on text description, including: a first extraction module, used to extract text features in the target text based on the text encoder in the preset model, and the target text has a corresponding target category; a second extraction module, used to extract image features in the original image based on the image encoder in the preset model, and the category of the original image is inconsistent with the target category; a fusion module, used to fuse the text features and the image features based on the semantic fusion module in the preset model to obtain fused semantic features; a generation module, used to generate an adversarial perturbation image from the semantic features based on the generator in the preset model; and an adding module, used to add the adversarial perturbation image to the original image to generate a corresponding adversarial image sample.
[0015] In an embodiment of the present application, text features in a target text are extracted based on a text encoder in a preset model, and the target text has a corresponding target category; image features in an original image are extracted based on an image encoder in the preset model, and the category of the original image is inconsistent with the target category; based on a semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features; based on a generator in the preset model, an adversarial perturbation image is generated from the semantic features; and the adversarial perturbation image is added to the original image to generate a corresponding adversarial image sample.
[0016] Among them, by extracting text features from the target text based on the text encoder in the preset model and extracting image features from the original image based on the image encoder in the preset model, multimodal feature extraction is achieved. Based on the semantic fusion module in the preset model, the text features and image features are fused, and based on the generator in the preset model, corresponding adversarial samples are obtained. Therefore, different target or multi-target adversarial perturbation images can be generated by dynamically combining text features and image features according to different target requirements, and applied to the input image to generate adversarial samples. There is no need to retrain the generator, thereby improving the efficiency and flexibility of adversarial sample generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A flowchart of a multi-target counterattack method based on text description according to an embodiment of the present application; Figure 2 A schematic diagram of a multi-target adversarial attack process based on text description according to an embodiment of the present application; Figure 3A Schematic diagram of the relationship between the change in the number of categories and the attack success rate in an embodiment of the present application; Figure 3B Schematic diagram of the relationship between the change in the number of categories and the attack success rate in an embodiment of the present application; Figure 4 A schematic diagram of a specific case of an embodiment of the present application; Figure 5 This is a schematic diagram of the framework of a multi-target anti-attack device based on text description according to an embodiment of the present application; Figure 6 This is a schematic diagram of the framework of a multi-target anti-attack device based on text description in an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0020] like Figure 1 As shown, the embodiment of the present application also provides a multi-target anti-attack method based on text description, and the method 100 includes: S101: Extract text features from the target text based on the text encoder in the preset model.
[0021] Among them, the target text has a corresponding target category.
[0022] S102: Extracting image features from the original image based on the image encoder in the preset model.
[0023] Among them, the category of the original image is inconsistent with the target category.
[0024] S103: Based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features.
[0025] S104: Based on the generator in the preset model, generate an adversarial perturbation image using the semantic features.
[0026] S105: Add the adversarial perturbation image to the original image to generate a corresponding adversarial image sample.
[0027] It should be noted that the method 100 can be applied to server components such as a server side, a service cluster side, or a cloud side.
[0028] The following is a detailed explanation of the above steps: S101: Extract text features from the target text based on the text encoder in the preset model.
[0029] The target text has a corresponding target category. For example, Figure 2 As shown, the target text may be "monitor" and its category may be "monitor".
[0030] The preset model can be a known model, a custom model, or a variant of a known model. Regardless of the model, it should have encoders for two or more modes, such as a text encoder and an image encoder. Encoders for other modes can also be added as needed.
[0031] The preset model includes a comparative language-image pre-training CLIP framework, which may include the above-mentioned encoders, such as a text encoder and an image encoder.
[0032] For example, as mentioned above, and as Figure 2 As shown in Figure 1, a text description of a target category (monitor) can be input into the CLIP text encoder to extract text features of the target category. The extracted text features are represented as a high-dimensional embedding vector of the target category, which is used to describe the semantic information of the target category.
[0033] S102: Extracting image features from the original image based on the image encoder in the preset model.
[0034] The original image category and the target category are inconsistent. If the original image could be a vehicle image, then the category should be vehicle, i.e., a vehicle image. However, the aforementioned text category (e.g., "monitor") does not correspond to the vehicle image. In other words, the target category of the text is used to interfere with the vehicle image. The goal is to ensure that the generated adversarial example is misclassified as the target category by the subsequent target classifier.
[0035] As mentioned above, and as Figure 2 As shown in Figure 2, at the same time, the input original image can also be fed into the CLIP picture encoder (i.e., image encoder) to extract the image's visual features, i.e., image features. Image features are represented as multi-dimensional embedding vectors of the input original image, which are used to describe the visual content.
[0036] Since this embodiment can process text and image modalities simultaneously, it extracts and fuses features of the input image with the text description of the target category, and dynamically generates adversarial perturbations adapted to the target category through a generative model. Therefore, through the multimodal capabilities of CLIP, it is possible to dynamically combine text descriptions and image features to generate multi-target adversarial perturbations, and apply them to the input image to generate adversarial samples.
[0037] S103: Based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features.
[0038] The semantic fusion module establishes a latent feature space mapping between the image's feature space and CLIP's text embedding space (i.e., the text feature space). Through this mapping, the model can generate image perturbations based on various semantic text cues, thereby enabling adversarial attacks.
[0039] According to the above, if Figure 2 As shown in FIG, the extracted text features and image features are input into the semantic module (i.e., the semantic fusion module) for fusion.
[0040] Specifically, the semantic module may include a semantic modulation module and a fusion module. The semantic modulation module receives the above features, and during the fusion process, the text features and the image features interact with each other through the semantic modulation mechanism in the semantic modulation module to form a unified feature representation.
[0041] The semantic modulation module uses linear mapping and a bidirectional attention mechanism to calculate attention maps in both the spatial and channel dimensions to refine the injection of semantic information. The fusion module then generates fused features that not only incorporate the visual features of the input image but also the textual semantic information of the target category, laying the foundation for the subsequent generation of adversarial perturbations.
[0042] Specifically, based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain the fused semantic features, including: using the linear transformation in the semantic fusion model, embedding the text features as the query, and using the latent encoding of the text features as the key and value to determine the attention map in the spatial and channel dimensions; based on the attention map, determining the parameters of the linear transformation, and using the parameters to map the text features to the latent subspace of the corresponding preset model, where the latent subspace is a subspace of the image features, so that the text features are fused with the image features.
[0043] The semantic fusion module uses a linear transformation to compute cross-attention between text embeddings (text features) as queries and latent codes as keys and values. This process is performed in the spatial and channel dimensions, generating two attention maps. Based on these attention maps, the parameters of the linear transformation (e.g., translation and scaling parameters) are calculated. These parameters are then used to map the text cue embeddings (text features) into the corresponding CLIP latent subspace, i.e., the space of image features, for feature fusion.
[0044] This module uses a two-dimensional attention mechanism (spatial and channel) to learn how to align visual and textual details. For example, in the spatial dimension, the scaling parameter is used to adjust the contribution of each spatial position, while in the channel dimension, the translation parameter of each channel is fine-tuned to accurately adjust the match in the latent space of the above image features.
[0045] It should be understood that the semantic fusion module can also be divided into a semantic modulation module and a fusion module.
[0046] S104: Based on the generator in the preset model, generate an adversarial perturbation image using the semantic features.
[0047] As described above, the fused semantic features are input into the generator to generate an adversarial perturbation image, which is added to the input original image to form an adversarial sample.
[0048] Among them, Figure 2 As shown in , the generator can include the above semantic module (i.e., semantic fusion module), or the semantic module can be decoupled from the generator. When the generator includes the above semantic module, the generator can directly generate adversarial perturbation images based on the semantic features, i.e., Figure 2 The adversarial perturbation p shown x,t The adversarial perturbation data may be image data.
[0049] Specifically, based on a generator in a preset model, semantic features are used to generate adversarial perturbation images, including: based on a decoder in the generator, the semantic features are decoded to obtain the adversarial perturbation images.
[0050] As mentioned above, in the process of generating adversarial perturbations, the semantic feature decoder, i.e., the generator decoder, is used to decode the semantic features to obtain the corresponding perturbations.
[0051] In order to generate perturbations more accurately, the generator needs to generate perturbations according to the following mechanism.
[0052] Specifically, based on the generator in the preset model, the semantic features are generated into an adversarial perturbation image, including: based on the text feature distribution of the target category, the generator generates an adversarial perturbation image from the semantic features, so as to offset the semantic features of the adversarial image sample to the feature center of the target category through the adversarial perturbation image.
[0053] As mentioned above, the generator calculates the feature offset of the input original image as the difference between the target category feature and the current category feature, so that the generated adversarial sample can be more accurately misclassified as the target category by the target classifier.
[0054] Among them, the target category is the target category of the above text, which serves as the interference category, and the current category is the category to which the original image belongs.
[0055] The generator can include two modules: a feature input module and a perturbation generation module. The feature input module is used to receive fused image features and text features, while the perturbation generation module generates adversarial perturbations through multi-layer convolution operations. The decoder can be set in the perturbation generation module to perform deconvolution operations.
[0056] In order to generate a perturbation that is more consistent with the specific original image feature input, the original image can also be used to perform image splicing of image features and semantic features to generate perturbations.
[0057] Specifically, the method 100 also includes: using the encoder in the generator to extract image features from the original image; after obtaining the fused semantic features, splicing the semantic features with the extracted image features; decoding the spliced image features through the decoder in the generator to obtain an adversarial perturbation image.
[0058] like Figure 2 As shown in the figure, according to the above description, the original image is input into the generator encoder in the generator to perform a convolution operation to obtain the image features of the original image, and the image features and semantic features are spliced to obtain the spliced image features, and then the deconvolution operation is performed according to the generator decoder to obtain the adversarial perturbation image.
[0059] It should be noted that the above Figure 2 The image features extracted by the image encoder are a 512-dimensional vector, while the generator encoder is a convolutional network that extracts spatial features. Furthermore, as mentioned above, it is also possible to bypass the generator encoder encoding of the original image and directly decode it using semantic features to obtain the adversarial perturbation image.
[0060] In addition, the generator is obtained through training, but the generator of this embodiment can meet the needs of multiple categories without multiple training.
[0061] Specifically, the method 100 further includes: generating a composite loss function based on the cross entropy loss function, the KL divergence loss function, and the regularization term; and generating a generator based on the composite loss function.
[0062] To achieve multi-target adversarial attacks, the generator uses a composite loss function to optimize the generation process. The loss function consists of the following three parts: Cross entropy loss function: ensures that the generated adversarial samples are misclassified as target categories by the subsequent target classifier. The specific formula is as follows: in, Predicting adversarial examples for the model Belong to category The probability distribution of C is the total number of categories, that is, the possible number of categories in the classification task, q c For the target category One-hot encoding of class c.
[0063] KL divergence loss function: optimizes the similarity between the distribution of adversarial samples and the characteristics of the target category. The specific formula is as follows: in, The function is used to calculate the similarity between two features. is the image feature of the adversarial sample, is the text feature of the target category, and KL is the KL divergence calculation function. Is to predict adversarial samples for the model Belong to category The probability distribution of .
[0064] Regularization term: constrains the magnitude of adversarial perturbations to ensure the imperceptibility of adversarial examples. The specific formula is as follows: Where x is the original image.
[0065] Combining the above three parts, the final composite loss function is: Among them, α, β, and γ are hyperparameters for adjusting the weights of each loss term.
[0066] S105: Add the adversarial perturbation image to the original image to generate a corresponding adversarial image sample.
[0067] Among them, the adversarial image samples are the adversarial samples mentioned above, which are used for the subsequent training of the target classifier.
[0068] According to the above, the generated adversarial perturbation is added to the original image to obtain the adversarial sample. The adversarial sample satisfies the following constraints: the perturbation amplitude is limited to Within the threshold range, the perturbation is imperceptible to the human eye while effectively misleading the target classifier. The resulting adversarial examples can target multiple target categories simultaneously, enabling dynamic switching of multi-target adversarial attacks.
[0069] like Figure 2 As shown, when generating adversarial samples x itj adv Finally, the adversarial example is fed into the white-box model (target classification model or target classifier) to obtain the white-box model's predicted category probability distribution for the adversarial example. The first part of the training loss is calculated by calculating the cross-entropy loss. The second part of the training loss is then calculated by calculating the KL loss, which is calculated by comparing the probability distribution obtained by calculating the similarity between the target category text features and the text features of all categories, with the category probability distribution predicted by the white-box model. This allows the white-box model to be iteratively updated using these two parts of the training loss.
[0070] like Figure 4 As shown in Figure 1, the final generated adversarial image samples include vehicle images with the "rooster" text category, traffic road section images with the "kangaroo" text category, and traffic signs with the "tiger" text category.
[0071] To verify the effectiveness of the method proposed in this example, experiments were conducted on multiple datasets, and ablation studies were performed on the impact of different modules and loss functions. The experimental results show that the method proposed in this example has significant performance advantages in multi-target adversarial attack tasks.
[0072] The experiments used the following public datasets: CIFAR10: Contains 60,000 32×32 color images divided into 10 categories, with 6,000 images per category. CIFAR100: An extension of CIFAR10, containing 100 categories, with 600 images per category. ImageNet-1k: Contains 1,000 categories, with at least 1,000 images per category, offering broad coverage and rich categories.
[0073] In addition, if Figure 3A As shown in the figure, in the CIFAR10 dataset, the relationship between the attack success rate and the number of categories can be seen from the figure. Three different lines can be seen from the figure. These three lines represent the attack success rate of the category with the highest attack success rate, the average attack success rate of the category, and the attack success rate of the category with the lowest attack success rate. From high to low, the attack success rate shows a downward trend with the number of categories. Figure 3BAs shown in the figure, the relationship between attack success rate and the number of categories in the CIFAR100 dataset is shown. Three different lines can be seen in the figure. These lines represent the attack success rate of the category with the highest attack success rate, the average attack success rate of the category, and the attack success rate of the category with the lowest attack success rate. From high to low, the attack success rate decreases with the number of categories. This shows that as the number of categories increases, the attack success rate decreases. The following experiments were conducted in the embodiments of this application to improve the attack success rate.
[0074] The experiment was conducted on a single NVIDIA 2080 GPU. The training process lasted about 10 hours, and the inference time for generating adversarial perturbations per image was 0.015 seconds. The CLIP framework used in the experiment was the pre-trained ViT-B / 32, and the weight parameters in the loss function were set to α=250, β=250, and γ=1. The training was conducted for a total of 120 epochs (training rounds), with an initial learning rate of , after 60 epochs, it is reduced to For different data sets, the perturbation amplitude limit They are set to 16 (CIFAR10), 32 (CIFAR100), and 64 (ImageNet-1k) respectively.
[0075] In the CIFAR10 and CIFAR100 datasets, the top 10 categories were selected as target attack categories. For the ImageNet-1k dataset, 10 target attack categories were randomly selected from the dataset due to the similar semantics of the top 10 categories. All experiments were trained and tested on the dataset.
[0076] To evaluate the effectiveness of the method proposed in this example, ablation experiments were conducted on the CIFAR10 and CIFAR100 datasets. First, the impact of the semantic module in the generator and the KL divergence term in the loss function were evaluated. The attack success rate (ASR) was used as an evaluation metric to measure the performance of the attack. The experimental results are shown in Table 1 below. As can be seen from the results, the introduction of the semantic module improves the average attack success rate of ten categories on the CIFAR10 and CIFAR100 datasets, indicating that the module effectively embeds semantic information into the image modality. Adding the KL divergence term to the loss function also slightly improves the attack success rate. Notably, the best experimental results were achieved when the semantic module and the KL divergence term were used simultaneously.
[0077] Overall, the embodiments of this application significantly improve the effectiveness of adversarial attacks. The introduction of the semantic module not only enhances the semantic consistency of the generated adversarial samples, but also provides clearer guidance for the generation of adversarial samples. At the same time, the introduction of the KL divergence term further optimizes the loss function structure of the generator, thereby obtaining adversarial samples that better match the characteristics of the target model.
[0078] Table 1 Here, ClipAdv refers to the CLIP framework mentioned above, and sem refers to the aforementioned semantic module.
[0079] The method proposed in this example is applicable to a variety of datasets (such as CIFAR10 and ImageNet-1k), and its performance and effectiveness can be verified in tasks such as classification and detection. Experiments show that this example can simultaneously implement efficient adversarial attacks on multiple target categories, demonstrating strong versatility and robustness.
[0080] In order to generate adversarial perturbations of different scales to better attack the image, this embodiment also adds a multi-scale fusion module. The multi-scale fusion module aims to effectively integrate feature information of different scales to improve feature expression ability and robustness.
[0081] Specifically, the method 100 also includes: receiving an image of a first preset size for carrying text features and an image of a first preset size for carrying image features through a multi-scale fusion module in a preset model, and generating feature images of three different scales corresponding to the text feature image and the image feature image respectively through an adaptive pooling operation; respectively for the text feature image and the image feature image, copying the corresponding feature images of three different scales along the spatial dimension to generate a feature image of the first preset size; receiving the original image, and using a lightweight network to extract features of the original image to determine the fusion weight of the multi-scale features, and acting on feature maps of different scales; respectively for the text feature image and the image feature image, performing weighted fusion on the corresponding feature images of three different scales and the corresponding feature image of the first preset size based on the fusion weight, to obtain a feature image of the second preset size, so as to fuse the text features and the image features.
[0082] It should be noted that the multi-scale fusion module can be set before the semantic fusion module.
[0083] For example, as described above, this module first receives text features and image features from a CLIP encoder, with the corresponding text features and image features belonging to an input feature map of size C×H×W. It then uses adaptive pooling to generate three feature maps of different scales: C×H / 2×W / 2, C×H / 4×W / 4, and C×H / 8×W / 8. It then restores the original size using a copy-and-concatenate method, replicating the three scaled feature maps along the spatial dimensions to form a single feature map of size C×H×W, i.e., a text feature map and an image feature map of that size. Furthermore, this module receives an additional original image as described above and extracts features from it using a lightweight network (e.g., a convolutional network). This module then learns the fusion weights for the multi-scale features, applying them to the four feature maps of different scales (C×H / 2×W / 2, C×H / 4×W / 4, C×H / 8×W / 8, and C×H×W) to control their contributions in the fusion process. Ultimately, after weighted fusion, a feature map (such as image features and text features) of size 4C × H × W is obtained. This map contains information from different scales and adaptively weights the image content, improving the model's ability to express multi-scale information. The weighted fused image features can be input into the semantic fusion module for further fusion.
[0084] To better incorporate semantic information into adversarial perturbations, this embodiment also incorporates adaptive instance normalization (AdaIn). AdaIN adjusts the mean and standard deviation by introducing external semantic features s, thereby integrating semantic information into the normalization process. Specifically, given content features x and external semantic features s, AdaIN calculates: in and are the mean and standard deviation of the content characteristics. and are the new mean and standard deviation calculated from the external semantic features.
[0085] Since s is a global semantic vector, this embodiment uses a multi-layer perceptron (MLP) to calculate the normalization parameter corresponding to s: in Used to scale feature distributions, Used to adjust the mean of features. MLP is a simple two-layer fully connected network.
[0086] Finally, the semantic feature s is calculated by MLP and , and then acts on the AdaIN normalization formula to dynamically adjust the content feature x. Here, the content feature x refers to the fused semantic feature, while the semantic feature s is the newly introduced semantic feature.
[0087] Generating perturbations with high transferability is a key challenge in adversarial attacks. To improve the universality of perturbations across different models and inputs, a high-frequency loss can be introduced to control the high-frequency components of the generated perturbations. High-frequency perturbations often contain details and local variations, and are prone to failure across different inputs and models. Therefore, penalizing the high-frequency components helps generate more universal perturbations and improves the transferability of the attack.
[0088] Specifically, the method 100 further includes: after generating the adversarial image sample, obtaining the high-frequency component of the image sample, and constraining the high-frequency component.
[0089] High-frequency loss refers to controlling the frequency components of the perturbation so that it is primarily concentrated in low-frequency regions. Low-frequency perturbations typically correspond to large-scale image changes and have strong mobility, while high-frequency components affect local details and are easily affected by model and data diversity. In the frequency domain, the low-frequency components of the perturbation are smoother and more widely applicable, while the high-frequency components are typically concentrated on details and noise.
[0090] To achieve this goal, after obtaining the final adversarial example, the generated perturbed image (perturbed sample) is first converted to the frequency domain via Fourier transform (Fourier Transform) or discrete cosine transform (DCT). The low-frequency components in the frequency domain contain macroscopic variations in the perturbation, while the high-frequency components correspond to details and rapid changes in the image. By applying penalties and constraints to the high-frequency components, the generation of detailed perturbations can be reduced, thereby making the perturbation more consistent across different images and models.
[0091] Among them, the realization formula of high-frequency loss is expressed as: in, represents the generated perturbation image, is the perturbation image The Fourier transform of Is a mask function used to extract low-frequency areas, i and j are the number of rows and columns of image pixels respectively. Here, the mask The low frequency region is set to 0 (low frequencies are preserved), while the high frequency portion is assigned a non-zero value, allowing the high frequency components to be penalized specifically. Specifically, It is a binary mask with a value of 0 in the low-frequency area and a value of 1 in the high-frequency area. This mask function can effectively isolate low-frequency and high-frequency components and impose a penalty on high-frequency components.
[0092] like Figure 5 As shown, the present invention provides a multi-target adversarial attack device 500 based on text description, including: a first extraction module 501, used to extract text features in a target text based on a text encoder in a preset model, where the target text has a target category, and the target category does not correspond to the target text; a second extraction module 502, used to extract image features in an original image based on an image encoder in a preset model; a fusion module 503, used to fuse text features and image features based on a semantic fusion module in a preset model to obtain fused semantic features; a generation module 504, used to generate an adversarial perturbation data image from the semantic features based on a generator in a preset model; and an adding module 505, used to add the adversarial perturbation data image to the original image to generate a corresponding adversarial image sample.
[0093] Specifically, the fusion module 503 includes: a transformation unit, which is used to use the linear transformation in the semantic fusion model to embed text features as queries, and use the latent encoding of text features as keys and values to determine the attention map in spatial and channel dimensions; a mapping unit, which is used to determine the parameters of the linear transformation based on the attention map, and use the parameters to map the text features to the latent subspace of the corresponding preset model, where the latent subspace is a subspace of the image features, so that the text features are fused with the image features.
[0094] Specifically, the generation module 504 is specifically used to: based on the text feature distribution of the target category, the generator generates an adversarial perturbation image using the semantic feature, so as to shift the semantic feature of the adversarial image sample to the feature center of the target category through the adversarial perturbation image.
[0095] Specifically, the generation module 504 is specifically configured to: decode the semantic features based on the decoder in the generator to obtain an adversarial perturbation image.
[0096] In addition, the device 500 also includes: a determination module, specifically used to generate a composite loss function based on the cross entropy loss function, the KL divergence loss function and the regularization term; and generate a generator based on the composite loss function.
[0097] Among them, the preset models include: comparative language-image pre-training CLIP framework.
[0098] In addition, the apparatus 500 further includes a third extraction module configured to extract image features from the original image using the encoder in the generator; a splicing module configured to splice the fused semantic features with the extracted image features after obtaining the fused semantic features; and a generation module 504 configured to decode the spliced image features using the decoder in the generator to obtain an adversarially perturbed image.
[0099] In addition, the device 500 also includes: a receiving module, which is used to receive an image of a first preset size for carrying text features and an image of a first preset size for carrying image features through a multi-scale fusion module in a preset model, and generate three feature images of different scales corresponding to the text feature image and the image feature image through an adaptive pooling operation; a copying module, which is used to copy the corresponding feature images of three different scales along the spatial dimension to generate a feature image of a first preset size for the text feature image and the image feature image respectively; a receiving module, which is used to receive the original image and use a lightweight network to extract features from the original image to determine the fusion weight of the multi-scale features, which acts on feature maps of different scales; a weighted fusion module, which is used to perform weighted fusion on the corresponding feature images of three different scales and the corresponding feature image of the first preset size based on the fusion weight for the text feature image and the image feature image respectively, to obtain a feature image of a second preset size to fuse the text features and the image features.
[0100] In addition, the apparatus 500 further includes: a constraint module, configured to obtain high-frequency components of the image samples after generating the adversarial image samples, and constrain the high-frequency components.
[0101] For other embodiments of the apparatus 500, reference may be made to the method embodiments described above, which correspond to the embodiment of the method 100 and are not described in detail here.
[0102] like Figure 6 As shown, the embodiment of the present application further provides a multi-target anti-attack device 600 based on text description, including: a memory 601 and a processor 602, the processor 602 reads a computer program in the memory 601, and is configured to perform the following operations: Based on the text encoder in the preset model, text features in the target text are extracted, where the target text has a target category, and the target category does not correspond to the target text; based on the image encoder in the preset model, image features in the original image are extracted; based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features; based on the generator in the preset model, the semantic features are used to generate an adversarial perturbation image; the adversarial perturbation image is added to the original image to generate a corresponding adversarial image sample.
[0103] For other embodiments of the device 600, reference may be made to the aforementioned embodiments of the apparatus 500 and the method 100, which correspond to the embodiments of the apparatus 500 and the method 100 and are not described in detail here.
[0104] It should be understood that the specific order or hierarchy of steps in the disclosed processes is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The accompanying method claims present elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.
[0105] While the above descriptions of embodiments and examples of the present invention are provided for the purpose of providing a more detailed and complete description of the present disclosure, they are not intended to be the only ways to implement or use the embodiments of the present invention. The embodiments cover features of various embodiments, as well as the method steps and sequences for constructing and operating these embodiments. However, other embodiments may be used to achieve the same or equivalent functionality and sequence of steps.
[0106] In the foregoing detailed description, various features are grouped together in a single embodiment to simplify the disclosure. This method of disclosure should not be interpreted as reflecting an intention that embodiments of the claimed subject matter require more features than are expressly recited in each claim. On the contrary, as reflected in the appended claims, the invention comprises less than all the features of any individual disclosed embodiment. The appended claims are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate preferred embodiment of the invention.
[0107] The above description of the disclosed embodiments is intended to enable any person skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments presented herein but is intended to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0108] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purposes of describing the above embodiments, but one of ordinary skill in the art will recognize that the various embodiments may be further combined and arranged. Therefore, the embodiments described herein are intended to encompass all such changes, modifications and variations that fall within the scope of the appended claims. Furthermore, to the extent the term "comprising" is used in the specification or claims, the term is intended to be encompassed in a manner similar to the term "including," as explained in terms of "including," used as a transitional word in the claims. Furthermore, any use of the term "or" in the specification of the claims is intended to mean a "non-exclusive or."
[0109] Those skilled in the art will also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of the two. To clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps mentioned above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.
[0110] The various illustrative logic blocks or units described in the embodiments of the present invention may be implemented or operated using a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor may be a microprocessor, or alternatively, any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented using a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0111] The steps of the methods or algorithms described in the embodiments of the present invention may be directly embedded in hardware, a software module executed by a processor, or a combination of the two. The software module may be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. For example, the storage medium may be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Alternatively, the storage medium may also be integrated into the processor. The processor and storage medium may be provided in an ASIC, which may be provided in a user terminal. Alternatively, the processor and storage medium may also be provided in different components in the user terminal.
[0112] In one or more exemplary designs, the functions described in the embodiments of the present invention can be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium or transmitted in the form of one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media that facilitate the transfer of computer programs from one location to another. Storage media can be any usable medium that can be accessed by a general-purpose or specialized computer. For example, such computer-readable media can include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions, data structures, and other forms readable by a general-purpose or specialized computer or processor. In addition, any connection can be appropriately defined as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote resource via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless methods such as infrared, wireless, and microwave, it is also included in the definition of computer-readable media. The aforementioned disks and discs include compact disks, laser disks, optical disks, DVDs, floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while discs typically reproduce data optically using lasers. Combinations of the above may also be included in computer-readable media.
[0113] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-target adversarial attack method based on text description, characterized in that: include: Extracting text features from a target text based on a text encoder in a preset model, wherein the target text has a corresponding target category; extracting image features from an original image based on an image encoder in the preset model, where a category of the original image is inconsistent with the target category; Based on the semantic fusion module in the preset model, the text features and the image features are fused to obtain fused semantic features; Based on the generator in the preset model, generating an adversarial perturbation image using the semantic features; The adversarial perturbation image is added to the original image to generate a corresponding adversarial image sample.
2. The method according to claim 1, characterized in that The semantic fusion module in the preset model is used to fuse the text features and the image features to obtain fused semantic features, including: Determine an attention map in spatial and channel dimensions using a linear transformation in a semantic fusion model, embedding text features as queries, and latent encodings of the text features as keys and values; Based on the attention map, parameters of the linear transformation are determined, and the parameters are used to map the text features to a latent subspace of a corresponding preset model, where the latent subspace is a subspace of the image features, so that the text features are fused with the image features.
3. The method according to claim 1, characterized in that The step of generating an adversarial perturbation feature image from the semantic feature based on the generator in the preset model includes: Based on the text feature distribution of the target category, the generator generates an adversarial perturbation image using the semantic feature, so as to shift the semantic feature of the adversarial image sample to the feature center of the target category through the adversarial perturbation image.
4. The method according to claim 1, wherein The step of generating an adversarial perturbation image from the semantic features based on a generator in the preset model includes: Based on the decoder in the generator, the semantic features are decoded to obtain an adversarial perturbation image.
5. The method according to claim 1, 3 or 4, characterized in that The method further comprises: Generate a composite loss function based on the cross entropy loss function, KL divergence loss function and regularization term; The generator is generated based on the composite loss function.
6. The method according to claim 1, characterized in that The preset model includes: a comparative language-image pre-training CLIP framework.
7. The method according to claim 1, characterized in that The method further comprises: Use the encoder in the generator to extract image features from the original image; After obtaining the fused semantic features, the semantic features are spliced with the extracted image features; The spliced image features are decoded by the decoder in the generator to obtain the adversarial perturbation image.
8. The method according to claim 1, characterized in that The method further comprises: An image of a first preset size for carrying the text feature and an image of a first preset size for carrying the image feature are received through a multi-scale fusion module in the preset model, and feature images of three different scales corresponding to the text feature image and the image feature image are generated respectively through an adaptive pooling operation; For each of the text feature image and the image feature image, copy the corresponding feature images of three different scales along the spatial dimension to generate a feature image of the first preset size; Receive the original image and perform feature extraction on the original image using a lightweight network to determine the fusion weights of multi-scale features, acting on feature maps of different scales; For the text feature image and the image feature image, respectively, the corresponding feature images of three different scales and the corresponding feature image of the first preset size are weightedly fused based on the fusion weight to obtain a feature image of the second preset size to fuse the text features and the image features.
9. The method according to claim 1, characterized in that The method further comprises: After generating the adversarial image sample, a high-frequency component of the image sample is obtained and constrained.
10. A multi-target anti-attack device based on text description, characterized in that: include: A first extraction module is configured to extract text features from a target text based on a text encoder in a preset model, wherein the target text has a corresponding target category; a second extraction module, configured to extract image features from an original image based on an image encoder in the preset model, wherein the category of the original image is inconsistent with the target category; A fusion module, configured to fuse the text features and the image features based on the semantic fusion module in the preset model to obtain fused semantic features; A generation module, configured to generate an adversarial perturbation image from the semantic features based on a generator in the preset model; An adding module is used to add the adversarial perturbation image to the original image to generate a corresponding adversarial image sample.
Citation Information
Patent Citations
Attack resisting method based on diffusion model
CN119047536A
Image generation method and device, electronic equipment and medium
CN119205988A
High mobility adversarial sample generation method based on multi-modal model
CN119851055A
Aerial time-sensitive target identification method based on large model
CN119888534A
Speed-limiting traffic sign camouflage sample generation method and system based on StyleGAN2 and CLIP
CN119964122A