Image classification method and system for improving adversarial robustness
The anchor point-guided optimization method improves the adversarial robustness of the visual-language model under the zero-sample strategy, solves the problem of model adversarial attack sensitivity, and achieves improving the robustness and classification accuracy of the model without increasing computing resources.
Patent Information
- Application Number
- CN202510256091.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-08-01
AI Technical Summary
Existing vision-language models are sensitive to adversarial attacks, relying on additional training data leads to risk of overfitting, and high computing resource consumption, affecting their application and stability in the security-sensitive field.
By using anchor point guidance optimization method under the zero-sample strategy, image features are moved from source point features to target distribution, and feature alignment is combined with vision-language model to improve model adversity robustness.
Without relying on additional training data, the model's robustness against attacks is significantly improved, the calculation overhead is reduced, the classification accuracy of adversarial samples is enhanced, and the classification accuracy is better migration and adaptability are also achieved.
Smart Images

Figure CN120411985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more particularly, to an image classification method and system for enhancing adversarial robustness. Background Art
[0002] With the rapid development of artificial intelligence technology, the research and application of Vision-Language Models (VLMs) in the fields of computer vision and natural language processing have gradually become a hot topic. The core idea of the VLMs model is to establish the relationship between images and text data through joint training, enabling the model to understand and generate text information related to images, or to identify and generate images based on text descriptions. In recent years, vision-language models pre-trained on large-scale image-text data, such as CLIP (Contrastive Language-Image Pretraining), ALBEF (Align Before Fuse), and BLIP (Bootstrapping Language-Image Pretraining), have achieved remarkable results and shown excellent performance in tasks such as image classification, image generation, visual question answering, and cross-modal retrieval.
[0003] Despite the breakthrough progress made by these VLMs models in multiple fields, they still face an important challenge - adversarial attacks. Adversarial attacks refer to the intentional perturbation of input data by making small changes to interfere with the normal operation of machine learning models. Research has found that VLMs models are very sensitive to such attacks, and maliciously modified input data may cause serious deviations in the model's output or even render the model completely ineffective. This not only limits their application in security-sensitive fields but also poses challenges to their reliability and stability in real-world scenarios.
[0004] Currently, the defense methods against adversarial attacks can be mainly divided into two categories: adversarial training and adversarial purification. Adversarial training enhances the model's robustness to attacks by introducing adversarial samples during the model training process, enabling the model to learn the characteristics of adversarial samples. The adversarial purification method reduces the impact of perturbations by transforming the input samples affected by adversarial perturbations into "clean" samples. Common purification methods include using generative adversarial networks (GANs) or diffusion models to remove adversarial perturbations. However, these purification methods still have some problems. For example, the effect of removing adversarial perturbations is limited, and additional computational resources are usually required for the generation process, resulting in increased computational overhead. In terms of improving the adversarial robustness of VLMs models, adversarial training needs to be carried out on an additional dataset to enhance the robustness of the image encoder. Therefore, an additional dataset is required to fine-tune the model. When the number of samples in the dataset is small, adversarial fine-tuning may lead to overfitting of the model or a decrease in its generalization ability on other datasets. Summary of the Invention
[0005] To overcome the defect that visual-language models rely on additional training data in zero-shot adversarial robustness training and have the risk of overfitting, the present invention provides an image classification method and system for enhancing adversarial robustness.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] In a first aspect, the present invention provides an image classification method for enhancing adversarial robustness, including the following steps:
[0008] Obtain any image sample as the source image, and respectively perform feature extraction on the source image through the image encoder and text encoder in the visual-language model to obtain source features and text features;
[0009] Add noise to the source image to obtain a perturbed sample; perform feature extraction on the perturbed sample through the image encoder, and take the average value of the image features of the perturbed sample as the anchor feature;
[0010] Based on the anchor-guided optimization method, move the source features towards the target distribution to obtain optimized image features;
[0011] Input the optimized image features and text features into the visual-language model to obtain the predicted class result.
[0012] As a preferred solution, the adding noise to the source image includes the following steps:
[0013] Randomly sample noise from a Gaussian distribution and add it to the source point image to obtain a number of perturbed samples;
[0014] Alternatively, attack the source point image through the PGD attack method to generate perturbed samples.
[0015] As a preferred solution, the anchor-guided optimization method moves the source point features towards the target distribution, including the following steps:
[0016] Generate a linear path between the source point features and the anchor point features based on linear interpolation;
[0017] After optimizing the interpolation coefficient α on the linear path, perform an anchor-guided linear movement to obtain optimized image features; its expression is:
[0018] f movement =(1 - α)f source +αf anchor
[0019] where, f movement represents the optimized image features; f source represents the source point features, f anchor represents the anchor point features; α is a hyperparameter used to control interpolation.
[0020] As a preferred solution, the optimization of the interpolation coefficient a on the linear path includes the following steps:
[0021] Use one of grid search, Bayesian optimization, or random search for hyperparameter tuning, or search and tune by manually setting several hyperparameter combinations.
[0022] As a preferred solution, the vision-language model includes the CLIP model; then, the input of the optimized image features and text features into the vision-language model includes the following steps:
[0023] Calculate the cosine similarity between the optimized image features and text features through the CLIP model, and select the category corresponding to the text prompt with the highest similarity as the classification result of the image; its expression is:
[0024]
[0025] where, represents the text features of the nth category, N is the total number of categories; sim(·,·) is the cosine similarity function; x c represents the input image sample; p(y = n|x c ) represents the probability that the prediction sample x c belongs to the nth category y; τ is the temperature parameter.
[0026] As a preferred solution, the method further includes the following steps:
[0027] Obtain any image sample and perform preprocessing on it to obtain a source point image; the preprocessing includes:
[0028] Convert the image sample into a size format suitable for input to the vision-language model;
[0029] And / or, perform normalization processing on the image sample.
[0030] In a second aspect, the present invention provides an image classification system for enhancing adversarial robustness, which applies the image classification method proposed by the present invention. Among them, the system includes:
[0031] A feature extraction module, configured to respectively perform feature extraction on the input source point image through an image encoder and a text encoder in the vision-language model to obtain source point features and text features;
[0032] A perturbation module, configured to add noise to the source point image to obtain a perturbation sample; and, configured to perform feature extraction on the perturbation sample through the image encoder, and take the average value of the image features of the perturbation sample as the anchor point feature;
[0033] A feature optimization module, configured to move the source point features towards the target distribution based on the anchor point guidance optimization method to obtain optimized image features;
[0034] A classification module, configured to input the optimized image features and text features into the vision-language model to obtain a predicted class result.
[0035] In a third aspect, the present invention also proposes a device, including a memory and a processor, where computer-readable instructions are stored in the memory. Among them, when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the image classification method for enhancing adversarial robustness as described in the present invention.
[0036] In a fourth aspect, the present invention also proposes a storage medium, on which computer-readable instructions are stored. Among them, when the computer-readable instructions are executed by a processor, all or part of the steps of the image classification method for enhancing adversarial robustness as described in the present invention are implemented.
[0037] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0038] The present invention does not rely on any additional training data or fine-tuning, and only enhances the adversarial robustness of the model through a zero-shot strategy. In practical applications, there is no need to retrain the model or provide an additional training set, thereby reducing the computational overhead and time cost;
[0039] The present invention performs feature alignment based on an anchor-guided optimization method, which can effectively guide the features of image samples back to the feature space of clean samples, greatly improving the robustness of the model against adversarial attacks and enhancing the classification accuracy of adversarial samples.
[0040] The present invention has good transferability and adaptability. In addition to image classification tasks, it can also adapt to other vision-language tasks, has broad application potential, and can meet the requirements of adversarial robustness under various tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of an image classification method for improving adversarial robustness shown according to an embodiment of the present invention.
[0042] Figure 2 It is a framework flowchart of an image classification method for improving adversarial robustness shown according to an embodiment of the present invention.
[0043] Figure 3 It is an architecture diagram of an image classification system for improving adversarial robustness shown according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0045] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0046] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0047] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0048] Embodiment 1
[0049] This embodiment is applied to improve the adversarial robustness of the vision-language model in zero-shot tasks. Taking the image classification task as an example, it can also be applied to other tasks, such as: image recognition, image-text matching and retrieval, personalized image generation, visual question answering tasks, and so on.
[0050] This embodiment proposes an image classification method for improving adversarial robustness, as Figure 1 、 2 shown, which is the flow chart and framework flow chart of the image classification method for improving adversarial robustness in this embodiment.
[0051] The image classification method for improving adversarial robustness proposed in this embodiment includes the following steps:
[0052] S1. Obtain any image sample as the source image, and perform feature extraction on the source image through the image encoder and text encoder in the vision-language model respectively to obtain source features and text features;
[0053] S2. Add noise to the source image to obtain a perturbed sample; perform feature extraction on the perturbed sample through the image encoder, and take the average value of the image features of the perturbed sample as the anchor feature;
[0054] S3. Move the source features towards the target distribution based on the anchor-guided optimization method to obtain optimized image features;
[0055] S4. Input the optimized image features and text features into the vision-language model to obtain the predicted class result.
[0056] The source image in this embodiment can be an adversarial sample or a clean sample, and the same processing is performed on any input sample.
[0057] During the processing, first perform feature extraction through the vision-language model without distinguishing between adversarial samples and clean samples. Since the distribution information of clean samples is lacking in zero-shot tasks, this embodiment generates a perturbed sample by adding noise to the source image, takes the average value of the image features of the perturbed sample as the anchor feature, uses the anchor feature as the feature distribution area of the clean sample, and then moves the source features towards the target distribution through the anchor-guided optimization method to obtain image features closer to the clean sample, thereby improving the adversarial robustness of the model and its adaptability to different data distributions. Further, classify the image by inputting the optimized image features into the vision-language model, that is, complete the image classification task.
[0058] This embodiment does not rely on any additional training data or fine-tuning, but enhances the adversarial robustness of the model only through the zero-shot strategy. In practical applications, there is no need to retrain the model or provide additional training sets, thus reducing the computational overhead and time cost.
[0059] This embodiment performs feature alignment based on the anchor-guided optimization method, which can effectively guide the features of image samples back to the feature space of clean samples, greatly improving the robustness of the model against adversarial attacks and enhancing the classification accuracy of adversarial samples.
[0060] This embodiment has good transferability and adaptability. In addition to image classification tasks, it can also adapt to other vision-language tasks, has broad application potential, and can meet the adversarial robustness requirements under various tasks.
[0061] Exemplarily, the vision-language model in this embodiment can optionally use vision-language models pre-trained on large-scale image-text data, such as CLIP, ALBEF, BLIP, FLAVA, ALIGN, DeCLIP, LLaVA, etc.
[0062] In an optional embodiment, the adding noise to the source image includes the following steps: randomly sampling noise from a Gaussian distribution and adding it to the source image to obtain a number of perturbed samples; or, attacking the source image through the PGD attack method to generate perturbed samples.
[0063] Among them, the process of adding Gaussian noise is realized by the following formula:
[0064]
[0065] In the formula, represents the i-th perturbed sample after adding Gaussian noise; x source represents the source image;
[0066] δ~N(0,σ 2 ) represents Gaussian noise sampled from the Gaussian distribution N(0,σ 2 ).
[0067] In the specific implementation process, the appropriate noise type can be optionally selected according to the task requirements to better simulate the real scenario and enhance the robustness of the model, such as salt-and-pepper noise, Poisson noise, Rayleigh noise, gamma noise, etc.
[0068] The PGD (Projected Gradient Descent) attack method is a powerful adversarial attack technique used to generate adversarial examples. It iteratively updates the input sample within a legitimate perturbation range, causing it to gradually deviate from the original class on the decision boundary of the model. The PGD attack method combines gradient information and projection operations, enabling the generation of high-quality adversarial examples and is widely used to evaluate and enhance the robustness of models.
[0069] During the process of adding noise to the source image, through multiple iterations, the gradient with respect to the current adversarial example is calculated based on the loss function, and the adversarial example is updated along the gradient direction until a preset number of iterations is reached or a specific condition is met, at which point the iteration stops, obtaining the perturbed sample.
[0070] In this embodiment, choosing to add noise to the source image to generate the perturbed sample and cooperating with the anchor-guided optimization method to complete the adversarial training can effectively avoid the complex adversarial training process and significantly reduce the consumption of computing resources.
[0071] Furthermore, the perturbed sample is subjected to feature extraction through an image encoder, and the average value of the image features of the perturbed sample is taken as the anchor feature, expressed as:
[0072]
[0073] where, f anchor represents the anchor feature, k is the number of perturbed samples, and F(·) represents the image encoder.
[0074] In adversarial training, there are usually significant differences in the feature distributions between clean samples and adversarial examples, and this difference is called covariate shift. In this embodiment, selecting the average value of the image features of the perturbed sample as the anchor feature can serve as a reference point for distinguishing clean samples and adversarial examples to obtain the target distribution of clean samples. Cooperating with the anchor-guided optimization method to move the source features towards the target distribution can effectively guide the features of adversarial examples back to the feature space of clean samples, thereby enhancing the robustness of the model against adversarial attacks.
[0075] In an alternative embodiment, the moving of the source features towards the target distribution by the anchor-guided optimization method includes the following steps:
[0076] Generate a linear path between the source features and the anchor features based on linear interpolation;
[0077] After optimizing the interpolation coefficient α on the linear path, perform an anchor-guided linear movement to obtain the optimized image features; its expression is:
[0078] f movement =(1 - α)f source+αf anchor
[0079] Among them, f movement represents the optimized image feature; f source represents the source point feature, which is extracted from the source point image x source through an image encoder and is denoted as f source = F(x source ); f anchor represents the anchor point feature; α ∈ [0, 1] is a hyperparameter used to control interpolation. When α = 0, the movement result is the same as the source point feature; when α = 1, the movement result is the same as the anchor point feature.
[0080] The Anchor-based Optimization Method (AOM) is a technique for feature alignment and distribution adjustment, which is widely used especially in the fields of multimodal learning and data augmentation. Its core idea is to guide the source point feature to move towards the target distribution through anchor points (i.e., clean samples in the target distribution), so as to achieve feature alignment and optimization, and improve the stability and generalization ability of the model.
[0081] Furthermore, in an optional embodiment, optimizing the interpolation coefficient α on the linear path includes the following steps:
[0082] Use one of grid search, Bayesian optimization, or random search for hyperparameter tuning, or search and tune by manually setting several hyperparameter combinations.
[0083] In an optional embodiment, the vision-language model includes the CLIP model.
[0084] Among them, CLIP (Contrastive Language-Image Pre-training) is a multimodal pre-training model that maps images and texts to the same feature space through contrastive learning, so as to achieve cross-modal semantic alignment between images and texts. The CLIP model contains a text encoder and an image encoder, and is especially suitable for zero-shot image classification.
[0085] Furthermore, inputting the optimized image feature and text feature into the vision-language model includes the following steps:
[0086] Calculate the cosine similarity between the optimized image feature and the text feature through the CLIP model, and select the category corresponding to the text prompt with the highest similarity as the classification result of the image; its expression is:
[0087]
[0088] Among them, represents the text feature of the nth category, where N is the total number of categories; sim(·,·) is the cosine similarity function; x c represents the input image sample; p(y = n|x c ) represents the probability that the predicted sample x c belongs to the nth category y; τ is the temperature parameter used to adjust the cosine similarity distribution between the image feature and the text feature.
[0089] In an optional embodiment, the method further includes the following steps: obtaining any image sample and preprocessing it to obtain a source image; the preprocessing includes: converting the image sample into a size format suitable for input to the vision-language model; and / or, normalizing the image sample.
[0090] This embodiment can be directly applied to the enhancement of adversarial robustness in vision-language models (VLMs), especially suitable for zero-shot testing without any training data.
[0091] As an exemplary illustration, the following implements the image classification method for enhancing adversarial robustness proposed in this embodiment when facing adversarial samples in specific image classification.
[0092] First, input the image dataset, select 16 datasets including TinyImageNet, CIFAR-10, CIFAR-100, STL-10, ImageNet, Caltech256, Flowers102, etc. as the benchmark datasets for testing, covering multiple different tasks and fields, and then convert the input images into a format suitable for model input (224×224×3) through preprocessing.
[0093] Meanwhile, this embodiment follows the test paradigm of the previous method (APT) and supplements 11 additional benchmark test sets including UCF101.
[0094] This embodiment uses various adversarial attack methods such as PGD-10, PGD-100, AutoAttack, and CW to generate adversarial samples for verifying the robustness of the model.
[0095] The zero-shot test is conducted using the image classification method for enhancing adversarial robustness proposed in this embodiment. By evaluating the performance of the model on each dataset, its classification accuracy and adversarial robustness are calculated. Further, in this embodiment, CLIP ViT-B / 32 is used as the backbone model for experiments. Under adversarial attacks, PGD-10 attacks are used to generate adversarial samples, and tests are conducted on 16 datasets. For the setting of hyperparameters, we set α = 1.2 and the variance σ of Gaussian distribution noise = 0.18. By comparing the results of the method of the present invention with existing methods (such as TeCoA, PMG-AFT, etc.), the effect of enhancing zero-shot robustness of the present invention can be verified. As shown in Tables 1 and 2 below, the comparison of clean sample accuracies of the present invention and different adversarial training methods on 16 datasets; Tables 3 and 4 are the comparison of zero-shot adversarial sample robustness of the present invention and different adversarial training methods on 16 datasets.
[0096] Table 1 Comparison of clean sample accuracies - 1
[0097]
[0098] Table 2 Comparison of clean sample accuracies - 2
[0099]
[0100] Table 3 Comparison of zero-shot adversarial sample robustness - 1
[0101]
[0102]
[0103] Table 4 Comparison of zero-shot adversarial sample robustness - 2
[0104]
[0105] As can be seen from Tables 1 and 2 above, compared with other methods, the method of the present invention shows better accuracy on most datasets without any fine-tuning. The average accuracy of the present invention on 16 datasets is 62.24%, which is better than TeCoA (46.99%) and PMG-AFT (55.71%). Tables 3 and 4 show the results of zero-shot robustness tests under PGD-10 attacks. Compared with other methods, the average robustness of the method of the present invention on 16 datasets is 41.72%, which is significantly better than TeCoA (26.96%) and PMG-AFT (31.95%). The robustness of the method of the present invention against adversarial attacks has increased by 24.34%, demonstrating its ability to effectively enhance adversarial robustness.
[0106] Further, by comparing the average test results of applying the method of this embodiment and different adversarial training methods to 11 datasets under AutoAttack and CW attacks, as shown in Table 5 below.
[0107] Table 5 Other Adversarial Attack Tests
[0108]
[0109]
[0110] As can be seen from Table 5, under the AutoAttack attack, the robustness of the method of the present invention is 15.83% higher than that of the existing optimal method PMG-AFT, and it exceeds APT (non-zero sample method) by 9.41% under the CW attack.
[0111] Embodiment 2
[0112] This embodiment applies the image classification method proposed in Embodiment 1 to propose an image classification system for enhancing adversarial robustness, as Figure 3 shown, which is the architecture diagram of the image classification system for enhancing adversarial robustness of this embodiment.
[0113] The image classification system for enhancing adversarial robustness proposed in this embodiment includes:
[0114] A feature extraction module, configured to respectively perform feature extraction on the input source point image through the image encoder and the text encoder in the vision-language model to obtain source point features and text features;
[0115] A perturbation module, configured to add noise to the source point image to obtain a perturbed sample; and, configured to perform feature extraction on the perturbed sample through the image encoder, and take the average value of the image features of the perturbed sample as the anchor point feature;
[0116] A feature optimization module, configured to move the source point features towards the target distribution based on the anchor point guidance optimization method to obtain optimized image features;
[0117] A classification module, configured to input the optimized image features and text features into the vision-language model to obtain a predicted class result.
[0118] In an optional embodiment, this embodiment further includes a preprocessing module.
[0119] The preprocessing module serves as the input end of the system and is configured to preprocess the input image samples to generate source point images; the preprocessing includes: converting the image samples into a size format suitable for input to the vision-language model; and / or, performing normalization processing on the image samples.
[0120] The preprocessing module outputs the source point image after preprocessing to the feature extraction module and the perturbation module for further processing.
[0121] It can be understood that the system in this embodiment corresponds to the method in Embodiment 1 above. The optional items in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.
[0122] Embodiment 3
[0123] This embodiment provides a computer device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the image classification method for enhancing adversarial robustness proposed in Embodiment 1.
[0124] Embodiment 4
[0125] This embodiment provides a storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, all or part of the steps of the image classification method for enhancing adversarial robustness proposed in Embodiment 1 are implemented.
[0126] Exemplarily, the storage medium includes but is not limited to various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0127] Exemplarily, the instructions, programs, code sets, or instruction sets can be implemented using conventional programming languages.
[0128] Exemplarily, the processor includes but is not limited to smartphones, personal computers, servers, network devices, etc., and is used to execute all or part of the steps of the image classification method for enhancing adversarial robustness described in Embodiment 1.
[0129] Each embodiment in the present invention is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can refer to the partial descriptions of the method embodiments. The device embodiments described above are merely exemplary. The modules described as separate components may or may not be physically separated. When implementing the solution of the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] Obviously, the above-mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or alterations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. An image classification method for enhancing adversarial robustness, characterized in that, It includes the following steps: Obtain any image sample as the source point image, and respectively perform feature extraction on the source point image through the image encoder and the text encoder in the vision-language model to obtain the source point feature and the text feature; Add noise to the source point image to obtain a perturbed sample; perform feature extraction on the perturbed sample through the image encoder, and take the average value of the image features of the perturbed sample as the anchor point feature; Based on the anchor-guided optimization method, move the source point feature towards the target distribution to obtain the optimized image feature; Input the optimized image feature and the text feature into the vision-language model to obtain the predicted class result.
2. The method for image classification for enhancing adversarial robustness according to claim 1, wherein The adding noise to the source point image includes the following steps: Randomly sample noise from the Gaussian distribution and add it to the source point image to obtain a number of perturbed samples; Or, attack the source point image through the PGD attack method to generate a perturbed sample.
3. The image classification method for enhancing adversarial robustness according to claim 1, wherein The moving the source point feature towards the target distribution based on the anchor-guided optimization method includes the following steps: Generate a linear path between the source point feature and the anchor point feature based on linear interpolation; After optimizing the interpolation coefficient α on the linear path, perform the anchor-guided linear movement to obtain the optimized image feature; its expression is: f movement = (1 - α)f source + αf anchor Among them, f movement represents the optimized image feature; f source represents the source point feature, and f anchor represents the anchor point feature; α is a hyperparameter used to control interpolation.
4. The method for image classification for enhancing adversarial robustness according to claim 3, wherein, The optimizing the interpolation coefficient α on the linear path includes the following steps: Use one of grid search, Bayesian optimization or random search for hyperparameter tuning, or search and tune by manually setting several hyperparameter combinations.
5. The method for image classification for enhancing adversarial robustness according to claim 1, wherein The vision-language model includes the CLIP model; then, the inputting the optimized image feature and the text feature into the vision-language model includes the following steps: Calculate the cosine similarity between the optimized image feature and the text feature through the CLIP model, and select the class corresponding to the text prompt with the highest similarity as the classification result of the image; its expression is: Among them, represents the text feature of the nth category, where N is the total number of categories; sim(·,·) is the cosine similarity function; x c represents the input image sample; p(y = n|x c ) represents the probability that the predicted sample x c belongs to the nth category y; τ is the temperature parameter.
6. The method for image classification for enhancing adversarial robustness according to any one of claims 1 to 5, wherein The method further includes the following steps: Obtain any image sample and perform preprocessing on it to obtain the source point image; the preprocessing includes: Convert the image sample into a size format suitable for input to the vision-language model; And / or, perform normalization processing on the image sample.
7. An image classification system for enhancing adversarial robustness, which applies the method for enhancing adversarial robustness of image classification according to any one of claims 1 to 6, characterized in that, It includes: A feature extraction module, which is used to respectively perform feature extraction on the input source point image through the image encoder and the text encoder in the vision-language model to obtain the source point feature and the text feature; A perturbation module, which is used to add noise to the source point image to obtain a perturbed sample; and is used to perform feature extraction on the perturbed sample through the image encoder, and take the average value of the image features of the perturbed sample as the anchor point feature; A feature optimization module, which is used to move the source point feature towards the target distribution based on the anchor-guided optimization method to obtain the optimized image feature; A classification module, which is used to input the optimized image feature and the text feature into the vision-language model to obtain the predicted class result.
8. The image classification system for enhancing adversarial robustness according to claim 7, characterized in that, The system further includes: A preprocessing module, which is used to perform preprocessing on the input image sample to generate the source point image; the preprocessing includes: converting the image sample into a size format suitable for input to the vision-language model; and / or, performing normalization processing on the image sample.
9. A computer device, including a memory and a processor, where computer-readable instructions are stored in the memory, characterized in that, When the computer-readable instructions are executed by the processor, the processor is caused to execute all or part of the steps of the image classification method for enhancing adversarial robustness according to any one of claims 1 to 6.
10. A computer storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, all or part of the steps of the image classification method for enhancing adversarial robustness according to any one of claims 1 to 6 are implemented.